AI HOT (Curated Pool)· aihot-apiZH00:00 · 07·21
→AgentDebugX: An open-source toolkit for failure observability, attribution, and recovery in LLM agents
LLM agent failures are hard to debug because the error step is rarely the root cause. AgentDebugX wraps debugging into a Detect-Attribute-Recover-Rerun loop, with DeepDebug doing multi-turn root-cause diagnosis via global trajectory understanding and cross-examination. On the Who and When benchmark, strict attribution accuracy hits 28.8% on qwen3.5-9b, 7.1 points above the strongest single-pass baseline. On GAIA, a single rerun fixes 13 of 73 failed tasks, lifting overall accuracy from 55.8% to 63.6%. The toolkit ships as a Python library, CLI, web console, and installable agentic skill, plus an opt-in Error Hub for sharing scrubbed failure-diagnosis-repair bundles.
#Kunlun Zhu#Xuyan Ye#Zhiguang Han
editor take
AgentDebugX debugs agent failures with multi-turn root-cause diagnosis, fixing 13 of 73 failed GAIA tasks in one rerun and lifting accuracy from 55.8% to 63.6%.
HKR breakdown
hook ✓knowledge ✓resonance —