AGENTFAILDB · RELIABILITY MAP
How multi-agent systems fail, measured locally.
750 recorded multi-agent runs on one consumer GPU, at zero API cost. This is a reliability map, not a ranking: the run count sits next to every rate.
| crewaillama3.1:8b | 250 | 83.2% | 68.0% | 80.0% | 74.0% | 100.0% | 94.0% | 0.11 |
| langgraphllama3.1:8b | 250 | 82.8% | 68.0% | 90.0% | 66.0% | 100.0% | 90.0% | 0.72 |
| autogenllama3.1:8b | 250 | 50.8% | 30.0% | 46.0% | 14.0% | 98.0% | 66.0% | 0.36 |
TASK SUCCESS BY FRAMEWORK · 750 RUNS
autogen
51%
127/250 runs · llama3.1:8b
crewai
83%
208/250 runs · llama3.1:8b
langgraph
83%
207/250 runs · llama3.1:8b
| Cascading hallucination | Delegation loop | Context degradation | Conflicting outputs | Role violation | Silent failure | Resource exhaustion | |
|---|---|---|---|---|---|---|---|
| autogen | |||||||
| crewai | |||||||
| langgraph |
Rule-based detectors, run-level labels; all 295 annotations have source=rule_based. Directional, not human-validated. Hover a cell for the run count.
| Code | Research | Data analysis | Debate & reasoning | Planning | |
|---|---|---|---|---|---|
| autogen | |||||
| crewai | |||||
| langgraph |
Success tracks task type, not difficulty: correctness-critical tasks (code, data analysis) fail most.
AGENTFAILDB · THE BENCHMARK BEHIND THIS PAGE
- Runs
- 750: 250 tasks × 3 frameworks (autogen, crewai, langgraph)
- Tasks
- 5 types × 50: research, code generation, debate & reasoning, planning, data analysis. Easy, medium, hard and adversarial.
- Model
- llama3.1:8b
- Hardware
- One 6 GB consumer GPU (RTX 3050). Fully local, $0 API cost.
- Recorded
- 2026-06-29 → 2026-06-30
- Traces
- Every message between agents, in order, per run.
- Labels
- Rule-based detectors, run-level labels; all 295 annotations have source=rule_based. Directional, not human-validated.
- Judge
- A local 8B LLM judge: directional, not ground truth. Two of the three replays above were judged passes although they failed.
- Prior work
- Complements MAST (Cemri et al., NeurIPS 2025), the human-validated 14-mode taxonomy of multi-agent failures. AgentFailDB doesn't replace it.
Every number on this page is computed from the benchmark's run data, and the replays are unedited traces. The three replays were picked because the failure is easy to see; the scorecard counts every run.