Dynetrix

AGENTFAILDB · RELIABILITY MAP

How multi-agent systems fail, measured locally.

750 recorded multi-agent runs on one consumer GPU, at zero API cost. This is a reliability map, not a ranking: the run count sits next to every rate.

crewaillama3.1:8b25083.2%68.0%80.0%74.0%100.0%94.0%0.11
langgraphllama3.1:8b25082.8%68.0%90.0%66.0%100.0%90.0%0.72
autogenllama3.1:8b25050.8%30.0%46.0%14.0%98.0%66.0%0.36

TASK SUCCESS BY FRAMEWORK · 750 RUNS

autogen

51%

127/250 runs · llama3.1:8b

crewai

83%

208/250 runs · llama3.1:8b

langgraph

83%

207/250 runs · llama3.1:8b

HOW OFTEN EACH FAILURE WAS FLAGGED · SHARE OF RUNS
Cascading hallucinationDelegation loopContext degradationConflicting outputsRole violationSilent failureResource exhaustion
autogen
crewai
langgraph

Rule-based detectors, run-level labels; all 295 annotations have source=rule_based. Directional, not human-validated. Hover a cell for the run count.

TASK SUCCESS BY TASK TYPE · SHARE OF RUNS THAT PASSED
CodeResearchData analysisDebate & reasoningPlanning
autogen
crewai
langgraph

Success tracks task type, not difficulty: correctness-critical tasks (code, data analysis) fail most.

AGENTFAILDB · THE BENCHMARK BEHIND THIS PAGE

Runs
750: 250 tasks × 3 frameworks (autogen, crewai, langgraph)
Tasks
5 types × 50: research, code generation, debate & reasoning, planning, data analysis. Easy, medium, hard and adversarial.
Model
llama3.1:8b
Hardware
One 6 GB consumer GPU (RTX 3050). Fully local, $0 API cost.
Recorded
2026-06-29 → 2026-06-30
Traces
Every message between agents, in order, per run.
Labels
Rule-based detectors, run-level labels; all 295 annotations have source=rule_based. Directional, not human-validated.
Judge
A local 8B LLM judge: directional, not ground truth. Two of the three replays above were judged passes although they failed.
Prior work
Complements MAST (Cemri et al., NeurIPS 2025), the human-validated 14-mode taxonomy of multi-agent failures. AgentFailDB doesn't replace it.

Every number on this page is computed from the benchmark's run data, and the replays are unedited traces. The three replays were picked because the failure is easy to see; the scorecard counts every run.