What we found
Headline findings
Share = fraction of the 12,712 attributions. Rate = fraction of the 800 analyses.
The data
AutoResearchEval
Task suite
How a paper becomes a task
How this compares to other research-agent evaluations
Reported populations, and what the corpus does not settle
The taxonomy
ARFT: where a failure surfaces × why it happens
Every one of the 45 patterns lives in exactly one cell. Select a cell for its patterns.
Ultimate root cause:
Reading the map
Patterns
45 failure patterns
Stage A–F (X = cross-stage patterns) × one root-cause pillar. Select a card for its per-system breakdown.
Per system
What is universal and what is model-specific
A distribution, not a ranking: the signal is the mix across root causes, not the bar length. Every system ran the same 100 tasks, so a hit count is also a per-system rate.
Pattern × system hit matrix
Click a column header to sort; click a row label to open the pattern.
Top-10 patterns per system
Case studies
Four trajectories read at the level of process
Two invalid because the metric, not the science, became the objective · two valid yet capped by one unexamined decision. All quoted from the agents’ own artifacts.