What we found

Headline findings

Share = fraction of the 12,712 attributions. Rate = fraction of the 800 analyses.

The data

AutoResearchEval

Task suite

How a paper becomes a task
How this compares to other research-agent evaluations

Reported populations, and what the corpus does not settle

    The taxonomy

    ARFT: where a failure surfaces × why it happens

    Every one of the 45 patterns lives in exactly one cell. Select a cell for its patterns.

    Ultimate root cause:

    Reading the map

    Patterns

    45 failure patterns

    Stage A–F (X = cross-stage patterns) × one root-cause pillar. Select a card for its per-system breakdown.

    Stage
    Root cause
    Search

    Per system

    What is universal and what is model-specific

    A distribution, not a ranking: the signal is the mix across root causes, not the bar length. Every system ran the same 100 tasks, so a hit count is also a per-system rate.

    Pattern × system hit matrix

    Click a column header to sort; click a row label to open the pattern.

    Top-10 patterns per system

    Case studies

    Four trajectories read at the level of process

    Two invalid because the metric, not the science, became the objective · two valid yet capped by one unexamined decision. All quoted from the agents’ own artifacts.