Browse with Rocco

We counted 264,224 agent turns. The causal labels are not ready to publish as causal findings.

A cross-framework workflow corpus has useful coverage, but regex labels, autonomous subturns, and missing turn-level bytes limit causal claims.

Photograph: NASA/JPL-Caltech

What exists in the checkout

The session summary contains 2,236 sessions and 264,224 assistant turns: 201,440 from Claude Code, 41,764 from Codex, 20,710 from OpenCode, and 310 from Gemini Antigravity. Recomputing the totals from multidimensional-session-summary.jsonl matches the report.

The full turn-level file is not present as data in this checkout. multidimensional-failure-corpus.jsonl is a 134-byte Git LFS pointer whose declared object size is 281,254,500 bytes. Any public analysis of turn text must fetch and verify that object first. Session-level aggregates can be inspected now; the underlying turn excerpts cannot be re-audited from this checkout.

Why 96.8% does not mean users are vague

The report assigns 255,674 turns to AUTONOMOUS_STEP_CONTINUATION, or 96.8% of all turns. The builder uses that label when human_prompt is empty. Empty prompts occur naturally when one user request produces many assistant tool subturns. The denominator is assistant turns, not distinct user instructions.

The report then classifies nearly all of those empty-prompt subturns as HUMAN_UNDERSPECIFIED_INPUT. This says more about transcript segmentation than human communication. A publishable estimate of user underspecification needs a user-turn denominator and session-aware pairing.

The labels are heuristic, not causal

The builder marks a prompt vague when it is short, lacks punctuation-like technical characters, or matches a small phrase list. It marks spiralling when an apology and a repeated edit co-occur, when a shell error is followed by a tool-name retry without a diagnostic tool, or when three edits touch a previously edited file. A retrieval miss is then inferred when a prompt contains a repository keyword and the same turn has a shell error or spiral flag.

Those are candidate signals. They do not establish that missing retrieval caused the failure. The output field causal_pattern overstates what the procedure measures. The clean public names are adjacent_pattern or heuristic_label until a human-coded sample provides precision, recall, and inter-rater agreement.

The current report also contains a denominator typo: the headline table totals 264,224, while a later sentence says 264,205. The distributions recompute to 264,224.

A publication-grade repair

First, reconstruct exchanges around user instructions rather than assistant subturns. Keep tool calls as events inside an exchange. Second, sample each proposed label by framework and register. Have two reviewers label the same blinded examples, report agreement, and adjudicate conflicts. Third, split temporal association from causal interpretation. A tool error followed by user anger is an observed sequence; "AI failure induced anger" requires a stronger annotation protocol.

Finally, publish privacy and exclusion rules before examples. This corpus comes from private work sessions. Aggregate counts do not authorize transcript publication. A useful research release can expose schemas, label definitions, synthetic fixtures, and validated aggregate results without releasing personal prompts or secrets.

Evidence ledger

  • Recomputed: session and assistant-turn totals from the available summary JSONL.

  • Read: the report and corpus-builder rules.

  • Unavailable: the full 281 MB turn-level Git LFS object in this checkout.

  • Heuristic only: vagueness, spiralling, retrieval-miss, and causal-pattern labels.

  • Privacy boundary: no private prompt excerpts are published here.

Evidence & limitations

This article references 2 supporting records. These files are not available as public downloads here. The article's findings should be read with its stated scope and limitations.

Ask about the methods or evidence

Keep reading

More research

All research

Contact · Toronto

ideas, built to launch

Describe the task, the systems involved, and what is getting in the way.