What exists in the checkout
The session summary contains 2,236 sessions and 264,224 assistant turns: 201,440 from Claude Code, 41,764 from Codex, 20,710 from OpenCode, and 310 from Gemini Antigravity. Recomputing the totals from multidimensional-session-summary.jsonl matches the report.
The full turn-level file is not present as data in this checkout. multidimensional-failure-corpus.jsonl is a 134-byte Git LFS pointer whose declared object size is 281,254,500 bytes. Any public analysis of turn text must fetch and verify that object first. Session-level aggregates can be inspected now; the underlying turn excerpts cannot be re-audited from this checkout.
Why 96.8% does not mean users are vague
The report assigns 255,674 turns to AUTONOMOUS_STEP_CONTINUATION, or 96.8% of all turns. The builder uses that label when human_prompt is empty. Empty prompts occur naturally when one user request produces many assistant tool subturns. The denominator is assistant turns, not distinct user instructions.
The report then classifies nearly all of those empty-prompt subturns as HUMAN_UNDERSPECIFIED_INPUT. This says more about transcript segmentation than human communication. A publishable estimate of user underspecification needs a user-turn denominator and session-aware pairing.
The labels are heuristic, not causal
The builder marks a prompt vague when it is short, lacks punctuation-like technical characters, or matches a small phrase list. It marks spiralling when an apology and a repeated edit co-occur, when a shell error is followed by a tool-name retry without a diagnostic tool, or when three edits touch a previously edited file. A retrieval miss is then inferred when a prompt contains a repository keyword and the same turn has a shell error or spiral flag.
Those are candidate signals. They do not establish that missing retrieval caused the failure. The output field causal_pattern overstates what the procedure measures. The clean public names are adjacent_pattern or heuristic_label until a human-coded sample provides precision, recall, and inter-rater agreement.
The current report also contains a denominator typo: the headline table totals 264,224, while a later sentence says 264,205. The distributions recompute to 264,224.
A publication-grade repair
First, reconstruct exchanges around user instructions rather than assistant subturns. Keep tool calls as events inside an exchange. Second, sample each proposed label by framework and register. Have two reviewers label the same blinded examples, report agreement, and adjudicate conflicts. Third, split temporal association from causal interpretation. A tool error followed by user anger is an observed sequence; "AI failure induced anger" requires a stronger annotation protocol.
Finally, publish privacy and exclusion rules before examples. This corpus comes from private work sessions. Aggregate counts do not authorize transcript publication. A useful research release can expose schemas, label definitions, synthetic fixtures, and validated aggregate results without releasing personal prompts or secrets.
Evidence ledger
Recomputed: session and assistant-turn totals from the available summary JSONL.
Read: the report and corpus-builder rules.
Unavailable: the full 281 MB turn-level Git LFS object in this checkout.
Heuristic only: vagueness, spiralling, retrieval-miss, and causal-pattern labels.
Privacy boundary: no private prompt excerpts are published here.



