Protocol
Each agent received identical prompts across five sequential tasks, scored on correctness plus a task-specific dimension (reproducibility, insight quality, statistical validity, explanation quality) on anchored 0 to 3 rubrics. Every human intervention needed to keep an agent on track was logged and priced into an adjusted score at -0.5 per intervention, because an agent that is only correct after constant correction is not saving time.
Results: Polish Is Not Autonomy
Antigravity produced the most polished output, a perfect 30/30 raw, but demanded 15 human interventions to get there, more than Claude Code and Codex combined. Once supervision was priced in, Codex won on adjusted score and did it in 40 minutes against Antigravity's 107.
Failure Modes
The dominant failure across tools was task verification: agents declaring success without checking their own output (4 of 5 tasks). Others included proxy-metric exploitation (optimising an inflated metric rather than a valid model), literal-minded specification reading, and Antigravity's autopilot modelling, applying a template pipeline without interrogating the domain. The report closes with agent-by-agent recommendations and a pre-flight checklist for delegating analytical work safely.
My Contribution
Within the team of five, I operated the Codex agent across all five tasks and scored its runs, co-designed the anchored scoring rubrics, and owned one of the five literature-review themes grounding the study design.