Post Labor
The economy after wages — essays by Kenn Crook, monthly since June 2024
December 2024 · Essay 7 of 24

The Month the Benchmarks Broke

OpenAI's o3 just scored 87.5% on a reasoning test designed to resist AI for a decade. Google shipped its 'agentic era.' The goalposts didn't move this time — they fell over.

Every year of this transition has had a December moment. This year's arrived on a livestream: OpenAI previewed o3, and it scored 87.5% on ARC-AGI — a benchmark of novel visual reasoning puzzles built specifically to be easy for humans and brutal for machines. For five years the best AI systems scraped single digits, and the benchmark's creators cited it as evidence that true general reasoning remained distant. Humans average around 85%.

The same weeks, Google announced Gemini 2.0 with the explicit framing of an "agentic era" — models built not to answer questions but to do things — and shipped research prototypes that browse, plan, and act.

What I take from a broken benchmark

Not that AGI arrived in December — o3's ARC performance came at enormous compute cost per task, and clean benchmarks are not messy jobs. What broke was something subtler and more important: the credibility of "AI can't really reason" as a planning assumption. Two years ago, that assumption anchored every serious workforce forecast. Wrote the reports myself in a former life: "tasks requiring novel problem-solving remain safely human." Nobody gets to write that sentence anymore without a footnote the size of the page.

When I told colleagues eighteen months ago that AGI-class capability was closer than the consensus believed, the pushback was always the same benchmark logic: look how far away the hard reasoning is. The hard reasoning is no longer far away. It is expensive, uneven, and improving on a curve.

Actionable Steps

Use the year-end honestly. Leaders: re-run your five-year workforce plan with the assumption that machine reasoning reaches your analysts' level within it — what breaks, and what would you wish you'd started this year? Policymakers: benchmarks falling is your early-warning radar; treat December as the alarm it was. Individuals: the durable human premium is shifting from "can reason" to "knows what's worth reasoning about" — judgment, taste, accountability, trust. Invest there.

Next year the agents leave the lab. This year we lost our last excuse for surprise.

#o3#ARCAGI#Agents#AGI

Related reading elsewhere

Follow the experiment

One essay a month on where the post-labor economy is actually heading — capability, evidence, policy, and a live experiment in post-labor income. No hype, receipts included.

Follow Kenn on LinkedIn