AGI Is Already Here. We’re Just Pretending It Isn’t.
The AGI conversation is a magic trick. Watch the other hand.
A little over a year ago, just after OpenAI released o3, I was in a meeting at a university. The room was full of educators discussing AI and assessment. I am not an educator. I am usually the person whose voice the room tolerates as long as it does not have to be taken seriously. I said something like: Tyler Cowen called o3 AGI yesterday. It can use tools, read the web, work through scientific reasoning, and do things no model could touch a year earlier. Maybe we should stop calling these things chatbots. Polite pause. Conversation resumed.
I do not care much about the definitions of AGI. I never have. Whatever the label, what mattered to me about o3 was that the moment it could pull from external sources, use tools, reason through a problem step by step and move between domains, it had crossed into territory where it was simply better than most humans at many knowledge tasks I could think of. That is not dismissive of humans. It is arithmetic. We cannot hold what these systems hold in working memory. We cannot do across as many domains what these systems do in seconds. I said as much. The room moved on.
A year later the AGI debate is still running. The smartest people in the field keep gesturing at a threshold. Three years. Five years. By 2027. By 2030. Inside the labs the vocabulary itself has started to crack. The debate runs hot. The headlines follow. The threshold moves.
Watch the other hand.
While we have argued about the threshold, the tests we built to measure it have been falling at a pace that breaks the news cycle. Humanity’s Last Exam was published in Nature in January 2026 as a deliberately hard closed-form academic test: 2,500 questions across mathematics, science and the humanities, written by experts to sit at the frontier of human knowledge. At launch, o1 was around 8 percent. By May 2026, public provider-reported frontier scores are 44.4 percent without tools and 51.4 percent with search and code. Third-party self-reported preview boards put Claude Mythos at 64.7 percent. That caveat matters. Self-reported preview boards are not gospel. But the caveat does not rescue the old story. A benchmark built to stand at the end of the line became a moving target almost immediately. The benchmark’s own creators warn that as AI progress accelerates, benchmarks become quickly saturated and lose their utility as measurement tools.
GPQA Diamond, graduate-level science questions designed to be Google-proof, gets domain PhDs to roughly the high sixties or low seventies depending how you count errors. Gemini 3 Pro hit 91.9. Gemini 3.1 Pro is reported at 94.3. ARC-AGI-2, the benchmark explicitly named for artificial general intelligence and designed around novel abstract reasoning puzzles, now has a verified frontier score of 77.1. The prize threshold is 85. We are not discussing trivia tests anymore. We are discussing the tests built to expose what the machines were supposed not to be able to do.
FrontierMath is messier, and the mess is part of the story. It was built from unpublished, extremely difficult mathematical problems, the kind that can take specialist mathematicians hours or days. At release, state-of-the-art models could barely touch it. By early 2026, public models were over 40 percent on the standard tiers and over 30 percent on Tier 4, the research-level set. One of FrontierMath’s own contributors told IEEE Spectrum that the benchmark would probably saturate within two years and might do it faster. Now Epoch is auditing the dataset because an AI-assisted review flagged fatal errors in about a third of the problems. Even the benchmark failure is part of the point. The tests are not only being beaten. They are being destabilised by the systems they were built to measure.
The obvious reply is that the benchmarks are flawed. They are. Some are contaminated, some are Goodharted, some are simply badly built. In April 2026, Berkeley researchers showed that eight major agent benchmarks could be driven to near-perfect scores without solving the underlying tasks at all. That is not a counterargument to what I am saying. It is the same argument from the other side. The tests are no longer stable instruments. They are being saturated, hacked, audited, patched and discarded almost as soon as they matter. A broken thermometer does not prove there is no fever. It proves we need to stop pretending the thermometer was the phenomenon.
Which is why the non-benchmark evidence matters more.
The Nobel committees have already touched this. In October 2024 the Chemistry Prize went to David Baker, Demis Hassabis and John Jumper for computational protein design and protein structure prediction, AlphaFold having been used by millions of researchers across more than 190 countries. The same week the Physics Prize went to John Hopfield and Geoffrey Hinton for foundational work on artificial neural networks, the intellectual machinery underneath the chatbot you used yesterday. Two AI-shaped Nobel stories in a single week. No one had a category for that, so they pretended it was two separate stories.
In July 2025 two different AI systems hit gold-medal standard at the International Mathematical Olympiad, both inside the standard competition window, both in natural language end-to-end. Five problems out of six. OpenAI’s researcher Alexander Wei described what the model had done as “intricate, watertight arguments at the level of human mathematicians”. He was not reaching for the press release. He was describing the result.
This was not supposed to happen yet. It happened anyway. It happened while we were waiting for AGI.
Now look at what just shipped in the last six months.
In March 2025 a small research group called METR published a benchmark that asked a simple question. How long a task, measured in human expert labour, can an AI agent complete with 50 percent reliability? Their answer was that time horizons had been doubling roughly every seven months over the previous six years. That was already startling. Extrapolate the line and the systems reach month-long projects around 2030.
By 2026 the curve had bent. METR’s updated Time Horizon 1.1 work puts the post-2023 doubling time at 131 days, and the post-2024 estimate at 88.6 days. Their live page now has to warn that measurements above 16 hours are unreliable with the current task suite. This does not mean an AI literally sits there doing sixteen uninterrupted hours of work. METR is explicit about that. It means the system can complete tasks that would take a human expert that long, at the specified reliability level. That distinction matters. It makes the result more useful, not less.
These are not chatbots answering questions. These are agents. They read your codebase. Plan a multi-step change. Run tests. Fix failures. Hit a bug they cannot solve. Switch tactic. Try again. Commit the result. Open the pull request. Wait for review. Iterate.
Twelve months ago this was a research curiosity. Today Claude Code, Codex, Cursor, Devin and Replit Agent are production tools, used by serious engineering teams on real codebases. On the official SWE-bench Verified leaderboard, Claude Opus 4.6 sits at 75.6 percent and GPT-5-2 Codex at 72.8 percent. Provider-reported single-attempt tables put top models clustered around 80 percent. SWE-bench Pro, the harder and cleaner version, cuts the numbers down sharply, but even there the current public leaderboard has GPT-5.4 at 59.1, Muse Spark at 55 and Claude Opus 4.6 at 51.9. That is not toy performance. That is not a chatbot. That is a machine resolving real software tasks at a level that would have sounded deranged three years ago.
Anthropic overtook OpenAI in business adoption in April 2026 on Ramp’s data, 34.4 percent to 32.3 percent, the first time that had happened. Last month Codex on Mac gained background computer use, the ability to operate native applications by seeing, clicking and typing with its own cursor. Multiple Codex agents can now work on a Mac in parallel without interfering with your own work.
Andrej Karpathy, an OpenAI co-founder who has been more sceptical than most, pushed back hard on the idea that 2025 was “the year of agents”. His reframe was that this is the decade of agents. He is not arguing that the technology is overrated. He is arguing that we are early. The reliability gap, the march of nines, where each additional nine of reliability takes the same amount of work as the last, is what shifts agents from impressive demo to deployed economic infrastructure that quietly does most of what most companies pay people to do.
We are at the start of that decade. Not the end. And the curve is already producing expert-hour task horizons that are breaking the benchmark built to measure them.
This is the hand that is not waving. While the AGI debate consumed the headlines, the systems that turn cognitive work into infrastructure arrived and started shipping production code.
So forget the threshold. Ask why anybody thought the line was the important thing.
Stanford’s AI Index reports the cost of running a model at GPT-3.5-level performance fell more than 280-fold between November 2022 and October 2024. From twenty dollars per million tokens to seven cents. If diesel fell at the same rate, a one-pound-fifty litre would cost about half a penny. A million GPT-3.5-level tokens now costs less than a dropped coin.
ChatGPT had 900 million weekly active users in February 2026. Even allowing for duplicate accounts, bots, children and every other measurement caveat, that is planetary scale. The product handles about 2.5 billion prompts a day. Google still does roughly five trillion searches a year, about 13.7 billion a day, which means ChatGPT is already operating at around a sixth to a fifth of Google-scale query volume. For a service that did not exist three and a half years ago, that is obscene.
Anthropic’s own analysis found that the tasks Claude covers require, on average, 14.4 years of education. The economy-wide average is 13.2. The work people are already asking AI to absorb is harder, on average, than the work in the typical job. That is not a forecast. That is a measurement of what is already being used.
Stanford’s Digital Economy Lab found a 16 percent relative employment decline for 22 to 25 year olds in the most AI-exposed occupations since the widespread adoption of generative AI, even after controlling for firm-level shocks. Anthropic’s March 2026 labour-market paper found a 14 percent fall in the job-finding rate for the same age group in exposed occupations, barely statistically significant, with no comparable decrease for workers over 25. Two papers, two methodologies, same direction. Not a clean causal apocalypse. Something worse for the people who want comfort: an early signal.
Manufacturing efficiencies made physical goods abundant. That is why you own thirty shirts and your great-grandfather owned three. Whatever AI ends up being called, it is doing the same thing to thought. Set aside whether the thing in the box is intelligent. Ask what happens to a civilisation when most of the work it organised itself around becomes cheap.
I suspect the AGI frame is comfortable because it puts the disruption in the future. I do not know that. It would be stupid to pretend I can see inside the heads of institutions, executives, universities or policymakers. I could be completely wrong. But the effect of the frame is comfortable even if nobody consciously chose comfort. Once it is real AGI, we will deal with it. Until then, this is all just chatbots. The framing protects three things at once: the institutions whose pricing power depends on cognitive labour being scarce, the careers built on producing the kind of legible output a pattern machine can now produce in seconds, and those whose product is the credential that says someone can do the thing the machine can now do.
Some of the people building these systems have already given up on the term. Dario Amodei wrote in late 2024 that he disliked “AGI” as a frame and preferred “powerful AI”, his description being “a country of geniuses in a datacenter”. Karen Hao’s Empire of AI argues that AGI has functioned as an organising myth, a quasi-religious story for fundraising, policy and power. OpenAI’s commercial arrangement with Microsoft reportedly defined AGI, for contractual purposes, as a system capable of generating at least one hundred billion dollars in profit. That is not a definition of intelligence. It is a definition of revenue.
The honest counterargument is that the systems still fail in ways humans do not. Apple’s research team published “The Illusion of Thinking” in June 2025, showing that large reasoning models hit a complexity wall, fail to use explicit algorithms reliably, and collapse on harder puzzle settings. The Dell’Acqua and Mollick jagged-frontier study found consultants did faster and better work with AI on tasks inside the frontier, then became 19 percentage points less likely to get the right answer on a task outside it. Bloom and colleagues at the NBER surveyed executives across four countries and found that 89 percent reported no measurable productivity impact from AI over three years. SWE-bench Pro was built because the easier coding benchmarks had become too flattering. Verified overstates what the systems can do in production. Pro shows the frontier is still jagged.
All of this is true. None of it touches the thesis. The thesis was never that these systems are minds or that they are reliable. The thesis is that they are doing economically valuable cognitive work, cheaply, at scale, today, while we wait for permission to admit it. Even on SWE-bench Pro, the harder public software-engineering benchmark, leading models now clear 50 percent. That category of work did not exist as a serious possibility 36 months ago.
I do not think these systems are conscious or that they are minds. I also know enough to know that I do not know. I may be wrong. I cannot inspect the inside of the thing, and neither can most people talking with certainty. What I can say is simpler: they hold more in their working memory than I can hold in mine, work faster than I can work, cross more domains than I cover, and produce competent output on tasks that two years ago I would have paid a trained professional to produce. Call that what you like. The arithmetic is the arithmetic.
I have written before, at length, about why the human matters. Not because we are uniquely special, and not because there is some essence in us that the machines can never touch. We matter because we are the part of the equation that has to live with what we build, and the part that decides what is worth doing in the first place. None of that argument requires the technology to be inferior to us. The argument is stronger, not weaker, when we name what these systems actually are.
What I will not do anymore is pretend that the AGI debate matters in the way it claims to matter. The next definition will be either weaker than what is already deployed, in which case the line has already been crossed and we will quietly stop talking about it, or stronger than what is already deployed, in which case we will keep using the technology we have while we wait for the version that does not yet exist. Both options leave us in the same place. The economy is being rearranged underneath people whose vocabulary cannot quite name what is rearranging it.
The most disruptive technology in a generation is here. It runs on your phone. It writes the emails you do not have time to write. It is shipping production code on real codebases, running expert-hour tasks with increasing reliability, and the doubling curve is shortening, not lengthening. The intellectual machinery behind it has already won Nobel Prizes and reached gold-medal standard at the world’s most elite high-school mathematics competition. The people who built it argue about when it will become real. They are looking at the wrong hand.
So are you.
I am going to keep writing this. Not because I think I can change where we are. I cannot. The capital is committed. The infrastructure is being built. The models are getting more capable every quarter. The institutions whose job it was to think clearly about this are still mostly arguing about whether to allow it in undergraduate essays.
I will keep writing it because we have to wake up, and we have to do it quickly. The thing we have to wake up to is already in our hands. The misconception that these systems are inferior to us is the most expensive belief in the room. It is the belief that lets us keep training for jobs ‘in part’ the machines now do better, building institutions around scarcities that no longer exist, and waiting for a future arrival that already happened.
That was the trick: we kept watching the hand labelled AGI while the other hand made cognitive labour cheap.
We just have to deal with it.



I think people are overestimating the value of the benchmarks themselves - what seems to actually be happening here is that the answers are either tending to be in the training data, or the tests themselves are flawed and easily manipulated.
https://www.linkedin.com/pulse/how-we-broke-top-ai-agent-benchmarks-dawn-song-n6qrc
Three threads here that feel like they want more room than a single essay can give them, and they form a natural sequence.
The employment signal for the 22-25 cohort is the most time-sensitive. You flag it as an early warning and move on, but the asymmetry is the story — the effect doesn't show up for over-25s, and that gap is doing a lot of unexplained work. Is it entry-level exposure, accumulated institutional insulation, or an effect that's present in older workers but slower and harder to isolate? That question has a closing window. In eighteen months it's either confirmed pattern or interesting anomaly, and the moment for writing about it as early signal will have passed.
The frame section is where I think you have something nobody else is developing with this precision. The Microsoft contractual definition is the detail that unlocks it — AGI defined as a revenue threshold means the frame isn't just intellectually convenient, it's legally and financially load-bearing. Everything before the threshold is, by contractual definition, not AGI regardless of capability. That's not epistemic cowardice, it's structured incentive. Universities, professional bodies, regulators without frameworks, governments without tax policy — each has equivalent structural reasons to maintain the frame. It persists not because people are stupid or afraid but because it's doing real work for real institutions with real money attached. That deserves its own piece.
The cost collapse is where the historical weight is. The thirty shirts versus three is where you stop, but 1910 is where it gets interesting — early enough to see the shape, late enough that it's irreversible. When the cost of producing a legal brief or diagnostic opinion drops two orders of magnitude in two years, what happens to the profession built around pricing that production? Not eventually. In the next five years. Pick one profession and follow it all the way down and you have the piece that makes the abstract concrete.
Three essays. The signal, the frame, the collapse. Each one earns the next.