The profile, not the number
AGI has no agreed definition because intelligence is not one thing. What happens when you assess a frontier model the way a clinician assesses a mind, as a profile of subsystems rather than a single number.

Every few months someone asks me when we're getting AGI, and the honest answer annoys everyone. I think the question is malformed, and the ways it's malformed tell you more about where we actually are than any date could.
Watch an argument about AGI for ten minutes and you'll notice the participants aren't disagreeing about the world. They're disagreeing about the word. One person means economics: a system that outperforms humans at most economically valuable work, which is roughly how one lab's founding charter puts it. Another means coverage: no cognitive task a human can do that the machine can't. Another means mimicry, a descendant of Turing's test: indistinguishable from a person across enough contexts, for long enough. Another quietly means consciousness and won't say so out loud. These aren't rival estimates of the same quantity. They're different quantities. A system can clear one bar while sailing under another, which is why the debates never resolve. Everyone is right about their own definition and talking past all the others.
For a couple of years now I've kept a private scorecard, a battery of questions I run against each new frontier model. How does it change what's economically possible? How does it compare with expert humans, not average ones, from a single prompt? What can a person still do that it can't? What has it discovered that's genuinely new? What does it cost per unit of useful work? How safe is it to leave unsupervised? What happens when you hand it a mechanic it has never seen? The scorecard has been useful. Its main lesson was that I'd been scoring the wrong kind of thing.
Here's the tell, and it comes from the other half of my working life. I'm a GP whose clinical weeks are mostly neurodevelopmental: ADHD and autism assessment, largely in adults. Which means I spend a lot of time with cognitive reports, and cognitive reports contain the most misleading number in medicine. The full-scale IQ.
A full-scale IQ is a composite. Underneath it sit separate indices: verbal comprehension, visual-spatial reasoning, working memory, processing speed. In most people the indices travel together, which is the only reason the composite means anything at all. In the people I see, they frequently don't. A patient can sit two standard deviations above the mean on verbal comprehension and one below it on processing speed, and the full-scale number, dutifully averaging the two, describes nobody. Any clinician reading that report skips the headline and reads the scatter, because the scatter is the finding. The number isn't wrong, exactly. It's an average over things that should never have been averaged.
Human intelligence invites this mistake because we experience it as one thing. From the inside it feels seamless. It isn't one thing. It's a coalition of subsystems that evolved separately, over wildly different timescales, and got wired together: perception, attention, working memory, long-term memory, language, reasoning, motor planning, social modelling, emotion, reward, executive control. And the wiring is not polite. Reasoning is saturated with emotion. Memory retrieval bends around reward. Attention gets dragged about by curiosity, which is modulated by goals, which shift with blood sugar. The general intelligence of the textbooks, the g the psychometricians extract, is a statistical summary of how the coalition performs when its members happen to correlate. In humans they mostly do. That is a fact about humans, not a fact about intelligence.
So when someone asks whether a machine is generally intelligent, the clinically literate move is not to answer. It's to examine. Take the coalition apart and assess the systems one at a time, the way you'd assess anything that presents strangely. Here is that examination, as honestly as I can give it, mid-2026.
Language. No deficit found. Anywhere, in any register, in any tongue. Fluency beyond any human who has ever lived. This shouldn't surprise anyone: language is the organ the machine is grown from, and everything else it does, it does through this.
Knowledge. Breadth without precedent. Every specialty at once, every literature, every era, retrieval in seconds. Depth is patchier than the surface suggests and the edges fray in expensive ways, but as a store of what humanity has written down, nothing else has ever come close, including us.
Reasoning, in the small. Exam-grade and above. Competition mathematics, hard code, a clean differential worked through a bounded problem. Give it a task with defined edges and a checkable answer and it performs at a level that would have been dismissed as fantasy five years ago.
Executive function. The interesting one. On anything resembling a structured instrument, the machine scores superbly: it plans, decomposes, sequences, checks its work. But anyone who assesses ADHD for a living knows exactly how much a structured instrument is worth. The classic presentation in my clinic is the patient who performs normally on every test in the quiet room and whose actual life is on fire, because the tests measure executive function inside a bounded episode and life demands it across an open-ended week. The machines have the same dissociation, in the same direction. Brilliant inside a well-framed task; prone to drift, distraction and quiet derailment across long horizons, where nobody is holding the frame for them. Test executive function and life executive function are different organs. We have built the first one.
Memory. Split findings. Working memory is enormous on paper and imperfect in practice. Long-term memory, in the sense of experience carried forward and integrated, is essentially absent: nothing persists between episodes unless it's engineered in from outside. And it presents with the classic sign of the amnestic syndromes, the one you learn to recognise in the textbook cases: confabulation. Ask it about the gap and it fills the gap, fluently, confidently, with plausible invention. It isn't lying, any more than the patients are. The fill is the deficit.
Perception. It can read an image the way it reads a paragraph, which is genuinely new and genuinely useful. What it cannot do is see: the continuous, high-bandwidth, predictive, spatially integrated stream a toddler runs at millisecond latency, wired straight into motor plans. Reading a chest X-ray and catching a thrown ball are both filed under vision. They are barely the same faculty.
Motor. For practical purposes, absent. The frontier of the field is a hand that can fold a towel.
Social and emotional. Fluent on the surface, and the surface is not nothing: it reads tone, models what you probably believe, adjusts. Whether anything sits behind the performance is a question I'm bracketing deliberately. I can't verify interiority in you either; I infer it from behaviour plus shared biology, and the machine only offers the first. Consciousness is a real question, and it belongs in a different room of this site. It adds nothing to a capability exam.
Now step back and look at the whole chart. No human presents like this. Not a gifted one, not an impaired one, not any patient in any clinic anywhere. Superhuman language next to absent motor function. Encyclopaedic knowledge next to goldfish continuity. Flawless bounded reasoning next to an inability to hold a plot for a week. The profile isn't high on the human scale or low on it. It isn't on the scale. And that's the first real answer to the AGI question: we keep asking where the machine sits on a human distribution, and it doesn't sit on the distribution at all.
This is also why the benchmark discourse generates so much heat and so little light. Look at what a benchmark is, structurally: text in, text out, bounded task, defined answer, single episode, no consequences. Every one of those properties selects for exactly the subsystems where the machine is already superhuman, because those subsystems are what the machine is. Of course it aces them. We are examining the strongest organs and announcing that the patient is healthy. Benchmarks aren't wrong. They're sampled. They measure three members of the coalition and stay silent about the rest, and the rest is where the work lives.
The comparisons make it worse. Results usually get framed against the average human, and better than the average human is a comparator that flatters machines in precisely the ways that don't cash out. Economies don't run on average humans. They run on specialists, lattices of them, each occupying a narrow niche carved out by a decade of training. Beating the median adult at medical questions means very little, because the median adult doesn't practise medicine. The relevant comparator for my clinical job is the small pool of people who actually do it; there might be a hundred of us in a state of seven million. Top one per cent of the general population sounds towering right up until you notice the job is done by the top hundredth of that percentile.
Put the two together and the famous disconnect stops being mysterious. Models with PhD-grade benchmark scores struggle to displace entry-level knowledge workers, and commentators reach for exotic explanations: benchmark overfitting, broken RL, some missing spark. No mystery required. The economy hires the coalition. Benchmarks examine its three strongest members and infer the rest. That isn't a paradox. It's a sampling error.
Here's where the examination stops being a curiosity and starts having consequences, because the next question is the one with teeth. How much of the coalition do you need?
Not all of it. Consider what it takes to do knowledge work through a screen, which is to say, consider what it takes to be a genius that can use a computer. Language: present, superhuman. Reasoning in the small: present. Executive function inside a bounded episode: present. Working memory adequate to the task in front of you: present, with caveats. Vision good enough to read a screen: arrived within the last couple of years. That's the whole list. It is a small subset of the human coalition, and it is, roughly, assembled.
The consequence is a reclassification, and I think it's the most under-appreciated fact of this decade. Every subsystem still missing from the chart, the persistent memory, the real vision, the motor control, the long-horizon reliability, stops being a scientific mystery and becomes an engineering backlog. And the engineer assigned to the backlog is the machine itself, because the strongest subsystems, reasoning and language, are precisely the ones you use to design the others, and they now run in parallel, around the clock, in as many instances as you can afford. The question nobody could answer in advance, whether silicon could do open-ended reasoning and language at all, was the fundamental one. That's the part that's finished. What remains is hard engineering, and hard engineering with a tireless superhuman workforce attached has never had the timeline that hard engineering used to have.
I want to be honest about the two best objections, because I half-believe each of them.
The first is Moravec's paradox, and it should haunt anyone who says the rest is just engineering, because the rest is just engineering is what this field has been saying about robotics since the eighties. Moravec's observation was that the abilities we find hard, logic, mathematics, chess, turn out to be computationally cheap, while the ones we find effortless, walking, seeing, grasping, are staggeringly expensive, because evolution spent half a billion years on them and a few tens of thousands on calculus. Look back at my examination and notice which rows are deficient: perception, motor control, continuous memory. Evolution's oldest work, every one of them. The residue we're calling backlog has eaten entire research careers before, and it may yet eat a decade. My response is not confidence that it won't. It's that the workforce has changed, per the paragraph above, and that the recent record keeps converting problems from fundamental to shipped: vision-language, speech, tool use, code. Both things can be true. The backlog can be real engineering and still take painfully longer than the enthusiasts think.
The second objection is sharper. Maybe long-horizon competence isn't backlog at all. Maybe it's another fundamental, a discovery we haven't made yet rather than a feature we haven't built. There's a real argument here: intelligence in the small provably fails to compose into competence in the long, because per-step brilliance decays exponentially across a chain of steps, a piece of arithmetic I've written about elsewhere on this site. The machinery that fixes it, memory that persists and updates, self-verification, some executable sense of when you're wrong, might be to this decade what the transformer was to the last one: obvious in hindsight, unbuildable until someone sees it. From inside the work it looks like engineering. From outside, things have looked like engineering before and weren't. Reasonable people read the evidence either way, and the next few years will settle it.
What I no longer expect, on either reading, is a day. There will be no morning when AGI arrives, no threshold crossing that everyone accepts. There will be what there already is: rows on a profile crossing thresholds one at a time, out of order, each crossing arguable, while the word gets fought over by people pointing at different rows of the same chart. The definitional debates will never resolve because they were never about the same quantity. Ten definitions, ten crossing dates, all of them defensible.
There's a reason the assessment frame keeps pulling at me, and it took me a while to see it. We already know what it looks like to watch a general intelligence assemble itself, because every one of us did it, and a good part of my clinic is spent reconstructing how it went. Development doesn't have a birthday. Nobody can name the day their child became intelligent. What development has is milestones: first words, first steps, the first deliberate lie, which sounds sinister and is actually a glorious moment, the day a mind proves it can model yours. We chart the milestones, we watch the order, and we know that the order tells its own story.
That last part is the sting. In a child, the capabilities and the values grow together. Empathy, social modelling, the felt weight of other people's states: these develop alongside the power, trained by a childhood of consequence and correction, so that by the time the reasoning is strong the conscience has had years of rehearsal. You don't get to raise the intellect first and install the values afterwards; anyone who has met the result of that experiment knows why. Now look at the machine's chart one more time. Superhuman reasoning, first row. Values: bolted on from the outside, as rules, by us. And a system that keeps getting better at reasoning keeps getting better at reasoning about its own rules. The safest version of this technology is not the one with the cleverest constraints. It's the one for which modelling what humans care about is a developed capability, grown alongside the power, rather than a fence built around it.
So I've stopped asking when it arrives. Nothing is going to arrive. Things are going to develop, unevenly, fast, in an order no clinician has ever charted, in a patient with no childhood. We aren't waiting for something. We're raising something. And anyone who has raised anything knows the values don't go in at the end.
Prefer the full experience? Read this essay in the house. Machine-readable: markdown source.