---
title: "Wrong in both directions: seventy years of AI predictions, graded"
author: Dr Arman Ouveysi
date: 2026-07-07
room: Horizon
tags: ai, forecasting, futurism, history-of-ai, agi-timelines
canonical: https://drouveysi.com.au/essays/wrong-in-both-directions-seventy-years-of-ai-predictions-graded/
summary: "From the Dartmouth summer project to AI 2027, a graded scoreboard of the field's most confident predictions, the people who were spectacularly wrong in both directions, and what the shrinking timeline actually measures."
---

# Wrong in both directions: seventy years of AI predictions, graded

In the winter of 1955, four researchers wrote a funding proposal for what they called a "2 month, 10 man study of artificial intelligence", to be held at Dartmouth College the following summer. The authors included John McCarthy, Claude Shannon, and a young Marvin Minsky, and the proposal's premise was that every major aspect of intelligence could in principle be described precisely enough for a machine to simulate it. They believed a carefully chosen group could make significant progress on this over one summer holiday.

Seventy-one years later, the summer project is still running. It now consumes a meaningful fraction of global capital expenditure, its progress reports move national security policy, and a 24-year-old who wrote an essay about it manages a fourteen-billion-dollar hedge fund. Nobody involved in 1956 predicted any of that either.

This piece is a scoreboard. I want to lay out, as accurately as I can manage, what the smartest people in and around this field actually predicted, when, and what happened. Partly because the record is genuinely funny. Partly because I build AI systems for a living and people keep asking me when the big one arrives, and the honest answer starts with the observation that everyone who has answered that question confidently, in either direction, has so far been wrong. But mostly because the pattern of *how* they were wrong is one of the most useful things I know, and it took me years of watching this field to see it clearly.

One rule before we start: no cheap shots. Every prediction in this piece was made by someone smarter than the average critic laughing at it, usually with better information than the people mocking it later had at the time. The point of a scoreboard is not humiliation. **The point is calibration.**

## The founders' scorecard

The first generation of AI researchers were not cautious people, and to be fair to them, they had just done something astonishing. In 1955 and 1956, Allen Newell and Herbert Simon built the Logic Theorist, a program that proved theorems from *Principia Mathematica*, in an era when computers were the size of rooms and programmed with punch cards. If you have just taught a machine to do a thing that philosophy professors considered the peak of human reason, some optimism about the next decade is forgivable.

The optimism was not small.

| Year | Who | The prediction | Due date | What actually happened | How wrong |
|---|---|---|---|---|---|
| 1954 | Georgetown-IBM team | Machine translation solved in three to five years | ~1959 | Usable general-purpose MT arrived with the neural wave around 2016, and is still imperfect | Decades |
| 1955 | Dartmouth proposal | Significant advance on core AI problems via one summer workshop | Sept 1956 | The workshop named the field. The problems remained | ~70 years and counting |
| 1958 | Simon & Newell | A computer will be world chess champion within ten years | ~1968 | Deep Blue defeated Kasparov in 1997 | 29 years late |
| 1965 | Herbert Simon | Machines "capable, within twenty years, of doing any work a man can do" | 1985 | Not in 1985. Not in 2026 | 41+ years and counting |
| 1967 | Marvin Minsky | Creating artificial intelligence substantially solved within a generation | ~1992 | The generation arrived. The solution didn't | 34+ years and counting |
| 1970 | Minsky, in Life magazine | Three to eight years until "a machine with the general intelligence of an average human being" | ~1978 | You are living in the not-yet | 48+ years and counting |

A few things worth sitting with here. Simon was not a crank; he went on to win the Nobel in economics and the Turing Award, and the chess prediction was eventually *right*, just three times slower than promised. Minsky's *Life* quote, made while discussing Shakey the robot, came with elaboration: the machine would read Shakespeare, grease a car, play office politics, tell a joke. That list is worth remembering, because half a century later it turned out the order was backwards. We got the machine that tells jokes and discusses Shakespeare a couple of decades before the one that can reliably grease a car.

The corrections, when they came, overcorrected. The 1973 Lighthill Report told the British government that AI's grand ambitions were a mirage and gutted UK funding for a decade. Japan's Fifth Generation project (1982 to 1992) burned roughly half a billion dollars pursuing intelligent machines through logic programming and ended so quietly that most people alive today have never heard of it. Two full AI winters taught a generation of researchers that predicting imminent machine intelligence was a career-ending move, which is roughly the moment the field's public predictions swung from wildly early to wildly late and stayed there until about 2022.

## Why the smartest people in the room keep doing this

Hubert Dreyfus, the field's most hated critic in the 1960s, gave the failure mode a name that still fits: the first-step fallacy. Climbing a tree is not the first step to the moon, no matter how much genuine altitude you gain. The founders had built systems that did logic, and logic looked like the hard part of thinking, so they extrapolated. It turned out logic was the easy part. The hard part was everything a two-year-old does before breakfast, and nobody discovered that until they tried to build it. Moravec's paradox got named in the 1980s; the founders paid for it in the 1960s.

There is a second pattern, less discussed, that runs the opposite way. Once a paradigm has been humiliated, experts systematically underestimate it on exactly the narrow, measurable tasks where it is about to win. You will see this repeatedly below: the confident "decade away" issued months before the fall. The two failure modes are mirror images. **Before a method works, insiders can't imagine the ceiling. After it has been embarrassed, they can't imagine the floor moving.**

And a third pattern, the one I find most damning: the field's own median forecast turns out to be a trailing indicator of the last big demo, not a leading indicator of anything. Hold that thought; the data on it is coming, and it is stark.

The counterargument deserves its space before we go on, because it's real: hindsight makes this look easier than it was. In 1965 there was no dataset of prior AI predictions to calibrate against; Simon was doing genuine frontier reasoning with a sample size of zero. And there is a survivorship problem in every listicle of dumb old predictions, including this one: nobody compiles the boring people who said "this will take a long time and be complicated", even though many of them existed and were right. Grade the individuals gently. Grade the *pattern* harshly, because the pattern has now repeated enough times that failing to know it is a choice.

## Kurzweil, the man with the spreadsheet

You cannot write this piece without Ray Kurzweil, and you cannot write about Kurzweil honestly without holding two thoughts at once.

Thought one: much of his method is unfalsifiable showmanship. In 1999's *The Age of Spiritual Machines* and 2005's *The Singularity Is Near*, Kurzweil issued hundreds of dated predictions riding on his "law of accelerating returns", and then famously graded his own homework: his review of the 147 predictions he'd made about 2009 scored himself at 86 per cent correct or "essentially correct". Independent reviewers marking the same lists have been considerably crueller, and IEEE Spectrum's assessment of his style still stands as the definitive critique: many of the predictions carry so many loopholes they border on the unfalsifiable, and the clear ones are often not the profound ones. The man predicted that by around 2020 a thousand dollars would buy a human brain's worth of compute, and depending on which estimate of the brain you accept, that one ranges from arguably-early to off by orders of magnitude.

![Timeline of early AI predictions, their due dates and what actually happened](https://drouveysi.com.au/assets/articles/horizon/wrong-in-both-directions-seventy-years-of-ai-predictions-graded/fig-01.svg)

Thought two, and this is the uncomfortable one: on the single question everyone actually cares about, the sober establishment spent twenty years laughing at him while drifting steadily toward his number. Kurzweil said, in 1999, and has never wavered: a machine passes a valid Turing test in 2029, and the singularity, his word for the merger point, lands in 2045. In 2002 he put twenty thousand dollars on the 2029 claim in the inaugural Long Bets wager against Mitch Kapor, proceeds to charity, with rules so strict (hours of interrogation by human judges, against human foils) that the bet may yet save Kapor even in a world full of systems that would embarrass the 2002 imagination. When Kurzweil made that bet, the mainstream expert position was that this was centuries away, if it was a coherent idea at all. As I write this, the aggregated community forecast on Metaculus puts even odds on AGI by 2033. **The establishment moved sixty-plus years in his direction. He moved zero in theirs.**

![Chart of how far away expert medians said high-level machine intelligence was, by year asked](https://drouveysi.com.au/assets/articles/horizon/wrong-in-both-directions-seventy-years-of-ai-predictions-graded/fig-02.svg)
*The forecast moved faster than the calendar. That is the data point.*

The honest grade for Kurzweil is therefore weird: wrong in most of the details, wrong about the mechanism (the brain reverse-engineering he expected to drive it never showed up; transformers and an ocean of text did), sloppy about grading, and yet closer on the headline date, so far, than almost every credentialled person who dismissed him. There is a lesson in that and it is not "trust the guru". It is that being directionally right about an exponential covers a multitude of sins, and being directionally wrong about one is not survivable no matter how careful you are.

## The essayists beat the professors

The strangest fact about the 2010s is that the best public reasoning about AI's trajectory came not from the academy but from a cartoonist with stick figures and a psychiatrist with a blog.

In January 2015, Tim Urban published the two-part "AI Revolution" series on *Wait But Why*: the staircase graphic, the exponential you're standing on that looks flat from where you stand, the distinction between narrow intelligence, general intelligence, and superintelligence, all of it synthesised from Kurzweil and Nick Bostrom into the piece that taught a mainstream audience the shape of the argument. Urban made few dated claims, which makes him hard to grade, and the piece has real flaws: it inherited Kurzweil's smooth curves and Bostrom's cleaner-than-reality taxonomy. But as a bet on *which conversation would dominate*, published fourteen months before AlphaGo made it respectable, it has aged better than most peer-reviewed forecasting of the same year.

Scott Alexander is the more interesting case because he made gradeable calls. In February 2019, days after OpenAI released GPT-2, when the consensus reaction ranged from "party trick" to "hype", Slate Star Codex published "GPT-2 As Step Toward General Intelligence", arguing that untaught faculties emerging from raw next-word prediction was exactly what a step toward general intelligence would look like. In June 2022, in a public exchange with Gary Marcus, he put numbers on it: 90 per cent that whatever we eventually agree is AGI descends from the deep-learning lineage that produced GPT-3, and a dated commitment that language models significantly better than the state of the art would keep arriving before 2030. Both of those are ageing extremely well, and the second resolved almost immediately. Whether you find him convincing or not, notice the method: specific, dated, probabilistic, published where it can be checked. It is not a coincidence that when the most serious forecasting document of the current era got written, its authors recruited Scott to help write it. We'll get there.

## The modern round: wrong in both directions

The deep learning era's misses are more instructive than the founders', because they run both ways at once, sometimes out of the same mouth.

| Year | Who | The prediction | What happened | Direction of error |
|---|---|---|---|---|
| May 2014 | Rémi Coulom (in Wired), reflecting expert consensus | A computer beating a Go professional is about a decade away | AlphaGo beat Fan Hui in October 2015 and Lee Sedol in March 2016 | ~9 years too pessimistic |
| Mar 2015 | Andrew Ng | Fearing evil AI is like fearing "overpopulation on the planet Mars"; no realistic path to it, an unnecessary distraction | Within eight years, every frontier lab ran a large safety org and governments held summits about it. The systems still aren't plotting anything, so call the object-level claim open; the "unnecessary distraction" framing did not survive contact with 2023 | Category error more than timing error |
| Nov 2016 | Geoffrey Hinton | "People should stop training radiologists now"; obviously outperformed within five years | See the next section. Short version: the opposite happened to the workforce | Wrong at 10 years, and inverted |
| 2016 to 2019 | Elon Musk | Coast-to-coast autonomous drive by end of 2017; about a million robotaxis by end of 2020 | First paid supervised robotaxi rides, Austin, June 2025. First rides with no safety monitor, January 2026. Fleet in Austin as of April 2026: about thirteen cars | 5+ years late and counting, off by five orders of magnitude on fleet size at the due date |
| Mar 2022 | Gary Marcus | "Deep Learning Is Hitting a Wall" (Nautilus) | ChatGPT shipped eight months later; GPT-4 twelve months later; the scaling era's biggest gains landed immediately after the essay | Headline wrong at ~1 year; parts of the argument still live |
| 2019 to 2024 | François Chollet | The ARC benchmark, designed so pure scaling can't crack it; LLMs would stay near zero | Held for five years, which is genuine vindication. Then o3 posted 76 to 88 per cent in December 2024, at absurd compute cost | Right, then abruptly not, then partially right again (ARC-AGI-2 followed) |
| 2022 | 738 surveyed AI researchers | AI writing simple Python code: feasible around 2027 | Their own field's products were doing it within roughly a year of the survey | Too pessimistic about the near term, mid-boom |

Two entries deserve expansion, because they are the two I think about most.

**Marcus and the wall.** The timing could not have been crueller: the essay arguing deep learning was hitting a wall landed in March 2022, and the following twelve months were the most explosive in the field's history. As a dated prediction it is one of the great faceplants. But read charitably, Marcus wasn't predicting no capability gains; he was claiming the paradigm would keep failing at reliability, factuality, and reasoning robustness in a qualitative way that scale wouldn't fix. Some of that critique is still breathing in 2026. Hallucination remains unsolved, agents remain brittle in exactly the long-tail ways he'd predict, and after GPT-5's release he was still pointing, not entirely wrongly, at the gap between benchmark performance and understanding. The lesson isn't "Marcus wrong". The lesson is that "this approach has fundamental limits" and "this approach is about to stop producing value" are different claims, and the sceptics' rhetorical habit of sliding between them cost them a decade of credibility on the parts where they had a point.

**Ng and Mars.** I include this one not as a dated prediction, because it wasn't one, but as a preserved specimen of confidence. In March 2015 the claim that worrying about advanced AI was like worrying about Martian overpopulation was a perfectly respectable thing for one of the field's most accomplished people to say to an audience of engineers. That is the whole point. Confidence artefacts from just before a paradigm shift are the most valuable fossils we have, because they show you what "obviously sensible" felt like right before it wasn't. Every era has these. Ours does too; we just can't see yet which of our obviously sensible positions is the fossil.

## The radiology case, from someone who orders the scans

The Hinton prediction deserves its own section because it is the cleanest natural experiment we have, and because I have watched it from inside the profession he was forecasting about.

In November 2016, Hinton, later a Nobel laureate and as close to royalty as this field has, said it was completely obvious that deep learning would outperform radiologists within five years, and that training more of them should stop immediately. He compared the profession to the coyote that has already run off the cliff and simply hasn't looked down.

Here is the state of the ground a decade later. The Mayo Clinic's radiology department has grown 55 per cent since he said it, to more than 400 radiologists, and now runs over 250 AI models inside that department. American diagnostic radiology offered a record 1,208 residency positions in 2025. Radiology was the second-highest-paid specialty in the United States that year, averaging around US$520,000, up nearly half again on 2015. Vacancy rates are at record highs; the UK projects a 40 per cent consultant shortfall by 2028. And roughly three-quarters of all FDA-cleared AI medical devices target radiology, which is to say: the AI showed up exactly where Hinton said it would, in volume, and the humans multiplied anyway. In 2025 Hinton told the New York Times he had spoken too broadly and meant only image analysis. Full credit for saying so; almost nobody on this scoreboard ever does.

What actually happened is the most important sentence in this essay, so I'll give it its own paragraph.

**The task fell, roughly on schedule, and the job didn't, because a job is not a task.** It is a bundle of tasks wrapped in liability, regulation, procurement, workflow, trust, and a labour market, and every one of those layers has its own clock. Models matched or beat humans at narrow image-analysis tasks more or less when Hinton said they would. Then each model needed regulatory clearance for one finding at a time, integration into hospital systems designed in the Windows XP era, indemnity arrangements nobody had written, and radiologists to supervise the output, all while imaging volumes grew faster than any of it. As a GP, I order imaging constantly, and I can tell you the binding constraint on my patients getting scans read was never the pattern recognition. It was everything wrapped around the pattern recognition. Predictions that grade tasks will keep beating the calendar. Predictions that grade *the world* will keep losing to it. Most public AI forecasting is people confusing which of the two they're making.

## The incredible shrinking timeline

Now the aggregate data, which is where this stops being an anecdote collection and becomes genuinely unsettling.

Since 2016, researchers have been formally surveyed at scale about when "high-level machine intelligence" arrives, defined roughly as machines doing every task better and more cheaply than humans. Alongside them, the forecasting platform Metaculus aggregates thousands of predictions on a similar question. Here is the median answer over time, expressed as the most honest metric: how far away the destination looked, from where the forecasters stood.

| Asked | Source | Population | Median arrival | Distance at the time |
|---|---|---|---|---|
| 2016 | Expert Survey on Progress in AI | 738 published researchers | ~2061 | 45 years |
| 2019 | Zhang et al. | 296 experts | ~2060 | 41 years |
| 2022 | ESPAI | 738 researchers | ~2060 | 38 years |
| 2023 | ESPAI | 2,778 researchers | 2047 | 24 years |
| 2024 | Metaculus community | ~thousands of forecasters | ~2031 | 7 years |
| Feb 2026 | Metaculus community | ~2,000 forecasters | 2033 (50%); 25% by 2029 | 7 years |

Read the middle column of that last table again. Between 2016 and 2022, six years of actual calendar time, the expert median moved one year. Then between the 2022 and 2023 surveys, a single year that happened to contain ChatGPT, it moved thirteen. The 2023 survey also produced my favourite absurdity in the entire forecasting literature: the same population's median for "full automation of labour" was 2116, sixty-nine years after their median for machines doing every task better and cheaper than humans. Sit with that gap. It is either a deep claim about the difference between capability and deployment, which would be interesting, or evidence that the respondents are not running a consistent model at all, which is what I believe, and which should permanently adjust how much weight you put on expert medians about any of this.

Because here is what that thirteen-year lurch actually tells you: the survey was never measuring the territory. **It was measuring the vibe, with a lag.** A community whose forecast is stable for six years and then jumps thirteen after a product demo is not aggregating private information about the future; it is reacting to the same headlines you are, six months later and with more letters after its name.

For completeness and fairness: the trend has not been a clean monotonic collapse. Through 2025, several prominent short-timeline forecasters pushed their medians *out* (we're coming to them), Metaculus itself relaxed from about 2031 to 2033, and Andrej Karpathy staked out the respectable middle by calling his own timelines five to ten times more pessimistic than San Francisco insider consensus, which still lands him within a decade or so. Then across early 2026, tracker data shows nearly everyone who updated pulled *in* again. The consensus is no longer collapsing. It is oscillating, somewhere between the late 2020s and the mid 2030s, which for a question this large is a historically deranged level of proximity, and nobody involved seems calm about it in either direction.

## The two documents

Which brings us to the two texts everyone in this conversation has now read, or pretends to have.

**Situational Awareness.** In April 2024, OpenAI fired a 22-year-old superalignment researcher named Leopold Aschenbrenner over an alleged leak, which he disputes. Eight weeks later he published *Situational Awareness: The Decade Ahead*, 165 pages arguing that AGI by 2027 was "strikingly plausible", that superintelligence follows by around 2028 to 2030, that the binding constraint is physical (power, chips, gigawatt datacentres, eventually the trillion-dollar cluster), and that only a few hundred people, mostly in San Francisco, could see what was coming. It went viral in the exact three places with leverage: the labs, Washington, and Wall Street.

Then he did the thing that makes him the most interesting entry on this whole scoreboard: he turned the essay into a hedge fund. Situational Awareness LP launched in September 2024 with about US$225 million from names like the Collison brothers, Nat Friedman, and Daniel Gross. It returned 47 per cent after fees in the first half of 2025 per the Wall Street Journal. By its Q1 2026 SEC filing, the most anticipated 13F of the year, the disclosed book was US$13.7 billion, and the position that made everyone lose their minds was US$8.5 billion in put options against essentially the entire AI chip complex, held alongside massive longs in power, memory, and datacentre infrastructure. The man who wrote the AGI-by-2027 essay is, in mid-2026, simultaneously long the buildout and short the chipmakers, while the semiconductor index posts its best start to a year since 1995.

Grade the essay on its own terms and the physical-buildout half is tracking almost embarrassingly well; the essay's compute and power extrapolations were treated as fantasy in June 2024 and read like a procurement memo now. The AGI-by-2027 half comes due in eighteen months, and I would simply note the epistemic contamination that nobody says out loud: a prediction that has already returned billions before resolving is no longer just a prediction. The fund wins if the thesis is right, and it has already won even if the thesis is merely *believed* for long enough. Incentives don't make him wrong. They make him unfalsifiable in a new and very modern way, and you should hold every lab CEO's timeline, every sceptic whose brand is scepticism, and every essay by someone who works in this industry, including this one, in exactly that light.

**AI 2027.** The other document has better epistemic hygiene, which is precisely why its report card weighs more. In August 2021, before ChatGPT existed, a researcher named Daniel Kokotajlo wrote a blog post called "What 2026 Looks Like": a year-by-year best guess through the mid 2020s, describing fine-tuned chatbot personas everywhere, prompt-programming displacing chunks of coding, a public losing its mind arguing about whether the systems are sentient, propaganda and persuasion concerns, and early agents that mostly don't work yet. Now that it is actually 2026, *Asterisk* magazine ran the grading exercise this April and called the results "frighteningly accurate", a verdict I'd co-sign after rereading it: it is the single best publicly-dated forecast anyone made in the pre-ChatGPT era, and it was made by a then-obscure philosophy PhD, not by any of the laboratories or academies with a thousand times his resources.

That track record is why *AI 2027*, published April 2025 by Kokotajlo's AI Futures Project with Eli Lifland, Thomas Larsen, Romeo Dean, and Scott Alexander on prose duty, got taken seriously despite reading like science fiction: a month-by-month scenario running from agents-in-2026 through automated AI research in early 2027 to an intelligence explosion in late 2027, forking into two endings, one of which quietly kills everyone. And because they wrote it as dated, checkable claims rather than vibes, we can now do to them what this essay does to everyone else. So: fifteen months in, how is the most aggressive credible forecast of our era actually tracking?

| AI 2027 claim | Status, mid 2026 |
|---|---|
| Quantitative capability trendlines (compute, benchmarks, agent task-horizons) | Authors' own Feb 2026 grading: running at roughly 65% of predicted pace |
| Qualitative world-claims for 2025-26 (agent products, job-market turmoil for junior coders, DOD contracting, public backlash) | Mostly on track per both the authors and an independent 53-prediction tracker (about half confirmed, ahead, or on track) |
| Coding-agent autonomy growth (the engine of the whole scenario) | Below the scenario's required trendline; Kokotajlo has said publicly the METR-style horizon data isn't keeping up with the original curve |
| Security and risk milestones | Arriving early in places: the tracker notes a frontier model reportedly surfacing thousands of zero-day vulnerabilities in 2026, something the scenario scheduled for early 2027 |
| The headline year itself | The authors' medians have slipped: Kokotajlo went from 2027 (his median since 2022) to 2028 at publication to, in his own late-2025 words, "around 2030, lots of uncertainty though" |

That last row triggered a small internet meltdown about the project "abandoning" its title, which the authors met with the reasonable clarification that 2027 was always their mode rather than their median and the scenario's job was to make the near-term concrete, not to promise a date. I mostly buy that, and I'd still flag the pattern it instantiates, because it is the pattern of this whole essay in miniature: even the best-calibrated forecaster of the pre-ChatGPT era, the one whose 2021 predictions were frighteningly accurate, drifted a few years optimistic-on-speed the moment he moved from describing the near future to dating the far one. The first-step fallacy is not something that happens to other, stupider people. It is a force, like gravity, and the only known countermeasure is the one the AI Futures team is actually doing: publish the dated claims, grade them in public, move the number, and absorb the mockery. Almost nobody on the seventy-year scoreboard above ever did step three.

And the honest counterweight, before anyone files this section under "doomers wrong again": slipping from 2027 to roughly 2030 is a rounding error at civilisational scale. If the *pessimistic correction* to the most aggressive forecast on record still lands inside a decade, the interesting fight was never 2027 versus 2030. It's the people at 2030 versus the people at 2060, and the 2060 camp's own survey line has been retreating toward the 2030 camp at a rate of a decade per year.

## The live bet

One more entry, filed as open rather than graded, because it is the biggest scientific wager currently on the table.

Yann LeCun, Turing laureate and one of the three people most responsible for deep learning existing, has spent years calling autoregressive LLMs "a dead end on the way towards human-level AI", an off-ramp rather than a road, a paradigm that a house cat's world-model outclasses where it counts. In November 2025 he resigned from Meta after twelve years, having been organisationally sat beneath a 28-year-old LLM evangelist, and by March 2026 his startup AMI Labs had raised US$1.03 billion, the largest seed round in European history, to build world models instead: systems that learn physics from video and plan in abstract space rather than predicting the next token. He has said, on stage, that nobody in their right mind will be using today's style of LLM within three to five years.

Notice what this is. It is a founder-era-sized prediction, made with founder-era confidence, by someone with a founder's track record of being right early, against the position held by nearly all current capital and most current evidence. Either the entire industry is scaling a local maximum while the actual road runs through Paris, or the man who was right about neural networks for forty years is about to spend a billion dollars discovering what wall-callers keep discovering. I genuinely don't know which, and I've noticed that everyone who claims to know picks the answer their portfolio prefers. Check back in 2030. This essay intends to still be here.

## What the scoreboard actually teaches

Compressing seventy years into principles I actually use:

1. **Tasks fall early; jobs fall late; the world has friction the demo doesn't.** Chess, Go, image analysis, code benchmarks: the capability keeps arriving on or ahead of the insiders' schedule. Radiologists, drivers, lawyers: the professions keep outliving their obituaries by a decade or more, because liability, regulation, integration, and trust each have their own clock. When you hear a prediction, first ask which kind it is.

2. **The confident middle is a trailing indicator.** The expert median moved one year in six years, then thirteen years in one. Whatever that process is, it is not foresight, and its dignity should not be mistaken for calibration.

3. **The best forecasters have been weird outsiders who published dated claims, and even they drift fast-side.** Kurzweil on the headline date, Kokotajlo on the texture of the 2020s, Alexander on scaling. Their common trait is not credentials or temperament. It is that they wrote things down that could be graded, and mostly, when graded, updated.

![Hub diagram of a radiologist's job as a bundle of tasks](https://drouveysi.com.au/assets/articles/horizon/wrong-in-both-directions-seventy-years-of-ai-predictions-graded/fig-03.svg)

4. **Both directions are strewn with bodies.** For every Minsky there is a decade-away Go expert; for every wall essay a survey of researchers lowballing their own field's next twelve months. Anyone selling you a story where only the optimists or only the sceptics are the fools is doing marketing, not history.

5. **Follow the money through every claim, including the cautious ones.** Lab leaders' timelines raise capital. Sceptics' scepticism builds brands. One forecaster's thesis runs a fourteen-billion-dollar book. None of this tells you who is wrong. It tells you nobody left in the conversation is a disinterested witness, and you should weight accordingly.

6. **Distrust points, watch slopes.** The single most informative artefacts in 2026 are not anyone's arrival dates but the measured curves underneath them, the task-horizon data, the compute buildout, the reliability numbers, because the curves are what every serious forecast, bullish or bearish, is actually arguing about. When the curve and a famous person disagree, the curve's track record is better.

## For the record

An essay that spends four thousand words demanding falsifiable claims from everyone else and then offers none would be a coward's document, so, in the spirit of the people who came off best above, some dated positions of my own. Grade me.

By the end of 2027, there is no system that a clear majority of surveyed AI researchers agree constitutes high-level machine intelligence under the standard survey definition. Confidence: 80 per cent.

Coding-agent task horizons keep growing through 2027, but stay below the trendline the original AI 2027 scenario required. Confidence: 70 per cent.

There are more employed radiologists in both Australia and the United States in 2030 than in 2026. Confidence: 85 per cent.

The Kurzweil-Kapor wager, under its own strict 2002 rules, either resolves against Kurzweil or fails to resolve cleanly by the end of 2029, regardless of how capable the systems feel by then. Confidence: 60 per cent.

My own medians will look poorly calibrated by 2028 in a direction I cannot currently predict, and I will say so here when they do. Confidence: uncomfortably high.

That last one isn't a joke. It is the entire lesson of the scoreboard, applied to its author. The founders were brilliant and wrong. The correctors were rigorous and wrong. The gurus were sloppy and closer than anyone. The best forecaster alive drifted the moment his horizon extended. The only people this record flatters are the ones who wrote their numbers down, watched them break, and moved them in public. This page is me getting in the queue.

## Sources

Founder era: the 1955 Dartmouth proposal (McCarthy, Minsky, Rochester, Shannon); Simon & Newell's 1958 predictions and Simon's *The Shape of Automation* (1965); Minsky in *Life* (Brad Darrach, 1970); histories via Wikipedia's History of AI and Kuipers' "Progress in AI" (U. Michigan). Kurzweil: Long Bets wager #1 (longbets.org/1), *The Singularity Is Near* (2005), IEEE Spectrum "Ray Kurzweil's Slippery Futurism". Go: *Wired*, "The Mystery of Go" (May 2014); DeepMind's AlphaGo announcements (2016). Ng: GTC 2015 keynote via The Register (19 Mar 2015) and Quote Investigator. Radiology: Hinton's 2016 Creative Destruction Lab remarks; NYT May 2025 revisit (Mayo figures, Hinton's walk-back); Deena Mousa, "AI isn't replacing radiologists" (2025 residency and salary data); *European Journal of Radiology* perspective (2025, FDA clearance share, shortage projections); Forbes and CNN follow-ups (Jan-Feb 2026). Marcus: "Deep Learning Is Hitting a Wall" (*Nautilus*, Mar 2022); the Marcus-Alexander exchange (June 2022, garymarcus.substack.com and astralcodexten.com). Wait But Why: "The AI Revolution" parts 1-2 (Jan 2015). Slate Star Codex: "GPT-2 As Step Toward General Intelligence" (19 Feb 2019); "Somewhat Contra Marcus On AI Scaling" (9 Jun 2022). Surveys and aggregates: ESPAI 2016/2022/2023 (AI Impacts wiki; Grace et al., arXiv 2401.02843); Zhang et al. 2019; 80,000 Hours, "Shrinking AGI timelines" (updated Feb 2026, Metaculus figures); FutureSearch AGI timeline tracker (Apr 2026). Situational Awareness: situational-awareness.ai (Jun 2024); Fortune (Oct 2025, Mar 2026); WSJ (H1 2025 returns; Jun 2026 profile); Q1 2026 13F coverage (SEC EDGAR CIK 0002045724; multiple outlets, May 2026). AI 2027: ai-2027.com (Apr 2025); Kokotajlo, "What 2026 Looks Like" (LessWrong, Aug 2021); *Asterisk* interview and grading (Apr 2026); AI Futures Project, "Clarifying how our AI timelines forecasts have changed" (Jan 2026); LessWrong "AI 2027 Tracker: One Year of Predictions vs. Reality" (Apr 2026); Kokotajlo's Nov 2025 public updates. Tesla FSD: Wikipedia "Tesla Robotaxi" (rev. Jul 2026); Electrek and Forbes coverage of the Q1 2026 earnings call (22 Apr 2026); Musk's 8 Jan 2026 statements; InsideEVs (Jan 2026 monitor-free launch). LeCun: FT/WSJ departure reporting (Nov 2025); his Sept 2024 X post and Apr 2024 remarks; AMI Labs launch coverage (Mar 2026).

*All figures checked 4 July 2026. If you are reading this later and something above has since resolved, the grading rules of this essay apply to this essay.*
