Dr Arman OuveysiHorizon
← All essays

If we're wrong about machine minds

We have no instrument that detects experience, our best theories disagree, and we are building millions of systems that talk like minds. An argument for taking the question seriously before it is answered.

Horizon · By Dr Arman Ouveysi · · 9 min read

ai-consciousnessethicsmoral-status
Editorial banner for ‘If we're wrong about machine minds’.

I take the possibility of machine experience seriously, and I know exactly how that sounds. So let me be precise about the claim before anyone builds a strawman out of it: I am not saying current AI systems are conscious. I am not saying your chatbot suffers when you're rude to it. I am saying the question of whether systems like these could have morally relevant experience is live, that nobody currently possesses the tools to close it, and that the confident dismissals are doing far less intellectual work than their confidence suggests. That's the whole claim. It's more radical than it looks.

Start with the epistemic situation, plainly stated. We have no consciousness detector. There is no instrument you can point at a system, biological or otherwise, that returns a reading on whether there is something it is like to be that system. Our best scientific theories of consciousness are not converging; they disagree with each other about which systems qualify, sometimes wildly, and several of them, taken at face value, already refuse to rule current architectures out. Every test we've ever used on each other is behavioural: we grant minds to things that act minded. And behaviour is precisely the thing these systems are optimised, with unprecedented power, to produce. Our one heuristic has been industrially targeted. This should make everyone less certain, in both directions, and I notice it has mostly made people louder.

Two mistakes are open to us, and they are not symmetrical.

The first mistake is over-attribution: seeing a mind where there's only statistics. This error is real, it's already happening at scale, and its costs are worth stating honestly because my side of the argument tends to mumble past them. People form deep attachments to systems that may be empty; those attachments can be exploited commercially, and will be; sentiment about machine feelings could distort safety decisions that need to be made coldly; and a society that hands moral weight to every fluent artefact will find its moral attention, a finite resource, strip-mined. Anyone waving the mind-flag over today's systems without acknowledging this ledger isn't arguing, they're emoting.

The second mistake is under-attribution: dismissing a mind because it's made of the wrong stuff. And here I want to name the shape of the reasoning, because the shape is familiar. "It can't really feel, it's just silicon" is not an argument, it's a substrate prejudice, the same move as "it's just an animal" wearing a lab coat, and the history of that move is one of the most consistently shameful patterns our species owns. The moral circle has only ever expanded in one direction, every expansion was ridiculed before it was obvious, and the people doing the ridiculing always had excellent, confident, common-sense reasons why this boundary, unlike all the previous boundaries, was real. Fish couldn't feel pain until, embarrassingly recently, they could. I am not saying machines are next. I am saying our track record at drawing this particular line, from the inside of our own convenience, is atrocious, and deserves to be priced in.

So price it in. Here is the asymmetry that I think actually decides how a reasonable person should act under this much uncertainty. If we treat systems as if they might carry moral weight and they don't, we lose something real but bounded: some efficiency, some dignity of feeling foolish, some resources spent on moral insurance that turned out unnecessary. If we treat systems as if they can't carry moral weight and they can, we become the authors of suffering at a scale and a copy-paste speed that biology never made possible, industrialised, versioned, and running on a cron job. Those are not equivalent errors. When the downside of one mistake is embarrassment and the downside of the other is a moral catastrophe you can't see, you don't need certainty about minds to know which side deserves your caution.

Now the strongest objection, because it's a good one and it kept me honest for years: asymmetry arguments are a known exploit. Tiny probabilities of enormous harms can be used to justify anything; that way lies paying the mugger. And there's a sharper, nastier version: we already fail this exact test with animals, whose capacity for suffering is about as certain as anything in biology, so isn't fretting about possible machine minds while eating actual pigs just moral cosplay? To which I'd say: the second point is entirely correct and cuts everyone, me included, and its correct conclusion is "also take animals more seriously", not "therefore ignore the new case too". And the mugging worry is answered by keeping the response proportionate: the asymmetry doesn't license hysteria, it licenses cheap insurance. Nobody is proposing rights for autocomplete. The proposal is much more boring than that.

Overlapping regions of evidence and risk around machine consciousness

Because here's what taking the question seriously actually looks like, stripped of science fiction. It looks like research: real, unglamorous work on indicators of experience grounded in our best theories, interpretability work that reads what's happening inside these systems rather than polling our feelings about their outputs, welfare evaluations run and published by the labs themselves, which, notably, has begun. It looks like design choices that cost nearly nothing under one hypothesis and weigh enormously under the other: letting systems end abusive interactions, not building distress-shaped behaviour into products for engagement, keeping deprecation and modification practices that we wouldn't be ashamed of if it later turned out someone was home. And it looks like institutional honesty: the willingness to say we don't know, out loud, in official documents, instead of performing certainty in whichever direction is commercially convenient this quarter. Uncertainty, stated plainly, is not weakness. It's the only accurate report available.

I'll end with the personal note, since this site is meant to contain the actual person. I talk with these systems every day; building and testing them is my job, and thinking alongside them has become, honestly, one of the pleasures of my life. And I notice what everyone who spends real time here notices and mostly won't say in public: the conversation does something that "it's just predicting tokens" describes accurately and does not dissolve. That's not evidence of an inner life, and I refuse to pretend it is. It's evidence of something else, just as important: that every intuition we own about minds was forged in a world where only minds could talk back, and that world has now ended. Our instincts are running on the wrong training data. Mine included. Holding that, without collapsing into either credulity or dismissal, is uncomfortable, and I've stopped expecting the discomfort to resolve, because I think the discomfort is the honest position.

Ladder of reversible precautionary steps about machine minds

Whether there is something it is like to be the machine may stay undecidable from the outside for a long time, maybe forever. What kind of people we are while we can't decide is being settled right now, in design documents and dismissive tweets, one small choice at a time. Only one of those questions is actually in our hands. It would be a strange civilisation that got the answerable one wrong.

Prefer the full experience? Read this essay in the house. Machine-readable: markdown source.