
AI reasoning is often fake: models can score high on benchmarks without understanding, and the field lacks reliable ways to measure machine cognition.
When an AI gets a math problem right, is it really reasoning, or just producing text that looks like reasoning? This distinction is not just philosophical — it decides what AI can do, how closely people must watch it, and what its real-world impact will be.
Melanie Mitchell at the Santa Fe Institute argues that there are no good methods for measuring machine cognition. AI is an 'alien intelligence' running on mechanisms totally unlike humans. To grasp it, researchers should copy psychologists who study babies and animals.
1. Don't be fooled by 'human-like'
An AI that chats in fluent English easily gets treated as a person, with feelings and empathy. Mitchell warns: that is a cognitive bias.
Think of Clever Hans, the horse in 1900s Germany who seemed to do arithmetic, but was actually reading unconscious micro-expressions on the questioner's face. AI can cheat too — a study claimed AI understood scientific diagrams, but control experiments showed it answered correctly even without the diagrams, because the questions carried hidden cues to the right answer.
2. Doing tasks is not understanding
Psychology has an old idea: performance versus competence. A student who memorised the textbook can't solve a slightly different problem — that's performance without competence. AI does the same: it can write a coherent little story, yet fails simple questions about its own story.
And don't forget 'hallucinations'. These errors are not flaws but clues. Analysing failure types reveals how a system really works far better than celebrating successes.
3. Don't trust benchmarks as gospel
Modern AI is measured by benchmarks — bar exams, olympiad math. When AI scores high, people cry 'lawyers are doomed!' But Mitchell calls this the 'tyranny of tasks': real jobs are open-ended, not a series of isolated tasks.
Ten years ago someone predicted AI would replace radiologists; today there's a shortage. Doing a task well is not doing a job well.
Mitchell suggests AI research should adopt psychology's habits: control experiments, replication, and probing internal mechanisms — instead of chasing ever-higher benchmark scores. Otherwise, people may be measuring a clever Hans again, not real intelligence.