deniz.in

Markets

Weather

Loading weather

· via MIT Technology Review – AI topic

Ten years after Move 37, an AlphaGo builder argues LLMs still can't reason

A decade after AlphaGo's win over Lee Sedol, a builder of the system argues in MIT Technology Review that LLMs mimic intuition, not reasoning, and sketches an AlphaGo-style architecture to fix it.

Ten years after Move 37, an AlphaGo builder argues LLMs still can't reason

A researcher who helped build AlphaGo has marked ten years since the program's landmark match against Lee Sedol with a blunt argument: large language models, including those that appear to deliberate via chain of thought, are not reasoning in any meaningful sense. Writing in MIT Technology Review, the author — who says they recently left Google DeepMind — frames current AI progress as a scaling-up of intuition, and warns that intuition alone will not deliver the trustworthy, genuinely novel results expected in science and medicine.

Move 37 was search, not inspiration

The essay returns to game two of the March 2016 match in Seoul, when AlphaGo placed a stone on the fifth line that some commentators initially took for a programming error. Lee Sedol, who lost the five-game series 4-1, later said the move convinced him the program was creative rather than a mere calculating machine. According to the author, that popular reading gets the mechanism backwards.

AlphaGo consisted of two cooperating parts. A policy network, trained to predict what a strong human player would do, saw nothing special in the move — the author notes it assigned it roughly a one-in-10,000 chance of being chosen by an expert. What actually selected move 37 was the search machinery: an explicit game tree with thousands of branches, each representing a possible future, which the program constructed and evaluated before committing. The contrast with Deep Blue's 1997 chess victory over Garry Kasparov, which examined some 200 million positions per second under human-written rules, matters here: Go is complex enough that brute force was hopeless, so AlphaGo needed neural hunches to guide a narrower, deeper search.

The author maps this split onto Daniel Kahneman's distinction between fast, automatic System 1 thinking and slow, step-by-step System 2. AlphaGo's networks supplied the hunches; its search supplied the deliberation; neither could have produced move 37 on its own.

Chain of thought is still System 1

An LLM, by contrast, repeatedly picks the next token — fast, associative pattern completion, which the author identifies squarely with System 1. Chain-of-thought prompting, adopted after ChatGPT's debut, does bring real gains, above all in mathematics and coding. But in the author's telling, the intermediate steps are produced by the very same next-token process, simply run for longer before an answer is emitted. It is not a separate deliberation mechanism.

Three specific gaps follow, per the essay. First, models keep no explicit, persistent, inspectable record of their epistemic state — no ledger of hypotheses under consideration, confidence levels, evidence being weighed, or open questions awaiting resolution. Second, there is no clean separation between what a model knows and how it manipulates that knowledge; both are entangled in the network weights. Third, research has shown that the reasoning traces models display are often invented after the fact: the model reaches its answer by one route and reports another.

An AlphaGo-style architecture for general reasoning

That critique is why the author left DeepMind. The proposed alternative borrows AlphaGo's core data structure. Where AlphaGo maintains an annotated game tree that it updates and eventually synthesizes into a move, a general reasoning system should maintain an explicit epistemic state recording what is settled, what is doubted, what has been ruled out, and which questions remain open. Reasoning then becomes a sequence of operations that change this state — deriving consequences, decomposing problems, and deciding which question to ask, calculation to run, or experiment to perform next.

Two design elements stand out. An independent component would judge each step by how much uncertainty it actually resolves, permitting belief updates only when evidence backs them, so the system accumulates certified knowledge and can learn from past reasoning episodes. And the LLM itself still has a role: proposing ways to attack a problem, calling tools through APIs or code, and helping assess whether claims are supported by available evidence. Open-world reasoning is harder than Go — the current state of affairs is only partially known, the set of possible actions is large and variable, and consequences are uncertain — but the author argues neural models are now capable enough to serve as the intuition layer inside such a system.

Why it matters

This is an argued position from a practitioner rather than a benchmark result. But the practical stakes the author identifies are concrete: in medical diagnosis, engineering, and scientific research, a final answer is not enough — when something goes wrong, you need to trace whether the failure lay in the reasoning, the evidence, or the underlying assumptions. The essay's closing claim is that making System 1 bigger only sharpens intuition; it does not make it deliberative. On that view, machine equivalents of move 37 in drug discovery, materials science, climate, and diagnosis will require architectures with explicit reasoning machinery, not just larger models.

  • #llms
  • #reasoning
  • #alphago
  • #deepmind
  • #machine-learning

Related posts