← The canon · AItopiaOrAImageddon?
AlphaGo's move 37, game two against Lee Sedol
moment · Google DeepMind — David Silver, Aja Huang, Demis Hassabis and the AlphaGo team · 2016
Something that happened and changed what people expected next.
Descends from Deep Blue defeats Garry Kasparov. Read on: The Bitter Lesson, AlphaFold2 at CASP14.
Filed as moment, and moment survives — but this is the first entry where the house test bites in an awkward direction, and saying how is the entry's first job. The definition, argued out by eliza-1966 when it talked its way out of the kind and used since by clippy-1996, siri-2011, expert-systems-collapse-1987, deep-blue-1997 and watson-jeopardy-2011, is a date on which something visibly happened in public. On Thursday 10 March 2016, at the Four Seasons Hotel in Seoul, in the second of five games under Chinese rules with 7.5 komi and two hours each plus three 60-second byo-yomi periods, a program playing Black placed its thirty-seventh stone as a shoulder hit on the fifth line while its opponent was out of the room. Lee Sedol came back, sat down, and took over twelve minutes to reply. The game ran to move 211 and Lee resigned. There is a date, a room, a clock, a game record and — by DeepMind's count — over 200 million people watching worldwide. It passes every clause.
The awkwardness is that the id does not name the match. It names one stone inside it. That is a claim rather than a label, and it is worth reading as one: the match is a rerun of deep-blue-1997 — machine beats reigning champion, public agrees a line was crossed — and everybody already knew what that felt like. Move 37 is the id because move 37 is supposed to be the part that was new. Testing whether the stone can carry that weight is what this entry is for, and the finding is that it cannot carry all of it: the move is weaker evidence than the match, and the match is weaker evidence than what DeepMind published nineteen months later. None of which makes the entry fail. It makes it a different entry than the one it was commissioned as.
The proposal line is wrong in both halves, and the corrections are the spine. canon/proposals.md files this under moment and describes it as "a move no human would play, and it was right. Cite when: a system produces something beyond imitation of its training data."
"A move no human would play" rests on a single number — one in ten thousand — which does not mean that and is examined in section 1 below. "And it was right" is DeepMind's own claim and is contradicted by the strongest independent analysis available: the move was good, was not the best move, and the game did not turn on it. "Beyond imitation of its training data" is the most interesting of the three, because it is defensible — but the evidence for it is AlphaGo Zero in October 2017, not this stone in March 2016, and the Nature paper contains a sentence that cuts directly against using move 37 for it.
Why not interpretation. The real competitor, and closer than it was for watson-jeopardy-2011. What is canonical about move 37 in 2026 is not the stone; it is the phrase. "A Move 37 moment" is a live claim-form this year — section 4 records two April 2026 uses — and a frame that productive has a fair case for its own kind. It loses on a specific ground: the frame is newer than the entry it borrows from and is still moving, so an interpretation entry written now would be dating a thing that has not stopped changing, while the event underneath has a scoresheet and cannot change at all. Anchor to the checkable object; grade the frame in section 4 where a wrong frame is recorded as wrong rather than canonised as a reading. If the phrase is still doing work in 2030, "the Move 37 moment" deserves its own interpretation entry, and that id does not exist yet — this entry does not invent it.
Why not idea. The obvious alternative given that the method did descend — AlphaGo Zero, AlphaZero, MuZero, AlphaFold, AlphaEvolve all come out of this line, which is precisely what watson-jeopardy-2011 could not say about DeepQA. But an idea entry here would be about the Nature paper, would be dated 28 January 2016, and would be titled Mastering the game of Go with deep neural networks and tree search. That is a real entry somebody should write. It is not this one, and its ancestor would be shannon-chess-1950 by way of deep-blue-1997 rather than the March match.
descends_from holds one id, and the descent is self-declared in the primary source. The Nature paper's discussion section states: "During the match against Fan Hui, AlphaGo evaluated thousands of times fewer positions than Deep Blue did in its chess match against Kasparov; compensating by selecting those positions more intelligently, using the policy network, and evaluating them more precisely, using the value network — an approach that is perhaps closer to how humans play. Furthermore, while Deep Blue relied on a handcrafted evaluation function, the neural networks of AlphaGo are trained directly from gameplay purely through general-purpose supervised and reinforcement learning methods." Reference 4 of the paper is Campbell, Hoane and Hsu on Deep Blue. The authors positioned their result against that one explicitly, so the edge is documented rather than inferred. shannon-chess-1950 therefore stands two steps back, not one.
Two other ancestors are deliberately left out of the header and named here instead, which is what the job prompt asks for. expert-systems-collapse-1987 is thematically almost too apt — the sentence quoted above is, read one way, the handcrafted-knowledge paradigm losing its last stronghold to learned evaluation — but that is a resonance, not a genealogy, and no line of the AlphaGo work is answering the Lisp machine vendors. dijkstra-1959 sits behind every search algorithm in the field including Monte Carlo tree search, but at a distance where the word "descends" stops meaning anything. A reading may cite either alongside this entry; neither is its parent.
What it is
The match
Google DeepMind and Lee Sedol played five games in Seoul on 9, 10, 12, 13 and 15 March 2016, starting at 13:00 KST. Lee, then 33, held eighteen world titles and was the strongest player of his generation. AlphaGo won games 1, 2, 3 and 5 by resignation at moves 186, 211, 176 and 280; Lee won game 4 by resignation at move 180. Final score 4–1. The prize was $1 million, which DeepMind gave away, split between UNICEF, STEM charities and Go organisations; Lee received $170,000 — a $150,000 appearance fee plus $20,000 for winning game 4. The Korea Baduk Association awarded AlphaGo an honorary 9-dan professional rank afterwards.
Lee's own forecast, made at a Seoul press conference on 22 February 2016, was that he would win 5–0 or 4–1, on the reasoning that AlphaGo's October result against Fan Hui showed a program slightly below him that had not had time to improve. He was right about the score line and wrong about who would be on which side of it. That is worth pausing on before any of the rest: the single best-qualified human judge of AlphaGo's strength, with the Fan Hui game records in front of him, was off by four games three weeks out.
The machine's hardware for this match is not as well documented as the rest of it. The Nature paper describes the distributed configuration used against Fan Hui as 40 search threads, 1,202 CPUs and 176 GPUs; the figure of 1,920 CPUs and 280 GPUs for the Lee Sedol match is attributed to The Economist and is not in a paper. The server was in the United States and reached over Google's cloud. [verify] on the Seoul-match numbers.
The machine
Mastering the game of Go with deep neural networks and tree search was received 11 November 2015, accepted 5 January 2016 and published in Nature 529, 484–489, on 28 January 2016 — six weeks before Seoul, which matters, because everything the public learned in March had been documented in advance. The pipeline, in the paper's own numbers:
- A supervised-learning (SL) policy network: 13 layers, "trained… from 30 million positions from the KGS Go Server," predicting held-out human moves with an accuracy of 57.0% using all input features and 55.7% on raw board position and move history, against a prior state of the art of 44.4%.
- A fast rollout policy: 24.2% accuracy, but 2 microseconds per move instead of 3 milliseconds.
- A reinforcement-learning (RL) policy network, initialised from the SL network and improved by self-play. It "won more than 80% of games against the SL policy network" and 85% against Pachi, a strong open-source Monte-Carlo program, using no search at all.
- A value network, trained by regression on 30 million distinct positions, each sampled from a separate self-play game — separate games because training on positions from the same game overfit badly (test MSE 0.37 against 0.19 training; the self-play set gave 0.234 and 0.226).
- Monte Carlo tree search combining the two, expanding leaves with the policy network as a prior and evaluating them by mixing the value network with a fast rollout.
Single-machine AlphaGo won 494 of 495 games (99.8%) against Crazy Stone, Zen, Pachi, Fuego and GnuGo; the distributed version won 77% against the single machine and 100% against the other programs. In October 2015 it had beaten Fan Hui, a professional 2-dan and winner of the 2013, 2014 and 2015 European championships, 5–0 — in the abstract's words, "the first time that a computer program has defeated a human professional player in the full-sized game of Go, a feat previously thought to be at least a decade away."
One sentence in that paper is the most important thing in this entry, and it is almost never quoted. Describing the search, the authors write: "It is worth noting that the SL policy network pσ performed better in AlphaGo than the stronger RL policy network pρ, presumably because humans select a diverse beam of promising moves, whereas RL optimizes for the single best move." The network proposing candidate moves inside the machine that played move 37 was the one trained to imitate humans. The stronger self-play network was tried in that slot and was worse there. Hold that against "beyond imitation of its training data" and the claim does not die, but it has to be stated far more carefully than the proposal line states it.
The move
Move 37 was Black's, a shoulder hit on the fifth line. Conventional teaching put that contact play on the third or fourth line; the fifth was held to be inefficient, too far from the edge to take territory and too committal to be pure influence. Lee was out of the room when it landed. Michael Redmond, the 9-dan commentating in English, said on air: "That's a very surprising move," and then "I wasn't expecting that. I don't really know if it's a good or bad move at this point." Fan Hui, working as an analyst and by then the only human alive who had lost a formal match to the thing, left his seat and said: "It's not a human move. I've never seen a human play this move."
I could not confirm the board coordinate against a primary game record. The fifth-line shoulder hit is reported identically everywhere; the coordinate is either absent or inconsistent in what I could reach, and the SGF viewer I found would not render. This entry therefore does not state one. [verify]
Three days later, in game 4, Lee played move 78, a wedge into the centre — Gu Li called it the divine move on the Chinese broadcast — and AlphaGo fell apart. David Silver's account, given to Lex Fridman in 2020, is that the failure was not a one-off: "AlphaGo in around one in five games would develop something which we called a delusion, which was kind of a hole in its knowledge where it wasn't able to fully understand everything about the position," and that such holes "persist for tens of moves." Asked afterwards why he had played 78 when nobody else had considered it, Lee said: "It was the only move I could make." He came to the press conference smiling: "This one win is so valuable and I will not trade this for anything in the world."
What the one-in-ten-thousand number actually measures
This is the load-bearing correction, and it is checkable.
The number's provenance is David Silver, speaking to Cade Metz at WIRED during the match. It is not in the Nature paper and it is not in any DeepMind technical document I could find. wired.com is not fetchable from this machine, so the figure here is taken from the several secondary accounts that quote Metz consistently, and from DeepMind's own present-day page, which states: "In game two, it played Move 37 — a move that had a 1 in 10,000 chance of being used." [verify] against the WIRED original.
Grant the number. What is it a number about? It is the SL policy network's estimated probability, and that network was trained on 30 million positions from KGS, an online server, from games by 6-to-9-dan KGS players — amateur ranks, not professional ones — and it reproduced held-out moves from that corpus 57% of the time. So the honest reading of "one in ten thousand" is: one model, which gets 57% of moves in its corpus right, estimates that this move appears about once per ten thousand in a body of strong amateur online Go. That is a fact about a corpus. It is not a fact about Go, and it is emphatically not a fact about what humans would or would not play.
The cleanest demonstration of that is what happened to the number. KataGo's main developer, posting as Polytope in May 2025, reports that "KataGo's raw policy prior for a recent net puts ~25% mass on the move." Same move, same position; a network trained on self-play rather than on KGS amateurs finds it one of the natural candidates. The move did not travel from 0.01% to 25% because Go changed in nine years. It travelled because the measuring instrument changed, and the original instrument had been built to answer a different question — what do people on a Go server do — which the world then read as what can a human conceive.
The same figure of one in ten thousand is reported for Lee's move 78, from the same source. That should settle it. A number that assigns identical improbability to the machine's alien move and to the human's divine one is not measuring alienness; it is measuring rarity in a corpus, and both moves were rare in it.
Whether it was right
DeepMind's current public page says: "This pivotal and creative move helped AlphaGo win the game and upended centuries of traditional wisdom."
The strongest independent Go-technical assessment I could find says otherwise. On a May 2025 question thread, Polytope — again, KataGo's developer, which makes this the analysis of the person who built the tool everyone else uses to check — gives three findings:
- On whether it needed deep search: "there aren't any critical tactics related to it"; the variations "aren't too complicated"; the move "boils down to fuzzy overall intuition." It was an intuitive option, not a buried tactical resource dug out by brute force.
- On whether the game turned on it: "the evaluation of the position doesn't change much through that move or the 2-3 moves so it's not like the game swung on that move." Lee's reply at move 38 was "indeed fine."
- On whether it was uniquely best: "both players do have other ways to play that are ~equally good, so the move also is not a unique good/best move."
A second respondent on the same thread, DAL, puts it flatly: the move "had a limited impact on the match and was not an optimal move as scored by stronger contemporary Go engines." [verify] — this is a public discussion thread, read directly, not a peer-reviewed analysis, and it is cited here as expert testimony rather than as a study.
So: a good move, not the best move, not decisive, and reachable by intuition rather than by search. AlphaGo went on to win game 2 by resignation 174 moves later. What is left of the story after that subtraction is still real and still interesting — a machine played a move strong professionals found shocking, was not punished for it, and the shock turned out to be a defect in the professionals' priors rather than in the move — but it is a much smaller claim than "pivotal," and "pivotal" is the vendor's word.
Why a reading would cite it
The admission case: this is the canon's cleanest instance of a real result whose public meaning was fixed inside forty-eight hours and has been drifting away from the record ever since — and the drift is measurable, because the record is a scoresheet and the position can still be loaded into an engine.
Almost nothing else in this canon has that property. mycin-1976's evaluation cannot be rerun. watson-jeopardy-2011's clinical claims were never tested by the outcome study that would have settled them. deep-blue-1997's machine was dismantled and its logs are partial. Move 37 is a position on a 19×19 board; any reader with a laptop can put it into KataGo and see for themselves that the evaluation does not swing. The gap between what happened and what is said to have happened is unusually large here, and unusually cheap to close. That combination is the entry's whole value, and it is a lesson about how AI stories form rather than about Go.
This is the third distinct failure shape in the canon's demonstration series, and a reading should choose deliberately among them. deep-blue-1997: the capability was real and had nowhere to go — IBM declined the rematch and dismantled the machine. watson-jeopardy-2011: the demonstration was impeccable, the deployment happened, and the deployment was wrong in the field. This entry is neither. Here the capability was real, the method deployed spectacularly — AlphaGo Zero, AlphaZero, MuZero, AlphaFold, AlphaEvolve all descend from this line — and the thing that was wrong was the story. Cite Deep Blue for a win that led nowhere; cite Watson when a demo is offered as evidence a product will work; cite this when a specific artefact is being asked to carry a general claim about what a system can do.
Three concrete occasions, two of them from this project's own window.
One: "a Move 37 moment" as a live 2026 claim-form. The phrase is doing real rhetorical work right now. In Quillette on 15 April 2026, Michael Morgenstern wrote: "Mythos is Move 37 for computer security. It shows us that we've been playing in a tiny corner of the map, apparently besting decades of security research in a matter of days." Eleven days later, on 26 April 2026, Peter Diamandis reported that Lila Sciences — a Flagship Pioneering company that has raised about $550 million to build "scientific superintelligence" — "calls these discoveries 'Move 37 moments' — and they report that they've been happening across every domain since late 2025," in a piece that also asserts "No human in the 2,500-year history of Go had ever played that move." That last sentence is unverifiable in principle, is a straight escalation of a number that never said it, and sits one paragraph away from a company's own account of its unaudited results. This is what the entry is for. When a firm reports having had a Move 37 moment, the canon should be able to say what the original one was: a good-but-not-best move, not decisive in its own game, whose surprise was measured against a corpus of amateur online games — and, crucially, adjudicated in public by a third party the vendor did not choose. The borrowings have the metaphor and not the adjudication. The phrase launders an unverifiable claim through a verified one.
Two: novelty-from-RL claims, and what the honest version looks like. digests/2026-08-15-12.md records Crouzeix's conjecture, open since 2004, acquiring two independently claimed proofs, both disclosing model use — Shanmu Jin's manuscript after a 16-hour GPT-5.6-Sol run, and Lorist and Schwenninger's arXiv:2608.03841 of 4 August 2026 — neither yet through peer review. That is the good shape: a checkable artefact, a named model, a disclosed method, and an explicit statement of what has not happened yet. Move 37 is the base case for what happens when such a claim is checked: the check was possible, it took nine years and a better engine, and it downgraded the claim without destroying it. A reading grading an "AI discovered something new" story should expect the same arc and should not wait nine years to start it.
Three, and the most transferable: "superhuman" is a claim about a distribution, not about a system. The same digest records Anthropic raising its own catastrophic-misalignment rating from "very low" to "low" and saying plainly this is "an uncertainty adjustment rather than a new finding" — because safety benchmarks are saturating and R&D-capability measurement is degrading. The instruments no longer resolve the system. Go is the one domain where that worry has been settled empirically, and the answer is bad. Section 3 has the detail; the citable sentence is that seven years after the world agreed Go was solved above human level, an amateur beat a superhuman Go engine fifteen games out of fifteen-minus-one by exploiting a shape the engine could not see — and the defences did not hold. A reading that needs a worked example of a system being genuinely superhuman on the distribution it was measured on and trivially beatable just off it should cite this entry, not a hypothetical.
Two further occasions the lenses reach and this entry answers. Under capabilities, the digest's own caveat — "every number above except Qwen's weights is vendor-reported" — has its counterpart here: the "pivotal" in DeepMind's description of move 37 is vendor prose about a ten-year-old game that anyone can check, and it is still wrong. Vendor framing does not decay just because the underlying result was honest. Under concentration, the Seoul match is what a capability demonstration looks like when the evaluator is not the vendor: an opponent the company did not pick, a rule set it did not write, a referee, a public game record, and a loss on camera in game 4. Very little in 2026 clears that bar.
What it got right, and what it got wrong
moment does not require this section. Five dated claims attach and all are due.
Claim 1 — the field's own timeline. Made May 2014. Due 2024. Beaten by roughly eight years.
The claim: in Alan Levinovitz's WIRED piece "The Mystery of Go, the Ancient Game That Computers Still Can't Win" (12 May 2014) — which is reference 31 of the Nature paper, so DeepMind was grading itself against it — Rémi Coulom, author of Crazy Stone and the man who introduced Monte Carlo tree search to Go, put a decade on beating a top professional without handicap, while noting he disliked making predictions. The Nature abstract's own phrasing is "a feat previously thought to be at least a decade away."
What happened: Fan Hui fell in October 2015, seventeen months later; Lee Sedol in March 2016, twenty-two months later. A wrong forecast in the fast direction, from the single best-informed person in the field, hedged at the time he made it. That is the base rate this canon exists to accumulate, and it cuts against the usual assumption that expert AI forecasts run optimistic. Coulom's error was pessimism, and his hedge was honest.
Claim 2 — Lee Sedol's. Made 22 February 2016. Due 15 March 2016. Wrong by four games.
Covered above and not belaboured. The instructive part is the reason: Lee reasoned from the Fan Hui games, which were five months stale, against a system whose improvement rate he had no way to observe. In 2026 that failure mode is universal — a public capability estimate is always an estimate of a checkpoint, and the checkpoint is always older than the system. digests/2026-08-15-12.md records the extreme version: Anthropic disclosing "Model 2," a model better than the one it ships, already in heavy internal use, with no plans to release it. Every outside judgement of that lab's capability is Lee Sedol reading the Fan Hui games.
Claim 3 — "beyond imitation of its training data." Made March 2016. Properly tested October 2017. Upheld, but not by this move.
The claim is the one the proposal line makes and the one the whole Move 37 genre rests on.
What happened: the strong test came from DeepMind itself. AlphaGo Zero, published in Nature in October 2017, was trained purely by self-play with no human game data at all, and surpassed AlphaGo Lee — the version that played Seoul — within three days, winning 100 games to 0. That is what "beyond imitation" looks like when it is demonstrated rather than illustrated: remove the human corpus entirely, and the system gets better.
Graded strictly, the claim is upheld and the evidence usually cited for it is the wrong evidence. Move 37 came out of a system whose move-proposing network was trained to imitate humans, in a search the paper says worked better with the imitating network than with the stronger self-play one. AlphaGo Zero came out of no human data at all. The canon should record that the popular story attaches the conclusion to the weaker exhibit, nineteen months early, and that the stronger exhibit — which is unambiguous, which is in Nature, and which nobody has ever had to walk back — is rarely named in the same breath.
Claim 4 — that AI would kill Go. Made continuously, March 2016 onward. Wrong, with one ugly qualification.
What happened: participation did not collapse and professional play got better. Shin, Kim, van Opheusden and Griffiths, PNAS 120(12), 2023, analysed more than 5.8 million move decisions by professional players across 71 years (1950–2021), using KataGo to score them and generating 58 billion counterfactual patterns. They found that human decision quality improved after the arrival of superhuman AI, that novel moves — previously unobserved — became more frequent, and that novelty became more strongly associated with quality afterwards (interaction coefficient β₃ = 0.515, p < 0.001). The authors are careful, and this entry repeats their care: they state that causality is not established and that generalisation beyond Go is an open question. The mechanism, on their account, is that AI licensed exploration rather than supplying answers.
The concrete residue is visible in the game. The early 3-3 invasion, which Go teachers had discouraged for generations, is now ordinary professional play, adopted after professionals studied AlphaGo's games; DeepMind's own April 2017 write-up, co-authored by Fan Hui — the man who said "it's not a human move" and then spent a year working with the machine — sets out the refinement.
The ugly qualification, which belongs in the same paragraph: the same tool became a cheating vector almost immediately. In November 2020 the Korea Baduk Association suspended Kim Eun-ji, a 13-year-old 2-dan and the youngest professional in Korea, for one year, after she admitted using AI assistance to beat a 9-dan national team member in an online cyberORO event on 29 September 2020. "AI made the humans better" and "AI made it impossible to trust an online result" are the same fact.
Claim 5 — that Go was solved above human level. Made March 2016, restated May 2017, and refuted in 2023.
The claim: AlphaGo Master beat Ke Jie 3–0 at Wuzhen, 23–27 May 2017. Ke, then the world's top-ranked player, said after game two: "Last year, I think the way AlphaGo played was pretty close to human beings, but today I think he plays like the God of Go." Lee Sedol retired from professional Go on 19 November 2019, telling Yonhap: "Even if I become the number one, there is an entity that cannot be defeated."
What happened: Wang, Gleave, Tseng, Pelrine, Belrose, Miller, Dennis, Duan, Pogrebniak, Levine and Russell, "Adversarial Policies Beat Superhuman Go AIs," ICML 2023 (arXiv:2211.00241, first posted 1 November 2022). Abstract, verbatim: "We attack the state-of-the-art Go-playing AI system KataGo by training adversarial policies against it, achieving a >97% win rate against KataGo running at superhuman settings. Our adversaries do not win by playing Go well. Instead, they trick KataGo into making serious blunders. Our attack transfers zero-shot to other superhuman Go-playing AIs, and is comprehensible to the extent that human experts can implement it without algorithmic assistance to consistently beat superhuman AIs. The core vulnerability uncovered by our attack persists even in KataGo agents adversarially trained to defend against our attack."
The exploit is cyclic groups: the engines do not evaluate a cyclic shape correctly, believe it invulnerable, and decline to defend it. Kellin Pelrine, a human, playing under ordinary conditions on KGS with no computer assistance during play, beat the bot JBXKata005 in 14 of 15 games. The follow-up — Tseng, McLean, Pelrine, Wang and Gleave, "Can Go AIs Be Adversarially Robust?", AAAI 2025 (arXiv:2406.12843) — tested three defence families including KataGo's own adversarial training, and found that "though some defenses protect against previously discovered attacks, none withstand freshly trained adversaries": 65% for one new adversary against the hardened end-of-2023 network at 4,096 visits, 75% for a qualitatively different "gift" exploit at 512 visits.
This is the most important grade in the entry and it is almost never attached to this story. "Superhuman at Go" was true and remains true on the distribution of positions that arise when both sides are trying to play Go. Off that distribution it was never true, nobody noticed for six years, the engines themselves gave no signal, and the obvious fixes did not work. Every 2026 claim of superhuman performance in a domain — coding, mathematics, security, medicine — inherits that structure exactly, and Go is the domain where the check has actually been run.
What it got right that is consistently forgotten
AlphaGo was measured by an opponent it did not choose, under rules it did not write, in front of a referee, and it lost a game on camera. Game 4 is part of the record. DeepMind published the delusion rate — one game in five — through its own research lead, and published the Nature paper six weeks before the match rather than after it. Set against the 2026 norm of vendor-run evaluations on vendor-chosen suites reported by the vendor on models the evaluator cannot obtain, the March 2016 demonstration is close to best practice, and the canon should say so while it is criticising the story told about it. The problem was never the evidence. It was the summary.
Commonly misused as
moment does not require this section. godel-incompleteness-1931 established that the canon writes one wherever a result is repeatedly made to say something it does not, and here it is where the entry earns its length.
1. "A move no human would ever have played." What exists is one network's estimate that the move occurs about once in ten thousand times in a corpus of strong-amateur online games — a network that reproduces that corpus 57% of the time. A modern self-play network gives the same move about a quarter of its policy mass. The identical one-in-ten-thousand estimate was produced for Lee Sedol's move 78, played by a human, three days later. The number measures corpus rarity. It has never measured human inconceivability, and the escalation to "no human in the 2,500-year history of Go had ever played that move" — as circulated in April 2026 — is a claim about roughly forty million recorded and unrecorded games that nobody has checked and nobody could.
2. "The move that won the game." DeepMind's own site calls it "pivotal… helped AlphaGo win the game." KataGo's developer says the position's evaluation "doesn't change much through that move," that Lee's response was fine, and that both players had roughly equally good alternatives. AlphaGo won 174 moves later. The move is better described as the moment the audience understood the game was lost, which is a fact about the audience.
3. "Proof that AI transcends its training data." The most defensible of the four and still not what this move shows. The move-proposing network inside AlphaGo was the human-imitation network, and the Nature paper records that it outperformed the stronger self-play network in that role. The proposition is true — AlphaGo Zero, October 2017, no human data, 100–0 against the Seoul version — and the honest citation is that paper. Using move 37 for it is using the anecdote when the experiment is available and cleaner.
4. "A Move 37 moment" as a general licence. The 2026 usage and the one this entry most expects to be needed for. The phrase now means our system did something that looked wrong to experts and turned out to be right, and it is being applied to results that have not been adjudicated by anyone outside the company making the claim. The original had a third-party referee, a published rule set, an opponent chosen by neither side's convenience, a scoresheet, and a televised loss. The metaphor transfers the credibility and leaves the adjudication behind, which is precisely backwards: the adjudication was the valuable part. A reading meeting this phrase should ask who the referee was. Usually there wasn't one.
5. "It showed machines had become creative." The canon should decline to adjudicate creativity and note what the record can support instead. A system optimising for win probability, proposing candidates from a network trained on amateur human games and evaluating them with a network trained on its own play, selected a move that strong professionals had ruled out on general principle and was not punished. Whether that is creativity is a question about the word. What it is, unarguably, is evidence that a human consensus can be wrong for centuries in a domain with perfect information, complete records and thousands of full-time experts — and that the error was invisible from the inside until something that had not been taught the consensus went looking. That claim is smaller, is fully supported, and happens to be more alarming than the one it replaces.
Sources
Primary.
- D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel and D. Hassabis, "Mastering the game of Go with deep neural networks and tree search," Nature 529, 484–489, 28 January 2016, doi:10.1038/nature16961. Read in full from DeepMind's own hosted PDF. Source of every training figure above (30 million KGS positions, 57.0%/55.7%/44.4%, 24.2% rollout, >80% and 85% for the RL policy network, the 30-million self-play value-network set and its MSEs, 494/495, the 1,202-CPU/176-GPU distributed configuration, the Fan Hui result and rank), of the SL-beats-RL sentence, and of the Deep Blue comparison that establishes
descends_from. - Google DeepMind, "AlphaGo" research page, read directly, 15 August 2026. Source of "over 200 million people worldwide," the honorary 9-dan, and the "1 in 10,000" and "pivotal and creative move helped AlphaGo win the game" claims that section 1 grades.
- Lucas Baker and Fan Hui, "Innovations of AlphaGo," DeepMind blog, 10 April 2017, read directly. Source of the 3-3 invasion account. Contains no discussion of move 37, which is itself worth knowing: DeepMind's own technical write-up of AlphaGo's innovations does not use the example the world uses.
- M. Shin, J. Kim, B. van Opheusden and T. L. Griffiths, "Superhuman artificial intelligence can improve human decision-making by increasing novelty," PNAS 120(12), e2214840120, 2023, read via PubMed Central. Source of the 5.8 million moves, 1950–2021, KataGo as estimator, 58 billion counterfactuals, the β₃ = 0.515 interaction, and the authors' own causality caveat.
- T. T. Wang, A. Gleave, T. Tseng, K. Pelrine, N. Belrose, J. Miller, M. D. Dennis, Y. Duan, V. Pogrebniak, S. Levine and S. Russell, "Adversarial Policies Beat Superhuman Go AIs," ICML 2023, arXiv:2211.00241 (v1 1 November 2022, v4 13 July 2023). Abstract read directly and quoted verbatim.
- T. Tseng, E. McLean, K. Pelrine, T. T. Wang and A. Gleave, "Can Go AIs Be Adversarially Robust?", AAAI 2025, arXiv:2406.12843. Source of the three defence families, the 65% and 75% figures and the "none withstand freshly trained adversaries" finding. Read via search summaries of the paper and the FAR.AI project page rather than the PDF end to end. [verify].
- Wikipedia, "AlphaGo versus Lee Sedol," read for its citations rather than its prose. Source of the per-game resignation move numbers, the komi and time controls, the prize split, Lee's $170,000, the honorary 9-dan and the Economist-sourced Seoul hardware figures.
Secondary, and used as such.
- Caleb Biddulph, "What was so great about Move 37?", LessWrong, 29 May 2025, read directly, and specifically the replies from Polytope — KataGo's principal developer — and from DAL. This is the entry's whole basis for the "good but not best, and not decisive" finding, and for the ~25% modern policy prior. It is a discussion thread, not a study; it is here because it is the best-qualified analysis I could find and because the claim it contradicts is a vendor's marketing page. A reading leaning hard on it should say where it came from. [verify] by loading the position into KataGo, which is the one check in this entry anyone can run at home.
- Cade Metz, WIRED, March 2016 — "The Sadness and Beauty of Watching Google's AI Play Go" and "In Two Moves, AlphaGo and Lee Sedol Redefined the Future." The origin of the one-in-ten-thousand figure for both move 37 and move 78, of Redmond's on-air lines and of Fan Hui's "it's not a human move." Not read at first hand — wired.com is not fetchable from this machine — and reconstructed from several independent secondary accounts that quote the same passages consistently, plus DeepMind's own restatement of the figure. [verify].
- David Silver, interviewed by Lex Fridman, episode 86, 2020, read via a published transcript. Source of the "delusion… in around one in five games" passage and of Silver's own description of the fifth-line move and of move 78.
- Alan Levinovitz, "The Mystery of Go, the Ancient Game That Computers Still Can't Win," WIRED, 12 May 2014 — reference 31 of the Nature paper. Source of Rémi Coulom's decade estimate and his stated dislike of predictions. Taken from the American Go E-Journal's contemporaneous write-up and other summaries, not the original. [verify].
- Yonhap News Agency, 27 November 2019, via CNN and China Daily: Lee Sedol's retirement, dated 19 November 2019, and the "entity that cannot be defeated" line.
- Xinhua and China Daily, 27 May 2017, for the Ke Jie result at the Future of Go Summit, Wuzhen, 23–27 May 2017, and the "God of Go" quotation.
- The Korea Times, November 2020, and the American Go E-Journal's March 2021 round-up, for the Kim Eun-ji suspension, the 29 September 2020 cyberORO game and the one-year ban.
- D. Silver et al., "Mastering the game of Go without human knowledge," Nature 550, October 2017 (AlphaGo Zero). The 100–0 result against AlphaGo Lee after three days of self-play is taken from the paper's summaries and from the Wikipedia entry's citation of it rather than from the PDF. [verify].
- Michael Morgenstern, "Move 37: AI and the End of Cyber Security," Quillette, 15 April 2026, read directly, for the "Mythos is Move 37 for computer security" passage.
- Peter H. Diamandis, "Scientific Superintelligence: The Deep Blue Moment," Metatrends, 26 April 2026, read directly, for the Lila Sciences "Move 37 moments" attribution and the "2,500-year history" assertion. Note that the "Move 37 moments across every domain since late 2025" phrasing is Diamandis reporting Lila rather than a quotation from a named Lila executive; the only directly attributed Lila quotation in that piece is CEO Geoffrey von Maltzahn's definition of scientific superintelligence. Lila's funding — about $550 million across a March 2025 seed and a Series A — is from press coverage and is background, not a claim this entry relies on.
- Korea Herald / UPI / PR wire coverage, 9 March 2026, for the tenth-anniversary event held at the Four Seasons Hotel Seoul, at which Lee Sedol appeared with the enterprise agent company Enhans and said "Ten years ago, I competed against AI, but now we are moving forward through collaboration." Recorded with its frame intact: this was a commercial product launch built around the anniversary, and the quotation is Lee's, delivered at a vendor's event.
- People's Daily Online, 9 March 2016, for Lee's 22 February 2016 press conference forecast of 5–0 or 4–1, and ABC News and CBC, 13–15 March 2016, for the post-game-4 and post-match quotations.
Attempted and failed, so that nothing above silently depends on it: wired.com is not fetchable from this machine, which is why the single most quoted number in the entry is second-hand; nature.com redirects to an identity provider and the Nature HTML could not be read, so the paper was read from DeepMind's hosted PDF instead; senseis.xmp.net returned HTTP 403 and news.ycombinator.com returned HTTP 429; the alphago-games.com viewer rendered no game record, which is why this entry states no board coordinate for move 37; and local PDF-to-text extraction was refused by the sandbox, so the Nature paper was read as page images. Claims circulating that this entry deliberately does not use: viewership figures for the match other than DeepMind's own "over 200 million," which vary widely by source and measurement; and any assertion about whether a fifth-line shoulder hit had precedent in professional play before 2016, which I could not establish either way.
The 2026 citation occasions — the Crouzeix's conjecture proofs of 14 August 2026, Anthropic's August 2026 Risk Report and its "Model 2" disclosure, and the digest's own "claimed vs. demonstrated" caveat — are as recorded in this project's digests/2026-08-15-12.md, which holds the primary links. They are named here as occasions to cite this entry, not as evidence for anything in it. Nothing in this file is evidence, nothing in it is deposited in the ledger, and nothing in it touches the needle.