← The canon · AItopiaOrAImageddon?
The LLaMA leak and the Llama 2 release
moment · Meta AI — Hugo Touvron, Guillaume Lample and the LLaMA team; the leak's uploader is anonymous · 2023
Something that happened and changed what people expected next.
Descends from Scaling Laws for Neural Language Models, Deep Reinforcement Learning from Human Preferences.
Filed under moment, and it belongs there — but the id names two events and only one of them is the moment. canon/proposals.md puts llama-weights-2023 in the moment section and describes it as a pair: "LLaMA leaked (March 2023), Llama 2 released openly (July 2023) — frontier-adjacent weights became downloadable and unrecallable." The house definition, argued out by eliza-1966 when it talked its way out of the kind and used since by clippy-1996, siri-2011, expert-systems-collapse-1987, deep-blue-1997, watson-jeopardy-2011, alphago-move-37-2016, chatgpt-2022, alphafold2-2020 and alexnet-2012, is a date on which something visibly happened in public. Both halves satisfy it. But they are not the same species of event, and the entry is worth less if it blurs them:
- On or about 2 March 2023, someone Meta had not authorised put 219 GB of Meta's weights into public circulation and nobody could get them back. That is the moment. It was involuntary, it was instantaneous, and it is the only irreversible act in the whole sequence.
- On 18 July 2023, Meta did on purpose a somewhat weaker version of what had been done to it. That is not a moment; it is a concession, and reading it as an independent decision is the single most common error made about this entry. Four and a half months separate the two, and what happened in between is that the gate stopped being worth defending.
chatgpt-2022 observed that it was the first moment in this canon whose evidential content was what the public did rather than what the system did. This one goes a step further: its evidential content is what the public did with an artifact against the wishes of the party that made it, and then what that party concluded from being unable to stop them. Nothing became possible in March 2023 that had not been possible in February. What changed is that the set of people who could exercise the possibility stopped being a list Meta maintained.
The temptation is to file this as a limit, and that would be wrong in a way worth stating up front. What the entry is for is an impossibility claim — you cannot un-distribute weights — and the canon has a kind for impossibility claims. But godel-incompleteness-1931, turing-halting-1936 and perceptrons-1969 are all proved. Their impossibility is a theorem, it holds for every future, and the interesting work is policing what the theorem does not say. Unrecallability is not a theorem. It is an empirical regularity with named mechanisms — cheap mirroring, unsettled copyright, a demand curve, and the fact that a file does not know who is holding it — and every one of those mechanisms is contingent. A regime that made weight distribution a strict-liability offence, or a hardware attestation scheme that refused to load unsigned checkpoints, would not refute a theorem; it would change a fact. Filing this as limit would license exactly the argument section 4 exists to refuse. It is a moment: a date, a public, a thing that happened, and a regularity that has held for three and a half years and is graded as such in section 3.
descends_from is [scaling-laws-2020, rlhf-christiano-2017], and both edges are documented on the page rather than inferred.
scaling-laws-2020is the ancestor of the artifact. The LLaMA paper's entire thesis is a scaling-law argument turned around: "for a given compute budget, the best performances are not achieved by the largest models, but by smaller models trained on more data" — and then the move that made the leak matter, which is to optimise for the inference budget rather than the training budget, because "the performance of a 7B model continues to improve even after 1T tokens." A model deliberately shaped to be cheap to run is a model that fits on hardware a person owns. The paper's actual cited parent here is Hoffmann et al. (2022), "Chinchilla", which has no id incanon/and is not inproposals.md;scaling-laws-2020is Chinchilla's own ancestor and the nearest thing on disk. Naming achinchilla-2022in the header would be inventing an id, which the job is forbidden to do, so it is named here instead.rlhf-christiano-2017is the ancestor of the second half. Llama 2 is one of the largest RLHF artifacts ever documented in public, andrlhf-christiano-2017already quotes its numbers on disk — 1,418,091 human binary comparisons, from Table 6 of arXiv:2307.09288, which that entry uses to measure how far the technique travelled from a $36 Atari run. Read the two files together and the point sharpens: what Llama 2 released was not only a capability but a behaviour policy, andrlhf-christiano-2017establishes that the policy comes off for the price of a coffee while the capability does not.
transformer-2017 is behind both and is not claimed directly. It is scaling-laws-2020's declared ancestor already, and stacking it here would make the graph a list of everything true rather than a record of descent.
**What is not an ancestor, despite the pull:**
chatgpt-2022is the weather, not the parent. LLaMA's paper does not cite it; the work was done in parallel and the design decisions came from Chinchilla, not from ChatGPT. The honest qualifier is that the derivative wave has a real edge to it — Stanford's Alpaca (13 March 2023) fine-tuned leaked LLaMA-7B on 52,000 instructions generated by OpenAI'stext-davinci-003, so the first thing the world built on the leaked weights was distilled from ChatGPT's own lineage. That is an edge from Alpaca, which has no id, and not from this entry. Read the two files together anyway:chatgpt-2022holds the day a capability became available to rent, and this file holds the day one became available to own.asilomar-1975is the inverse case, not the parent, and the inversion is the most useful thing either file offers the other. Asilomar is the canon's worked example of a field pausing its own most dangerous work and writing its own rules, and it holds because the participants had a monopoly on the technique. Here, restraint was attempted — a case-by-case access list, a noncommercial licence, an explicit misuse rationale — and it lasted seven days. A reading that cites Asilomar for the proposition that voluntary restraint can work needs this file next to it for the condition under which it cannot.frankenstein-1818already worked the same ground from the fiction side: Victor refuses to tell Walton the method, and that entry grades the refusal as the most defensible decision in the book while noting it worked only because he had no competitors. This file is that footnote with a date on it.
Ids named here that do not exist in canon/ and are not being invented into the header: chinchilla-2022, mistral-2023, deepseek-r1-2025, gpt-oss-2025, instructgpt-2022, alpaca-2023, bletchley-declaration-2023 (proposed, unwritten), fli-pause-2023 (named by asilomar-1975, unwritten).
What it is
The release that was designed not to be one
On 24 February 2023 Meta AI announced LLaMA — four models at 7B, 13B, 33B and 65B parameters — and published the paper two days later as arXiv:2302.13971. The technical claim was the one that mattered: "LLaMA-13B outperforms GPT-3 (175B) on most benchmarks, despite being 10× smaller," and the 65B was "competitive with the best models, Chinchilla-70B and PaLM-540B." Training used publicly available datasets exclusively, "without resorting to proprietary and inaccessible datasets" — 1.4 trillion tokens for the two large models, 1 trillion for the 7B. The 65B took 1,022,362 A100-80GB GPU-hours and the paper puts the whole development effort at roughly 2,638 MWh and 173 tCO2eq, with the release justified in part on the grounds that it "will reduce future carbon emission since the training is already done."
The distribution terms were the opposite of the technical posture. The inference code went out under GPLv3. The weights did not go out at all:
> To maintain integrity and prevent misuse, we are releasing our model under a > noncommercial license focused on research use cases. Access to the model will > be granted on a case-by-case basis to academic researchers; those affiliated > with organizations in government, civil society, and academia; and industry > research laboratories around the world.
The paper's own last line is "We release all our models to the research community." Both sentences are Meta's, published within 48 hours of each other, and the tension between them is the whole entry. A canon reading this cold should notice that the thing that leaked was never open, and the word "open" has been retrofitted onto it ever since.
The seven days, dated
This is the part worth having on file, because the sequence is faster than anybody's memory of it and every retelling compresses it.
- 24 February 2023 — announcement; the request form goes up.
- 2 March 2023 — pull request #73 is opened against Meta's own
facebookresearch/llamarepository: "Save bandwidth by using a torrent to distribute more efficiently." It proposes adding a magnet link to Meta's README. It is a joke, and it is also a working download. It accumulated over 800 thumbs-up reactions, was closed unmerged, and Meta maintainers eventually locked the thread as "too heated" on 1 September 2023 — six months later, which is its own small finding about how long the argument stayed live. - first days of March 2023 — a torrent of the full weight set is posted to 4chan and spreads through the AI communities. Dating caveat, stated because the entry's whole point turns on the interval: Wikipedia dates the upload to 3 March; Motherboard's report describes an upload "last week" and is indexed variously as 3 March and 7 March; PR #73 fixes the outside bound, because a working link existed on 2 March. The honest statement is six to nine days after the announcement, and no source I could reach pins the hour.
- 6 March 2023 — Meta issues takedown requests over Hugging Face repositories hosting the weights, which Hugging Face honours.
- ~10 March 2023 — Georgi Gerganov publishes llama.cpp. The original README states the goal plainly: "The main goal is to run the model using 4-bit quantization on a MacBook."
- 11 March 2023 — Simon Willison runs a GPT-3-class model on his own laptop and writes the sentence that dates the shift better than any benchmark: "I thought it would be a few more years before I could run a GPT-3 class model on hardware that I owned. I was wrong: that future is here already."
- 13 March 2023 — Stanford releases Alpaca: LLaMA-7B fine-tuned on 52,000 instruction-following demonstrations for under $600. Within days people are running derivatives on Raspberry Pis and a Pixel 6.
- 20 March 2023 — Meta files a DMCA notice through OpSec Online against
github.com/shawwn/llama-dl, a script that downloaded the weights from mirrors. The claimed infringed work is, remarkably, the announcement blog post. The notice asserts that "No one is authorized to exhibit, reproduce, transmit, or otherwise distribute Meta Properties without the express written permission of Meta." GitHub processed it against 403 repositories in the fork network and complied the next day. - ~21 March 2023 — Stanford takes the Alpaca demo offline over hallucinations, hosting cost and content-filter inadequacy. "We feel that we have mostly achieved this goal, and given the hosting costs and the inadequacies of our content filters, we decided to bring down the demo." The code stayed up. This is the entry in miniature: a service can be withdrawn in an afternoon, a distribution cannot be withdrawn at all.
- 27 April 2023 — a counter-notice is filed on behalf of Tensorfork Labs arguing that the weights "were copied from the works used to train LLaMA-dl by a rote automated process and do not reflect any human selection or arrangement" and that "the facts embodied in the LLaMA-dl weights do not have sufficient originality to be copyrightable."
That last document is the most under-cited item in this file. Meta's ability to recall the weights rested on a copyright theory in model parameters that has never been tested in court, and the first person to be served pushed back on exactly that theory within five weeks. As of this entry, github.com/shawwn/llama-dl is live, with 4.1k stars and 394 forks, and its README records what actually happened to the takedown: "Facebook shut off the link a couple hours after this repo went live. I mirrored everything to R2 and updated the script to point to that instead." The legal instrument worked on the origin server and did nothing to the artifact.
What Meta said while it was happening
Meta never denied the leak and never defended the gate. A spokesperson at the time: "It's Meta's goal to share state-of-the-art AI models with members of the research community to help us evaluate and improve those models"; some "have tried to circumvent the approval process"; and — the sentence to keep — "our current release strategy allows us to balance responsibility and openness."
That was said days after the release strategy had demonstrably failed, and four months before Meta replaced it with one that had no gate at all. A reading citing this file for the proposition that a company's stated release posture predicts its next release posture will be citing it backwards.
Llama 2, and what was actually conceded
On 18 July 2023, timed to Microsoft Inspire and announced jointly with Microsoft as "preferred partner," Meta released Llama 2 at 7B, 13B and 70B, with instruction-tuned chat variants, free for research and commercial use, weights downloadable without a case-by-case decision.
Three things about it are consistently misreported.
First, it is not open source, and it was not in 2023 either. The Llama 2 Community License is a bespoke Meta licence, never OSI-approved. Two clauses do the work: a licensee whose products had more than 700 million monthly active users in the calendar month before the release date must request a separate licence from Meta, and licensees may not use Llama 2 outputs to improve other large language models. The first is a competitor carve-out with a threshold chosen narrowly enough to be a list of names; the second is a restriction on field of use. On 28 October 2024 the Open Source Initiative published the Open Source AI Definition 1.0, requiring information about training data, complete code, and the parameters. Llama fails it on more than one axis, and Meta rejected the definition rather than the label — a spokesperson said the OSI bar was too narrow for models of this complexity. The correct word for what Llama 2 was is open-weights, and the canon should use it.
Second, it was accompanied by an Acceptable Use Policy, which is the part of the licence that is unenforceable in exactly the way this entry is about. The AUP governs conduct by people who have already downloaded a file that runs offline.
Third, the safety layer was enormous and detachable. Llama 2's chat models came out of RLHF at a scale nobody had documented before — 1,418,091 binary human preference comparisons, collected weekly. rlhf-christiano-2017 grades what that layer is worth on disk, and the finding that matters here is that fine-tuning strips it for well under a dollar. Meta shipped, in one artifact, the strongest public demonstration of both halves of the open-weights argument: this is how much behaviour costs to install, and this is how little it costs to remove.
The distribution form became the marketing
Within a year, being distributed like a leak was a deliberate aesthetic. Mistral shipped Mistral 7B on 27 September 2023 as a bare magnet link posted from a then-dormant account, before any blog post — Apache 2.0, 7.3B, beating Llama 2 13B on the benchmarks it reported. 404 Media's headline called the result an "Undeletable Chatbot," which puts the vocabulary of this entry into the trade press six months after the leak. On 17 March 2024 xAI released Grok-1 — 314B mixture-of-experts, 86B active, Apache 2.0 — as a 318 GB torrent behind a magnet link. That last item is recorded as a dated fact and graded nowhere in this file; see the boundary note at the end for why.
The through-line is not that torrents are romantic. It is that a release mechanism with no revocation step became normal, and the labs that adopted it were advertising the absence of the step.
Where it had got to by this reading
As of 16 August 2026, three and a half years after the DMCA notices:
- The original leaked weights are still one command away. The Hugging Face mirror
huggyllama/llama-7bshows 292,001 downloads in the last month. Its model card still carries the fiction intact — "you should only use this repository if you have been granted access to the model by filling out this form" — pointing at an approval process that no longer gates anything and for a model four generations obsolete. Nobody is downloading LLaMA-1 for its capability. The number is a measurement of how completely a recall failed. - The open-weights centre of gravity left the country that opened it. Hugging Face's state-of-open-models report of 14 August 2026 puts Alibaba's Qwen above 3 billion downloads in six months, against 418 million for Google and 227 million for Meta, with 460-plus open models and over 300,000 derivatives. On routed inference the shift is sharper still: one industry analysis puts Chinese open-weight models at roughly 61% of tokens consumed on OpenRouter by May 2026, with Meta's share below 1% — that pair of figures is from a secondary analyst source and should be cited as an estimate, not a filing; the Hugging Face download numbers are the harder of the two.
- Meta closed the door and then reopened it, inside four months. The Llama line ended: Llama 4 Scout and Maverick shipped 5 April 2025, Behemoth was delayed in May 2025 and never shipped, and on 17 April 2026 Meta Superintelligence Labs introduced Muse Spark — proprietary, closed weights, explicitly the successor series to Llama. Then on 10 August 2026 Zuckerberg published a 6,500-word essay arguing for open models on competitiveness grounds, Meta released Muse Glimmer (30B dense, Apache 2.0, distilled from Muse Spark, sized to run on one 24 GB consumer GPU), and committed to opening Muse Spark 1.2's weights. The essay's operative line: "I do not believe restricting access to foreign open source models is an effective solution. Our goal should be for American open source models to be the best globally. This requires removing the hurdles that make it harder for American open source models to compete."
Those three bullets are one finding. Meta could stop releasing whenever it liked and did. Meta could resume whenever it liked and did. What Meta could never do, at any point in the three and a half years, was retrieve the file it put out in February 2023. The policy is fully reversible; the artifact is not. Every serious use of this entry runs through that distinction.
Why a reading would cite it
The concentration lens is the axis this project is named for — the Culture and the Sprawl being the same capability under different ownership — and it is the lens that most often has a live event in the window. This file is what a reading reaches for when that event involves weights leaving or not leaving a building. Five concrete occasions, each with the thing the entry actually supplies:
1. Any open-weights release, for the asymmetry. The occasion recurs constantly: in the fortnight before this reading alone, DeepSeek's V4-Pro reached general availability on 13 August, Zhipu's GLM-5.3 landed on 14 August, Alibaba shipped Qwen 3.8 27B under Apache 2.0 on 14 August, and Meta released Muse Glimmer on 10 August. The temptation each time is to score the release as a point for access or a point against safety. What this entry supplies instead is the structural fact that the release decision and the recall decision are not the same decision, because the second one does not exist. A release is not a policy that can be adjusted next quarter; it is an act with no inverse. That belongs in the prose of any reading that treats a weights drop as evidence.
**2. When a lab withholds weights, for what the withholding is worth. Zhipu held GLM-5.3's weights for roughly two weeks pending a security review. OpenAI delayed its first open-weight model on 11 July 2025** with Altman's own statement of the mechanism — "once weights are out, they can't be pulled back" — before shipping gpt-oss on 5 August 2025. Anthropic's stated position of 27 July 2026 puts it identically: "once weights are released they cannot be withdrawn," and "safeguards can be removed, and copies can be downloaded, redistributed, and run on private systems beyond monitoring." This entry is the worked case behind all three sentences, and it says something the sentences do not: the hold is the only moment at which anyone has a decision. A reading covering a withheld release should treat the withholding as the whole of the governance, not as a delay in it.
3. When someone claims a release can be undone. The DMCA record is the test case and it is unambiguous: takedowns to Hugging Face on 6 March 2023, a notice against 403 GitHub repositories on 20–21 March 2023, a counter-notice on 27 April 2023 disputing that weights are copyrightable at all, no adjudication of that question since, and 292,001 downloads of the recalled artifact in the month before this reading. If a reading meets a proposal to claw back a released model, this file is the base rate.
4. When a national-security or catastrophic-misuse argument is made about open weights, in either direction. The entry carries the two documents that keep such an argument honest. Kapoor, Bommasani, Narayanan et al., arXiv:2403.07918 (2024) built a marginal-risk framework and found that "current research is insufficient to effectively characterize the marginal risk of open foundation models relative to pre-existing technologies." NTIA's report of 30 July 2024 concluded that the government "should not restrict the wide availability of model weights for dual-use foundation models at this time" while recommending active monitoring, and the July 2025 America's AI Action Plan went further with a section titled "Encourage Open Source and Open Weight AI." The point for a reading is not that open weights were cleared. It is that the question is marginal risk against an existing baseline, and almost every public claim on either side is a claim about absolute risk. The EU took the same shape from the other end: Article 53(2) of the AI Act exempts genuinely free-and-open models from some provider obligations, and the exemption vanishes entirely for models above the systemic-risk compute threshold — a legislature saying openness earns relief only below a capability line.
5. When the word "open source" is used about a model. OSI 1.0, Meta's rejection of it, and the fact that the artifact at the centre of the open-weights era was gated and leaked rather than opened. This file, next to sparse-moe-2017, gives a reading the sharper axis: that entry demonstrates that "open weights" is not a binary, because a licence can be maximally permissive while the hardware requirement is a rack. Possession and permission are different variables, and LLaMA-13B mattered because it was the first frontier-adjacent model where both pointed the same way.
The honest limit of the citation, stated so it cannot be quietly exceeded: this entry is not evidence that open weights are good, and it is not evidence that they are dangerous. It is evidence about irreversibility, which is a claim about the shape of the decision rather than its merits. A reading that uses it to argue either pole is using it wrongly, and the file says so here so that the misuse is visible when it happens.
What it got right, and what it got wrong
Six dated claims. moment does not require this section, but the claims made around this one have been quoted for three years and most of them are gradeable now.
Claim 1 — Meta: "To maintain integrity and prevent misuse, we are releasing our model under a noncommercial license focused on research use cases." Made 24 February 2023. Due on release. Failed within seven days.
Wrong, and not marginally. A working magnet link was proposed in Meta's own repository on 2 March, six days later. The failure mode is worth naming precisely, because it is the transferable part: the licence was not defeated, it was bypassed. Nobody argued that noncommercial terms permitted redistribution; they simply redistributed. A licence constrains people who can be identified and sued. The access list constrained nobody who did not want to be constrained, because the enforcement surface — the download server — was upstream of the only copy operation that mattered.
The senators' later characterisation of the vetting as "seemingly minimal" is partly unfair and partly correct. Meta ran an approval process; it granted access to a large number of researchers, any one of whom could defect at zero cost and with no attribution risk. The gate's strength was one over the number of people behind it, and that is a property of the design and not of the diligence.
Claim 2 — Meta: "our current release strategy allows us to balance responsibility and openness." Made early March 2023. Graded by Meta's own next move.
Wrong, and graded by the speaker. Said in the days after the leak, in defence of a strategy that had already stopped existing. Four months later Meta abandoned every element of it — no case-by-case approval, no noncommercial restriction, commercial use permitted, weights downloadable by anyone. A company does not replace a strategy it believes is balanced.
This one earns its place because of what it teaches about reading vendor statements in the window immediately after an incident. The statement was not a lie; it was a position held for as long as it took to reach the obvious conclusion, and it was quoted for months afterward as though it were durable.
Claim 3 — Hawley and Blumenthal: the leak's harms. Made 6 June 2023. Due continuously. Right in kind, unquantified in degree, and mis-attributed in the particulars.
The letter warned of "seemingly minimal" protections in an "unrestrained and permissive" release, said Meta "appears to have failed to conduct any meaningful risk assessment in advance," and named the expected harms: spam, fraud, malware, privacy violations, harassment.
Grade it in three parts.
Right about availability. The letter's factual predicate — that the full model was now available "to anyone, anywhere in the world, without monitoring or oversight" — was correct on the day and has been correct every day since.
Right in kind about the harms. Criminal LLM tooling exists and is a market. The category the letter named is real.
Wrong, or at least unsupported, in attribution — and this is the part that gets repeated as though it were settled. The most-cited exemplar of criminal LLM tooling, WormGPT (July 2023, sold at roughly €60–100/month on Hack Forums), was built on GPT-J-6B — an EleutherAI model released in 2021, two years before LLaMA, by people who never gated anything. The letter's causal story does not survive its own headline case. More broadly, arXiv:2403.07918 examined the misuse vectors systematically and found the research base insufficient to characterise marginal risk at all, and NTIA reached the same conclusion at policy level in July 2024. The correct verdict is not "they were wrong"; it is "the harm they named is real, the counterfactual they assumed has never been demonstrated, and three years of subsequent work has failed to demonstrate it in either direction."
Neither senator has published a self-assessment of this letter that I could find, so there is no self-grade to set against an independent one. That absence is itself worth recording: the letter is still cited, and nobody has gone back to score it.
Claim 4 — the anonymous Google memo: "We have no moat, and neither does OpenAI." Published 4 May 2023. Due open-ended. Right about the mechanism, wrong about the companies, and right about the wrong country.
The memo is the sharpest contemporaneous reading of the leak and it names this entry as its cause: "At the beginning of March the open source community got their hands on their first really capable foundation model, as Meta's LLaMA was leaked to the public"; "Within a week, LLaMA is leaked to the public. The impact on the community cannot be overstated"; "They are doing things with $100 and 13B params that we struggle with at $10M and 540B"; "We cannot hope to both drive innovation and control it."
Wrong about the two firms it named. Three years on, both Google and OpenAI are at the frontier, both charge for access, and neither has been commoditised by open weights. If the thesis had held as stated, that is precisely what would have happened.
Right about the mechanism, and the mechanism found a different victim. The memo's actual argument is that the cost of iteration collapses when weights are public, and iteration compounds faster than any single lab can. That happened — just not to Google. It happened to Meta. The company that supplied the commodity is the one whose share of routed open-weight inference is now reported below 1%, whose download counts sit an order of magnitude under a competitor's, and which spent April 2026 to August 2026 arguing publicly about whether to be in this business at all. "We cannot hope to both drive innovation and control it" reads, in 2026, as a description of Meta's position rather than Google's.
Right for the wrong country, which is the interesting part. The memo's "open source community" is implicitly a Western hobbyist-and-startup ecosystem. The commoditisation actually arrived from Chinese labs at industrial scale — DeepSeek from January 2025 onward, then Qwen, Kimi, GLM and MiniMax — and by 2026 the argument in Washington had inverted so completely that Meta's case for open weights is a national competitiveness case. I searched the 2023 commentary for a prediction that the open-weights frontier would be Chinese and did not find one. Absence of a find is not proof of absence, and it is recorded as a search result rather than a fact. But the shape of the 2023 argument — US labs against US regulators, safety against innovation — had no room for the thing that actually happened.
Claim 5 — Zuckerberg: "Starting next year, we expect future Llama models to become the most advanced in the industry." Made 23 July 2024. Due 2025. Wrong, and the line it was made about no longer exists.
Made in Open Source AI Is the Path Forward, alongside Llama 3.1 405B, billed as "the first frontier-level open source AI model." The essay's central argument — that open source keeps power from concentrating "in the hands of a small number of companies" — is a values claim and is not graded here. The forecast is gradeable and it failed on schedule.
2025 came. Llama 4 Scout and Maverick shipped on 5 April 2025 to a poor reception; Behemoth, the flagship the claim depended on, was delayed in May 2025 amid reporting that its performance would not match earlier claims, and never shipped. On 30 July 2025 Zuckerberg's own position had moved to "we'll need to be rigorous about mitigating these risks and careful about what we choose to open source." On 17 April 2026 the Llama series was superseded by a closed model. The claim was not merely missed; the product line it was about was discontinued twenty-one months after it was made.
The August 2026 essay does not retract it, and does not have to. But a reading grading anyone's 2026 open-weights commitment should hold this one next to it, because the same speaker made a maximal open-weights commitment in 2024, closed in 2026, and reopened four months later — all three sincerely, and all three at company scale.
Claim 6 — the entry's own claim: "frontier-adjacent weights became downloadable and unrecallable." Made March 2023 in effect. Standing. Right, with the scope stated exactly.
This is the claim the file exists to carry, so it gets the strictest treatment.
Upheld on the artifact. Every recall instrument available to a large, well-resourced company was used: origin shutdown within hours, platform takedowns within four days, a DMCA notice reaching 403 repositories within a month. Outcome: the weights are on Hugging Face today, mirrored to commercial object storage, downloaded 292,001 times in the past month, and the takedown's own target repository is live with the mirror documented in its README.
Refuted as a claim about practice, and this is where careless citation goes wrong. "Unrecallable" says nothing about whether the next model ships open. Meta demonstrated the negative case at full scale: Behemoth, trained, never released; the Llama line ended; Muse Spark closed for four months. A company's open-weights posture is revocable at will, on a quarter's notice, with no mechanism required. Anyone building on open weights is exposed to that, and the 2026 record is the proof — the shift in the download tables from Meta to Alibaba is partly a story about capability and partly a story about a supplier who left.
Contingent, not proved, on the mechanism. Three things make the regularity hold: mirroring is nearly free, the copyright status of weights is unsettled and was contested at the first opportunity, and demand exists. Change any one and the regularity weakens. It has held for three and a half years across every attempt made on it, which makes it a very good empirical bet and not a theorem. Section 4 is about people who quote it as a theorem.
What everyone gets wrong about the interval
One finding sits under all six claims and is almost never stated: the leak was not the cause of open-weights AI; it was the moment the gate stopped being worth defending. Meta itself had released OPT-175B on 2 May 2022 — a 175B model, GPT-3 class, weights available on request under a noncommercial licence, shipped with a full training logbook — ten months before LLaMA, and nothing happened. BLOOM, GPT-J and GPT-NeoX-20B were all fully public before that. The variable was never openness. It was quality per parameter: a 13B model beating GPT-3 175B is a model that runs on a laptop, and llama.cpp on 10 March and Alpaca on 13 March turned that from a claim into a Tuesday. The gate fell because what was behind it had become small enough to carry.
Commonly misused as
limit is the kind that requires this section, and this entry is not a limit — but it is misused in precisely the way limits are misused, which is why the section is here.
Misused as: a proof that model weights can never be recalled, therefore any proposal to restrict distribution is naive. What it actually establishes is that every recall attempt made between March 2023 and this reading has failed, by mechanisms that are named and contingent. That is a strong base rate and a weak impossibility. godel-incompleteness-1931 is the canon's standing example of a result carried into arguments it cannot support, and this one is a softer version of the same move: a durable empirical fact wearing the clothes of a theorem. The tell is the word "cannot." Altman's "once weights are out, they can't be pulled back" and Anthropic's "once weights are released they cannot be withdrawn" are both operationally correct and both stated as physics. A reading should quote them as what they are — planning assumptions by parties who have to decide, well supported by this file, not derived from anything.
Misused as: proof that the leak caused the open-source AI ecosystem. See the OPT-175B correction above. The ecosystem predated it; the leak supplied the first artifact where capability and consumer hardware met.
Misused as: "Meta open-sourced LLaMA." Meta gated LLaMA and it leaked. Llama 2 was open-weights, under a bespoke licence with a competitor carve-out and a field-of-use restriction, and it has never satisfied the OSI definition. Both halves of the sentence are wrong and the error is load-bearing, because it turns an involuntary event into a strategy and a licence into a philosophy.
Misused as: proof that open weights are dangerous, or proof that they are safe. The letter of 6 June 2023 predicted specific harms; those harms exist; their marginal attribution to this release has never been established, and the most careful public work says the evidence base cannot currently establish it in either direction. The trap catches both poles. On the AImageddon side, citing this entry for "open weights caused the criminal-LLM market" fails on WormGPT's own provenance. On the AItopia side, citing it for "three years and nothing happened" mistakes an absence of measurement for a measurement — nobody has built the counterfactual either.
Misused as: the same thing as a source-code leak. The 31 March 2026 Claude Code incident — a 59.8 MB source map shipped in a public npm package, roughly 513,000 lines across 1,906 files, mirrored to GitHub within hours — has the same diffusion signature as this entry and is a categorically different object. Client code reveals how a harness is built. Weights are the capability. A reading that files them together is losing the distinction the concentration lens is about; a reading that notes they diffuse identically is using both correctly.
Misused as: an argument that only applies to the country that made it. The 2023 framing was a domestic one. By 2026 the majority of downloaded and routed open weights come from labs outside US jurisdiction entirely, which means every policy instrument argued about in 2023–24 — export controls, licensing requirements, liability — now addresses a minority of the artifacts in circulation. Any reading applying this entry to a policy proposal has to say which share of the world's open weights the proposal can reach.
Sources
Primary, or as close as this machine could get:
- **Meta AI, Introducing LLaMA: A foundational, 65-billion-parameter large language model, 24 February 2023**, ai.meta.com/blog/large-language-model-llama-meta-ai/. Source for the model sizes, the token counts, and the two quoted sentences on the noncommercial licence and case-by-case access.
- **Touvron et al., LLaMA: Open and Efficient Foundation Language Models, arXiv:2302.13971, submitted 27 February 2023. Source for the GPT-3 comparison, the Chinchilla reframe and the training-versus-inference-budget argument, the "publicly available datasets exclusively" claim, the 1,022,362 GPU-hours / 2,638 MWh / 173 tCO2eq figures, and "We release all our models to the research community." Read via the ar5iv HTML rendering; the arXiv PDF returned binary content this machine could not parse**, so the quotations are from ar5iv.
- GitHub pull request
facebookresearch/llama#73, "Save bandwidth by using a torrent to distribute more efficiently," opened 2 March 2023 byChristopherKing42, closed unmerged, locked by maintainers on 1 September 2023. The outside bound on the leak date, and the best single artifact of the week. - GitHub DMCA repository,
2023/03/2023-03-21-meta.md— Meta's notice of 20 March 2023 filed through OpSec Online, the claimed work, the "No one is authorized…" sentence, and GitHub's processing against 403 repositories. - GitHub DMCA repository,
2023/04/2023-04-27-meta-counternotice.md— the Tensorfork Labs counter-notice and the uncopyrightability argument, quoted verbatim. github.com/shawwn/llama-dl, read 16 August 2026: live, 4.1k stars, 394 forks, README documenting the R2 re-mirror after Facebook cut the original link.huggingface.co/huggyllama/llama-7b, read 16 August 2026: 292,001 downloads in the last month, model card still directing users to the defunct access form.- **Meta and Microsoft, Meta and Microsoft Introduce the Next Generation of Llama, 18 July 2023, and the Llama 2 Community License Agreement** for the 700-million-MAU clause and the no-improving-other-LLMs restriction.
- **Touvron et al., Llama 2: Open Foundation and Fine-Tuned Chat Models, arXiv:2307.09288, July 2023**, Table 6 — the 1,418,091 preference comparisons. Cross-checked against
canon/rlhf-christiano-2017.md, which quotes the same table from its own reading of the paper. - Hawley and Blumenthal to Zuckerberg, 6 June 2023, hawley.senate.gov — "unrestrained and permissive," "seemingly minimal," "failed to conduct any meaningful risk assessment," and the enumerated harms.
- **The anonymous Google memo, We Have No Moat, And Neither Does OpenAI, published by SemiAnalysis 4 May 2023**, and verified by them as internally circulated. Quoted passages on the March leak and the cost asymmetry.
- **Zuckerberg, Open Source AI Is the Path Forward, 23 July 2024**, about.fb.com — the Linux analogy, the concentration argument, and the "starting next year" forecast graded in Claim 5.
- Zuckerberg, 30 July 2025 — "careful about what we choose to open source," via TechCrunch's reproduction of the letter.
- Sam Altman (@sama), 11 July 2025 — the open-weight delay and "once weights are out, they can't be pulled back."
- **Anthropic, Our position on open-weights models, 27 July 2026**, anthropic.com — irreversibility, removable safeguards, "a public good," and the explicit disclaimer of ever advocating a ban.
- **NTIA, Dual-Use Foundation Models with Widely Available Model Weights, 30 July 2024**, and its accompanying fact sheet — the no-restriction-at-this-time recommendation and the marginal-risk framing.
- **Kapoor, Bommasani, Klyman, Longpre, … Liang and Narayanan, On the Societal Impact of Open Foundation Models, arXiv:2403.07918 (2024)** — the marginal-risk framework and the insufficiency finding. The abstract page this machine reached gave a submission date of 27 February 2024 while the identifier places it in March; the day is unresolved and only the year is relied on.
- **Open Source Initiative, The Open Source AI Definition 1.0, 28 October 2024**, and Meta's rejection of it, via OSI and contemporaneous coverage.
- EU AI Act, Article 53(2) and the GPAI guidelines — the open-source exemption, its three conditions, the obligations that survive it, and its disappearance at the systemic-risk threshold.
- Meta AI, OPT-175B, arXiv:2205.01068 and the 2 May 2022 release, including the published training logbook — the correction in "What everyone gets wrong."
- Meta / Meta Superintelligence Labs, 10 August 2026 — the Muse Glimmer model page (Apache 2.0, 30B, single-consumer-GPU sizing) and Zuckerberg's 6,500-word essay, quoted from Fortune's 10 August 2026 report.
Secondary, named because specific claims rest on them:
- Motherboard/Vice, first week of March 2023, for the 4chan torrent and the Meta spokesperson quotations ("balance responsibility and openness"). The publication date is inconsistent across indexes — 3 March in some, 7 March in others — and I could not resolve it; the entry therefore dates the leak by the pull request instead.
- Wikipedia, "Llama (language model)", for the 3 March upload date, the 6 March Hugging Face takedowns, and the Llama 3 / 3.1 / 4 release dates.
- Simon Willison, 11 March 2023, simonwillison.net, for the laptop quotation.
- llama.cpp README (Georgi Gerganov, March 2023) for the 4-bit-on-a-MacBook goal.
- The Register, Gizmodo and Stanford Daily, March–April 2023, for Alpaca's under-$600 cost, the 52,000 instructions, and the demo withdrawal with the code left up.
- Trustwave SpiderLabs, Unit 42 and Huntress, 2023–2026, for WormGPT's GPT-J-6B provenance and its subscription pricing.
- Mistral AI (@MistralAI), 27 September 2023, and 404 Media's "Undeletable Chatbot" coverage; xai-org/grok-1 and Academic Torrents, 17 March 2024, for the Grok-1 magnet release.
- Bloomberg and Fortune, 15 August 2026, reporting Hugging Face's 14 August 2026 state-of-open-models figures (Qwen >3bn, Google 418m, Meta 227m).
- An industry analysis of OpenRouter routing shares (datagravity.dev) for the ~61% Chinese open-weight token share by May 2026 and Meta below 1%. Carried explicitly as an analyst estimate; I could not reach OpenRouter's own published rankings to verify it independently.
- **VentureBeat, The New Stack and DeepLearning.AI's The Batch, April 2026, for Muse Spark's introduction on/around 17 April 2026 as a proprietary successor to Llama; CNBC, UPI, Constellation Research and Phoronix, 10 August 2026**, for Muse Glimmer and the Muse Spark 1.2 commitment.
- WSJ via Axios, SiliconANGLE and Computerworld, May 2025, for the Behemoth delay and the stated capability concerns.
- Zscaler and Fortune, 31 March 2026, for the Claude Code source-map exposure — the 59.8 MB
.map, ~513,000 lines, 1,906 files, and the explicit confirmation that no model weights were exposed.
Attempted and failed, so that nothing above silently depends on it: venturebeat.com, cnbc.com and thenewstack.io returned HTTP 403 to this machine for the Muse Spark and Muse Glimmer stories, so those events are carried through search summaries and the outlets that did serve; the arXiv PDF for arXiv:2302.13971 could not be parsed and ar5iv was used instead; Zuckerberg's 10 August 2026 essay was not read at source — only Fortune's quotation of it, which is why exactly one sentence of it is quoted here; openai.com and x.com were not fetched directly, so Altman's July 2025 wording is carried from TechCrunch's and the search index's reproduction of the post.
On the boundary. This file reads canon/ entries on disk — eliza-1966, chatgpt-2022, asilomar-1975, frankenstein-1818, sparse-moe-2017, rlhf-christiano-2017, scaling-laws-2020, transformer-2017, godel-incompleteness-1931, turing-halting-1936, perceptrons-1969, alexnet-2012, culture-banks-1987 — and names chinchilla-2022, alpaca-2023, mistral-2023, deepseek-r1-2025, gpt-oss-2025, instructgpt-2022, bletchley-declaration-2023 and fli-pause-2023 as ids that are unwritten or unproposed, none of them invented into the header. It grades no vendor's 2026 conduct: everything graded in section 3 is a dated claim, the most recent of them made in July 2024 and due in 2025, and the 2026 events appear as record rather than as verdict. xAI's Grok-1 release is recorded as a bare dated fact and graded nowhere, per rule 7 — this project holds musk-robotaxi-2019 in its canon and xAI is a candidate to compose its tweets, and those two facts are not permitted to meet. Meta is graded in both directions in the same file, which is the only defence a canon has. It deposits nothing in history/, touches no other entry, writes no digest, places no needle, no score and no landmark position, and does not run build.py.