← The canon · AItopiaOrAImageddon?
"Computer-Based Medical Consultations: MYCIN"
idea · Edward H. Shortliffe, with Bruce G. Buchanan and Stanley N. Cohen, Stanford University · 1976
A framing later work is built out of rather than argued about. Cited when the thing being watched descends from it and the descent explains its shape.
Read on: The collapse of the Lisp machine market.
Filed as idea and idea is right, though the proposal line in canon/proposals.md describes the entry as an event — "expert system that matched or beat infectious-disease faculty in a blinded 1979 evaluation and was never used on a patient" — which reads like an argument for moment. It is not, and the reason matters more than the filing does.
The wrinkle, and it is the same shape as eliza-1966: what descends from MYCIN into 2026 is not the technology. Production rules with certainty factors are a dead line; nothing inside a transformer works that way, and Shortliffe himself co-authored the paper that retired the uncertainty calculus. Two other things descend, and both are alive. The first is the evaluation genre — the blinded comparison of a machine's recommendations against named clinicians', scored by independent expert raters who do not know which prescriber is the computer. The 1979 paper closes by proposing exactly that generalisation, in its own words, and the design it proposed is the design of the Science paper published on 30 April 2026. The second is the negative result the genre cannot see: MYCIN won that comparison and was never used in patient care, and its builders wrote down why, in detail, while it was still fresh. A reading in 2026 is not citing a 1970s antibiotic program. It is citing the first well-run instance of a study design that is still being run, together with the fifty-year-old record of what happened next.
Why not moment. The strongest alternative, and it fails on dates. The moment entries — Clippy shipping, Siri shipping, Deep Blue — are days on which something visibly happened in public. MYCIN has no such day. The program was begun in 1972, the dissertation was finished in October 1975, the book appeared in 1976, the knowledge base was "laid to rest" in 1978, the blinded evaluation was published in 1979, and the retrospective that explains the non-adoption came in 1984. The thing this entry is for — that demonstrated performance did not produce use — is by construction not an event. It is the absence of one, and an absence cannot be filed to a date.
Why not limit. Tempting, and it would be the same error section 4 of godel-incompleteness-1931 exists to catch. "Performance does not produce adoption" is a robust empirical regularity about institutions, not a theorem. It has no proof, no scope condition, and no statement of the conditions under which it fails — and it plainly does fail sometimes, or nothing would ever be adopted. Filing it beside Gödel and Turing would lend it an authority it has not earned, in the act of warning against exactly that move.
descends_from is empty, and it is a checked fact. MYCIN's real ancestor is DENDRAL — Lederberg, Feigenbaum, Djerassi and Buchanan's mass-spectrometry program, from which Shortliffe took the idea of encoding expert knowledge in production rules and Feigenbaum's "knowledge is power" hypothesis. DENDRAL is not in canon/, is not in canon/proposals.md, and is not among the candidates named at the proposal list's cap; if it is ever written, this entry should list it. The other ancestor is negative: Ledley and Lusted's 1959 Science paper on Bayesian probability in medical diagnosis, which MYCIN was in part a reaction against — also not in the canon. Note what is not an ancestor despite sitting next to it in every popular history: eliza-1966. ELIZA and MYCIN are both 1960s–70s medical-adjacent programs at the two poles of the field, and the ELIZA entry already points here for "the adjacent question of why beating clinicians in a study does not change care." But Weizenbaum's line runs through Colby and psychiatry, Shortliffe's through DENDRAL and mass spectrometry, and descends_from means came out of.
What it is
A rule-based consultation program that advised physicians on the selection of antimicrobial therapy for patients with severe infections — first bacteremia, later acute infectious meningitis. It began in 1972 as the doctoral work of Edward Shortliffe, who had come to Stanford in 1970 to study medicine and was simultaneously pursuing a PhD in what would now be called biomedical informatics. Stanley Cohen, then Chief of Clinical Pharmacology, was his dissertation advisor; Stanton Axline supplied the infectious-disease expertise; Bruce Buchanan, out of DENDRAL, was the principal computer-science colleague. The dissertation was completed in October 1975 and published by Elsevier in 1976 as Computer-Based Medical Consultations: MYCIN, which is the title and date this entry carries. The name is from the -mycin suffix that ends a great many antibiotics.
It ran in Interlisp on SUMEX-AIM — a DEC PDP-10 installed on the Stanford medical campus in 1973 as a shared national resource for AI-in-medicine research, and granted one of the last available ARPANET connections, the first host on that network not funded by the Department of Defense.
The mechanism is simple enough to state completely. Knowledge lived in IF/THEN production rules, kept strictly separate from the program that interpreted them. A rule looked like this, in the English translation MYCIN could generate from the LISP:
> IF: (1) the gramstain of the organism is grampos, (2) the morphology of the > organism is coccus, and (3) the growth conformation of the organism is clumps, > THEN: there is suggestive evidence (.7) that the identity of the organism is > staphylococcus.
That trailing number is a certainty factor, running from −1 (false) through 0 (unknown) to +1 (true), combined across rules by a small non-Bayesian calculus. The program reasoned backward from goals, which meant it controlled the dialogue and could therefore accept one- or two-word answers — a design choice made, as Chapter 32 says outright, specifically to avoid having to understand free text. A consultation asked "around 50 or 60" questions. The program held 120 organism names on its list of possible identities. Reported rule counts vary by era and by source, from "some 350" in the contemporaneous 1978 account to the 450–600 figures usually quoted later; the knowledge base grew throughout the decade and no single number is correct for all of it.
Three subsystems: a Consultation Program that gathered data and advised, an Explanation Program that answered WHY (why are you asking this?) and HOW (how was that established?) by unwinding the goal stack in English, and a Rule-Acquisition Program that let an expert add or edit rules and immediately re-run the case. The second of those was not decoration. Buchanan and Shortliffe state the working assumption plainly: "physicians would not ask a computer program for advice if they had to treat the program as an unexaminable source of expertise."
Because the rules were separable from the interpreter, removing them left a domain-independent engine — "empty MYCIN" or "essential MYCIN", EMYCIN (van Melle), the first expert-system shell and the direct ancestor of the commercial shells that drove the 1980s expert-systems industry, and with it the boom that ended in the second AI winter. MYCIN also seeded a dense family of Stanford dissertations: TEIRESIAS (Davis, knowledge acquisition), GUIDON (Clancey, tutoring) and its successor NEOMYCIN, BAOBAB, VM, and through EMYCIN the PUFF, CENTAUR, SACON, WHEEZE, GRAVIDA, CLOT and DART systems, plus ONCOCIN, the oncology program Shortliffe's group built next precisely to overcome the barriers that had stopped MYCIN.
The evaluations. An earlier study of MYCIN's bacteremia advice (Yu et al., 1979a) gave encouraging results but, as the team itself wrote, was "difficult to interpret because of the potential bias in an unblinded study." So they ran it again properly, and the design is the reason this entry exists.
Ten patients with infectious meningitis were selected by a physician unacquainted with MYCIN's methods or its knowledge base, from a county hospital affiliated with Stanford, by retrospective chart review. The cases were chosen to be diagnostically challenging and deliberately diverse: no more than three viral, and at least one each of tuberculous, fungal, viral and bacterial, including at least one with a positive CSF gram stain and at least one negative. Each case was reduced to a clinical summary; only what was in the summary went to MYCIN, and the program was not modified.
The same summaries went to five faculty members of Stanford's Divisions of Infectious Diseases in Medicine and Pediatrics, one senior postdoctoral ID fellow, one senior medical resident and one senior medical student. Textbooks were allowed; there was no time limit. That gave ten prescribers per case: MYCIN, the eight humans, and the therapy the patients had actually received.
Then eight infectious-disease specialists at institutions other than Stanford, all of whom had published on the management of meningitis, were sent the clinical summaries and the ten prescriptions per case, in random order and in a standardised format. They did not know the identity of any prescriber and did not know that one of them was a computer program. Each rated all 100 prescriptions as equivalent to their own choice, an acceptable alternative, or not acceptable — 800 assessments in all.
The results, from Table 31-1:
| Prescriber | Rated acceptable (n=80) | Majority-acceptable cases (n=10) | Cases failing to cover a treatable pathogen | |---|---|---|---| | MYCIN | 52 (65%) | 7 (70%) | 0 | | Faculty-1 | 50 (62.5%) | 5 (50%) | 1 | | Faculty-2 | 48 (60%) | 5 (50%) | 1 | | ID fellow | 48 (60%) | 5 (50%) | 1 | | Faculty-3 | 46 (57.5%) | 4 (40%) | 0 | | Actual therapy given | 46 (57.5%) | 7 (70%) | 0 | | Faculty-4 | 44 (55%) | 5 (50%) | 0 | | Resident | 36 (45%) | 3 (30%) | 1 | | Faculty-5 | 34 (42.5%) | 3 (30%) | 0 | | Student | 24 (30%) | 1 (10%) | 3 |
MYCIN was top of the list on the item-level measure — 65% against a five-faculty mean of 55.5% — and the difference among prescribers was significant (F = 3.29, 9 and 70 df, p < 0.01). It never failed to cover a treatable pathogen, and it did so while prescribing fewer drugs than the physicians who had actually treated the patients, who in eight of ten cases prescribed two or three antimicrobials where one or none would have sufficed. The authors' own stated primary limitation: "the small number of cases studied."
And the closing sentence of the paper, which is the thing that descends:
> The methodology of the evaluation is of interest because it was developed in > an attempt to analyze clinical decisions for which there is no clear right or > wrong choice. Since most areas of medicine are characterized by a variety of > acceptable approaches, even among experts, the technique used here may be > generally useful in assessing the quality of decision making by other computer > programs.
And then nothing happened. From Chapter 36, in the authors' own words: "For example, although it was explicitly our initial intention to implement and test MYCIN on the hospital wards, this experiment was never undertaken. Instead the infectious disease knowledge base was laid to rest in 1978 despite studies demonstrating its excellent decision-making performance."
Chapter 32 gives the reasons, and they are worth reading in full because almost every popular retelling substitutes a different set. "MYCIN was never used routinely in patient-care settings." What stopped the ward trial "was the simple fact that we knew the program was likely to be unacceptable, for mundane reasons quite separate from its excellent decision-making performance." Those reasons, in their order:
- No recognised need. "Although there was a demonstrated need for a system like MYCIN..., we did not feel there was a recognized need on the part of individual practitioners. Most physicians seem to be quite satisfied with their criteria for antibiotic selection, and we were unconvinced that they would be highly motivated to seek advice from MYCIN."
- No workflow integration. "Our second concern was our inability to integrate MYCIN naturally into the daily activities of practitioners." A physician had to find a free terminal, log on, and answer a series of questions "many of which were simply transcriptions of lab results already known to be available on other computers at Stanford." Linking SUMEX to the Stanford lab machines was considered and rejected.
- Latency. "When the machine was heavily loaded, annoying pauses between MYCIN's questions were inevitable, and a total consultation could have required as long as 30 minutes or an hour. This was clearly unacceptable and would have led to rejection of the system despite its other strong features."
- Friction. "Slight annoyances, such as the requirement that the physicians type their answers, would have further alienated users."
- Ordinary institutional decay. The need to extend the knowledge base to other infectious diseases just as Axline and Yu were both leaving Stanford; the difficulty of funding knowledge-base expansion for a program that already looked finished; and, stated without flinching, "our own lack of enthusiasm for implementation studies once we had come to identify some of the computer science inadequacies in MYCIN's design and preferred to work on those in a new environment."
Within five years the knowledge base was obsolete anyway: third-generation cephalosporins arrived and changed antibiotic selection for a range of common problems. The authors drew the general lesson themselves — "This point emphasizes the need for knowledge base maintenance mechanisms once expert systems are introduced for routine use in dynamic environments, where knowledge may change rapidly over time."
Why a reading would cite it
The occasion is live, dated, and almost embarrassingly exact.
On 30 April 2026, Science published Brodeur, Buckley, Kanjee et al., "Performance of a large language model on the reasoning tasks of a physician" (392:524–527, DOI 10.1126/science.adz4433), from Harvard Medical School and Beth Israel Deaconess, with Arjun Manrai and Adam Rodman as co-senior authors. It compares OpenAI's o1 series against hundreds of physicians across six clinical reasoning experiments, including real emergency-department records from a Massachusetts hospital presented exactly as they appeared in the EHR, without pre-processing, at sequential decision points from triage onward. The model matched or exceeded attending physicians in diagnostic accuracy at the early ED decision points, and eclipsed both prior models and the physician baselines across the reasoning tasks.
That is the 1979 study with better instruments: a machine's recommendations, real cases the machine's builders did not choose, human clinicians as the comparison arm, and expert judgement as the scoring rule. The genre is the one Yu, Fagan and colleagues proposed in the last paragraph of their paper.
What MYCIN supplies is the other half — the half a comparison study structurally cannot contain. And the 2026 authors say so themselves, which is what makes the citation an amplification rather than a rebuttal. Manrai: "this does not mean AI will necessarily improve care—how and where it should be deployed remain understudied, and we desperately need rigorous prospective trials to evaluate the impact of AI on clinical practice." Brodeur: "a model might get the top diagnosis right but also suggest unnecessary testing that could expose a patient to harm," and "humans should be the ultimate baseline when it comes to evaluating performance and safety." A reading covering this result, or the next one, should be able to say: this is the fifth decade of this finding, the first instance is documented down to the reason the ward trial never ran, and the reasons were not about accuracy.
The middle term between them is already on the record too. Goh et al., "Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial" (JAMA Network Open, October 2024), randomised 50 US-licensed physicians in family, internal and emergency medicine to conventional resources or conventional resources plus GPT-4. Median score per case: 76% with the model, 74% without — a two-point difference (95% CI −4 to 8, p = 0.60). The model on its own scored significantly higher than either group of doctors. That is the MYCIN structure with the outcome measured rather than assumed: superior standalone performance, no improvement in the hands of the people who would actually use it.
The narrower, recurring occasions:
- Knowledge decay. MYCIN's knowledge base was overtaken by a new drug class inside five years, and its authors turned that into a general warning about maintenance. When a reading covers a training cutoff, a model's medical currency, or a guideline change that a deployed system has not absorbed, this is the precedent and it comes with the lesson already drawn.
- Explanation as a precondition. MYCIN's designers built the WHY/HOW facility on the belief that clinicians would refuse an unexaminable oracle, and their 1980 attitude survey of physicians found that a program's ability to explain its reasoning was judged the single most important requirement for a medical advice-giving system. Cite it when transparency obligations, model cards or "explainable AI" requirements are at issue — including the healthcare AI bills AB 1979 and AB 2575, which cleared California's Senate Appropriations 5–2 each on 13–14 August 2026 as recorded in this project's own digest. The 1970s answer to that requirement was a real derivation trace. The 2026 answer is generated text, which is a different object.
- The claimed-versus-demonstrated discipline. The Science paper's own preprint (arXiv:2412.10849, December 2024) was titled "Superhuman performance of a large language model on the reasoning tasks of a physician." The word did not survive to publication. That is a small, checkable instance of a claim being tempered on its way through review, and it is the kind of thing a reading tracking the gap between announcement and record should notice without editorialising.
What it got right, and what it got wrong
Not required for idea, but MYCIN carries a lot of dated claims and the canon is worth less if only prediction entries get graded. Claim dates are given per item; where the claim is the authors' own retrospective it is dated to the 1984 book.
Right — the performance claim, and it survived being made harder. Claim: a rule-based program could select antimicrobial therapy as well as an infectious-disease specialist. Made: 1976. Due: 1979, on their own initiative, under blinding and with external raters they did not choose the outcome of. Graded: correct, at the size of the study. This is worth stating plainly because the rest of this section is critical, and it should not be read as a claim that MYCIN did not work. It worked.
Right — knowledge separated from inference generalises. Claim: keeping the rules out of the program means the same engine can advise in a different domain. Made: 1976. Due: within the decade. Graded: correct and then some. EMYCIN produced systems for pulmonary function, structural analysis, and oncology, and the shell became a product category. This is the entry's clearest technical descent, and it is also where the entry's caution belongs: it descended into an industry that busted. Shortliffe and Shah's own account of why is brittleness and maintenance — "companies often found that the systems were expensive to maintain and difficult to update. They generally had no machine learning component, so maintenance was crucial."
Right — the evaluation methodology would generalise. Claim, 1979: "the technique used here may be generally useful in assessing the quality of decision making by other computer programs." Due: unspecified. Graded: correct beyond anything the authors can have had in mind. Forty-seven years later the design is still the standard, and the 2026 Science study is a direct instance of it.
Right — knowledge bases rot in dynamic domains. Claim, 1984, after watching it happen. Due: continuously. Graded: correct, and now structural rather than incidental: a model's knowledge is fixed at training and the maintenance mechanism is retraining, which is expensive and infrequent, so the failure mode did not go away when the representation changed.
Right — the non-adoption diagnosis. Claim, 1984: a decision aid that cannot be integrated into a clinician's existing workflow will not be used, however good its advice. Due: continuously. Graded: correct, and it generalised so thoroughly that it acquired a name — Miller and Masarie's 1990 "demise of the Greek Oracle model," the observation that the consultation paradigm itself, in which a clinician stops what they are doing to interrogate an advisor, was the thing that failed. Shortliffe's group had already acted on this in the early 1980s by rebuilding ONCOCIN around it; the title of his 1982 editorial on the subject was "Good Advice is Not Enough."
Wrong, and graded false by the authors themselves — that a knowledge base good enough to perform is good enough to explain. Claim, implicit in the Explanation Program and explicit in the GUIDON work: "We originally assumed that a knowledge base that is sufficient for high-performance problem solving would also be sufficient for tutoring." Made: 1972–76. Due: by 1984. Graded, in their words: "This assumption turned out to be false, and this negative result spawned revisions in our thinking about the underlying representation of MYCIN's knowledge." Clancey's 1983 analysis found MYCIN's uniform rules silently compiled together several different kinds of knowledge — chance associations, statistical correlations, heuristics, causal relations, definitions, structural and taxonomic facts — so that a rule A → B was often a collapsed chain A → A₁ → A₂ → B whose intermediate steps a student needed and the program did not have. The conclusion in the 1984 book: "MYCIN's knowledge is, in a sense, compiled knowledge. It performs well but is not very comprehensible to students without the concepts that have been left out." That is the ancestor of every present-day finding that a system's competence does not come with an account of itself, and it was established, published and acted on by the people who had the most to lose by it.
Wrong — that physicians would require an examinable system. Claim: "Our working assumption was that physicians would not ask a computer program for advice if they had to treat the program as an unexaminable source of expertise." Made: 1970s, reaffirmed 1984, and supported by their own 1980 survey in which explanation was rated the top requirement. Due: continuously. Graded: false as a prediction about behaviour, whatever its merits as a design principle. The systems now being evaluated against physicians do not expose a derivation; the model in the 2026 Science study belongs to a series whose internal reasoning is deliberately not shown to users, and what a user sees when a model "explains" is generated text about the answer rather than a trace of how it was produced. The requirement MYCIN's designers treated as a precondition for adoption turned out not to be one. Whether it should have been is a live argument and this entry does not settle it; what the entry establishes is that the assumption was made explicitly, by people who tested it, and that events went the other way.
Wrong — certainty factors, and the refutation has Shortliffe's name on it. Claim: uncertainty in a rule-based medical system can be handled by a simple modular calculus that avoids the impractical demands of Bayesian methods. Made: 1975–76. Due: unspecified. Graded: superseded. Heckerman and Shortliffe, "From certainty factors to belief networks" (Artificial Intelligence in Medicine 4:35–52, 1992), traced the model's flaws to its imposing on uncertain rules the same modularity that logical rules have — the rule "if ALARM then BURGLARY" licenses you to increase belief in a burglary no matter how your belief in the alarm rose, which is not how evidence works. Belief networks, grounded in probability, replaced it. Note also the honest wrinkle recorded in the 1984 book: the team knowingly built redundant, overlapping rules for robustness, which "is at odds with the independence assumptions of the CF model," and their own fixes were "pragmatic, and not entirely satisfactory."
Unresolved rather than graded — the 1979 paper's forward-looking sentence. "Because MYCIN compared favorably with infectious disease experts in this study, we believe that it could be a valuable resource for the practicing physician whose clinical experience for specific infectious diseases may be limited." Made: 1979. Due: never specified. What happened: MYCIN never became that resource, and the same paper's next sentences named exactly what would have to be settled first — "questions concerning the program's acceptability to practicing physicians and its impact on patient care, as well as issues of cost and legal implications, remain to be answered." Those questions are still open in 2026 for the current generation of systems, which is the whole point of the entry; the 2024 randomised trial is one of the few direct attempts at the second of them and it came back null. A reading should carry that as an open question with a long history, not as a verdict.
Commonly misused as
Not required for idea. Included because the sentence "MYCIN beat the doctors and was never used" is now a stock move in AI-and-medicine commentary, and most of its uses are wrong in one of the following ways.
- "A study showing a model beating doctors means care is about to change." The live one, and the reason this entry was proposed. MYCIN is the cleanest counter-instance available: it won, on a hard case mix, blinded, judged by outsiders, and the ward trial was never run — for reasons that were about terminals, typing, latency, funding, staff turnover and the absence of felt need among the people who would have used it. None of those is an argument about accuracy, and none of them would have been detected by a better study of accuracy. The correct inference from a benchmark win is that a deployment question has become worth asking, not that it has been answered.
- "MYCIN outperformed the doctors." True on one measure and worth stating precisely. On the item-level rating MYCIN was top of ten prescribers (65% versus a faculty mean of 55.5%). On the majority-of-evaluators measure it tied with the therapy the patients had actually been given (7 cases out of 10 each), and the actual treating physicians also never failed to cover a treatable pathogen — their failing was over-prescribing, not missing. The study had ten cases, chosen to be diagnostically challenging and deliberately spread across four aetiologies; six had negative CSF smears. The authors named the case count as their primary limitation, and observed that "if more routine cases had been selected, there would have been greater consensus among evaluators." It is a real result and a small one, and it is not evidence that MYCIN was better than infectious-disease specialists at practising medicine.
- "MYCIN was never used on a patient." Close enough to be worth tightening rather than repeating. MYCIN was never used in patient care and never guided anyone's treatment; the planned ward trial was never undertaken. It was run on real patients' records — the ten evaluation cases were real people at a county hospital, and Chapter 31 notes several hundred further retrospective and prospective cases used in testing. The accurate sentence is "no patient was ever treated on MYCIN's advice."
- "It was killed by liability" / "by the FDA" / "by doctors' egos." These appear constantly in secondary retellings and none of them is what the developers wrote. Buchanan and Shortliffe's list is above; legal responsibility appears in the surrounding literature and in their own note that physicians are sensitive to guidelines against prescribing drugs they do not understand, but it is not what they say stopped them. The mundane version is both true and more useful: the thing that killed MYCIN was that using it would have taken up to an hour on a loaded time-sharing machine, required typing, and duplicated lab results that were already sitting on another computer in the same institution.
- "MYCIN proves AI in medicine doesn't work." It proves that a demonstrated performance advantage is not sufficient for adoption — a claim about deployment, not capability. The capability claim held up, and the architecture generalised so successfully that it produced a commercial industry. Reaching for MYCIN to deflate a capability result is the mirror image of reaching for it to inflate one.
- "MYCIN was an early chatbot." It printed English and accepted English, and it had no language understanding whatsoever. The backward-chaining control structure was chosen in large part so that it would not need any: the program asked the questions, so answers could be one or two words. Their one attempt to accept volunteered free text, BAOBAB, "never functioned at a performance level sufficiently high to justify its incorporation into MYCIN."
- "Certainty factors were Bayesian probabilities." They were explicitly not, and were introduced because the developers judged the Bayesian requirements — independence assumptions, or elicitation of conditional probabilities nobody had — unusable in that setting. Heckerman later showed the CF model admits a probabilistic reading, but only under assumptions that are usually false, which is a different and less flattering claim than "it was Bayes all along."
- "The 1980s expert-systems bust shows knowledge engineering was a dead end." The people who ran it draw a narrower lesson: the systems were brittle, had no learning component, and were expensive to maintain — a maintenance and acquisition failure, not a refutation of the idea that domain knowledge matters. Shortliffe and Shah's 2022 verdict is that "the knowledge is power aphorism has been somewhat forgotten in today's AI research and application communities—arguably to their detriment," which is a position a reading may cite as a position, not as a settled finding.
- "MYCIN shows today's clinical AI will also fail." It shows nothing of the kind, and nothing in this file is evidence. The barriers Buchanan and Shortliffe named were specific and several have genuinely fallen: there is no terminal to find, no logging on, no typing, no 30-minute wait, and lab results are in a record the system can read. Whether the remaining barriers — felt need, integration, liability, and demonstrated effect on care — fall too is an open empirical question that prospective trials are supposed to answer and mostly have not. This entry supplies the history and the questions, not the verdict.
Sources
Primary, read directly. Victor L. Yu, Lawrence M. Fagan, Sharon Wraith Bennett, William J. Clancey, A. Carlisle Scott, John F. Hannigan, Robert L. Blum, Bruce G. Buchanan and Stanley N. Cohen, "An Evaluation of MYCIN's Advice," Chapter 31 of Buchanan and Shortliffe (eds.), Rule-Based Expert Systems: The MYCIN Experiments of the Stanford Heuristic Programming Project (Addison-Wesley, 1984), pp. 589–596 — read in full, and the source of the study design, Table 31-1 in its entirety, the significance test, the authors' stated limitations, and the closing methodology paragraph. It is an edited version of Yu et al., "Antimicrobial selection by a computer: a blinded evaluation by infectious diseases experts," JAMA 242(12):1279–1282 (1979), PMID 480542.
Bruce G. Buchanan and Edward H. Shortliffe, "Human Engineering of Medical Expert Systems," Chapter 32 of the same volume, pp. 599–603 — read in full, and the source of every quotation about why MYCIN was never moved to the wards, the 30-minutes-to-an-hour figure, the 1978 freeze, the cephalosporin obsolescence, the physicians' attitude survey, and the "Good Advice is Not Enough" editorial (Shortliffe, 1982b). "Major Lessons from This Work," Chapter 36 of the same volume, pp. 669–691 — read in part, and the source of the "laid to rest in 1978" sentence, the 120-organism figure, the 50-or-60 questions figure, the unexaminable-oracle assumption, the tutoring assumption and its refutation, the CF-independence admission, and the list of knowledge types that MYCIN's uniform rules could not distinguish. The full book is posted openly by Shortliffe at people.dbmi.columbia.edu/~ehs7001/Buchanan-Shortliffe-1984/.
Edward H. Shortliffe and Nigam H. Shah, "AI in Medicine: Some Pertinent History," Chapter 2 of Cohen et al. (eds.), Intelligent Systems in Medicine and Health (Springer, 2022), DOI 10.1007/978-3-031-09108-7_2 — read in part, for the DENDRAL lineage, SUMEX-AIM and its ARPANET connection, the MYCIN rule example quoted above, the descendant diagram (Fig. 2.5), INTERNIST-1's NEJM-CPC result and Miller's move to QMR, the "Greek Oracle" citation (Miller and Masarie, 1990, named there and not read here), the two AI winters, and the knowledge is power verdict.
Shortliffe's dissertation was completed at Stanford in October 1975 and published as Computer-Based Medical Consultations: MYCIN (Elsevier, 1976), which supplies this entry's title and year; the book itself was not read for this entry. Rule counts are reported as a range for the reason given in the text; the "some 350" figure is from the contemporaneous "MYCIN: a knowledge-based consultation program for infectious disease diagnosis" (International Journal of Man-Machine Studies, 1978), seen in abstract only. The earlier unblinded bacteremia evaluation is cited here as Chapter 31 cites it, "Yu et al., 1979a," and was not read.
For the grading. David E. Heckerman and Edward H. Shortliffe, "From certainty factors to belief networks," Artificial Intelligence in Medicine 4:35–52 (1992), and Heckerman's "The Certainty-Factor Model," both openly posted by Microsoft Research — the ALARM/BURGLARY modularity example is from the former. William J. Clancey, "The epistemology of a rule-based expert system — a framework for explanation," Artificial Intelligence 20(3):215–251 (1983), for the compiled-knowledge analysis; consulted through abstracts and secondary description, not read in full.
For the citation occasion. P. G. Brodeur, T. A. Buckley, Z. Kanjee et al., "Performance of a large language model on the reasoning tasks of a physician," Science 392:524–527, 30 April 2026, DOI 10.1126/science.adz4433, with Arjun Manrai and Adam Rodman as co-senior authors; the publisher's page returned HTTP 403 and the paper itself was not read — the design details, the o1-series identification, the ED-record finding and all direct quotations from the authors are taken from the Harvard Medical School / Beth Israel press materials (EurekAlert releases 1125790 and 1126008) and Medical Xpress's 4 May 2026 report, and a reading that leans on the numbers should go to the paper. The preprint is arXiv:2412.10849 (December 2024), titled "Superhuman performance of a large language model on the reasoning tasks of a physician." Ethan Goh et al., "Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial," JAMA Network Open, October 2024, PMID 39466245, for the 50-physician trial and the 76%/74% result; abstract and secondary coverage, not the full paper.
Attempted and failed: the open-access PMC copy of Shortliffe's "Artificial Intelligence in Medicine: Weighing the Accomplishments, Hype, and Promise" (Yearbook of Medical Informatics, 2019) returned a CAPTCHA wall; Forbes's MYCIN piece returned HTTP 403; the Stanford Digital Repository scan of "Computer-Based Consultations in Clinical Therapeutics" was image-only and unreadable. A quotation widely attributed to Shortliffe — that "adoption of a new tool is not based solely on demonstrated need coupled with demonstrated high performance of the tool. In retrospect, that was naive. Acceptability is different from high performance" — circulates without a verifiable citation, and I could not establish its origin. It is named here as unverified and is not relied on anywhere above; the Chapter 32 and Chapter 36 passages say the same thing, in print, and can be checked.
The California healthcare-AI bill votes of 13–14 August 2026 are as recorded in this project's own digests/2026-08-15-12.md, which holds the primary links. They are named here as a citation occasion. Nothing in this file is evidence, nothing in it is deposited in the ledger, and nothing in it touches the needle.