← The canon · AItopiaOrAImageddon?
The DARPA Grand Challenge
moment · DARPA, under director Tony Tether — won by the Stanford Racing Team, led by Sebastian Thrun · 2005
Something that happened and changed what people expected next.
Read on: "The Future of Employment: How Susceptible Are Jobs to Computerisation?", Tesla Autonomy Investor Day.
Filed as moment, and moment holds — but the id names a year and the entry's content is an interval, so the filing has to be argued rather than assumed. The house definition, settled by eliza-1966 when it talked its way out of the kind and used since by clippy-1996, siri-2011, expert-systems-collapse-1987, deep-blue-1997, watson-jeopardy-2011 and alphago-move-37-2016, is a date on which something visibly happened in public. On Saturday 8 October 2005, at 6:40 a.m., on a 132-mile course through the Mojave desert starting and finishing at Primm, Nevada, twenty-three robotic vehicles were launched at five-minute intervals by DARPA personnel, tracked from a forty-person operations center, followed by government chase drivers holding emergency stops, across terrain closed to every competitor since 29 July. Five finished. There is a date, a clock, a course, a scoresheet, a government agency that owned none of the entrants, and a Report to Congress filed in March 2006 that names every team and the exact distance it covered. It passes every clause of the test, and it passes them with better documentation than any other entry in this canon.
The awkwardness is that the year in the id is not the year the finding lives in. What makes this canonical is not that five vehicles finished in 2005; it is that none finished in 2004 and five finished 574 days later. The 2004 race is not background to this entry. It is the control condition, and the entry carries both dates for the same reason an experiment carries both arms.
The proposal line, corrected
canon/proposals.md reads: "DARPA Grand Challenge (2004: no finisher; 2005: five finished) — eighteen months between total failure and routine success. Cite when: demonstrated-under-controlled-conditions versus actually-deployed, especially in robotics and vehicles."
The "cite when" clause is exactly right, is the reason this entry earns admission, and section 2 is built on it. The description in front of it is wrong in one important way and imprecise in three small ones, and correcting all four is the entry's first job.
"Routine success" is the load-bearing error. It is not a small overstatement; it inverts the shape of the result. DARPA's own Table 1 gives the distance completed by all twenty-three finalists, and the distribution is not a field that had learned to drive:
| Distance completed | Vehicles | |---|---| | 132 miles (the whole course) | 5 | | 66–81 miles | 2 | | 22–44 miles | 7 | | 1–17 miles | 9 |
The median vehicle covered 26 of 132 miles — about a fifth of the course. Sixteen of twenty-three did not reach halfway. The last-placed finalist, MITRE Meteorites, managed one mile. And these twenty-three were not the field; they were the survivors of a four-stage cull — 195 applications, 136 videos, 118 site visits, 43 semifinalists — and at the qualification event that produced them, only 5 of 43 teams completed all three test runs, and 20 of 43 never completed even one. What happened in October 2005 was not a field reaching competence. It was a very heavy tail with five survivors at the end of it, and the entry's whole value depends on not smoothing that over.
"Five finished" is the headline number and "four" is the qualified one, and DARPA's own report uses both. Section 5 of the report says "five teams completed the Grand Challenge course"; the conclusion on page 14 says teams progressed to "four vehicles finishing the course within the 10-hour limit." Both are true. Oshkosh's TerraMax was stopped by DARPA at sunset about 80 miles in for the safety of its chase crew, sat in the desert overnight in autonomous mode with its engine running, resumed at sunrise on 9 October, and finished at an average of 10.2 mph for a running time of about 12 hours 51 minutes — outside the limit, ineligible for the prize, and by DARPA's account "the first recorded event in which a ground vehicle operated in autonomous operations for more than 24 hours without any human intervention other than to command the vehicle to stop and resume at sunrise and the addition of 5 gallons of diesel fuel." Five completed the route; four beat the clock. Say which.
"Eighteen months" rounds down and DARPA's "19 months" rounds up. From 13 March 2004 to 8 October 2005 is 574 days — 18 months and 25 days. Neither rounding is wrong and the difference matters to nobody, but a canon that polices magnitudes should know that the two available figures come from the proposal and from DARPA's own conclusion respectively, and that the true number sits between them.
"No finisher" in 2004 understates it. Thrun and his co-authors put the 2004 result precisely in the paper of record: of fifteen robots that started, "none of the participating robots navigated more than 5% of the entire course." The best, Carnegie Mellon's Sandstorm, went about 7.4 miles of 142 and beached itself on an embankment after taking a hairpin too fast. "No finisher" makes it sound close. It was not close.
Why not the other kinds
Why not idea. The strongest rival. What plausibly descends from here into 2026 is not a race but an instrument: the publicly announced, independently judged, pass/fail inducement prize with a withheld test and an open field. That instrument has a documented paper trail — the National Academy of Engineering's 1999 report Concerning Federally Sponsored Inducement Prizes in Engineering and Science, the prize authority Congress granted at 10 U.S.C. § 2374a, and DARPA's 2003 decision to spend it on ground vehicles. An idea entry on inducement prizes would be legitimate and this canon should probably have one. But it would be a different entry with a different id, and this id names a race. Races are events; events are moments.
Why not prediction. Tempting, because this date is load-bearing for an unusual number of dated, gradeable claims — a congressional goal, an agency's self-assessment, and a decade of industry timelines — and section 3 grades them all. But the event is not a claim. It is the thing several claims were made about, and DARPA's own forecasts about it are graded in section 3 rather than promoted to the file's kind.
Why not limit. Nothing was proved, and it is worth saying because "the Grand Challenge showed off-road autonomy works" is deployed as though a result had been established. Five vehicles traversed one course, once, with no moving obstacles, following a route file they were handed. That is a demonstration.
descends_from is empty, and that is a finding rather than a gap
Nothing in canon/ documentably precedes this. The ancestry that actually produced the Grand Challenge is legislative and administrative: the NAE's 1999 inducement-prize report, § 2374a, and Section 220 of the Floyd D. Spence National Defense Authorization Act for Fiscal Year 2001 (Public Law 106-398, enacted 30 October 2000). None of those is in the canon and I am not inventing ids for them. Per the house rule, they are named here instead.
deep-blue-1997 is the entry I most want to link to, and it is a sibling, not an ancestor. I found no evidence that DARPA's 2003 decision took anything from IBM's 1997 match; its report cites the NAE study and consultations with the Marine Corps and Army TRADOC, not a chess game. The relationship is convergent rather than descended — two institutions independently reaching for the scored public contest as a way to settle a capability question — and the contrast between them is this entry's single most useful contribution, which section 2 develops. alexnet-2012 is the third instance of the same convergence and is likewise not an ancestor. Three entries, three independent arrivals at the same instrument, is a better fact than a fabricated lineage would have been.
One edge I would want and cannot yet have: musk-robotaxi-2019 is proposed and unwritten. It is the natural descendant of this entry on the forecasting side, and section 3 deliberately does not grade Musk's robotaxi claims, because that is that entry's whole job and doing it here would either duplicate it or — worse — grade one vendor's timeline inside a file whose other subjects are graded by different standards. Rule 7 is easier to keep when the grading of a given predictor lives in one place.
What it is
The mandate, which nobody remembers and which is the entry's spine
This was not a research project about cars. DARPA states the origin in its own report, quoting the statute:
> This decision supported the Congressional mandate stated in section 220 of the > Floyd D. Spence National Defense Authorization Act for Fiscal Year 2001 that > "It shall be a goal of the Armed Forces to achieve the fielding of unmanned, > remotely controlled technology such that . . . by 2015, one-third of the > operational ground combat vehicles are unmanned."
The pressure behind it was resupply convoys taking casualties to roadside bombs. The prize money came from Program Element 0603764E, Land Warfare Technology — the same line that funded the Future Combat Systems program, which the Army terminated in 2009. The authority capped DARPA at $10 million of prizes per fiscal year and required the Under Secretary of Defense for Acquisition, Technology and Logistics to personally approve any prize over $1 million, which is why the increase from $1M to $2M has a named signatory.
Every popular account of this event is a story about Silicon Valley. The document trail says it was a procurement problem with a 2015 deadline. Section 3 grades that deadline.
2004: the control condition
Announced 2003; run 13 March 2004; a 142-mile route from Barstow, California to Primm, Nevada; $1 million prize; 106 applications, 15 vehicles at the start. Nobody finished. Nobody got past 5%. The prize went unclaimed.
Scientific American, giving the race one of its 2004 awards and quoted approvingly by DARPA in the report to Congress, is worth reading for how a failure gets narrated in real time:
> Of the 15 vehicles that started the Grand Challenge . . . not one completed the > 227 kilometer course. One crashed into a fence, another went into reverse after > encountering some sagebrush, and some moved not an inch. The best performer, the > Carnegie Mellon entry, got 12 kilometers before taking a hairpin turn a little > too fast. The $1-million prize went unclaimed. In short, the race was a > resounding success.
That last sentence, written about a total failure four months after it happened, is either the best or the worst thing in the file depending on what happened next.
2005: the funnel, the course, the rules
Announced 8 June 2004. By the 11 February 2005 deadline DARPA had 195 applications from 36 states and three foreign countries — an 84% increase over 2004's 106. The cull was four-stage and is documented: 136 five-minute videos scored by at least two DARPA staff each; 118 site visits in May 2005, two government personnel per team, three runs on a 200-metre course; 40 semifinalists plus 3 late alternates = 43 at the National Qualification Event, held 28 September – 5 October at the California Speedway in Fontana; and 23 finalists announced on 5 October. The cap of 25 vehicles on the course was set not by DARPA's judgment but by "the practicalities of race operations as well as the Bureau of Land Management event permit."
The course was 132 miles, and DARPA's description of it is specific: a dry lake bed, cattle guard gates, narrow roads, highway and railroad underpasses, broken pavement, gravel utility roads and off-road trails, more than 50 turns of at least 90 degrees, more than 50 utility poles beside the road, and a finish through Beer Bottle Pass — "a steep, narrow downslope with a sheer drop-off on the side." Speed limits set by DARPA ran from 10 mph to 40 mph, and DARPA notes that "completing the 132-mile route required approximately 6 hours at the defined course speeds." The winning time was 6 hours 53 minutes. The finishers were not racing each other so much as running down a clock DARPA had already mostly set.
Vehicles received the route as an RDDF — a route definition data file of longitudes, latitudes, corridor widths and speed limits — on CD-ROM two hours before their start. The 2005 file held 2,935 waypoints; corridor width varied from 3 to 30 metres.
The four properties that make this a good test, and they are the point
deep-blue-1997 exists in this canon largely to record what a vendor-run evaluation looks like: IBM owned the venue, the machine, the operators, the records and the share price. alexnet-2012 exists partly to record the opposite — withheld test labels scored by people who did not build the systems. The 2005 Grand Challenge is the strongest instance of the second kind that this canon holds, and a reading reaching for it should know exactly why:
1. The scorer had no entrant. DARPA set the course, ran the clock, held the emergency stops and published the results. It had a stake in someone finishing — section 3 grades that bias — but none whatsoever in which team it was. Stanford, Carnegie Mellon and Oshkosh were graded by the same stopwatch. 2. The test was physically withheld. The route was revealed two hours before the start, and "the general area surrounding the route was closed to teams starting on July 29, 2005, to ensure no participant had advance access to the actual route area." That is a contamination control enforced with a land closure rather than a promise, and it is stricter than anything a 2026 model evaluation does. 3. Pass/fail, not self-scored. The metric was: did the vehicle arrive, and when. There is no benchmark to select, no prompt to tune, no subset to report. 4. Every entrant published its method. DARPA required a technical paper from each team and published the set — "openly published to enable technical interchange among teams and with others in the robotics community." The winner's paper is Appendix B of the report to Congress. The loser's papers are in the same archive.
Hold those four properties, then read the rest of this entry, because the entry's real finding is that the test had all four and still did not tell you when the capability would ship. That is a harder and more useful lesson than "vendor benchmarks are unreliable," and it is what this file has that deep-blue-1997 does not.
What the test did not contain
The winners say it, in the peer-reviewed paper, more plainly than any critic has:
> While the DARPA Grand Challenge was a milestone in the quest for self-driving > cars, it left open a number of important problems. Most important among those > was the fact that the race environment was static. Stanley is unable to > navigate in traffic. . . . Even within the domain of driving in static > environments, Stanley's software can only handle limited types of obstacles. > For example, the present software would be unable to distinguish tall grass > from rocks.
Three structural exclusions sit underneath that, each documented by Thrun et al.:
- No moving objects at all. "When a faster robot overtook a slower one, the slower robot was paused by DARPA officials, allowing the second robot to pass the first as if it were a static obstacle. This eliminated the need for robots to handle the case of dynamic passing." The only vehicle Stanley ever met on course was H1ghlander, held stationary for it at Mile 101.5.
- No route-finding. "The RDDF defined the approximate route that robots would take, so no global path planning was required. As a result, the race was primarily a test of high-speed road finding, obstacle detection, and avoidance in desert terrain."
- No unexpected platform. "Before both the 2004 and 2005 Grand Challenges, DARPA revealed to the competitors that a stock 4WD pickup truck would be physically capable of traversing the entire course."
So the honest one-line statement of what was demonstrated is: a vehicle handed its route in advance, in a world containing nothing that moves, can stay on a desert trail at about 19 mph for seven hours. That is a real and difficult result. It is also about a quarter of the problem, and the people who did it said so in print the following year.
The winner, and the part of it that is actually AI
Stanley was a stock 2004 Volkswagen Touareg R5 TDI, chosen for fuel efficiency, with skid plates and a reinforced bumper. Sensors: five SICK laser range finders on the roof rack, forward-pointing at different tilt angles, useful to about 25 m; one colour camera for long-range road perception; two 24 GHz radars out to 200 m; a GPS receiver, a two-antenna GPS compass and a 6-DOF IMU. Compute: six Pentium M blade computers in the trunk, Linux, about 500 W drawn from the stock alternator, roughly 30 software modules running in parallel. The radar and a stereo rig were both tested and both dropped before the race — the radar for a flaky USB driver and because "the probability of encountering large frontal obstacles was small in high-speed zones." The car that won drove on lasers, one camera, and dead reckoning.
The team was about 50 people, formed in July 2004, sponsored by Volkswagen's Electronics Research Lab, Mohr Davidow Ventures, Red Bull and Android — the company Google bought two months before the race — with personal donations from David Cheriton and Vint Cerf.
What makes this an entry in an AI canon rather than a robotics one is where the knowledge came from. Three components were learned rather than written:
- Terrain classification from human driving. "The specific functions involved in detecting obstacles are determined through a machine learning algorithm, which relies on human driving to acquire 'training examples' of drivable terrain." A person drove; the labels came from where the car had safely been.
- Self-supervised vision, online. The camera module classified terrain by colour and texture out beyond laser range, and it was trained during the race from the lasers themselves — "using near-range data classified by the lasers to determine the current best model of the road surface," adapting at 8 Hz. The cheap, short-range, reliable sensor labels data for the expensive, long-range, unreliable one, continuously, with no human in the loop.
- Speed policy from human data. The transfer function setting velocity against terrain roughness "emulates human driving characteristics, and is learned from data gathered through human driving."
Localisation was an unscented Kalman filter at up to 100 Hz. The design principle the team states outright: "Treat autonomous navigation as a software problem."
DARPA saw this clearly and said so in March 2006, which section 3 grades as the best call in the document: "The competition was in large measure a software race that tested the ability of teams to define and implement robust software systems able to adapt and 'learn' the sensor signature of navigable versus impassable terrain through repeated exposure."
Three months before it won, the winner failed six times in one run
This is the fact that should travel with the result, and it is in the Stanford Racing Team's own technical paper — a milestone list, published before the race, in which the team grades itself:
> July 1, 2005: Autonomous traversal of the entire 2004 DARPA Grand Challenge > Course, with the exception of public roads (partially achieved July 16, 2005; > the team encountered a total of six failures, each at a level that would have > been fatal in a actual race).
Fourteen weeks before winning a race that permitted zero failures, the winning vehicle recorded six fatal failures in a single traversal of an easier course. Seven weeks before the race, on 20 August, it managed 140 uninterrupted miles. By suspension of development it had logged about 1,200 autonomous miles and lateral accuracy of roughly 30 cm.
And it was still failing on the day. Thrun et al. report that Stanley's laser data stream stalled 17 times during the race, for 300–1,100 ms each, almost all between Mile 22 and Mile 35, inserting phantom obstacles into the map. Four caused significant swerves: at Mile 22.37 Stanley drove briefly onto the berm; at Mile 34.69 it swerved across an open lake bed with no obstacle on it. The authors' verdict — "None of these incidents led to a collision or an unsafe driving situation during the race" — is accurate and is also the sound of a system that got away with it. During 4.7% of the course the GPS reported 60 cm of error or worse, and at Beer Bottle Pass the DARPA corridor "aligns poorly with the map data," such that "a robot that followed the GPS via points blindly would likely have failed to traverse this narrow mountain pass" and driving "66 cm further to the left would have been fatal in many places."
The winning margin over Carnegie Mellon was about eleven minutes. The team attributes it to one module: without the vision system Stanley would have been capped at 25 mph and finished around 7 h 05 min, "possibly behind CMU's Sandstorm."
The runner-up lost to a component nobody identified for twelve years
H1ghlander started first, led early, then began losing power on climbs and could not reach speed even on the flat. It lost more than forty minutes and was paused at Mile 101.5 to let Stanley by. Carnegie Mellon could not work out why — not that day, and not for over a decade.
On 16 October 2017, at an event marking the tenth anniversary of the Urban Challenge, H1ghlander's engine bay was being cleaned with the engine running when Spencer Spiker, CMU's operations lead, leaned against it with his knee and the engine started to die. The culprit was a small electronic filter between the engine control module and the fuel injectors: touching it dropped power, pressing it killed the engine. Twelve years of desert vibration explained at once. Red Whittaker's line to Chris Urmson, who had built the perception system and had spent a decade wondering: "You're off the hook!"
This is the entry's sharpest single item and a reading should use it exactly as it stands. The most-cited scoreboard in the history of autonomous vehicles — the one that made Stanford the winner, put Thrun on the path to Google, and is narrated as the moment the better approach won — was decided by an intermittent hardware fault in the runner-up that no one could identify for twelve years. The result was correctly recorded. It measured something other than what everyone took it to measure.
The most consequential thing at the 2005 race finished twelfth
David Hall of Team DAD had run stereo vision in 2004, concluded it was inadequate, and by September 2005 had a prototype bolted to his vehicle. DARPA photographed it and put it in the report to Congress as a subsystem note:
> Figure 10 shows a novel 64-sensor configuration, developed and demonstrated by > Team DAD from Morgan Hill, California. A rotating LIDAR system was designed to > create a low-cost system capable of full azimuthal coverage operating at an > update rate needed by a moving vehicle.
Team DAD placed twelfth, at 26 of 132 miles. That prototype became the Velodyne HDL-64E, and by the 2007 Urban Challenge it was on top of five of the six vehicles that finished. It went on to sit on Google's cars and to define the sensor architecture of the entire industry for a decade at roughly $75,000 a unit.
The winning vehicle's sensor suite — five fixed lasers with a 25 m useful range — was obsolete within two years, beaten by hardware built by a team that covered a fifth of the course. Any reading tempted to treat a leaderboard as a ranking of which approach matters should hold those two facts together.
What it cost
DARPA's accounting, filed with Congress: approximately $7.8 million, plus the $2 million prize. Around 100 government personnel, more than 200 staff in the final two weeks working 18-hour days, a 40-person operations center, four hilltop communications towers covering 500 square miles, and 30 leased pickup trucks as chase vehicles. The prize was 20% of what the event cost to run. Teams' own spending is not in the report and I did not find a credible total; the 43 semifinalist teams comprised, by DARPA's count, "more than 1,000 innovators who committed a significant portion of their personal time."
Why a reading would cite it
The occasion is standing, not singular, and this entry earns admission on a sentence in the project's own most recent reading. digests/2026-08-16-00.md closes the robotics lens by rejecting its only candidate: a Wired feature on Unitree's G1 whose "own subheading asks whether the robot 'can ever hold down a real job'," dismissed because "it does not meet this lens's demonstrated-versus-deployed test." LENSES.md states that test in the standing lens: "What was demonstrated under controlled conditions and what is actually deployed and working are different findings; say which."
The 2005 Grand Challenge is the founding measured instance of that gap, and what makes it worth citing is that the demonstration was honest. Deep Blue's lesson is about a test the vendor owned. This one is harder, because there is nothing wrong with the test:
- The demonstration was clean and the deployment took fifteen years. The course was withheld, the judge was disinterested, every method was published, and the result held up. And Waymo opened a fully driverless public ride service in Phoenix on 8 October 2020 — fifteen years to the day. (A coincidence, not a cause; but it is a checkable one, and it is the cleanest measured interval this canon has between "demonstrated in a controlled test" and "a member of the public can buy it.") A 2026 reading meeting a clean, independently-run capability demonstration should not conclude anything at all about schedule from the cleanness of the test.
- The winners' own limitations section is the template. "The race environment was static." "Stanley is unable to navigate in traffic." "Unable to distinguish tall grass from rocks." When a 2026 lab publishes an agentic-capability result, the Grand Challenge question is not "is the number real" but what did the environment not contain — no other agents, no adversary, no distribution shift, the route supplied.
digests/2026-08-15-12.mdrecords Anthropic marking its own catastrophic-misalignment risk up on the strength of "measurement limits in R&D-capability evaluation"; that is a lab reporting the same category of problem, and this entry is the dated precedent for taking the limitation more seriously than the score. - The scoreboard can be correct and still measure the wrong thing. H1ghlander lost to a pressure-sensitive electronic filter. Nobody knew for twelve years. When a 2026 reading meets a benchmark margin of a few points between two frontier systems, this is the instance where an eleven-minute margin in a cleanly-run contest turned out to be a hardware fault, and where the field built a decade of narrative on the ranking anyway.
- The subsystem beats the winner. The durable output was a lidar from a team that covered 26 miles. A reading that reports only the top of a leaderboard is reporting the part of the event least likely to matter in ten years.
- The variance is the finding, not the mean. The winner failed six fatal failures in July, swerved onto a berm at Mile 22, and won by eleven minutes. Five of twenty-three finished; the median did a fifth of the course. Where a 2026 capability claim rests on a single successful run, the Grand Challenge is the case where the full distribution was published and the tail was everything.
- Off-road autonomy is the control experiment that is still running. This is the item most useful to the robotics lens and least known. In June 2024, twenty years on, an Army official assessing its own off-road autonomy software said: "The good news is we are moving forward in that area. The bad news is industry is nowhere near where people think in terms of off-road autonomy. There's still a lot of development to do." Off-road autonomy at militarily relevant speeds is precisely and exactly what the 2005 race demonstrated. Two decades later the customer that commissioned the demonstration says the capability is not there.
What it should not be cited for. Any general rate of AI progress — see section 4, first item. Nothing in this file is evidence, nothing in it deposits into the ledger, and nothing in it touches the needle.
What it got right, and what it got wrong
Not required for moment. Included because the founding claim was a dated statutory goal, because the agency graded itself in writing twice, and because the base rate this project is assembling is worth more with them than without.
Wrong, completely, and by the party that commissioned the whole thing — Congress, made 30 October 2000, due 2015. Section 220 of the FY2001 NDAA: "It shall be a goal of the Armed Forces to achieve the fielding of unmanned, remotely controlled technology such that . . . by 2015, one-third of the operational ground combat vehicles are unmanned." Graded eleven years past due: not approached, and the program that was supposed to get there has stopped. The Congressional Research Service, updated 20 May 2025, records that on 1 May 2025 the Secretary of the Army and the Chief of Staff published a letter under the Army Transformation Initiative directing cancellation of programs that "deliver dated, late-to-need, overpriced, or difficult-to-maintain capabilities," and that the Army planned to "halt work on its embattled Robotic Combat Vehicle (RCV) program." The Army's reported positions: "The RCV will stop development. The future of the robotic software development is unknown." CRS's own summary sentence is the grade: "it can be argued past efforts were less than fully successful."
This is the most valuable item in the file and the least cited anywhere. A legislature set a numeric capability target twenty-five years out, an agency built a genuinely excellent measurement instrument to chase it, the instrument produced a genuine result inside two years, the result was celebrated worldwide — and the target was missed so completely that the successor program was cancelled a decade after the deadline. Every stage of that chain worked except the one that mattered. A reading meeting a 2026 government AI target with a date on it has a fully closed prior case.
Overclaimed, and hedged by its own next sentence — DARPA, March 2006. The report to Congress: "the achievements of the vehicles that completed the route can be said to demonstrate conclusively that autonomous vehicle are able to travel over rugged terrain at militarily relevant distances and speeds." Graded: "conclusively" is not available from n=5, one course, one run, no traffic, and no replication. The sentence immediately following is the honest one — "An autonomous vehicle that can operate safely in all environments remains a challenge, however, as future military missions will require unmanned vehicles that can operate closely with mounted and dismounted personnel in complex environments such as urban terrain" — and the two sentences are almost never quoted together. When they are, DARPA looks careful. When only the first is quoted, DARPA looks like a press release. Both halves or neither, which is the discipline deep-blue-1997 imposes on Drew McDermott and which applies here with the roles reversed.
Right, and it is the best call in the document — DARPA on machine learning, March 2006. From the same conclusion: "the success of the learning-based approach in this real-world context will impact other domains of machine learning and cognition, immediately and in the future." Graded twenty years later: correct, and more so than the vehicle claim it sits beside. The agency looked at a desert race, identified the transferable finding as learning from data rather than hand-coded rules, and said so in a procurement document six years before alexnet-2012. Note carefully what this does and does not license: DARPA was right about the method and wrong about the timetable, in the same paragraph, which is exactly the split deep-blue-1997 extracts from Simon and Newell — **right about what, wrong about when, and the two errors independent enough that grading them together destroys the information in both.**
Partly right, self-graded, and now due — DARPA's own legacy claim, made on its website, due 31 December 2025. DARPA writes that the challenges "helped to create a mindset and research community that a decade later would render fleets of autonomous cars and other ground vehicles a near certainty for the first quarter of the 21st century." The first quarter of the 21st century ended seven and a half months before this entry was written, so the claim is gradeable rather than pending, and the house rule about recording both the predictor's grade and an independent one applies directly, because this is a self-assessment published by the party being assessed.
Graded on the record, splitting the claim in two because it splits cleanly:
- "Fleets of autonomous cars" — substantially right, and this canon should say so plainly. As of March 2026 Waymo reports 220.6 million rider-only miles driven without a human driver across Phoenix, the San Francisco Bay Area, Los Angeles, Austin and Atlanta, at roughly half a million paid rides a week across about ten US cities, targeting a million a week by the end of 2026. That is a fleet, it is autonomous, and members of the public pay for it.
- "And other ground vehicles" — wrong, per the Army grade above.
- "Near certainty" — wrong as of when it mattered, and the intervening record is the reason. Ford and Volkswagen shut Argo AI on 26 October 2022 after roughly $3.6bn of combined investment and a $2.7bn Ford write-down, with CEO Jim Farley saying "profitable, fully autonomous vehicles at scale are a long way off." GM shut Cruise in December 2024 after more than $10bn of losses, following the 2 October 2023 incident in San Francisco in which a Cruise vehicle struck a pedestrian and then pulled over dragging them about 20 feet. On 18 March 2018 an Uber test vehicle killed Elaine Herzberg in Tempe; the NTSB found the system detected her about six seconds out, never classified her correctly, had no model for a pedestrian outside a crosswalk, and had its automatic emergency braking disabled. Nothing about that decade was a near certainty.
An independent grader reaching a different answer, and named because the house rule requires it: IEEE Spectrum's twenty-year retrospective by Allison Marsh notes autonomous vehicles confined to "specific applications" and closes, of the self-driving car, "I don't expect to find one under my carport anytime soon." DARPA's self-grade and the outside grade differ, they differ in the direction the canon should expect, and the honest resolution is that DARPA was right about one clause of three and the clause it was right about took twenty years.
Wrong by an order of magnitude, in four years, by the same person — Chris Urmson, made 2015, revised 2019. Urmson built Carnegie Mellon's perception system in 2004, 2005 and 2007, then led Google's self-driving programme. At TED in 2015 he said he wanted his eleven-year-old son never to need a driving test — a five-year horizon, due about 2020. At SXSW in 2016 he was already hedging: "If you read the papers, you see maybe it's three years, maybe it's 30 years." By 2019, as CEO of Aurora, he put integration onto roads "over the next 30 to 50 years." Graded: the same expert, on the same technology, moved his own forecast by roughly a factor of ten in four years, and the move was in the direction of longer. This is the single most useful forecasting item here, because Urmson is not a hype merchant — he is one of the two or three best-informed people alive on the subject, he was inside the programme, and he was still wrong by an order of magnitude about a technology he was personally building. Set it against deep-blue-1997's finding that Murray Campbell's engineering forecast for Deep Blue was off by only 20% and the pattern sharpens: people forecasting delivery of a system they are building are roughly calibrated; the same people forecasting when a capability becomes an industry are not, and Urmson made both kinds of forecast about the same career.
Right in structure, and the intervals are the artifact — the prize instrument itself. DARPA spent $9.8 million total and got: a 2004 failure, a 2005 result 19 months later, the 2007 Urban Challenge (Victorville, 3 November 2007; 89 teams down to 11 finalists, six finished; CMU's Boss first, Stanford's Junior second, Virginia Tech's Odin third), the Velodyne HDL-64E, and the founding personnel of Waymo, Aurora, Nuro, Zoox and Argo. Graded the way deep-blue-1997 grades the Fredkin Prize — as a measured rate rather than an asserted one — the inducement prize did what it says on the tin: three well-posed thresholds produced a three-year public record of the actual pace of a capability, which no amount of forecasting produced. What it did not produce was any information about the pace of the fifteen years after it stopped, and DARPA ran no further ground-vehicle grand challenge.
Right, and cheaper to say than to believe — Tony Tether's own framing, October 2005. Tether compared the result to "the Wright Brothers' first flight in 1903 in Kitty Hawk, N.C., proving it could be done," and said "these vehicles haven't just achieved world records, they've made history." Graded: the analogy is better than it was meant to be, and it cuts against the optimism it was deployed for. Kitty Hawk in December 1903 was followed by roughly a decade before the first scheduled commercial passenger service and about three decades before aviation was an ordinary way to travel. If 2005 was Kitty Hawk, then a fifteen- year wait for a paying public passenger and a still-unfinished rollout in 2026 is the analogy working, not failing. A person reaching for "this is the Kitty Hawk moment" in 2026 about any AI capability is, if they mean it, forecasting decades.
Defensible on people and sensors, weaker on method — Sebastian Thrun's retrospective. "It was truly the birth moment of the modern self-driving car." Graded: the personnel lineage is real and documented — Thrun, Montemerlo, Urmson, Levandowski, Salesky, Ferguson and Levinson all trace through these races into Google, Waymo, Aurora, Argo, Nuro and Zoox — and the sensor lineage is real via Velodyne. The method lineage is thinner than the phrase implies: Stanley's five fixed lasers, hand-tuned UKF and 30-module pipeline are not what a 2026 driving stack looks like, and the intervening rebuild through HD maps, learned perception and end-to-end models is most of the actual work. It is a birth moment for an industry and a community. It is not the ancestor of the software.
A rare case where the vendor number survives independent checking, and it should be recorded because this canon usually finds the opposite. Waymo reports 82% fewer any-injury-reported crashes and 94% fewer serious-or-worse injury crashes than human benchmarks over 220.6 million rider-only miles through March 2026; its methodology is published in Traffic Injury Prevention and its page names IIHS, UMTRI, VTTI and Swiss Re as independent reviewers. An independent 2026 IIHS analysis of overlapping data reportedly found 81% fewer injury crashes against the vendor's 82%. Graded: the numbers agree to within a point. The lesson is not that vendor numbers are fine; it is that vendor-reported numbers can be checked when the data is releasable and the methodology is published, and that a reading should demand that condition rather than assume either honesty or dishonesty. The 2026 frontier-model equivalent — digests/2026-08-15-12.md: "every number above except Qwen's weights is vendor-reported. No independent evaluator has published on this window's releases" — is the case where the condition is absent.
Commonly misused as
Not required for moment. Included because the "eighteen months" framing is this event's main afterlife in AI commentary, and stopping it is likely to be this file's most common service.
- "Eighteen months from total failure to complete success — that's how fast AI moves." The load-bearing misuse, and the reason to be careful is that the interval is true and the inference from it is not. Four things have to hold before that number means what it is used to mean, and none of them is checked by the people using it. (a) Success was not routine: the median vehicle did a fifth of the course. (b) The 2004 field was not the 2005 field — 195 applications against 106, an 84% increase, with 160 teams that had never entered before; a large part of the delta is a bigger, better-funded, better-informed population, not eighteen months of learning by fixed participants. (c) The courses differ in both directions and are not cleanly comparable — 2005 was ten miles shorter with more than 50 turns of at least 90 degrees, while 2004 had more elevation gain and the Daggett Ridge switchbacks — and I found no source that establishes which was harder. (d) Everyone had watched everyone else fail, in public, with published technical papers. The eighteen months is a real measurement of a specific contest under those four conditions. It is not a rate, and it is certainly not a rate for anything other than desert driving.
- "The moment self-driving cars were solved." Corrected at length above. The environment was static, the route was supplied, and the winner's authors wrote that it could not drive in traffic or tell tall grass from a rock.
- "DARPA created the self-driving car industry." Substantially true on personnel and on the sensor, and it should still be said with the two qualifications that make it interesting: the decisive hardware came from a team that covered 26 of 132 miles, and the winning software architecture was superseded well before the industry existed. Founding an industry and being the ancestor of its technology are different claims.
- "Prizes are a cheap way to buy research." The $2 million prize cost DARPA $7.8 million to administer — about 100 government personnel, 200+ staff in the last fortnight, four communications towers, 30 chase trucks. The prize was a fifth of the bill. The leverage argument for inducement prizes is real and rests on unpaid team spending, not on the ratio people usually quote.
- "The military got what it paid for." It did not, and this is the least-known fact about the most-known race. See section 3: the statutory goal was missed and the successor program stopped in 2025. What the Defense Department demonstrably got was a civilian industry it does not control and a recruiting effect DARPA itself claimed ("applicants to participant universities are specifically citing the Grand Challenge in their engineering graduate school applications").
- "A well-designed independent evaluation tells you when a technology will arrive." The core misuse this entry exists to stop, and the reason it is worth more than another vendor-benchmark cautionary tale. This test was withheld, disinterested, pass/fail, and fully published — four properties no 2026 model evaluation has all of — and the honest forecast it supported in October 2005 was still off by fifteen years. Test quality bounds how much you can trust the measurement. It does not bound how much the measurement tells you about deployment.
- And the mirror: "demos never mean anything." Also wrong. Waymo is real, paid for, at scale, and demonstrably safer than the human baseline on the strongest data available. The demonstration was not fake and the capability did arrive. It arrived on a schedule nobody in the desert would have accepted.
- "Stanley won because Stanford had the better approach." Partly. It also won by eleven minutes over a vehicle with an intermittent electronic fault that cost forty minutes and was not diagnosed until 2017. Both are true. A canon that lets the first stand alone is repeating a scoreboard as though it were a measurement.
Sources
A note on how this file was researched, because deep-blue-1997 asked for it to be fixed and this is the fix. That entry warned that almost none of its web material was read directly, and that "a canon that polices misquotation cannot be built out of summaries." Four documents here were read at first hand, as page images, and they carry nearly every number above:
- **DARPA, Report to Congress: DARPA Prize Authority — Fiscal Year 2005 report in accordance with 10 U.S.C. § 2374a, March 2006, approved for public release, retrieved from grandchallenge.org; pages 1–15 and Appendix A read. The primary source for this entry. Origin of: the verbatim Section 220 quotation and DARPA's statement that the prize decision "supported" it; the NAE 1999 report as the deliberative basis; the 195/136/118/43/23 selection funnel and the 84% increase over 106; the BLM permit as the cap on the field; the route description, the 29 July land closure and the two-hour reveal; Table 1, the full distance-completed table for all 23 finalists**, from which the median of 26 miles is computed here; the TerraMax overnight account; the $7.8M-plus-$2M cost, the ~100 government personnel and the 30 chase trucks; Program Element 0603764E and the FCS link; the Scientific American block quotation; the "demonstrate conclusively" claim and its hedge; the learning-based-approach claim; Figure 10 and the Team DAD rotating lidar; and the § 2374a statutory text.
- **Thrun, S., Montemerlo, M., Dahlkamp, H., Stavens, D., Aron, A., Diebel, J., Fong, P., Gale, J., Halpenny, M., Hoffmann, G., Lau, K., Oakley, C., Palatucci, M., Pratt, V., Stang, P., et al., "Stanley: The Robot that Won the DARPA Grand Challenge," Journal of Field Robotics 23(9), 661–692 (2006), received 13 April 2006, accepted 27 June 2006; retrieved from robots.stanford.edu; pages 661–666 and 686–692 read. Origin of: the 2004 "none . . . more than 5%" finding and 107 registered/15 raced; 195 registered/23 raced for 2005 and the 6:53:58 time; the 2,935 waypoints and 3–30 m corridors; the three structural exclusions (static environment, DARPA-paused passing, no global path planning, stock-4WD announcement); the sensor and compute description and the dropped radar and stereo; the NQE run statistics; the 17 laser stalls, the Mile 22.37 berm and the Mile 34.69 phantom swerve; the 4.7% GPS-error figure and the Beer Bottle Pass 66 cm analysis; the vision module's contribution to the winning margin; and the discussion section's limitations**, quoted at length above.
- Stanford Racing Team, "Stanford Racing Team's Entry In The 2005 DARPA Grand Challenge," the team's required technical paper (reproduced as an appendix to the DARPA report); all 18 pages read. Origin of: the ~50-person team formed July 2004 and the four working groups; the sponsor list including Android and the Cheriton/Cerf donations; "Treat autonomous navigation as a software problem" and "the DARPA Grand Challenge is largely a software competition"; the three machine-learning components quoted verbatim above; and the milestone list with the six fatal failures of 16 July 2005 and the 140-mile run of 20 August 2005.
- Congressional Research Service, "The Army's Robotic Combat Vehicle (RCV) Program," IF11876, version 15, updated 20 May 2025, by Andrew Feickert; all 3 pages read. Origin of: the June 2024 Army off-road autonomy assessment quotation, the 1 May 2025 Army Transformation Initiative letter and the halt to RCV development, the "RCV will stop development" line, and CRS's own "less than fully successful" summary.
Read directly, in the browser. Waymo's Safety Impact page, figures as of March 2026: 220.6M rider-only miles, 94%/82%/93% reductions with absolute counts, the five-city mileage breakdown, the Traffic Injury Prevention peer review, and the four named independent reviewers.
A discrepancy in the winning time, recorded rather than resolved. DARPA's own Report to Congress states Stanley's time as "6 hours, 53 minutes, 8 seconds" (p. 10). Thrun et al. in the Journal of Field Robotics, DoD's own DVIDS report and later retrospectives give 6 hours, 53 minutes, 58 seconds. The report's adjacent claim that Stanley was "approximately 11 minutes faster than the next vehicle" is arithmetically the better fit for 6:53:58 against Sandstorm's 7:04:50, which suggests the report contains a typo rather than a competing measurement — but that is an inference, not a finding, and the official government document and the winners' peer-reviewed paper disagree about the most quoted number in the event. Neither figure changes anything, which is exactly why it is worth recording: this is what it looks like when a canon checks a number it did not need to check.
Via search summaries, and flagged as such — these arrive at one remove and should be verified against originals before a reading leans on them: the H1ghlander diagnosis (IEEE Spectrum, "Carnegie Mellon Solves 12-Year-Old DARPA Grand Challenge Mystery," 2017 — the 16 October 2017 date, Spencer Spiker, the filter between engine control module and fuel injectors, the 40-minute cost and Red Whittaker's "You're off the hook!"); Chris Urmson's 2015 TED, 2016 SXSW and 2019 Aurora statements; Tony Tether's Kitty Hawk and "made history" quotations (DVIDS, October 2005); DARPA's innovation-timeline legacy claim; the 2007 Urban Challenge results; the Velodyne HDL-64E history and its presence on five of six 2007 finishers; the Argo AI shutdown figures and Jim Farley's quotation; the Cruise incident and shutdown; the NTSB findings on the Herzberg collision; the IIHS 81% figure; IEEE Spectrum's twenty-year retrospective by Allison Marsh; Thrun's "birth moment" quotation and Stanley's presence in the Smithsonian's National Museum of American History; and the 2004/2005 course-difficulty comparison, which I have deliberately left unresolved in section 4 because no source I found settles it.
Attempted and failed, or not attempted. I did not obtain team-level spending figures for any 2005 entrant from a primary source, and have therefore made no claim about them. I did not read the DARPA 2005 rulebook itself, only the sections DARPA quotes in its report. Section 220's operative text is quoted here as DARPA quotes it in a document filed with Congress rather than from the statute, which I could not extract from the Public Law text; the ellipsis is DARPA's. The claim that the 2004 race also used an RDDF with a two-hour reveal is supported only by secondary summaries and is not relied on above.
The 2026 citation occasions — the robotics lens rejecting its only candidate on the demonstrated-versus-deployed test on 16 August 2026, the vendor-reported-benchmark note of 15 August, and Anthropic's citation of measurement limits in R&D-capability evaluation — are as recorded in this project's own digests/2026-08-16-00.md and digests/2026-08-15-12.md, which hold the primary links. They are named here as occasions to cite this entry, not as evidence for anything in it. Nothing in this file is evidence, nothing in it is deposited in the ledger, and nothing in it touches the needle.