Measured evidence instead of momentum
Keep ThinkingAI · Markets · Society
The ledger

The receipts.

Notable claims about AI, kept on file. What was promised, when, by whom, and what actually happened. True claims stay in the ledger next to the false ones; an honest scoreboard has both.

Confirmed · 3
Refuted · 8
Mixed · 6
Abandoned · 1
Pending · 1

August 2025

Sam Altman, OpenAI

Capability

Launching GPT-5, Altman said: "GPT-5 is the first time that it really feels like talking to an expert in any topic, like a PhD-level expert." OpenAI billed the model as smarter, faster, and more accurate, with a lower hallucination rate than its predecessors.

What happened · August 2025: GPT-5 posted state-of-the-art scores on several benchmarks at launch, including 74.9% on SWE-bench Verified, but reception fell well short of the framing. A broken model router made the system look worse than it was on day one, users pushed back on the removal of GPT-4o, and OpenAI restored it for paying subscribers within a week. Altman conceded to The Verge: "I think we totally screwed up some things on the rollout," while noting API traffic had doubled in 48 hours. Independent reviewers generally found a strong model, not a PhD-level expert in every topic.

The benchmarks were genuinely strong and the launch was still received as a disappointment, which is what a PhD-level promise does to expectations. The framing held best on narrow technical tasks and worst in everyday use. Capability claims at this scale now get graded by hundreds of millions of users within days.

Mixed

March 2025

Dario Amodei, Anthropic

Prediction

At a Council on Foreign Relations event, the Anthropic CEO said: "I think we will be there in three to six months, where AI is writing 90% of the code. And then in 12 months, we may be in a world where AI is writing essentially all of the code."

What happened · October 2025: Inside Anthropic the prediction roughly arrived on schedule: by October 2025 Amodei said Claude was writing about 90% of the code for many teams at the company, while adding that this required just as many software engineers as before, or more. Industry-wide the figure did not hold, with large employers such as Microsoft and Google reporting roughly a quarter to a third of new code as AI-generated through 2025, and the 12-month "essentially all of the code" horizon did not arrive.

The prediction was truer than critics allowed and narrower than it sounded: it came close inside the frontier lab that builds the tools, and missed the median company. Amodei also supplied the correct footnote himself, which is that AI writing 90% of the code did not mean 90% fewer engineers. Both the hit and the qualifier belong on the record.

Mixed

December 2024

Marc Benioff, CEO of Salesforce

Capability

"We're not adding any more software engineers next year because we have increased the productivity this year with Agentforce and with other AI technology that we're using for engineering teams by more than 30%." Benioff also said Salesforce would have fewer support engineers and would add 1,000 to 2,000 salespeople instead.

What happened · September 2025: The freeze was real and held through 2025, and Benioff later extended it, saying he was not hiring more engineers in fiscal 2026 either because of coding agents. In September 2025 he said Agentforce had let him cut customer support from 9,000 heads to about 5,000. Salesforce kept hiring in sales and AI roles, and by 2026 Benioff was saying publicly that AI cannot replace software engineers yet.

One of the few AI workforce claims the claimant largely followed through on, and that deserves credit. The 30% productivity figure was self-reported and never audited, and the same CEO who froze engineering hiring later conceded AI cannot replace software engineers yet. The freeze and the caveat both belong in the record.

Mixed

September 2024

Sam Altman, OpenAI

Prediction

It is possible that we will have superintelligence in a few thousand days (!); it may take longer, but I'm confident we'll get there. Altman made the claim in his essay The Intelligence Age, arguing that deep learning worked, got predictably better with scale, and that the path forward is paved with compute, energy, and human will.

A few thousand days from September 2024 lands somewhere in the early 2030s, so this receipt stays open for years by construction. The hedge, that it may take longer, does real work and makes the claim hard to falsify until the outer window closes. Worth noting that the essay is partly a fundraising document for the compute buildout it describes, which does not make it wrong but does make it interested.

Pending

June 2024

Apple

Product

At WWDC on June 10, 2024, Apple announced a more personal Siri as the centerpiece of Apple Intelligence: onscreen awareness, understanding of personal context, and the ability to take hundreds of new actions within and across apps. Apple said the features would roll out over the course of the next year, and it marketed the iPhone 16 around them.

What happened · May 2026: None of the personalized Siri features shipped in that cycle. In March 2025 Apple said delivery was taking "longer than we thought," pushed the rollout into the coming year, and pulled an iPhone 16 ad built around the features. By May 2026 they still had not shipped, and Apple agreed to pay $250 million to settle a class action alleging it had advertised AI capabilities that did not exist.

Apple's mistake was announcing a demo as a product and selling phones against it for months. The company has not abandoned the features and still says they are coming in 2026, which is why this reads as refuted rather than abandoned: the original claim failed publicly, expensively, and on the record.

Refuted

February 2024

Klarna

Product

Our AI assistant, built with OpenAI, does the equivalent work of 700 full-time customer service agents. It handled 2.3 million conversations in its first month, cutting resolution time from 11 minutes to under 2.

What happened · May 2025: The deployment was real and the numbers held. But by May 2025, CEO Sebastian Siemiatkowski said the company had cut costs too aggressively at the expense of quality, and Klarna began hiring human agents again so customers could always reach a person. The AI stayed; the full-replacement story did not.

The claim was true and the reversal was also true, which is what makes this the model receipt. Automation worked on volume, and the company itself decided that full replacement was the wrong product. Both halves deserve to be quoted, and usually only the first one is.

Mixed

January 2024

Rabbit (Jesse Lyu)

Product

At CES on January 9, 2024, Rabbit founder Jesse Lyu unveiled the R1, a $199 handheld powered by what the company called a Large Action Model that learns app interfaces and carries out tasks on your behalf, doing just about everything your phone can do, only faster. The first batch of 10,000 units sold out within a day.

What happened · May 2024: When the R1 shipped in April 2024, the LAM connected to only four services (Uber, DoorDash, Midjourney, and Spotify), several basic features were missing or broken, and more than 100,000 preorder customers received what The Verge scored a 3 out of 10, calling it an unfinished, unhelpful AI gadget with basically no evidence of a LAM at work. Wired later counted it among the three biggest hardware flops of 2024.

The R1 is a useful calibration for demo-driven claims: the keynote showed an agent that acts on apps, and the shipping product was a thin wrapper around four logged-in websites. Rabbit priced the device fairly and kept shipping updates, but the gap between the promised Large Action Model and what reviewers could actually test was never closed.

Refuted

November 2023

Humane (Imran Chaudhri and Bethany Bongiorno)

Product

On November 9, 2023, Humane launched the AI Pin, a $699 screenless wearable plus a $24 monthly subscription, pitched as the first device of a post-smartphone era of ambient computing. The company, founded by two former Apple executives and backed by more than $230 million in funding, promised an assistant that would replace apps and screens with voice, gestures, and a laser display projected onto your palm.

What happened · February 2025: The Pin shipped in April 2024 to scathing reviews; The Verge concluded it just doesn't work, and Marques Brownlee called it the worst product he had ever reviewed. By August 2024 daily returns were outpacing sales. On February 18, 2025, Humane sold most of the company to HP for $116 million, and on February 28 its servers shut down, leaving every sold Pin without calling, messaging, or AI features.

The ambition was genuine and so was some of the engineering; HP paid $116 million largely for the team, the patents, and the operating system. But the product shipped before it worked, and a device whose features live in the cloud dies with the company behind it. The bricking of every sold unit about ten months after launch is the part of this receipt most worth remembering.

Refuted

May 2023

Arvind Krishna, CEO of IBM

Prediction

IBM would pause or slow hiring for back-office roles such as human resources, a group of roughly 26,000 workers. "I could easily see 30% of that getting replaced by AI and automation over a five-year period," Krishna told Bloomberg, about 7,800 jobs.

What happened · May 2025: The automation half came true: by 2025 IBM's AskHR handled 94% of routine HR inquiries and the company had replaced several hundred HR workers with AI agents. The replacement half did not: Krishna told the Wall Street Journal in May 2025 that "our total employment has actually gone up," because AI freed investment to hire more programmers, marketers and salespeople. In 2026 IBM announced it would triple US entry-level hiring, including for roles it had once said AI could do.

This was a five-year forecast about roles, widely misreported as 7,800 people being laid off on the spot. Two years in, IBM's own account is that automation removed tasks while total headcount grew, though thousands of HR positions were genuinely eliminated along the way. A hiring pause and a workforce reduction are different events, and this receipt is often cited as the second when it was only ever the first.

Mixed

March 2023

OpenAI

Capability

GPT-4 "passes a simulated bar exam with a score around the top 10% of test takers; in contrast, GPT-3.5's score was around the bottom 10%."

What happened · 2024: The pass was real and was independently confirmed: researchers estimated GPT-4 at around 297 out of 400 on the Uniform Bar Exam, well above the passing threshold in most jurisdictions. A 2024 peer-reviewed re-evaluation by Eric Martinez found the top-10% framing inflated, because OpenAI compared against a February test pool heavy on repeat takers; against first-time July takers GPT-4 landed around the 62nd percentile overall, which is still a clear pass.

The result itself held up, and within two years it went from headline to baseline that every new model clears. The disputed part was the percentile, not the pass: the comparison pool did real work in the marketing. A true claim can still be framed at its most flattering edge, and this one was.

Confirmed

November 2022

CNET (Red Ventures)

Product

Starting around November 2022, CNET quietly published dozens of AI-generated personal finance explainers under a "CNET Money Staff" byline. The disclosure, visible only after clicking the byline, read: "This article was generated using automation technology, and thoroughly edited and fact-checked by an editor on our editorial staff."

What happened · January 2023: In January 2023 Futurism exposed the program and documented basic factual errors, including a botched compound interest calculation. CNET ended up issuing corrections on 41 of the 77 AI-written stories, some for phrasing that was not entirely original, and parent company Red Ventures told staff it was pausing AI-generated content across all of its sites.

The failure was more editorial than technical: the byline hid the authorship behind a click, and the promised human fact-check did not catch errors a careful reader would. CNET corrected most of the stories and said it would keep experimenting with AI tools, which makes this a governance receipt rather than a verdict on the technology itself.

Refuted

November 2022

Meta AI

Product

Meta AI released Galactica, a large language model trained on 48 million scientific papers, textbooks, and reference material, and presented it as a system that can store, combine and reason about scientific knowledge. A public demo invited users to generate summaries, literature reviews, and answers to scientific questions.

What happened · November 2022: Within hours of the November 15, 2022 launch, researchers showed the demo producing authoritative-sounding but fabricated science, including fake papers, wrong biographies, and plausible nonsense. Meta withdrew the public demo on November 18, three days after launch, saying the models remain available for researchers. The paper and model weights stayed public, but the product never returned.

Galactica was a serious research effort and its benchmark results were real, which made the public demo worse, not better: the model's fluency was exactly the hazard when applied to science. Pulling it after three days was the right call, though it happened only after public pressure. The episode became the standard example of why hallucination matters more in high-trust domains than in chat.

Abandoned

February 2021

Zillow

Capability

Zillow made the Zestimate itself the initial cash offer for eligible homes in more than 20 cities through its Zillow Offers home-buying service. "This exciting advancement demonstrates the confidence we have in the Zestimate," said COO Jeremy Wacksman, crediting advances in machine learning and AI.

What happened · November 2021: Nine months later Zillow announced it was shutting down Zillow Offers, cutting about 25% of its workforce and expecting more than $500 million in second-half 2021 losses on homes it had overpaid for. CEO Rich Barton said the unpredictability in forecasting home prices "far exceeds what we anticipated."

The Zestimate was a decent retrospective valuation tool with a published 1.9% median error on listed homes; the failure was promoting it into a forward pricing engine in a fast-moving market. The model did roughly what it was built to do, and the business decision to bet a balance sheet on it is what failed. Model accuracy and model risk are different things.

Refuted

November 2020

Google DeepMind

Capability

At the CASP14 blind assessment, DeepMind announced that AlphaFold 2 had achieved a median accuracy of 92.4 GDT across all protein targets, a level CASP's organizers regard as competitive with experimental methods, and said the 50-year-old grand challenge of predicting protein structure from amino acid sequence had effectively been solved.

What happened · October 2024: The claim held. DeepMind published the peer-reviewed Nature paper in July 2021 and opened a structure database that grew past 200 million predicted proteins. On October 9, 2024, Demis Hassabis and John Jumper received the Nobel Prize in Chemistry for protein structure prediction, sharing the prize with David Baker, who was honored for computational protein design.

A rare capability claim that never needed revision. CASP is a blind assessment judged against unpublished experimental structures, so the 92.4 figure was earned in the open, and the field's own organizers declared the problem solved before any prize was involved. This is the control case for what a verified AI breakthrough looks like.

Confirmed

April 2019

Elon Musk, Tesla

Prediction

We will have more than one million robotaxis on the road. A year from now, we'll have over a million cars with full self-driving, software... everything. Musk made the prediction at Tesla's Autonomy Day for investors, promising Level 5 autonomy with no geofence and a robotaxi network in parts of the US by 2020.

What happened · June 2025: End of 2020 came and went with zero Tesla robotaxis in service. Tesla did eventually launch a robotaxi pilot in Austin on June 22, 2025, but it was a limited trial: roughly 10 to 20 Model Y vehicles, invite-only riders, a geofenced area, and a safety monitor in the front passenger seat.

The deadline and the number were both wrong by an order of magnitude, and the 2025 pilot, while real, was supervised and geofenced rather than the open Level 5 network described in 2019. The fair reading is that the direction was right and the timeline was not. Six years late with safety monitors is a different claim than one million driverless cars in eighteen months.

Refuted

November 2016

Geoffrey Hinton, University of Toronto

Prediction

People should stop training radiologists now. It's just completely obvious that within five years, deep learning is going to do better than radiologists. It might be ten years, but we've got plenty of radiologists already. Hinton made the remarks at the Machine Learning and the Market for Intelligence conference in Toronto, comparing radiologists to the coyote that has run off the cliff but not yet looked down.

What happened · July 2026: A decade later, no radiologists have been replaced by AI. The profession faces a persistent shortage, residency positions are at record highs, and average US radiologist pay reached about $571,000 in 2025. More than 1,000 radiology AI devices have FDA clearance, but they work as tools inside radiology practices rather than as replacements for the radiologists using them.

Hinton was right that deep learning would get very good at reading images, and it did. He was wrong that task-level performance would translate into job replacement, because a radiologist's work is much more than image classification and demand for imaging rose as it got cheaper to analyze. Hinton himself later softened the prediction, which makes this a cleaner miss than the mythology around it suggests.

Refuted

June 2016

Jeff Bezos, CEO of Amazon

Prediction

The technology behind Alexa could become the "fourth pillar" of Amazon's business alongside its retail marketplace, Prime and AWS, Bezos said at Recode's Code Conference, revealing that more than 1,000 people were already working on it.

What happened · July 2024: Alexa reached enormous scale, but people used it mostly for free tasks like timers, weather and music rather than the voice shopping Amazon had bet on. The Wall Street Journal reported in July 2024 that Amazon had lost tens of billions of dollars on its devices business, including more than $25 billion between 2017 and 2021, and that the unit had no profit timeline. Amazon has kept investing, launching an AI-upgraded Alexa Plus in 2025.

The adoption half of the bet worked and the monetization half never did. One caveat: the reported losses cover Amazon's whole devices unit, not Alexa alone, so no exact Alexa figure is public. Even with that discount, the fourth pillar has not arrived, and the claim stays open only because Amazon keeps funding it.

Mixed

October 2013

IBM, with MD Anderson Cancer Center

Product

IBM and MD Anderson announced that Watson would power the cancer center's Moon Shots program, aimed at ending cancer, starting with leukemia. The pitch, repeated across IBM's marketing in the early 2010s, was that Watson's cognitive computing would read the medical literature, reason over patient records, and revolutionize cancer treatment.

What happened · January 2022: MD Anderson shelved its Oncology Expert Advisor project in 2017 after spending about $62 million; a University of Texas audit found the system had not been used on a single patient. IBM kept marketing Watson for Oncology elsewhere, but in January 2022 it sold the Watson Health data and analytics assets to private equity firm Francisco Partners, which folded them into a new company called Merative.

Watson's Jeopardy win was genuine engineering, and the idea that a machine could digest the oncology literature was not foolish in 2013. The failure was in the gap between a question-answering demo and a clinical system that had to integrate with real hospital data and real workflows. The episode is now the standard cautionary tale for selling general AI capability as a finished medical product.

Refuted

April 2009

IBM

Capability

IBM announced it was in the final stages of building a question-answering supercomputer, Watson, designed to parse natural-language clues well enough to take on Jeopardy's two greatest champions, Ken Jennings and Brad Rutter, and beat them on television.

What happened · February 2011: Over three televised games on February 14-16, 2011, Watson defeated Jennings and Rutter with a final total of $77,147 against $24,000 and $21,600, taking the $1 million top prize, which IBM donated to charity.

The win was real and it was earned on a genuine natural-language problem under broadcast conditions. The instructive half is what it did not generalize into: IBM spent years trying to convert the brand into products, most visibly Watson Health, whose assets were sold off in 2022. Game shows are closed worlds with clean answers; most markets are not.

Confirmed

Grading: confirmed means happened as claimed · refuted means did not happen, or failed publicly · mixed means true in part, reversed or narrowed · abandoned means quietly dropped · pending means not yet evaluable. Suggest a claim for the ledger via the contact page.

Newsletter

New papers, when they are ready. One email per analysis, nothing else.