Gary Marcus: The skeptic the industry keeps proving right
He sold an AI company to Uber, wrote the book that predicted the LLM era's failure modes, and testified to the Senate next to Sam Altman. The distrust wave has an intellectual author, and it is him.
The core position
Large language models are powerful pattern matchers, not paths to general intelligence. Real progress requires hybrid systems that combine learning with explicit reasoning, and real safety requires verification instead of vibes.
The lab's read
Marcus is the distrust wave's intellectual infrastructure. Long before procurement departments started asking how would we know, he was cataloguing the failure modes that made the question necessary. Whether or not neurosymbolic AI is the full answer, the market now prices the skepticism he spent a decade arguing for.
The insider who doubted
Gary Marcus is not an outsider throwing stones. He trained as a cognitive scientist under Steven Pinker at MIT, built a career studying how children actually learn language, and in 2015 founded Geometric Intelligence, a machine learning company focused on learning from small amounts of data. Uber acquired it in 2016 to anchor its AI labs, and Marcus briefly led the effort from inside one of the industry's most aggressive adopters.
That biography matters for reading everything he writes afterward. When Marcus argues that deep learning alone cannot get us to robust intelligence, he is not rejecting a technology he failed to master. He is describing the ceiling he hit while building with it.
Rebooting AI, before it was fashionable
In 2019, with computer scientist Ernest Davis, Marcus published Rebooting AI, a systematic argument that the field was confusing narrow competence with general understanding. The book's central claims were specific: systems trained on correlation would fail at compositionality, would hallucinate structure that was not there, and would break in ways their creators could not predict when pushed outside the training distribution.
At the time, the argument read as pessimism. The GPT era then proceeded to demonstrate each failure mode at planetary scale, in production, with invoices attached. Hallucination, brittle reasoning, unpredictable jailbreaks: the book reads less like prophecy than like a bug tracker written early.
The LLM years
When ChatGPT arrived in late 2022, Marcus became the most cited internal critic of the scaling thesis. His Substack became a running audit of LLM failure modes, written against a news cycle that announced a new miracle every week. In May 2023 he testified before the US Senate Judiciary subcommittee at the same hearing as Sam Altman, arguing for independent evaluation of AI systems before deployment, a position that sounded cautious then and reads as obvious now.
His public debates, with Yoshua Bengio and later with defenders of pure scaling, established the shape of the field's central argument: whether intelligence emerges from scale alone or whether something structurally different is required. Marcus's wager is that it is structural, and that the missing piece is explicit reasoning over representations, the thing classical AI was good at and deep learning discarded.
What he proposes instead
The constructive half of the critique is neurosymbolic AI: systems that combine neural networks' pattern recognition with symbolic machinery for logic, rules, and structured knowledge. Marcus has argued for this architecture since before it had a respectable constituency, and he founded Robust.AI to build robotics on principles closer to it.
His policy positions follow the same logic. If systems cannot verify themselves, verification must come from outside: independent audits, pre-deployment evaluation, liability for harms. The audit category now forming in enterprise budgets is, in large part, his agenda arriving on schedule.
Where the criticism lands
The fair reading of Marcus is not that he predicted everything correctly. He has been wrong about timelines and occasionally wrong about what LLMs cannot do at all. The fair reading is that he identified the right questions early: how would we know if the system understands, who checks, and what happens to the people affected when nobody does.
For a lab documenting the collision of AI claims and AI reality, Marcus is less a subject than a methodological ancestor. The distrust of unchecked capability claims is not anti-AI. It is the precondition for using the technology where it actually works. That sentence could sit above everything he has written.
What to take seriously
Skepticism is a research position, not a mood
Marcus's critique is specific and falsifiable: compositionality, grounding, reliable reasoning. Vague unease is cheap. Named failure modes are useful.
Beware capability claims without verification
His consistent demand for independent evaluation is now standard procurement practice. Ask how would we know before asking what can it do.
LLMs are a tool, not the toolbox
The neurosymbolic argument reduces to a practical point: different tasks need different machinery. Language models for language, reasoning systems for reasoning, rules for rules.
Being early looks like being wrong
Rebooting AI was dismissed in 2019 and reads as descriptive in 2026. Measured criticism ages better than enthusiasm.
Insiders make the best critics
Marcus's authority comes from having built and sold an AI company. Critique from inside the craft carries weight that commentary from outside cannot.
Sources & further reading
Keep Thinking
Independent analysis: no reselling, no vendor commissions. On the side, we help a small number of companies implement what we write about.
Work with the lab