Background: From Logic to Transformers
Łukasz Kaiser studied mathematics at the University of Wrocław, earned a PhD in logic and automata theory at RWTH Aachen, and worked as a CNRS researcher in Paris before joining Google Brain in 2013. Four years later he co-authored *Attention Is All You Need*. Of the eight Transformer co-authors, some founded Character.AI (Noam Shazeer), Cohere (Aidan Gomez), and Essential AI (Ashish Vaswani); Kaiser joined OpenAI in 2021 and contributed to GPT-4, o1, o3, and GPT-5.
His central claim from recent interviews (Jon Hernandez, Matt Turck's MAD Podcast, OpenAI Forum, October–November 2025): the "predict the next token" paradigm has reached its ceiling.
Why Next-Token Prediction Is Not Enough
Kaiser cites OpenAI's 2021 paper *Training Verifiers to Solve Math Word Problems* (Cobbe et al.), which estimated that solving high-school-level math word problems by scaling alone would require hundreds of billions of parameters — far beyond GPT-3's 175B. His conclusion: pure autgressive prediction is fundamentally memory and pattern matching. It learns statistical associations from internet text but lacks explicit intermediate reasoning steps, so it guesses or loops on multi-step problems.
What Makes Reasoning Models Different
With the o1 series (September 2024), Kaiser says AI entered a new paradigm:
- Tiny data: reasoning models train via reinforcement learning on verifiable tasks using a tiny amount of data compared to internet-scale pretraining.
- Chain of thought as a core mechanism: models generate hidden intermediate reasoning steps, like scratch paper, spending more compute before answering.
- Better generalization: math problems that once seemed to need hundreds of billions of parameters can now be solved by small reasoning models — not because models got bigger, but because the training paradigm changed.
- GPT-5 is natively trained on images and audio, can generate both, with video coming.
- GPT-5 Pro runs multiple chains of thought in parallel, then synthesizes a final answer — which explains its higher cost.
- Retraining (e.g., GPT-4o) is mainly for cost reduction, not capability: GPT-4o is not stronger than GPT-4, but much cheaper.
- Pretraining remains part of the pipeline, but post-training, RL, and data filtering contribute more to capability gains.
- Kaiser feels limited competitive pressure: researchers move between labs and no secret stays buried long.
- On business models: OpenAI has a strong internal culture against maximizing engagement for advertising. ChatGPT shopping recommendations explicitly cannot affect model outputs — biasing a model toward certain rankings in post-training produces strange results.
- He worries more about AI weapons than AI-generated slop: "You cannot control how it will be used."
- Cogniscendo: "2025-10-29: Some confirmations from OpenAI" (Jon Hernandez interview breakdown)
- The MAD Podcast with Matt Turck: "OpenAI's Łukasz Kaiser" (Nov 2025)
- OpenAI Forum: "Learning Powerful Models: From Transformers to Reasoners and Beyond" (Oct 2025)
- Cobbe et al., 2021: *Training Verifiers to Solve Math Word Problems*
No AI Winter
Kaiser rejects "AI winter" narratives for three reasons:
1. The reasoning paradigm is very early, like Transformer in 2017. 2. Plenty of low-hanging engineering optimizations remain unexploited at OpenAI. 3. AI progress will extend like Moore's law — exhausting each S-curve and finding the next one.
GPT-5 and Future Models
A Contrarian Take on Video Data
Kaiser is skeptical that large volumes of video data improve tasks like mathematics. Language already encodes humanity's abstract worlds; video captures the physical one. He believes video matters for robotics and physical simulation, but video-based "world models" will not generalize well, because the world in our heads is largely abstract, not physical.
Competition, Advertising, and Values
His Research Direction
Kaiser's current focus is learning reasoning from arbitrary data — beyond verifiable domains like math and code, toward the fuzzy, subjective, real-world decisions where most human reasoning happens. He sees this as the next major breakthrough.
Key Takeaways
1. The paradigm shift has happened: from bigger models to models that "think." 2. Data is not the bottleneck; reasoning via RL + chain of thought is the key. 3. Multimodality is necessary but not sufficient — true intelligence draws on the abstract world encoded in text. 4. No AI winter: early paradigm, huge optimization headroom, new S-curves ahead. 5. Values matter: don't optimize for engagement; fear misuse (weapons), not slop.
Kaiser's 2025 Global ML Tech Conference talk was titled "Reasoning Models: Past, Present and Future." The creator of Transformer is now trying to create the next paradigm — toward a model that "learns to become a great programmer, a great conversational agent, capable of both vision and language."