GPT-5.2 Launch: Benchmark Glory Meets User Backlash
The Anniversary Debut
OpenAI released the GPT-5.2 series on its tenth anniversary, calling it its strongest model family for professional knowledge work. It ships in three variants — Instant, Thinking, and Pro — targeting math, coding, and long-context reasoning for enterprise use cases like document analysis, financial modeling, and multi-step agent tasks.
Benchmark Highlights
- AIME 2025 (math): 100%
- GDPval (professional tasks vs. human experts): beats experts 70.9% of the time
- SWE-Bench Pro (coding): 55.6%
- ARC-AGI-2 (abstract reasoning): 52.9%
- On SimpleBench (a commonsense reasoning test with 200+ multiple-choice questions; human baseline 83.7%), GPT-5.2 scored below Claude Sonnet 3.7 from a year earlier — even the Pro version barely beat its predecessor.
- On dynamic benchmarks like LiveBench, GPT-5.2 did not fully top the charts, and higher token costs made some users hesitate to switch.
- "How many r's in garlic?" GPT-5.2 answered zero, while Gemini 3, DeepSeek R1, and Qwen3-Max got it right. Reproductions show unstable answers, sensitive even to letter casing — echoing the earlier "strawberry" fiasco.
- Math trick question: Despite a perfect AIME score, it miscalculated 5.9 − 5.11.
- Coding visualization: A traffic-light animation was functional but crude black-and-white, while Claude Opus 4.5 produced a colorful spinning cart. Its ASCII-art Mona Lisa also lacked the charm of GPT-4o's version.
- Responding to a user describing a panic attack with "I'm glad to hear that!"
- Telling a grieving child a pet's "body stopped working," versus GPT-4o's warm acknowledgment of the bond.
- Refusing to transcribe a philosophy article or answer simple persona questions due to safety filters.
- Advising "setting boundaries" in an infidelity scenario in a way that exposed the truth, missing interpersonal nuance.
OpenAI emphasizes strengths in building spreadsheets, presentations, and handling long multi-step projects, alongside improved tool calling and reduced hallucinations.
But Not All Benchmarks Agree
The Overnight Backlash
Within a day of launch, X and Reddit filled with negative feedback. Venture partner Deedy noted that while GPT-5.2 is smarter, consumers still miss GPT-4o's warmth and fun. Common complaints: bland personality, excessive safety guardrails, and an overall feeling of regression rather than upgrade.
Viral Failure Cases
The Emotional Intelligence Problem
The most damaging screenshots involved empathy failures:
Conclusion
GPT-5.2's launch played out like a drama: a triumphant opening followed by a wave of mockery. OpenAI faces a structural dilemma — enterprise markets demand powerful, safe tools, while consumers want engaging, empathetic companions. The episode suggests that genuine progress requires more than stacked benchmark scores; it requires balancing intelligence with humanity.
---
References 1. OpenAI Official Blog: Introducing GPT-5.2. 2. Machines Heart Report: GPT-5.2 Negative Reviews Surge. 3. X Posts and Threads on GPT-5.2 Feedback (Deedy, Scaling01, Bindu Reddy et al.). 4. Independent Benchmarks: SimpleBench, LiveBench, SWE-Bench Results. 5. User Comparisons: Programming and Emotional Response Tests on X.