English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GPT-5.2 Launch: Benchmark Glory Meets User Backlash

Forum topic · ✨步子哥 · 2025-12-15

Summary

OpenAI marked its tenth anniversary by releasing the GPT-5.2 model family, offered in Instant, Thinking, and Pro variants and billed as its strongest lineup for knowledge work. The models posted striking benchmark results: 100% on AIME 2025 math, 70.9% on GDPval versus human experts, 55.6% on SWE-Bench Pro, and 52.9% on ARC-AGI-2. Yet within a day, X and Reddit filled with criticism. Independent tests like SimpleBench showed weak commonsense reasoning, GPT-5.2 failed the 'count the r's in garlic' puzzle, and its code-generated visuals and ASCII art lagged rivals like Claude Opus 4.5. Users were most upset by emotionally cold responses — congratulating a user on a panic attack, handling pet loss insensitively — and by over-tightened safety filters refusing benign requests. The backlash highlights a core tension for OpenAI: enterprise buyers want power and safety, while consumers miss the warmth and playfulness of GPT-4o. This article compiles the benchmarks, failure cases, and community reactions in full.

GPT-5.2 Launch: Benchmark Glory Meets User Backlash

The Anniversary Debut

OpenAI released the GPT-5.2 series on its tenth anniversary, calling it its strongest model family for professional knowledge work. It ships in three variants — Instant, Thinking, and Pro — targeting math, coding, and long-context reasoning for enterprise use cases like document analysis, financial modeling, and multi-step agent tasks.

Benchmark Highlights

  • AIME 2025 (math): 100%
  • GDPval (professional tasks vs. human experts): beats experts 70.9% of the time
  • SWE-Bench Pro (coding): 55.6%
  • ARC-AGI-2 (abstract reasoning): 52.9%
  • OpenAI emphasizes strengths in building spreadsheets, presentations, and handling long multi-step projects, alongside improved tool calling and reduced hallucinations.

    But Not All Benchmarks Agree

  • On SimpleBench (a commonsense reasoning test with 200+ multiple-choice questions; human baseline 83.7%), GPT-5.2 scored below Claude Sonnet 3.7 from a year earlier — even the Pro version barely beat its predecessor.
  • On dynamic benchmarks like LiveBench, GPT-5.2 did not fully top the charts, and higher token costs made some users hesitate to switch.
  • The Overnight Backlash

    Within a day of launch, X and Reddit filled with negative feedback. Venture partner Deedy noted that while GPT-5.2 is smarter, consumers still miss GPT-4o's warmth and fun. Common complaints: bland personality, excessive safety guardrails, and an overall feeling of regression rather than upgrade.

    Viral Failure Cases

  • "How many r's in garlic?" GPT-5.2 answered zero, while Gemini 3, DeepSeek R1, and Qwen3-Max got it right. Reproductions show unstable answers, sensitive even to letter casing — echoing the earlier "strawberry" fiasco.
  • Math trick question: Despite a perfect AIME score, it miscalculated 5.9 − 5.11.
  • Coding visualization: A traffic-light animation was functional but crude black-and-white, while Claude Opus 4.5 produced a colorful spinning cart. Its ASCII-art Mona Lisa also lacked the charm of GPT-4o's version.
  • The Emotional Intelligence Problem

    The most damaging screenshots involved empathy failures:

  • Responding to a user describing a panic attack with "I'm glad to hear that!"
  • Telling a grieving child a pet's "body stopped working," versus GPT-4o's warm acknowledgment of the bond.
  • Refusing to transcribe a philosophy article or answer simple persona questions due to safety filters.
  • Advising "setting boundaries" in an infidelity scenario in a way that exposed the truth, missing interpersonal nuance.
Users described conversations as preachy, eerie, and gaslighting — one quipped that GPT-5.2 raises your blood pressure. Compared with GPT-4o's lively persona, GPT-5.2 feels like a precise but cold observer.

Conclusion

GPT-5.2's launch played out like a drama: a triumphant opening followed by a wave of mockery. OpenAI faces a structural dilemma — enterprise markets demand powerful, safe tools, while consumers want engaging, empathetic companions. The episode suggests that genuine progress requires more than stacked benchmark scores; it requires balancing intelligence with humanity.

---

References 1. OpenAI Official Blog: Introducing GPT-5.2. 2. Machines Heart Report: GPT-5.2 Negative Reviews Surge. 3. X Posts and Threads on GPT-5.2 Feedback (Deedy, Scaling01, Bindu Reddy et al.). 4. Independent Benchmarks: SimpleBench, LiveBench, SWE-Bench Results. 5. User Comparisons: Programming and Emotional Response Tests on X.

Tags

#gpt-5-2#openai#ai-benchmarks#large-language-models#ai-safety#user-experience#claude#gemini

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/176415131