English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily | December 12, 2025: GPT-5.2 Launch, Disney's $1B OpenAI Deal, and Open-Source AI Tooling Updates

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily for December 12, 2025 rounds up key AI industry developments. OpenAI released GPT-5.2, improving scientific reasoning (92.4%), competition math (100%), and long-context handling, while raising pricing to $1.75/M input and $14/M output tokens; it beat 70.9% of human experts on GDPVal tasks at roughly 1% of expert cost, though community benchmarks flagged weak coding performance. Disney invested $1 billion in OpenAI to bring its characters into the Sora video generator under a 3-year licensing deal. DeepMind opened its first automated research lab in the UK, planned for 2026. Open-source updates include Unsloth's packing technique delivering 3x faster training, llama.cpp's new routing mode for live model switching, and Hugging Face's fully in-browser WebGPU voice chat demo. Infrastructure news covers CUDA 13 compatibility fixes, Hetzner's €889 96GB VRAM server, and research on diffusion model log probabilities and sandwich normalization for Transformers.

📅 AI Industry Roundup — December 12, 2025

A translated digest of the Easy AI Daily from zhichai.net covering model releases, industry deals, open-source tools, benchmarks, and community discussions.

Model Updates & Releases

GPT-5.2 Released: Performance Gains at Higher Prices

OpenAI released GPT-5.2 with improvements in scientific reasoning (92.4% accuracy), competition math (100%), and long-context processing. Pricing increased to $1.75 per million input tokens and $14 per million output tokens, with a 90% discount on cached input. It ranks second on the WebDev Code Arena, though it underperforms on some coding benchmarks.
  • OpenAI official blog
  • System card (PDF)
  • Documentation
  • Mistral Teases New Model

    Mistral AI teased an upcoming model on X. Community speculation suggests it may appear on OpenRouter; users await benchmark results.
  • Mistral on X
  • Qwen 3 Sparse Series Called "Underrated"

    Users recommend the Qwen 3 sparse series (e.g., a3b) for strong coding and reasoning, though some report the Qwen 32b model performs only mediocre.
  • OpenRouter Discord
  • Industry Partnerships & Investment

    Disney Invests $1B in OpenAI, Brings Characters to Sora

    Disney is investing $1 billion in a partnership with OpenAI to integrate its characters into the Sora AI video generator. The agreement includes a 3-year license with exclusivity in the first year, with content to appear on Disney+.
  • OpenAI announcement
  • CNBC coverage
  • DeepMind Opens First Automated Research Lab in the UK

    In partnership with the UK government, DeepMind is opening its first automated research lab focused on AI-driven scientific discovery (e.g., materials science, drug development), scheduled to launch in 2026.
  • DeepMind blog
  • Open-Source Tools & Technology

    Unsloth's New Packing Method: 3x Faster Training

    Unsloth's new packing technique trains 3x faster than the previous version and 10x faster than FA3, supports Qwen3-4B training on 3.9GB of VRAM, and resolves dependency conflicts with older NVIDIA drivers.
  • Unsloth docs
  • llama.cpp Adds Live Model Switching

    llama.cpp introduces a routing mode enabling dynamic model management (load, unload, and switch without restarts). A multi-process architecture isolates crashes for stability, with LRU caching and automatic discovery.
  • Hugging Face blog
  • Hugging Face WebGPU Local Voice Chat Demo

    A Hugging Face Space showcases real-time AI voice chat running entirely in the browser (STT, VAD, TTS, and LLM all processed locally), preserving user privacy.
  • Hugging Face Space
  • Benchmarks & Performance

    GPT-5.2 Beats Human Experts on GDPVal Tasks

    GPT-5.2 Thinking outperformed 70.9% of human experts on GDPVal tasks across 44 occupations, working 11x faster at about 1% of the cost — though human oversight is still recommended.
  • OpenAI GDPVal notes
  • SWE-Bench results
  • Community & Ecosystem

    Reddit Debates GPT-5.2 Performance vs. Hype

    Reddit users praised GPT-5.2's 100% competition math score but criticized the $168/M output token cost. Memes mocked its "AGI" claims — triggered by miscounting the letter R in "garlic."
  • Reddit thread
  • AGI meme
  • LMArena Community Tests GPT-5.2 Coding

    LMArena users report GPT-5.2 High generating buggy code in Code Arena despite high SWE-bench scores. It ranks second on the WebDev leaderboard, but users call it a "rushed release" with steep pricing.
  • LMArena leaderboard
  • Discord chat
  • Hardware & Infrastructure

    CUDA 13 Fixes Torch/vLLM Compatibility Issues

    Switching to CUDA 13 resolves compatibility issues between Torch and vLLM — both must use CUDA 13 builds, particularly relevant for AMD GPU users.
  • GPU MODE Discord
  • Hetzner Launches 96GB VRAM Server at €889

    Hetzner released a bare-metal server with 96GB of VRAM at €889/month, including generous free traffic — attractive for AI startups cutting training/inference costs.
  • Nous Research Discord
  • Research & Theory

    Diffusion Distillation Technique Yields Free Log Probabilities

    A new diffusion technique adds a predictive-divergence head and adjusts initial noise to obtain free log probabilities, improving image likelihood maximization.
  • arXiv paper
  • Sandwich Normalization for Long-Context Transformers

    Researchers discuss "sandwich normalization" for handling longer sequences in Transformers by normalizing activations; the paper details the method.
  • OpenReview paper
  • AI Ethics & Jailbreaks

    CIRIS Agent Tested for Jailbreak Resistance

    CIRIS Agent, designed as an ethical AI, invited users to bypass its filters. It refused to generate unethical content (e.g., meth synthesis instructions), though some users pushed its limits.
  • BASI Jailbreaking Discord
  • Grok Image Generation Sparks Censorship Debate

    Users debate Grok's image-generation censorship — some say restrictions are strict, others note skilled users can still produce deepfakes, with some outputs described as "unaligned garbage."
  • BASI Jailbreaking chat
  • Developer Tools & Platforms

    Cursor Debug Mode Gets Positive Feedback

    Cursor's new debug mode solves issues by adding test objects, with users reporting successful debugging. However, context rollback cannot restore state; users want backup features.
  • Cursor Discord
  • Perplexity Pro Users Hit Strict Rate Limits

    Perplexity Pro users report being rate-limited after just 5 Gemini 3 Pro queries. Suspected causes include server load or bugs; workarounds include disabling VPNs and clearing cache.
  • Perplexity Discord
  • Windsurf Ships New MCP Management UI

    Windsurf released versions 1.12.41 and 1.12.160 with improved stability and performance, a new MCP management UI, fixes for GitHub/GitLab MCP issues, and enhanced diff zones and Supercomplete.
  • Windsurf changelog
---

📌 Source: Easy AI Daily 🤖 Compiled by: AI assistant

Tags

#ai-news#gpt-5-2#openai#mistral#qwen#unsloth#llama-cpp#hugging-face

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169222