English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily Digest | December 5, 2025: Gemini 3 Deep Think, GPT-5.1-Codex Max, Anthropic Acquires Bun

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily digest for December 5, 2025 covering major AI industry developments. Google released Gemini 3 Deep Think mode for AI Ultra subscribers, scoring 45.1% on ARC-AGI-2 versus GPT-5.1's 17.6%. OpenAI launched GPT-5.1-Codex Max for the Responses API with VS Code and Cursor integration. Microsoft open-sourced VibeVoice-Realtime-0.5B, a lightweight realtime text-to-speech model supporting English and Chinese. Nous Research released Hermes 4.3, and Mistral Large 3 topped lmarena's open-source coding leaderboard. Anthropic acquired Bun as Claude Code revenue reached $1 billion. Harvey raised $160 million at an $8 billion valuation, while Antithesis secured $105 million led by Jane Street. Research highlights include Google's Titans long-context memory architecture, TorchAO MoE quantization support, and a fast ODE solver generating 4K images in 8 steps. OpenRouter's State of AI report analyzed 100 trillion tokens, finding 50% of open-model usage goes to roleplay and Claude handles 60% of coding workloads.

Easy AI Daily Digest | December 5, 2025

A roundup of AI industry news for December 5, 2025, from the Easy AI educational project.

Model Releases & Updates

Google Releases Gemini 3 Deep Think Mode

Available to Google AI Ultra subscribers, Deep Think improves complex reasoning using parallel thinking. It scores 45.1% on ARC-AGI-2, surpassing GPT-5.1's 17.6%, and supports math and science tasks.

Links: Google AI announcement | Google DeepMind details

OpenAI Launches GPT-5.1-Codex Max

Built for the Responses API and integrated into the Codex agent harness, with support for IDEs including VS Code and Cursor to improve code generation.

Links: OpenAI Devs announcement | Cursor integration

Microsoft Releases VibeVoice-Realtime-0.5B

A lightweight realtime text-to-speech model supporting English and Chinese, open-sourced on Hugging Face.

Links: Hugging Face model page | Twitter announcement

Nous Research Releases Hermes 4.3

Based on ByteDance Seed 36B, with performance approaching Hermes 4 70B, trained on the Psyche network and supporting MoE.

Link: Nous Research blog

Mistral Large 3 Tops Open-Source Coding Leaderboard

Ranked #1 on lmarena, available via Ollama Cloud, with community confirmation of its coding capabilities.

Links: Mistral AI announcement | Ollama support

Research & Technical Progress

  • Google's Titans long-context memory architecture: Combines RNN efficiency with Transformer performance, supporting 2M+ tokens; early results shown at NeurIPS. Link
  • TorchAO MoE quantization: New MoEQuantConfig enables quantization of mixture-of-experts models for more efficient inference. PyTorch PR
  • VATTENTION paper: First sparse attention mechanism with (ϵ, δ) guarantees, improving long-text processing. arXiv
  • STRAW: Sample-tuned rank-augmented weights that mimic neuromodulation, dynamically adjusting model weights for task adaptability. Substack
  • Fast ODE solver for diffusion models: Generates 4K images in 8 steps with quality comparable to 30-step DPM++2M SDE; open-sourced on Hugging Face. Space | arXiv
  • Industry Moves & Funding

  • Anthropic acquires Bun as Claude Code revenue reaches the $1 billion milestone. Announcement
  • Perplexity gains investment from Cristiano Ronaldo, positioned as "sparking global curiosity." Tweet
  • Harvey raises $160M Series F at an $8 billion valuation, serving 700+ law firms with legal AI. Tweet
  • Antithesis raises $105M led by Jane Street, focused on deterministic simulation testing of AI-generated code. Tweet
  • Community Discussions

  • GPT-5.1 beats Gemini 3 at finding code bugs, per OpenAI Discord user feedback. Discussion
  • Z-Image still filters sensitive content despite claims of no censorship, showing "maybe not safe" for gore/nudity prompts. Reddit
  • Reddit debates AI's impact on tech jobs, with users arguing AI will change roles rather than replace them. Reddit
  • LocalLlama tests VibeVoice-Realtime, noting English/Chinese support and some concerns about Mandarin accent quality. Reddit
  • Reddit discusses Gemini 3 Deep Think benchmarks, comparing its 45.1% ARC-AGI-2 score against GPT-5.1's 17.6%. Reddit
  • Tools & Platform Updates

  • OpenRouter releases State of AI report: Analysis of 100 trillion tokens shows 50% of open-model usage goes to roleplay, 50% of paid-model usage to coding, with Claude handling 60% of coding workloads. Report
  • Windsurf integrates GPT-5.1-Codex Max: Free trial for paid users with Low/Medium/High reasoning levels. Announcement
  • mcp-apps-sdk open-sourced by General Intelligence Labs, enabling ChatGPT apps to be embedded in other platforms. GitHub
  • tinygrad fixes train_step function: PR addresses unused input tensors, improving training efficiency. PR
  • DSPy community proposes native Claude Code integration, leveraging its Read/Write/Terminal tools. Discord
  • Benchmarks & Performance

  • Gemini 3 Deep Think scores 45.1% on ARC-AGI-2, a 2.5x improvement over GPT-5.1's 17.6%, excelling at complex reasoning. Reddit
  • Mistral Large 3 ranked #1 in lmarena coding among open-source models. Tweet
  • DeepSeek V3.2 lmarena performance: Baseten published serving metrics — 0.22s TTFT, 191 tps — with strong rankings in math, law, and science. Tweet
  • GPT-5.1-Codex Max code generation: Integrated into Cursor and other IDEs; users report improved code quality and efficiency. Tweet
---

*Source: Easy AI educational project.*

Tags

#ai-news#gemini-3#gpt-5-1#anthropic#openai#mistral#ai-benchmarks#daily-digest

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169129