English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily News Digest | December 2, 2025

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily for December 2, 2025 covers a busy day of AI model releases and industry news. Mistral AI launched the Mistral 3 family including the 675B MoE Large 3 and Apache 2.0 Ministral 3B/8B/14B models. Apple released CLaRa-7B-Instruct on Hugging Face, Runway previewed Gen-4.5, DeepSeek shipped V3.2 with strong reasoning at low pricing, and Amazon unveiled the Nova 2.0 family with Sonic 2.0 speech-to-speech. Anthropic acquired the MIT-licensed Bun JavaScript runtime to boost Claude Code, launched Claude for Nonprofits, and published engineering survey findings. OpenAI reportedly developed a 'Garlic' model outperforming GPT-4.5. Research highlights include test-time compute scaling studies and OPPO's FINDER deep-research agent benchmark. Tooling news spans LangSmith Agent Builder beta, LlamaIndex's LlamaAgents, Hugging Face Skills, and Perplexity's open-source BrowseSafe prompt-injection defense. Hardware topics cover $5/hr B300 GPU rentals in Mongolia and 400GB VRAM rigs.

Easy AI Daily | December 2, 2025

A digest of AI industry news compiled from the Easy AI teaching project, covering model releases, company moves, research, tooling, community discussions, and hardware.

Model Releases and Updates

Mistral 3 Family (Large 3 and Ministral 3B/8B/14B)

Mistral AI released the Mistral 3 family, including the 675B MoE Mistral Large 3 (ranked #6 among open models) and Apache 2.0-licensed Ministral 3B/8B/14B. Ecosystem tools such as vLLM and llama.cpp already support the models; early evaluations show strong coding performance.
  • Mistral blog
  • Arena leaderboard
  • vLLM support
  • Apple Releases CLaRa-7B-Instruct

    Apple published the CLaRa-7B-Instruct model on Hugging Face. (Source)

    Runway Previews Gen-4.5

    Runway previewed Gen-4.5, improving cinematic realism, with early access opened. (Source)

    DeepSeek V3.2 Released

    DeepSeek V3.2 (including the Speciale variant) delivers strong reasoning performance at low pricing; Fireworks already offers API access. It scores highly on LisanBench.
  • Benchmarks
  • API availability
  • Amazon Nova 2.0 Family

    Amazon launched Nova 2.0 Pro (reasoning), Lite (speed), Omni (multimodal), and Sonic 2.0 (speech-to-speech). Nova 2.0 Pro hits 93% on τ²-Bench Telecom, and Sonic 2.0 ranks #2 in audio reasoning.
  • Nova 2.0 analysis
  • Sonic 2.0
  • Company News

    Anthropic Acquires Bun Runtime

    Anthropic acquired the MIT-licensed Bun JS/TS runtime to enhance Claude Code. The Bun team joins Anthropic; Claude Code reportedly reached a $1B run rate within six months.
  • Anthropic announcement
  • Bun announcement
  • Claude for Nonprofits

    Anthropic partnered with GivingTuesday to offer discounted plans, new integrations, and training for nonprofit organizations. (Source)

    Anthropic AI Work Impact Survey

    A survey of 132 engineers and 200,000 Claude Code sessions shows engineers increasingly prioritize Claude for problem-solving, changing team dynamics.
  • Survey thread
  • Follow-up
  • OpenAI "Garlic" Leak and GPT-5.1

    The Information reports OpenAI's "Garlic" model outperforms GPT-4.5 on coding and reasoning. OpenAI also released a GPT-5.1 Instant podcast covering reasoning and personality control.
  • Leak report
  • GPT-5.1 podcast
  • Research and Benchmarks

  • Test-time compute scaling: Large-scale research shows test-time compute strategies can improve complex reasoning without retraining; effectiveness depends on allocation strategy rather than raw compute. (Summary, Paper)
  • OPPO FINDER benchmark: OPPO's FINDER benchmark (100 tasks) and DEFT taxonomy show agents fail at evidence integration, verification, and planning. (Overview)
  • Neel Nanda on interpretability: Nanda argues for studying CoT in practical interpretability, pushing back on "interpretability has failed" narratives. (Clarification, Technical)
  • Gradium raises $70M seed: Paris-based Gradium exited stealth with a $70M seed round, launching transcription/synthesis APIs supporting 5 European languages. (Announcement)
  • Agents and Tooling

  • LangSmith Agent Builder public beta: No-code agent builder supporting prompts, tools, triggers, MCP, and memory/summarization. (Launch)
  • LlamaIndex LlamaAgents and LlamaSheets: Workflow templates and spreadsheet parsing, plus community office hours. (Recap)
  • Hugging Face Skills: Universal agent context compatible with Cursor, Claude Code, and Gemini CLI, using Claude's skills specification. (Source)
  • Perplexity open-sources BrowseSafe: BrowseSafe and BrowseSafe-Bench defend against prompt injection, outperforming safety classifiers. (Announcement)
  • Community Highlights (Reddit and Discord)

  • Reddit: /r/LocalLLaMA discusses Mistral 3's lineup and gaps, plus Mongolian GPU rentals (B300 at $5/hr with InfiniBand) compared to CoreWeave/Lambda. Non-technical subreddits discuss OpenAI's Code Red memo, possible ads in paid ChatGPT, the "dead internet" theory, and AI's impact on college education.
  • Discord: Model release talk (Mistral 3, Arcee Trinity, Flux 2 Pro rankings); kernel optimization (PyTorch conv3D slowdown, CUDA syncwarp race conditions, NVIDIA nvfp4_gemm leaderboard); developer tooling issues (Manus.im auth instability, OpenRouter DeepSeek errors, Cursor subagent problems); security (RawChat GPT-4o jailbreaks, SEED Framework at 99.4% jailbreak resistance, Gemini 3 Pro jailbreak attempts); industry chatter (OpenAI Alert Red memo, 400GB VRAM rigs, Gradium funding).
  • PyTorch issue
  • LMArena leaderboard
  • Flux leaderboard
  • Arcee Trinity manifesto
  • Hardware and Infrastructure

  • Mongolia GPU rentals: Fibo Cloud offers B300 Blackwell Ultra GPU rental at $5/hour with 3.2 Tb/s InfiniBand and pre-installed PyTorch/SLURM. (Landing page)
  • 400GB VRAM rigs: Users built 6x RTX 3090 rigs (400GB VRAM) using MCIO adapters and PSU synchronization for running DeepSeek 3.2-class models.
  • NVIDIA nvfp4_gemm competition: Users submitted kernels to NVIDIA's leaderboard, reducing latency with eval_better_bench.py; CPU queue bottlenecks discussed.
---

*Source: Easy AI teaching project*

Tags

#ai-news#mistral-3#anthropic#openai#deepseek#amazon-nova#claude-code#daily-digest

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169131