Easy AI Daily News Digest — December 2, 2025
A roundup of AI industry developments curated by the Easy AI educational project.
Model Releases and Updates
Mistral 3 Model Family (Large 3, Ministral 3B/8B/14B)
Mistral AI released the Mistral 3 family, including Mistral Large 3 (a 675B MoE model ranked #6 among open models) and the Apache 2.0-licensed Ministral 3B/8B/14B. Ecosystem tools like vLLM and llama.cpp already support them; early benchmarks show strong coding performance.- Mistral blog | Arena leaderboard | vLLM support
- Twitter announcement
- Twitter announcement
- Benchmarks | API availability
- Nova 2.0 analysis | Sonic 2.0
- Anthropic announcement | Bun announcement
- Twitter announcement
- Survey thread | Follow-up
- Leak report | GPT-5.1 podcast
- Test-time compute scaling: Large-scale studies show test-time compute strategies improve complex reasoning without retraining; effectiveness depends on allocation strategy rather than raw compute. (Summary | Paper)
- OPPO FINDER benchmark: OPPO's FINDER benchmark (100 tasks) and DEFT taxonomy reveal agent failures in evidence integration, verification, and planning. (Overview)
- Neel Nanda on interpretability: Nanda argues for studying chain-of-thought in practical interpretability, countering "interpretability failed" hype. (Clarification | Technical)
- Gradium raises $70M seed: Paris-based Gradium exited stealth with a $70M seed round, launching transcription/synthesis APIs supporting 5 European languages. (Announcement | Founder thread)
- LangSmith Agent Builder (beta): No-code agent builder supporting prompts, tools, triggers, MCP, and memory/summarization. (Release | Video)
- LlamaIndex LlamaAgents and LlamaSheets: Workflow templates and spreadsheet parsing, plus community office hours. (Recap | Invite)
- Hugging Face Skills: Universal agent context compatible with Cursor, Claude Code, and Gemini CLI, using Claude's skills specification. (Twitter)
- Perplexity open-sources BrowseSafe: BrowseSafe and BrowseSafe-Bench defend against prompt injection, outperforming safety classifiers. (Announcement)
- /r/LocalLLaMA discussions on Mistral 3's open models, Large 3's 675B MoE, and lineup gaps. (Thread 1 | Thread 2)
- Mongolian GPU rentals (B300, $5/hr, InfiniBand) vs. CoreWeave/Lambda. (Thread)
- Non-technical subs on OpenAI's Code Red memo, GPT-5.1, and possible ads in paid tiers. (Thread 1 | Thread 2)
- Debates on "dead internet," the "adpocalypse," and AI's impact on college education. (Dead internet | Adpocalypse | College)
- Model releases: Mistral 3, Arcee Trinity, Flux 2 Pro rankings. (LMArena | Arcee blog | Flux leaderboard)
- Kernel optimization: PyTorch conv3D slowdown, CUDA syncwarp race conditions, NVIDIA nvfp4_gemm leaderboard. (PyTorch issue)
- Developer tools: Manus.im instability/auth issues, OpenRouter DeepSeek errors, Cursor sub-agent and DeepSeek integration issues.
- Safety: RawChat stealth mode (GPT4o jailbreaks), SEED Framework (99.4% jailbreak resistance), Gemini 3 Pro jailbreak attempts. (UltraBr3aks)
- Industry: OpenAI Alert Red memo, 400GB VRAM rigs, Gradium's $70M round.
- Mongolian GPU rentals: Fibo Cloud offers B300 Blackwell Ultra rentals in Mongolia at $5/hour with 3.2 Tb/s InfiniBand and preinstalled PyTorch/SLURM. (Landing page)
- 400GB VRAM rigs: Users built 6x 3090 rigs using MCIO adapters and synchronized PSUs for models like DeepSeek 3.2. (Rig image | PSU sync)
- NVIDIA nvfp4_gemm competition: Community submissions to NVIDIA's kernel leaderboard reduced latency via eval_better_bench.py; CPU queue bottlenecks discussed. (Leaderboard)
Apple Releases CLaRa-7B-Instruct
Apple published the CLaRa-7B-Instruct model on Hugging Face.Runway Previews Gen-4.5
Runway previewed Gen-4.5 with improved cinematic realism; early access is open.DeepSeek V3.2 Released
DeepSeek V3.2 (including Speciale) shows strong reasoning performance at low prices; Fireworks offers API access. High scores on the LisanBench benchmark.Amazon Nova 2.0 Family
Amazon launched Nova 2.0 Pro (reasoning), Lite (speed), Omni (multimodal), and Sonic 2.0 (speech-to-speech). Pro scored 93% on τ²-Bench Telecom; Sonic 2.0 ranks #2 in audio reasoning.Business News
Anthropic Acquires Bun Runtime
Anthropic acquired the MIT-licensed Bun JS/TS runtime to strengthen Claude Code. The Bun team joins Anthropic; Claude Code reportedly reached a $1 billion run rate within six months.Claude for Nonprofits
Anthropic partnered with GivingTuesday to offer discounted plans, new integrations, and training for nonprofit organizations.Anthropic AI Work Impact Survey
A survey of 132 engineers and 200,000 Claude Code sessions shows engineers prioritize using Claude to solve problems, reshaping team dynamics.OpenAI "Garlic" Leak and GPT-5.1
The Information reports OpenAI's "Garlic" model outperforms GPT-4.5 on coding and reasoning. OpenAI also released a GPT-5.1 Instant podcast covering reasoning and personality control.Research and Benchmarks
Agents and Tooling
Community Highlights
Discord
Hardware and Infrastructure
*Source: Easy AI educational project*