Easy AI Daily Digest | December 2, 2025
📅 AI Industry News — December 2, 2025
Model Releases and Updates
Mistral 3 Model Family (Large 3 + Ministral 3B/8B/14B)
Mistral AI released the Mistral 3 family, including Mistral Large 3 (a 675B MoE model, ranked #6 among open models) and Ministral 3B/8B/14B under the Apache 2.0 license. Ecosystem tools such as vLLM and llama.cpp already support them; early evaluations show strong coding performance.Links: Mistral blog | Arena leaderboard | vLLM support
Apple Releases CLaRa-7B-Instruct
Apple published the CLaRa-7B-Instruct model on Hugging Face.Link: Twitter
Runway Previews Gen-4.5
Runway previewed the Gen-4.5 model with improved cinematic realism; early access is open.Link: Twitter
DeepSeek V3.2 Released
DeepSeek V3.2 (including the Speciale variant) shows strong reasoning performance at low pricing; Fireworks already offers an API. High scores on LisanBench.Links: Benchmarks | API availability
Amazon Nova 2.0 Family
Amazon launched Nova 2.0 Pro (reasoning), Lite (speed), Omni (multimodal), and Sonic 2.0 (speech-to-speech). Pro scores 93% on τ²-Bench Telecom; Sonic 2.0 ranks #2 in audio reasoning.Links: Nova 2.0 analysis | Sonic 2.0
Business Moves
Anthropic Acquires Bun Runtime
Anthropic acquired the MIT-licensed Bun JS/TS runtime to strengthen Claude Code. The Bun team joined Anthropic; Claude Code reportedly reached a $1B run rate within 6 months.Links: Anthropic announcement | Bun announcement
Claude for Nonprofits
Anthropic partnered with GivingTuesday to offer discounted plans, new integrations, and training for nonprofits.Link: Twitter
Anthropic Survey on AI's Impact on Work
A survey of 132 engineers and 200k Claude Code sessions shows engineers prioritize using Claude for problem-solving, changing team dynamics.Links: Survey thread | Follow-up
OpenAI "Garlic" Leak and GPT-5.1
The Information reports OpenAI's "Garlic" model outperforms GPT-4.5 on coding and reasoning. OpenAI also released a GPT-5.1 Instant podcast covering reasoning and personality control.Links: Leak report | GPT-5.1 podcast
Research and Benchmarks
- Test-time compute scaling: A large-scale study shows test-time compute strategies can improve complex reasoning without retraining; effectiveness depends on allocation strategy rather than raw compute. Summary | Paper
- OPPO FINDER benchmark: OPPO's FINDER benchmark (100 tasks) and DEFT taxonomy show agents fail at evidence integration, verification, and planning. Overview
- Neel Nanda on interpretability: Nanda argues for studying CoT in practical interpretability, pushing back on "interpretability is failing" hype. Clarification | Technical
- Gradium raises $70M seed: Paris-based Gradium exited stealth with a $70M seed round, launching transcription/synthesis APIs supporting 5 European languages. Announcement | Founder thread
- LangSmith Agent Builder (beta): No-code agent builder supporting prompts, tools, triggers, MCP, and memory/summarization. Release | Video
- LlamaIndex: Launched LlamaAgents (workflow templates) and LlamaSheets (spreadsheet parsing), plus community office hours. Recap | Invite
- Hugging Face Skills: Universal agent context compatible with Cursor, Claude Code, and Gemini CLI, using Claude's skills spec. Twitter
- Perplexity BrowseSafe: Open-sourced BrowseSafe and BrowseSafe-Bench to defend against prompt injection, outperforming safety classifiers. Announcement | Results
- /r/LocalLLaMA on Mistral 3: Discussion of the open-source 3B/8B/14B models, Large 3's 675B MoE, and gaps in the model lineup. Thread 1 | Thread 2
- /r/LocalLLaMA on Mongolia GPU rentals: B300 GPUs at $5/hr with InfiniBand, compared to CoreWeave/Lambda. Thread
- OpenAI Code Red: Non-technical subreddits discuss OpenAI's Code Red memo, GPT-5.1, and possible ads in paid tiers. Thread 1 | Thread 2
- Internet challenges: Discussions of the "dead internet" (AI-generated content), the "adpocalypse" (ChatGPT adding ads), and college education gaps. Dead internet | Adpocalypse | College
- Model releases: Mistral 3 (Large 3, Ministral), Arcee Trinity, Flux 2 Pro rankings. LMArena | Arcee blog | Flux leaderboard
- Kernel optimization: PyTorch conv3D slowdown, CUDA syncwarp race conditions, NVIDIA nvfp4_gemm leaderboard. PyTorch issue | NVIDIA leaderboard
- Developer tools: Manus.im instability and auth issues, OpenRouter DeepSeek errors, Cursor sub-agent and DeepSeek integration issues.
- Security: RawChat stealth mode (GPT-4o jailbreak), SEED Framework (99.4% jailbreak resistance), Gemini 3 Pro jailbreak attempts. UltraBr3aks
- Industry news: OpenAI Alert Red memo, 400GB VRAM rigs, Gradium's $70M funding.
- Mongolia GPU rental market: Fibo Cloud offers B300 Blackwell Ultra GPU rentals at $5/hr with 3.2 Tb/s InfiniBand and preinstalled PyTorch/SLURM. Landing page
- 400GB VRAM rigs: Users build 6x RTX 3090 rigs with MCIO adapters and synchronized older PSUs, targeting models like DeepSeek 3.2. Rig image
- NVIDIA nvfp4_gemm competition: Users submit nvfp4_gemm kernels to NVIDIA's leaderboard; eval_better_bench.py reduces latency, with discussion of CPU queue bottlenecks. Leaderboard
Agents and Tooling
Community and Platforms — Reddit
Community and Platforms — Discord
Hardware and Infrastructure
*Source: Easy AI Educational Project*