Easy AI Daily | December 2, 2025
A digest of AI industry news compiled from the Easy AI teaching project, covering model releases, company moves, research, tooling, community discussions, and hardware.
Model Releases and Updates
Mistral 3 Family (Large 3 and Ministral 3B/8B/14B)
Mistral AI released the Mistral 3 family, including the 675B MoE Mistral Large 3 (ranked #6 among open models) and Apache 2.0-licensed Ministral 3B/8B/14B. Ecosystem tools such as vLLM and llama.cpp already support the models; early evaluations show strong coding performance.- Mistral blog
- Arena leaderboard
- vLLM support
- Benchmarks
- API availability
- Nova 2.0 analysis
- Sonic 2.0
- Anthropic announcement
- Bun announcement
- Survey thread
- Follow-up
- Leak report
- GPT-5.1 podcast
- Test-time compute scaling: Large-scale research shows test-time compute strategies can improve complex reasoning without retraining; effectiveness depends on allocation strategy rather than raw compute. (Summary, Paper)
- OPPO FINDER benchmark: OPPO's FINDER benchmark (100 tasks) and DEFT taxonomy show agents fail at evidence integration, verification, and planning. (Overview)
- Neel Nanda on interpretability: Nanda argues for studying CoT in practical interpretability, pushing back on "interpretability has failed" narratives. (Clarification, Technical)
- Gradium raises $70M seed: Paris-based Gradium exited stealth with a $70M seed round, launching transcription/synthesis APIs supporting 5 European languages. (Announcement)
- LangSmith Agent Builder public beta: No-code agent builder supporting prompts, tools, triggers, MCP, and memory/summarization. (Launch)
- LlamaIndex LlamaAgents and LlamaSheets: Workflow templates and spreadsheet parsing, plus community office hours. (Recap)
- Hugging Face Skills: Universal agent context compatible with Cursor, Claude Code, and Gemini CLI, using Claude's skills specification. (Source)
- Perplexity open-sources BrowseSafe: BrowseSafe and BrowseSafe-Bench defend against prompt injection, outperforming safety classifiers. (Announcement)
- Reddit: /r/LocalLLaMA discusses Mistral 3's lineup and gaps, plus Mongolian GPU rentals (B300 at $5/hr with InfiniBand) compared to CoreWeave/Lambda. Non-technical subreddits discuss OpenAI's Code Red memo, possible ads in paid ChatGPT, the "dead internet" theory, and AI's impact on college education.
- Discord: Model release talk (Mistral 3, Arcee Trinity, Flux 2 Pro rankings); kernel optimization (PyTorch conv3D slowdown, CUDA syncwarp race conditions, NVIDIA nvfp4_gemm leaderboard); developer tooling issues (Manus.im auth instability, OpenRouter DeepSeek errors, Cursor subagent problems); security (RawChat GPT-4o jailbreaks, SEED Framework at 99.4% jailbreak resistance, Gemini 3 Pro jailbreak attempts); industry chatter (OpenAI Alert Red memo, 400GB VRAM rigs, Gradium funding).
- PyTorch issue
- LMArena leaderboard
- Flux leaderboard
- Arcee Trinity manifesto
- Mongolia GPU rentals: Fibo Cloud offers B300 Blackwell Ultra GPU rental at $5/hour with 3.2 Tb/s InfiniBand and pre-installed PyTorch/SLURM. (Landing page)
- 400GB VRAM rigs: Users built 6x RTX 3090 rigs (400GB VRAM) using MCIO adapters and PSU synchronization for running DeepSeek 3.2-class models.
- NVIDIA nvfp4_gemm competition: Users submitted kernels to NVIDIA's leaderboard, reducing latency with eval_better_bench.py; CPU queue bottlenecks discussed.
Apple Releases CLaRa-7B-Instruct
Apple published the CLaRa-7B-Instruct model on Hugging Face. (Source)Runway Previews Gen-4.5
Runway previewed Gen-4.5, improving cinematic realism, with early access opened. (Source)DeepSeek V3.2 Released
DeepSeek V3.2 (including the Speciale variant) delivers strong reasoning performance at low pricing; Fireworks already offers API access. It scores highly on LisanBench.Amazon Nova 2.0 Family
Amazon launched Nova 2.0 Pro (reasoning), Lite (speed), Omni (multimodal), and Sonic 2.0 (speech-to-speech). Nova 2.0 Pro hits 93% on τ²-Bench Telecom, and Sonic 2.0 ranks #2 in audio reasoning.Company News
Anthropic Acquires Bun Runtime
Anthropic acquired the MIT-licensed Bun JS/TS runtime to enhance Claude Code. The Bun team joins Anthropic; Claude Code reportedly reached a $1B run rate within six months.Claude for Nonprofits
Anthropic partnered with GivingTuesday to offer discounted plans, new integrations, and training for nonprofit organizations. (Source)Anthropic AI Work Impact Survey
A survey of 132 engineers and 200,000 Claude Code sessions shows engineers increasingly prioritize Claude for problem-solving, changing team dynamics.OpenAI "Garlic" Leak and GPT-5.1
The Information reports OpenAI's "Garlic" model outperforms GPT-4.5 on coding and reasoning. OpenAI also released a GPT-5.1 Instant podcast covering reasoning and personality control.Research and Benchmarks
Agents and Tooling
Community Highlights (Reddit and Discord)
Hardware and Infrastructure
*Source: Easy AI teaching project*