Easy AI Daily News Digest – November 4, 2025
Model Updates and Evaluation
- Minimax M2 tops WebDev Leaderboard – Ranked 4th overall and 1st among open-source models, excelling in coding, reasoning, and agent tasks at low cost and high speed. WebDev Leaderboard
- Qwen hallucination and instruction-following – Evaluations show Qwen models hallucinate on rare facts at twice the rate of Llama, but Qwen3 8B performs best at instruction-following continuation, beating the larger GPT OSS 20B. LLM-Propensity-Evals
- DeepMind releases IMO-Bench – A math reasoning benchmark suite: IMO-AnswerBench (answers), IMO-ProofBench (proofs), and IMO-GradingBench (grading). Gemini DeepThink scored 89.0% on the basic set and 65.7% on the advanced set. Twitter
- Qwen3-VL update – Integrated into Jan with thinking-capability API. Users report that conversion frameworks (Ollama vs MLX) affect accuracy; testing your target stack is recommended. Alibaba Qwen Twitter
- MCP solidifies as the core agent tool protocol – LitServe converts models/agents into MCP servers in 10 lines of code; Reka released a search/fact-checking MCP server for VS Code; Anthropic shared code-execution guidance. LitServe | Reka | Anthropic
- Fenic integrates OpenRouter – Fenic's dataframe API supports multi-provider AI workflows with scalable batch processing and no code changes; suitable for LLM ETL and context engineering. GitHub
- Windsurf launches Codemaps – Powered by SWE-1.5 and Sonnet 4.5, generating interactive visual maps of codebases to reduce confusion and boost productivity. Codemaps
- ComfyUI + LM Studio – Community discussion on connecting ComfyUI with LM Studio for a local Gemini Storybook alternative, using five text prompts and sampler-split story image generation. Discord
- llama.cpp releases official WebUI – Supports 150k+ GGUF models, PDF/image ingestion, conversation branching, and constrained JSON generation; praised as a milestone for local AI UX. GitHub
- Tinybox Pro v2 unveiled – George Hotz's new workstation: 8x RTX 5090 GPUs, 5U rackmount, priced at $50,000, shipping in 4–12 weeks; debate over cost vs cloud computing. Tinycorp Shop
- GPU shortage drives price increases – New cloud providers charge around $2/GPU-hour while hyperscalers reach $7; community discusses value and recommends local AMD cards. Discord
- MLX-Swift adds continuous batching – Enables multi-stream local inference, automatically upgrading single-request streams to batches for higher throughput. Twitter
- Google Project Suncatcher: TPUs in space – Google prototyped an orbital ML compute system; Trillium TPUs passed particle accelerator radiation testing, with two prototype satellites planned for 2027 in partnership with Planet. Sundar Pichai Twitter
- China subsidizes AI data center electricity – 50% electricity subsidies announced; Huawei plans gigawatt-scale SuperPoDs by 2027, focused on DeepSeek models. Twitter
- Epoch launches Frontier Data Centers Hub – Open-source tracker of 1GW+ AI data centers using satellite imagery and public records, freely available. Twitter
- Deutsche Telekom and NVIDIA build Munich data center – $1.1B investment equipping 10k GPUs (DGX B200 + RTX Pro) to expand European AI compute. Twitter
- Vidu Q2 ranks on Artificial Analysis – 8th place; supports multi-reference image conditioning, generating 8-second 1080p videos, API pricing between Hailuo 02 Pro and Veo 3.1. Twitter
- MotionStream: real-time interactive video generation – 29 FPS with 0.4s latency on an H100, supporting drag-and-gesture control for long videos. Twitter
- Generalist AI releases GEN-0 – A 10B+ parameter robotics foundation model trained on 270k+ hours of dexterous manipulation data, emphasizing physical common sense (grasping, stability, placement). Twitter
- Coca-Cola's 2025 Christmas ad is AI-generated – Reduced human involvement again this year; the company calls AI an irreversible trend, with improved ad quality. Twitter
- Fox News airs AI-generated protest footage by mistake – The network aired AI-generated footage of a food-stamp protest, later issuing a correction; raising concerns about AI content verification in media. Reddit
- Reddit debates the Qwen ecosystem – Users compare Qwen with GPT-OSS; one reports Qwen outperforming GPT-OSS-20B on an RTX 3060. Reddit
- Getty Images largely loses UK AI image lawsuit – The landmark ruling sparks further discussion of AI content copyright. Reuters
- Teachers worry about student AI dependence – Educators report students relying on AI for homework, raising concerns about loss of critical thinking and motivation, and how education systems should adapt. Reddit
- Cache-to-Cache: direct LLM semantic communication – A research paradigm where LLMs share semantic information directly via caches, bypassing text, improving accuracy and latency; auditability concerns raised. Reddit
- Context engineering blueprint released – A 41-page document covering agents, query enhancement, retrieval, prompting, memory, and tools, emphasizing the shift from prompt engineering to context engineering. Twitter
Agent and Tool Ecosystem
Local Inference and Hardware
AI Industry News
Multimodal and Robotics
Community and Media
Education and Research
*Source: Easy AI teaching project.*