Easy AI Daily | October 30, 2025
A roundup of the day's AI industry news, model releases, and community discussions, aggregated by the Easy AI education project.
Key Model Releases & Research
- Kimi Linear (KDA) by Moonshot AI: A hybrid architecture combining Kimi Delta Attention (KDA) with MLA. Delivers 75% KV cache reduction and 6x decoding throughput improvement, with strong long-context and RL performance. Open-source CUDA kernels and vLLM integration available. Model on Hugging Face
- MiniMax M2: Switched from hybrid to full attention architecture. Supports 200k context at 100 TPS, with free trial access. Community discussion notes full attention may be superior for multi-hop reasoning compared to linear attention approaches.
- Looped LLMs (ByteDance / Princeton / Mila): Recurrent latent reasoning allows 1.4B/2.6B parameter models to match standard 4B/8B models with better data efficiency; potentially combinable with MoE scaling.
- Anthropic introspection research: A paper claims LLMs can detect modifications to their internal activations; community debates whether this reflects genuine introspection or pattern recognition. Paper
- OpenAI Aardvark (GPT-5): An agentic security researcher in private beta that reads code, writes tests, and proposes vulnerability patches.
- Cognition Computer Use: Public beta enabling Devin to operate desktop and mobile tools; also launched SWE-1.5 at 950 tok/s—6x faster than Haiku and 13x faster than Sonnet, running on Cerebras hardware.
- Voyage voyage-3-large: Tops the Hugging Face RTEB leaderboard, supports INT8 quantization to reduce vector database costs; strong in finance, legal, and medical retrieval.
- Cartesia Sonic-3 TTS: SSM-based architecture achieving under 250ms real-time latency across 42 languages (including 9 Indian languages).
- Perplexity Patents & Discover: New patent research tool plus Discover and finance features (e.g., politician stock holdings).
- Hugging Face Smol Training Playbook: A 200+ page guide covering pretraining, fine-tuning, and infrastructure, emphasizing ablation studies and practical strategies. Playbook
- HKUST Toolathlon benchmark: 32 applications and 600+ tools; Claude Sonnet 4.5 achieved only 38.6% correctness, highlighting tool-use capability gaps.
- Qwen 3 VL merged into llama.cpp (PR #16780); currently MLX (Mac) only, with Q6 quantization reported to work well.
- Claude Skills community roundup: 10 workflow-changing skills including Rube MCP (500+ app integrations), Superpowers, and Document Suite. Skills repo
- George R.R. Martin v. OpenAI: A judge allowed the copyright lawsuit to proceed, with authors alleging ChatGPT generates content similar to their works.
- Udio removed wav downloads for subscribers, sparking user backlash and calls for open-source AI music alternatives.
- LMArena: ReCaptcha complaints; hailuo-2.3-fast added to the video leaderboard.
- Cursor: Debates over Composer 1 speed/accuracy, pricing, and cache usage; Cursor 2.0 bugs reported.
- Unsloth: RTX 8000 VRAM discussion; Qwen3 4B GRPO fine-tuning OOM fixes via 4-bit loading and batch size tuning.
- OpenRouter: Launched Sonar Pro Search with Perplexity; Sora 2 bias complaints.
- Modular Mojo: MAX performance rivaling NVIDIA and exceeding JAX on some ML tasks; early scikit-learn alternative ("Scijo") benchmarks show speed gains.
- GPU MODE: CUDA scan algorithm optimization vs CUB; FP8 quantization with TorchAO and GemLite benchmarks.
- Latent Space: ScaleAI's RLI benchmark shows Manus agent at only 2.5% automation rate. Leaderboard
- Nous Research: Hack The Box hosting an MCP-only CTF focused on AI security, free on November 20. Signup
- MCP Contributors: Model Context Protocol RFC delayed pending tangible implementations.
Product Launches
Open Source & Local AI
Legal & Industry
Community Highlights (Discord)
*Source: Easy AI education project.*