Easy AI Daily News | December 6, 2025
Model & Inference Infrastructure
vLLM 0.12.0 Released with DeepSeek Optimizations
- vLLM 0.12.0 introduces the experimental GPU Model Runner V2 and Prefill Context Parallel support.
- Optimized support for DeepSeek-V3.2's "thinking" mode, including tokenizer and tool-call parsers; also adds EAGLE decoding and NVFP4 quantization.
- Links: vLLM release | DeepSeek support details
- cuTile is a Python compiler library targeting TileIR, shipped with CUDA 13.1; the programming guide was rewritten.
- PTX 9.1 adds SIMD conversions and async "sharp+tma" operations. cuTile does not yet support mxfp/nvfp, but fp4 is planned.
- Links: cuTile repo | CUDA 13.1 docs
- Adds
AutoModelForMultimodalLMand an any-to-any pipeline supporting 2+ inputs/outputs (e.g., Gemma3n multimodal-to-text, Qwen3-Omni text + audio). - Link: release notes
- LangChain: new content moderation middleware (screens inputs, outputs, and tool results) and cost tracking with custom tool/API costs. DeepAgents CLI scored 42.7% on Terminal Bench 2.0, comparable to Claude Code. (moderation | cost tracking)
- Together AI + Meta: production-grade TorchForge RL support via the Together platform for long-horizon agentic workflows. (announcement)
- SonarSource: SonarQube MCP server brings enterprise static analysis into Claude Code/Cursor for more accurate AI code generation. (release)
- Kimi CLI: integrates with JetBrains IDEs via ACP. (details)
- Kling Video 2.6: native synchronized audio (speech, sound effects, ambient sound); new "Element/Subject Library" for Kling O1 enables persistent subject memory and consistency. (release | audio)
- Runway Gen 4.5 "Whisper Thunder": finer aesthetic control for world-building. (release)
- Alibaba Cloud Qwen3-TTS: 49+ voices, 10 languages and dialects, natural prosody; real-time and offline APIs; demos on HF/ModelScope. (release)
- Google Gemini 3 Pro: complex document derendering to HTML/LaTeX, screen understanding, spatial trajectory generation (robotics/XR), and a "thinking" mode for high-FPS video analysis. (details)
- FLUX.2 [dev] (Black Forest Labs): #1 open text-to-image on Artificial Analysis Image Arena, #2 in editing. FLUX.2 [klein] uses Apache-2.0 for commercial use. (analysis)
- Meituan LongCat-Image / LongCat-Image-Edit: image editing model under Apache-2.0, with demo. (release)
- MixtureVitae: licensed pretraining dataset targeting math/code, avoiding Books2 copyright risk. (details)
- Intel SignRoundV2: advances in extreme low-bit PTQ (e.g., 4-bit) for LLMs, improving quantization accuracy. (details)
- OpenRouter + a16z report: analysis of 100 trillion tokens shows reasoning models exceed 50% of usage, heavy Chinese closed-model traffic, and coding as a key use case. (report)
- NeurIPS 2025: Yejin Choi's keynote covered reasoning work including EPO; Sakana AI presented the "Continuous Thought Machine" (Neural ODE-based test-time compute scaling). (Sakana AI)
- OpenAI Residency applications open for engineers with basic ML experience; Google Gemini 3 Vibe Coding Hackathon offers $500K in API credits. (Residency | Hackathon)
- Trending this week: Google Gemini hackathon, Amanda Askell's AI ethics AMA, Qwen3-TTS launch, OpenAI Residency, and Cloudflare outage affecting tools like Claude. (AMA)
- Basketball AI: community system using RF-DETR for player/jersey detection, SAM2 tracking, SmolVLM2 jersey-number recognition, plus SigLIP, UMAP, and K-Means for team clustering, with trajectory correction and shot detection. (discussion)
- Anthropic survey: of 1,250 professionals, 86% say AI boosts productivity, but 69% feel stigma about using it. (study | discussion)
- Image generation: SteadyDancer vs. Wan2.2 Animate comparison (SteadyDancer maintains 100% identity match); Detail Daemon + ZIT combo for high-quality fantasy art. (SteadyDancer | Detail Daemon)
- CUDA Tile & GPU programming: debates on NVIDIA cuTile, PTX 9.1 features, and CUDA-L2 surpassing cuBLAS via RL optimization.
- Gemini 3 vs Opus 4.5: community benchmarks found Gemini more expensive with lower SWE-Bench scores; GPT-5.1-High performed better in bug-finding tests. (comparison sheet)
- Model-agnostic tool orchestrator: HuggingFace user released a production tool orchestrator based on Anthropic's Programmatic Tool Calling, letting any LLM write Rhai scripts to orchestrate tools, claiming 97-99% token reduction. (repo)
- MCP token usage analysis: tokenization is model-dependent — OpenAI uses tiktoken, Claude uses the count_tokens API; Claude 3 no longer provides a local tokenizer. (tiktoken | Claude API)
NVIDIA Releases cuTile Library and CUDA 13.1
Hugging Face Transformers v5 RC
Agents & Tooling Ecosystem
Multimodal Models & Generation Tools
Open Models & Datasets
Community & Industry
Reddit Highlights
Discord Discussions
*Source: Easy AI Education Project*