English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily News | December 6, 2025: vLLM 0.12.0, NVIDIA cuTile, Transformers v5 RC, Kling 2.6, and More

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily News for December 6, 2025 covers major AI industry updates. vLLM 0.12.0 ships experimental GPU Model Runner V2 with Prefill Context Parallel and optimized DeepSeek-V3.2 thinking-mode support. NVIDIA releases the cuTile Python compiler library targeting TileIR alongside CUDA 13.1 with PTX 9.1. Hugging Face launches Transformers v5 RC with AutoModelForMultimodalLM and any-to-any pipelines. LangChain adds content moderation middleware and cost tracking; Together AI partners with Meta on TorchForge RL support; SonarSource ships a SonarQube MCP server. On multimodal: Kling Video 2.6 adds native audio, Runway releases Gen 4.5 'Whisper Thunder', Alibaba Cloud launches Qwen3-TTS with 49+ voices in 10 languages, and Gemini 3 Pro gains document derendering and spatial trajectory generation. Open models include FLUX.2 [dev], Meituan's LongCat-Image, the MixtureVitae licensed dataset, and Intel's SignRoundV2 quantization. OpenRouter and a16z publish a report analyzing 100 trillion tokens, finding reasoning models now exceed 50% of usage.

Easy AI Daily News | December 6, 2025

Model & Inference Infrastructure

vLLM 0.12.0 Released with DeepSeek Optimizations

  • vLLM 0.12.0 introduces the experimental GPU Model Runner V2 and Prefill Context Parallel support.
  • Optimized support for DeepSeek-V3.2's "thinking" mode, including tokenizer and tool-call parsers; also adds EAGLE decoding and NVFP4 quantization.
  • Links: vLLM release | DeepSeek support details
  • NVIDIA Releases cuTile Library and CUDA 13.1

  • cuTile is a Python compiler library targeting TileIR, shipped with CUDA 13.1; the programming guide was rewritten.
  • PTX 9.1 adds SIMD conversions and async "sharp+tma" operations. cuTile does not yet support mxfp/nvfp, but fp4 is planned.
  • Links: cuTile repo | CUDA 13.1 docs
  • Hugging Face Transformers v5 RC

  • Adds AutoModelForMultimodalLM and an any-to-any pipeline supporting 2+ inputs/outputs (e.g., Gemma3n multimodal-to-text, Qwen3-Omni text + audio).
  • Link: release notes
  • Agents & Tooling Ecosystem

  • LangChain: new content moderation middleware (screens inputs, outputs, and tool results) and cost tracking with custom tool/API costs. DeepAgents CLI scored 42.7% on Terminal Bench 2.0, comparable to Claude Code. (moderation | cost tracking)
  • Together AI + Meta: production-grade TorchForge RL support via the Together platform for long-horizon agentic workflows. (announcement)
  • SonarSource: SonarQube MCP server brings enterprise static analysis into Claude Code/Cursor for more accurate AI code generation. (release)
  • Kimi CLI: integrates with JetBrains IDEs via ACP. (details)
  • Multimodal Models & Generation Tools

  • Kling Video 2.6: native synchronized audio (speech, sound effects, ambient sound); new "Element/Subject Library" for Kling O1 enables persistent subject memory and consistency. (release | audio)
  • Runway Gen 4.5 "Whisper Thunder": finer aesthetic control for world-building. (release)
  • Alibaba Cloud Qwen3-TTS: 49+ voices, 10 languages and dialects, natural prosody; real-time and offline APIs; demos on HF/ModelScope. (release)
  • Google Gemini 3 Pro: complex document derendering to HTML/LaTeX, screen understanding, spatial trajectory generation (robotics/XR), and a "thinking" mode for high-FPS video analysis. (details)
  • Open Models & Datasets

  • FLUX.2 [dev] (Black Forest Labs): #1 open text-to-image on Artificial Analysis Image Arena, #2 in editing. FLUX.2 [klein] uses Apache-2.0 for commercial use. (analysis)
  • Meituan LongCat-Image / LongCat-Image-Edit: image editing model under Apache-2.0, with demo. (release)
  • MixtureVitae: licensed pretraining dataset targeting math/code, avoiding Books2 copyright risk. (details)
  • Intel SignRoundV2: advances in extreme low-bit PTQ (e.g., 4-bit) for LLMs, improving quantization accuracy. (details)
  • Community & Industry

  • OpenRouter + a16z report: analysis of 100 trillion tokens shows reasoning models exceed 50% of usage, heavy Chinese closed-model traffic, and coding as a key use case. (report)
  • NeurIPS 2025: Yejin Choi's keynote covered reasoning work including EPO; Sakana AI presented the "Continuous Thought Machine" (Neural ODE-based test-time compute scaling). (Sakana AI)
  • OpenAI Residency applications open for engineers with basic ML experience; Google Gemini 3 Vibe Coding Hackathon offers $500K in API credits. (Residency | Hackathon)
  • Trending this week: Google Gemini hackathon, Amanda Askell's AI ethics AMA, Qwen3-TTS launch, OpenAI Residency, and Cloudflare outage affecting tools like Claude. (AMA)
  • Reddit Highlights

  • Basketball AI: community system using RF-DETR for player/jersey detection, SAM2 tracking, SmolVLM2 jersey-number recognition, plus SigLIP, UMAP, and K-Means for team clustering, with trajectory correction and shot detection. (discussion)
  • Anthropic survey: of 1,250 professionals, 86% say AI boosts productivity, but 69% feel stigma about using it. (study | discussion)
  • Image generation: SteadyDancer vs. Wan2.2 Animate comparison (SteadyDancer maintains 100% identity match); Detail Daemon + ZIT combo for high-quality fantasy art. (SteadyDancer | Detail Daemon)
  • Discord Discussions

  • CUDA Tile & GPU programming: debates on NVIDIA cuTile, PTX 9.1 features, and CUDA-L2 surpassing cuBLAS via RL optimization.
  • Gemini 3 vs Opus 4.5: community benchmarks found Gemini more expensive with lower SWE-Bench scores; GPT-5.1-High performed better in bug-finding tests. (comparison sheet)
  • Model-agnostic tool orchestrator: HuggingFace user released a production tool orchestrator based on Anthropic's Programmatic Tool Calling, letting any LLM write Rhai scripts to orchestrate tools, claiming 97-99% token reduction. (repo)
  • MCP token usage analysis: tokenization is model-dependent — OpenAI uses tiktoken, Claude uses the count_tokens API; Claude 3 no longer provides a local tokenizer. (tiktoken | Claude API)
---

*Source: Easy AI Education Project*

Tags

#ai-news#vllm#nvidia-cuda#hugging-face-transformers#qwen3-tts#flux2#gemini-3#open-source-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169092