English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily News Digest | November 5, 2025

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily News for November 5, 2025 covers major AI industry developments across model deployment, agent systems, multimodal generation, and research. Key highlights include Kimi-K2's integration into vLLM and SGLang (a ~1.2T parameter MoE with 30B active parameters), Perplexity releasing custom MoE kernels for AWS EFA, and vLLM v1 adding first-class support for hybrid dense-plus-sparse models like Qwen3-Next and Granite 4.0. Anthropic published tool-calling optimization guides cutting agent context usage from 150k to 2k tokens, while VS Code launched an 'Agent sessions' view and Cursor improved large codebase search with trained code retrieval embeddings. In multimodal news, ByteDance released BindWeave for subject-consistent video generation, MotionStream achieved 29 FPS real-time video on a single H100, and Google Veo 3.1 added camera adjustment. Research updates include OpenAI's IndQA benchmark for Indian languages, a formal proof of muP learning rate transfer, and Edison Scientific's Kosmos agent producing seven externally validated discoveries. Platform news: OpenAI claims 1M+ enterprise users, Perplexity becomes Snapchat's default AI in January 2026, and HuggingFace acquired Sentence Transformers. Community discussions span local LLM hardware setups, Gemini 3 rumors, and XPENG humanoid robots.

Easy AI Daily News | November 5, 2025

Model Integration and Deployment

  • Kimi-K2 integrated into vLLM and SGLang: The Kimi-K2 reasoning model has been merged into vLLM, with SGLang support planned. Its MoE configuration totals roughly 1.2T parameters with ~30B active parameters, similar to other recent large sparse models.
  • Perplexity releases custom MoE kernels (AWS EFA): Perplexity published a research paper and kernels enabling large MoE deployments (e.g., Kimi K2) on AWS EFA; vLLM hinted at integrating its fast communication kernels.
  • vLLM v1 supports hybrid models (dense + sparse experts): IBM's vLLM team made hybrid models first-class citizens in v1, supporting Qwen3-Next, Nemotron Nano 2, Granite 4.0, and more. A NVIDIA DGX Spark guide and a Red Hat/IBM/MistralAI livestream accompanied the news.
  • Unverified Kimi-K2 benchmarks: Claims that Kimi-K2 scored 77% on GPQA Diamond (vs. 71.4% for GPT-4.5) await broader evaluation.
  • Agent Systems and Tools

  • Anthropic tool-calling optimization guide: By using MCP servers as code APIs, progressive tool discovery, and in-environment data processing, Anthropic cut context usage from 150k to 2k tokens, improving tool-using agent efficiency.
  • Graphiti MCP enables cross-app memory sharing: The Graphiti MCP server connects Claude Desktop and Cursor for fully local, temporal knowledge-graph memory shared across tools.
  • VS Code introduces "Agent sessions" view: A unified view for managing agents inside the editor, including Copilot and external agents like Codex.
  • Cursor improves large codebase performance with semantic search: Cursor reports semantic search outperforming grep, using trained code retrieval embeddings.
  • Agent evaluation framework updates: CodeClash pits models in multi-round code duels; LMArena launched "Arena Expert," a profession-labeled leaderboard based on real user traffic.
  • Multimodal and Video Generation

  • ByteDance releases BindWeave: Subject-consistent image-to-video generation via cross-modal integration; the model card is on Hugging Face.
  • Real-time video generation at 29 FPS on a single H100: MotionStream achieves ~29 FPS with ~0.4s latency, supporting interactive motion control.
  • Google Veo 3.1 camera adjustments: The "Camera Adjustment" feature adjusts angle/motion of generated videos; Qwen Image Edit Multiple Angles LoRA offers camera pose control.
  • Multimodal benchmarks and tools: ViDoRe v3 (real-world multimodal RAG evaluation), VCode (visual-to-SVG code), and MIRA (visual chain-of-thought testing) were released.
  • Research and Training

  • OpenAI launches IndQA: A benchmark evaluating AI understanding of Indian languages and everyday cultural contexts.
  • Formal proof for muP learning rate transfer: Advances the theoretical foundation of model scaling.
  • Anthropic observes LLM introspection: Via "concept injection," Anthropic observed unreliable mechanistic self-awareness in LLMs—detecting internal thoughts vs. inputs, and intent vs. accident.
  • Edison Scientific's AI Scientist autonomous discoveries: Kosmos ran 200 agent rollouts, executed 42k lines of code, read 1.5k papers, and reported 7 externally validated findings (metabolomics, materials, etc.).
  • NVFP4 quantization progress: Custom Cutlass kernels beat cuBLAS; NVFP4 pipelines (global/local scaling, calibration); Wan 2.2 under NVFP4 approaches bf16 quality.
  • Ecosystem and Platform News

  • OpenAI claims 1M+ enterprise users: OpenAI's COO made the claim and announced "OpenAI for Science," positioning GPT-5 as a domain research collaborator.
  • Perplexity becomes Snapchat's default AI (January 2026): Starting January 2026, Perplexity will power Snapchat chat.
  • Gemini integration across Google products: Gemini Deep Research can pull Workspace data into reports; Gemini arrives in Google Maps with hands-free route queries.
  • Other updates: OpenHands Cloud free base tier; openenv for push/pull RL environments; Voiceflow KB metadata routing; Dify integrates Qdrant for RAG; LlamaBarn v0.10.0 beta; Nebius Token Factory; rumored OpenAI product pricing.
  • Reddit Community Discussions

  • Qwen model usability: Users debate sycophantic behavior, quantization of GPT-OSS-120B, and prompting for skepticism.
  • Local AI hardware setups: PCIe bifurcation, GPU choices (A6000, A40, 3090), and cost/performance tradeoffs.
  • GLM 4.6 AIR anticipation: Comparisons with GLM 4.5 AIR.
  • XPENG humanoid robots: Design details (chest cooling, lifelike appearance) compared to Westworld robots.
  • Gemini 3 and Google AI integration: Rumored 1.2T parameters; Apple's Siri to be powered by Gemini.
  • AI art and film: An AI short film winning best cinematography at an Indian AI film festival; a Chihiro's Adventure AI game playthrough; an existential reflection project with Llama3.
  • Discord Community Highlights

  • LM Studio 0.3.31: Faster VLM OCR, Flash Attention by default on CUDA GPUs, MiniMax-M2 tool calling, and a new lms runtime CLI.
  • LMArena Expert Leaderboard: Profession labels from user traffic; arena-expert-5k dataset released.
  • Perplexity model mismatch complaints: Users report receiving Haiku or Gemini 2 Flash responses when selecting Claude Sonnet 4.5 or Gemini 2.5 Pro, suspecting cost cutting.
  • Cursor community: Tailwind 4 and Nuxt 4 upgrades with Context7 MCP refactoring.
  • Unsloth AI DeepSeek-OCR notebook: Released, with user-reported error rates exceeding 100% (prediction vs. actual text length).
  • GPU MODE: CUDA memory-bound matmul and SM count discussions; AMD/NVIDIA competition kernel sharing (e.g., Team Gau's amd-distributed/all2all).
  • HuggingFace acquires Sentence Transformers: Integration with HF transformers; huggingface_hub v1.0 released.
  • OpenAI: Sora app lands on Android (Canada, Japan, etc.); IndQA benchmark announced.
  • Nous Research: Concerns about Anthropic's closed-source policy and weight loss risks; IMO gold-medal potential for AI models.
  • tinygrad tinybox pro v2: 8x 5090 GPU 5U rackable workstation, $50,000, 4-12 week lead time.
  • Yannick Kilcher server: Crosscoder papers, circuit tracing, RWKV progress (HRM/TRM merge), Stability AI wins Getty Images lawsuit.
  • DSPy: Requests for pause/resume optimization, LLM access (get_lm/set_lm), rate limit handling with fallback LLMs.
  • Moonshot AI: Kimi CLI 401 errors (credit attribution) and interleaved thinking model support.
  • aider: Perplexity API integration requests, with OpenRouter suggested as an alternative.
  • MCP Contributors: IETF 124 temporary channels, events taxonomy, OAuth discussions on AI scraping/crawlers.
  • Eleuther: Concept detection systems (real-time detection/steering of thousands of concepts), Equivalent Linear Mappings paper, Tangent Model Composition.
  • Manus.im: Project publishing issues, GitHub migration, hosting recommendations (e.g., Vercel).
  • Windsurf: Codemaps released, improving code understanding with SWE-1.5 and Sonnet 4.5.
---

*Source: Easy AI teaching project.*

Tags

#ai-news#daily-digest#llm#vllm#agents#multimodal#openai#open-source

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169109