Easy AI Daily News | November 5, 2025
Model Integration and Deployment
- Kimi-K2 integrated into vLLM and SGLang: The Kimi-K2 reasoning model has been merged into vLLM, with SGLang support planned. Its MoE configuration totals roughly 1.2T parameters with ~30B active parameters, similar to other recent large sparse models.
- Perplexity releases custom MoE kernels (AWS EFA): Perplexity published a research paper and kernels enabling large MoE deployments (e.g., Kimi K2) on AWS EFA; vLLM hinted at integrating its fast communication kernels.
- vLLM v1 supports hybrid models (dense + sparse experts): IBM's vLLM team made hybrid models first-class citizens in v1, supporting Qwen3-Next, Nemotron Nano 2, Granite 4.0, and more. A NVIDIA DGX Spark guide and a Red Hat/IBM/MistralAI livestream accompanied the news.
- Unverified Kimi-K2 benchmarks: Claims that Kimi-K2 scored 77% on GPQA Diamond (vs. 71.4% for GPT-4.5) await broader evaluation.
- Anthropic tool-calling optimization guide: By using MCP servers as code APIs, progressive tool discovery, and in-environment data processing, Anthropic cut context usage from 150k to 2k tokens, improving tool-using agent efficiency.
- Graphiti MCP enables cross-app memory sharing: The Graphiti MCP server connects Claude Desktop and Cursor for fully local, temporal knowledge-graph memory shared across tools.
- VS Code introduces "Agent sessions" view: A unified view for managing agents inside the editor, including Copilot and external agents like Codex.
- Cursor improves large codebase performance with semantic search: Cursor reports semantic search outperforming grep, using trained code retrieval embeddings.
- Agent evaluation framework updates: CodeClash pits models in multi-round code duels; LMArena launched "Arena Expert," a profession-labeled leaderboard based on real user traffic.
- ByteDance releases BindWeave: Subject-consistent image-to-video generation via cross-modal integration; the model card is on Hugging Face.
- Real-time video generation at 29 FPS on a single H100: MotionStream achieves ~29 FPS with ~0.4s latency, supporting interactive motion control.
- Google Veo 3.1 camera adjustments: The "Camera Adjustment" feature adjusts angle/motion of generated videos; Qwen Image Edit Multiple Angles LoRA offers camera pose control.
- Multimodal benchmarks and tools: ViDoRe v3 (real-world multimodal RAG evaluation), VCode (visual-to-SVG code), and MIRA (visual chain-of-thought testing) were released.
- OpenAI launches IndQA: A benchmark evaluating AI understanding of Indian languages and everyday cultural contexts.
- Formal proof for muP learning rate transfer: Advances the theoretical foundation of model scaling.
- Anthropic observes LLM introspection: Via "concept injection," Anthropic observed unreliable mechanistic self-awareness in LLMs—detecting internal thoughts vs. inputs, and intent vs. accident.
- Edison Scientific's AI Scientist autonomous discoveries: Kosmos ran 200 agent rollouts, executed 42k lines of code, read 1.5k papers, and reported 7 externally validated findings (metabolomics, materials, etc.).
- NVFP4 quantization progress: Custom Cutlass kernels beat cuBLAS; NVFP4 pipelines (global/local scaling, calibration); Wan 2.2 under NVFP4 approaches bf16 quality.
- OpenAI claims 1M+ enterprise users: OpenAI's COO made the claim and announced "OpenAI for Science," positioning GPT-5 as a domain research collaborator.
- Perplexity becomes Snapchat's default AI (January 2026): Starting January 2026, Perplexity will power Snapchat chat.
- Gemini integration across Google products: Gemini Deep Research can pull Workspace data into reports; Gemini arrives in Google Maps with hands-free route queries.
- Other updates: OpenHands Cloud free base tier; openenv for push/pull RL environments; Voiceflow KB metadata routing; Dify integrates Qdrant for RAG; LlamaBarn v0.10.0 beta; Nebius Token Factory; rumored OpenAI product pricing.
- Qwen model usability: Users debate sycophantic behavior, quantization of GPT-OSS-120B, and prompting for skepticism.
- Local AI hardware setups: PCIe bifurcation, GPU choices (A6000, A40, 3090), and cost/performance tradeoffs.
- GLM 4.6 AIR anticipation: Comparisons with GLM 4.5 AIR.
- XPENG humanoid robots: Design details (chest cooling, lifelike appearance) compared to Westworld robots.
- Gemini 3 and Google AI integration: Rumored 1.2T parameters; Apple's Siri to be powered by Gemini.
- AI art and film: An AI short film winning best cinematography at an Indian AI film festival; a Chihiro's Adventure AI game playthrough; an existential reflection project with Llama3.
- LM Studio 0.3.31: Faster VLM OCR, Flash Attention by default on CUDA GPUs, MiniMax-M2 tool calling, and a new
lms runtimeCLI. - LMArena Expert Leaderboard: Profession labels from user traffic; arena-expert-5k dataset released.
- Perplexity model mismatch complaints: Users report receiving Haiku or Gemini 2 Flash responses when selecting Claude Sonnet 4.5 or Gemini 2.5 Pro, suspecting cost cutting.
- Cursor community: Tailwind 4 and Nuxt 4 upgrades with Context7 MCP refactoring.
- Unsloth AI DeepSeek-OCR notebook: Released, with user-reported error rates exceeding 100% (prediction vs. actual text length).
- GPU MODE: CUDA memory-bound matmul and SM count discussions; AMD/NVIDIA competition kernel sharing (e.g., Team Gau's amd-distributed/all2all).
- HuggingFace acquires Sentence Transformers: Integration with HF transformers; huggingface_hub v1.0 released.
- OpenAI: Sora app lands on Android (Canada, Japan, etc.); IndQA benchmark announced.
- Nous Research: Concerns about Anthropic's closed-source policy and weight loss risks; IMO gold-medal potential for AI models.
- tinygrad tinybox pro v2: 8x 5090 GPU 5U rackable workstation, $50,000, 4-12 week lead time.
- Yannick Kilcher server: Crosscoder papers, circuit tracing, RWKV progress (HRM/TRM merge), Stability AI wins Getty Images lawsuit.
- DSPy: Requests for pause/resume optimization, LLM access (get_lm/set_lm), rate limit handling with fallback LLMs.
- Moonshot AI: Kimi CLI 401 errors (credit attribution) and interleaved thinking model support.
- aider: Perplexity API integration requests, with OpenRouter suggested as an alternative.
- MCP Contributors: IETF 124 temporary channels, events taxonomy, OAuth discussions on AI scraping/crawlers.
- Eleuther: Concept detection systems (real-time detection/steering of thousands of concepts), Equivalent Linear Mappings paper, Tangent Model Composition.
- Manus.im: Project publishing issues, GitHub migration, hosting recommendations (e.g., Vercel).
- Windsurf: Codemaps released, improving code understanding with SWE-1.5 and Sonnet 4.5.
Agent Systems and Tools
Multimodal and Video Generation
Research and Training
Ecosystem and Platform News
Reddit Community Discussions
Discord Community Highlights
*Source: Easy AI teaching project.*