📅 AI Industry Roundup — November 11, 2025
Model Updates & Performance
#### Moonshot AI releases Kimi K2 Thinking; AMA highlights, evaluations, and INT4 design
- Architecture: KDA (Kimi Delta Attention) + NoPE MLA hybrid attention stack; Muon optimizer supporting ~1T parameters
- Training: H800 GPUs, native INT4 QAT (quantization-aware training)
- Roadmap: vision capabilities coming soon; K3 planned to use KDA or hybrid attention
- Evaluations: Rank #7 on LisanBench; #2 on LMArena Text leaderboard (#2 among open models); supports agentic workflows with 200–300 tool calls
- Inference tips: use the official "kimi-k2-thinking-turbo" endpoint, enable streaming, temp=1.0
- Omnilingual ASR: open-source, covering 1600+ languages (including 500 underserved ones) with an accompanying corpus
- Gelato-30B-A3B: strong GUI operation performance — ScreenSpot-Pro 63.8%, OS-World-G 69.1%, outperforming larger models like Qwen3-VL-235B
- SYNTH synthetic dataset released along with Baguettotron model (200B tokens trained, SOTA on non-code tasks)
- Discussion of curriculum learning and RLVR scaling issues
- Fei-Fei Li published an essay on world models and spatial intelligence
- AMD Instinct MI355X: 2.2× performance improvement
- NVIDIA TensorRT-LLM adds Wide Expert Parallelism support
- Epoch AI predicts gigawatt-scale datacenters coming online in 2026
- H100/H200 spot prices expected to rise; Siemens launches a vLLM optimization platform; Baseten pushes "own your weights" training infrastructure
- Web auth is ill-suited for agent workflows; MCP standard focuses on tool discovery
- GEPA self-evolving agents support reflective learning
- Weave launches hallucination detection tools; FlowAgent for LangChain; Together AI publishes benchmarking guidance
- Strix Halo networking: InfiniBand 50Gbps performs similarly to Thunderbolt 10Gbps — network bandwidth is not the bottleneck
- Qwen3-VL OCR: outperforms Gemini 2.5 Pro, Claude Opus 4, and others
- BERT Chatbot with dLLM: open-sourced, turning any BERT into a chatbot with discrete diffusion
- Kimi K2 reportedly trained for ~$4.6M with near-GPT-5 performance
- Memes criticizing ChatGPT reliability (e.g., poisonous berry identification)
- Senators using AI-generated charts; OpenAI's Sora could cost up to $15M/day; Google introduces Nested Learning to address catastrophic forgetting
- LMArena: Sora 2 Pro account-sharing debate; criticism of OpenAI rule enforcement (Spotify, Meta); anticipation for Gemma 3 coding; Nano Banana 2 takedown theories
- Perplexity AI: Comet Browser YouTube search/playback issues; YouTube adblock clash (Chromium updates); referral program fraud bans; context window limits
- LM Studio: Gemma 4B context retention issues; Qwen3-VL OCR; NPU LLM performance (slower than GPU)
- Cursor Community: Sonnet 4.5 costs up to $1.02 NZD/min; Composer-1 disconnects; student verification errors; OpenRouter API key issues
- HuggingFace: NexusAI ComfyUI pro workflows; Maya1 open-source voice AI (3B params, 20 emotions); Ploke Rust AI interface for native project parsing
- GPU MODE: GMP-verified INT8×INT8→INT32 GEMM kernel (300.26 T-ops/s on A100); Blackwell microbenchmarking; NVSHMEM low-latency communication kernels
- OpenRouter: Kimi K2 prompt-induced crashloop fixed; Orchid AI ETA 2–48 months; Gemini 2.5 Flash token consumption (800k tokens for a 24-second video)
- OpenAI: Sora 2 quality complaints (static people, poor audio); GPT-5.1 Pro release speculation; AI censorship concerns
- Unsloth AI: AgentRL + Qwen2.5 7B integration delays; UD Quants performance drop (1.5 vs 4 tk/s); Kimi K2 Thinking GGUF looping issues in LM Studio
- Nous Research: Kimi's tone beats ChatGPT but weaker tracking; DeepSeek V3.2 low cost ($0.42/M tokens); Palantir identity debate
- Moonshot AI: Kimi K2 outperforms GLM 4.6; Unsloth reports Kimi-K2-Thinking issues; Kimi-for-coding quota drains fast ($19 plan lasts 1.5–2.5 days)
- Modular (Mojo): Mojo try-except outperforms Rust's Result; MAX beats TensorRT on B200; goal is a systems language with affine/linear types
- Yannick Kilcher: Qwen3-VL thinks it's text-only (Ollama's fault); Extropic talk grifty but interesting; Google's Nested Learning for catastrophic forgetting
- Latent Space: Terminal-Bench 2.0 released (89 tasks) with Harbor framework; Kimi K2 beats GPT-5 on Tau2 Bench; Meta's EdgeTAM real-time segmentation tracker (16 FPS on iPhone 15 Pro Max)
- Eleuther: Weave as WandB reporting alternative; NeurIPS chat; SAE nonlinear feature relationships paper accepted to AAAI 26
- tinygrad: big 3090→4090 gains, 5090 only marginal; migration to pyproject.toml; custom backward function kernel issues
- DSPy: Planner for multi-agent tool sprawl; TOON Adapter PR (perf concerns); first-class Agent CLI support aligned with Agent Client Protocol
- aider: Kimi models smarter with less verbose prompting; development moving to aider-ce branch; OpenRouter recommended for MoonshotAI K2 API
- MCP Contributors: 2025-11-25 spec release (freeze Nov 14); SEP-1330 SDK review; PII interception validation issues (Cursor, Claude)
- Manus.im: VEO3 connection loss blocks video creation; cancellations over token rates ($99 plan lasts hours); new engineers introduced (workflow automation, LLM integration)
> Links: AMA highlights | Evaluations | Agentic tool use
---
Speech & Computer Interaction Models
#### Meta releases multilingual ASR and Gelato-30B-A3B computer-use grounding model
> Links: Meta Omnilingual ASR | Gelato-30B-A3B
---
Data & Pretraining
#### SYNTH dataset and curriculum learning
> Links: SYNTH dataset | Fei-Fei Li essay
---
Hardware & Infrastructure
#### GPU and datacenter progress
> Links: AMD performance | GW datacenter forecast
---
Agents & Evaluation Tools
> Links: Auth for agents | GEPA agents | Weave detection
---
Reddit Discussions
#### /r/LocalLLaMA
> Links: Strix Halo test | Qwen3-VL OCR | BERT Chatbot
#### Less technical AI subreddits
> Links: Kimi K2 cost | ChatGPT memes | Sora cost
---
Discord Community Highlights
---
*Source: Easy AI Teaching Project*