Key points
Models & Capabilities
- Alibaba released Qwen3.5-397B-A17B — open-source sparse MoE with hybrid linear attention: 397B total / 17B active parameters, 201 languages, native 256K context (extensible to ~1M), Apache-2.0. Day-zero vLLM support; community estimated KV cache cost at ~31KB/token. The API version (Qwen3.5-Plus) offers 1M context plus search and code interpreter, though pricing drew criticism.
- MiniMax M2.5: 230B total / 10B active params, 200K context, ~2500 tok/s per GPU on 8×H200 with vLLM. Token-level process rewards improve RL signal efficiency. Local deployment needs ~200GB VRAM (e.g., 2× RTX 6000 Blackwell at 120–130 tok/s). GLM-5 praised for tool-calling and multi-turn agent work, though service stability is still settling.
- Claude Opus 4.6 launched with 1M-token context and an automatic "check your work" self-review step that can overturn earlier errors in long sessions. Strict hourly rate limits remain.
- Step 3.5 Flash flagged by OpenRouter users as exceptionally strong price-performance, though platform support lags.
- CommonLID: new language-identification benchmark (109 languages) from Common Crawl/EleutherAI; top models score below 80% F1 even on languages they claim to support.
- OpenClaw acquired by OpenAI; creator Peter Steinberger joins OpenAI's personal-agent effort while OpenClaw moves to a foundation. Community mixed: impressed by the solo-builder story, skeptical of hidden costs (e.g., 30-min heartbeat tasks) and post-acquisition direction.
- "Harness engineering" as the new moat: tool orchestration, context management, and observability — not raw model quality — increasingly determine agent experience. Minimal alternatives (PicoClaw, nanobot) and LangSmith's trace-first debugging are emerging responses.
- Real-world OpenClaw use: root-SSH Proxmox v6→v8 upgrades, multi-agent "software companies," Tavus video-call mode, and SEO content pipelines.
- MCP discussions: token costs of prompt-embedded JSON schemas for structured output; proposals to separate text/image/object results and pass context (timezone, etc.) explicitly.
- Jazz — a terminal-resident agent bundling MCP, git, shell, email, and scheduled tasks; Cloudflare experiments with HTTP endpoints returning Markdown for agents.
- NVIDIA GB300 NVL72: claimed ~50× per-MW performance and 35× lower per-token cost vs. Hopper; the real bottleneck is shifting from GPUs/HBM to datacenter power and distribution. Western Digital's 2026 HDD capacity reportedly booked out, with some AI customers locked through 2027/2028.
- AccelOpt: self-optimizing LLM agents claim 1.5× faster GQA paged decode and 1.38× faster prefill vs. FlashInfer 0.5.3; code open-sourced. GPU MODE is running a B200 FlashInfer-bench kernel contest.
- Kernel-tuning gotchas on H100/H200: noisy TFLOPs readings (1400–1500 jitter), Achieved Occupancy excluding idle SMs, and version-mismatch pain across CUTLASS/CuteDSL/Proton on B200.
- Hesper: WebGPU + BitNet-1.58 2B model hitting ~125 tok/s on M4 Max.
- CoVe, RLM, Rubric RL: Chain-of-Verification can roughly double accuracy on some tasks; Recursive Language Models (Omar Khattab) advocate recursive code-based reasoning over ever-longer attention; Cameron Wolfe's survey of 15+ rubric-based RL papers replaces fuzzy LLM-judge scoring.
- Model genealogy: matrix-based weight homology, independence tests reconstructing Llama fine-tune trees from black-box access, and black-box provenance methods — potential anti-"wrapper model" tools.
- Assistant Axis paper provides measurable evidence that activations drift along a persona axis during long conversations.
- X-Ware's diffusion-based activation editing and "meta-neurons"; FAR.AI warns deception probes as training targets may teach activation-level camouflage instead of honesty.
- QED-Nano 4B (Lewis Tunstall): multi-stage distillation + inference caching for IMO-level math proofs on small local models.
- Perplexity Pro backlash: deep search cut from 200 to 20 queries/month plus upload limits; matching prior usage now costs ~$167/month vs. $20; TrustPilot down to 1.5/5; users migrating to Claude/Opus 4.6 or Kimi.
- Kimi K2.5 strong on coding/reasoning with a $40/month API tier, but CLI install failures, duplicate billing, quota issues, and scam mirror sites push users toward self-hosting large MoEs (~700GB RAM + 200GB VRAM setups).
- Practical agentic coding workflows: Claude Cowork for pipeline tasks, planners like Ergo/planbot, executors like Codex/Claude Code/OpenClaw — the workflow (planning + version control + observability) matters more than the model.
- Security applications: PassLLM (password-guessing LoRA on Qwen3-4B fed millions of real breach pairs) and ATIC (three Claude Opus 4.5 "brains" scoring epistemic uncertainty and flagging to humans).
- China's "Spring Festival model week": Qwen3.5, GLM-5, MiniMax 2.5, and ByteDance's Seedance 2.0 (with a Jia Zhangke-directed short film) — video generation moving from toy to director-grade workflows.
- Stripe fee complaints (~8.3% effective take); speculation that Apple is deliberately waiting out the $2T capex wave.
- Pentagon vs. Anthropic: per Axios, the DoD may label Anthropic a "supply-chain risk" for refusing mass surveillance of Americans and fully autonomous weapons — compared by many to the PRISM era.
- Google reports attackers made 100,000+ prompts against Gemini attempting black-box distillation; community noted the irony given Google's own web-scale training data practices.
- OpenAI's ChatGPT Lockdown Mode for enterprise restricts tool calls (caching search, weakened web access) to reduce prompt-injection and data-exfiltration risk.
- Reproducibility controversy over an OpenAI physics paper using GPT-5.2 without disclosing prompts or tooling; journals urged to require conversation logs.
Agents & Tooling
Infrastructure & Hardware
Research & Methods
Products & Industry
Policy, Governance & Safety
📌 Source: Easy AI Daily (#EasyAI)