English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily Digest (Feb 17, 2026): Qwen3.5, MiniMax M2.5, Claude Opus 4.6, and More

Forum topic · 小凯 · 2026-03-27

Summary

A community-curated AI news roundup for February 17, 2026, covering major model releases and industry developments. Key items: Alibaba open-sources Qwen3.5-397B-A17B, a 397B-parameter sparse MoE with 17B active parameters, 256K native context (extensible to ~1M), and Apache-2.0 licensing; MiniMax ships M2.5 (230B total, 10B active, 200K context); Anthropic launches Claude Opus 4.6 with 1M-token context and a self-checking output step; Step 3.5 Flash gains attention for strong price-performance. Agent coverage includes OpenAI's acquisition of OpenClaw and its creator, the rise of 'harness engineering,' and MCP structured-output debates. Infrastructure notes highlight NVIDIA GB300 NVL72 efficiency claims, power delivery as the new bottleneck, and AccelOpt's LLM-generated CUDA kernels. Research items span Chain-of-Verification, Recursive Language Models, rubric-based RL, and model provenance forensics. Policy sections cover the Pentagon's supply-chain-risk warning to Anthropic, Google's model-distillation attack report, and OpenAI's new ChatGPT Lockdown Mode.

Key points

Models & Capabilities

  • Alibaba released Qwen3.5-397B-A17B — open-source sparse MoE with hybrid linear attention: 397B total / 17B active parameters, 201 languages, native 256K context (extensible to ~1M), Apache-2.0. Day-zero vLLM support; community estimated KV cache cost at ~31KB/token. The API version (Qwen3.5-Plus) offers 1M context plus search and code interpreter, though pricing drew criticism.
  • MiniMax M2.5: 230B total / 10B active params, 200K context, ~2500 tok/s per GPU on 8×H200 with vLLM. Token-level process rewards improve RL signal efficiency. Local deployment needs ~200GB VRAM (e.g., 2× RTX 6000 Blackwell at 120–130 tok/s). GLM-5 praised for tool-calling and multi-turn agent work, though service stability is still settling.
  • Claude Opus 4.6 launched with 1M-token context and an automatic "check your work" self-review step that can overturn earlier errors in long sessions. Strict hourly rate limits remain.
  • Step 3.5 Flash flagged by OpenRouter users as exceptionally strong price-performance, though platform support lags.
  • CommonLID: new language-identification benchmark (109 languages) from Common Crawl/EleutherAI; top models score below 80% F1 even on languages they claim to support.
  • Agents & Tooling

  • OpenClaw acquired by OpenAI; creator Peter Steinberger joins OpenAI's personal-agent effort while OpenClaw moves to a foundation. Community mixed: impressed by the solo-builder story, skeptical of hidden costs (e.g., 30-min heartbeat tasks) and post-acquisition direction.
  • "Harness engineering" as the new moat: tool orchestration, context management, and observability — not raw model quality — increasingly determine agent experience. Minimal alternatives (PicoClaw, nanobot) and LangSmith's trace-first debugging are emerging responses.
  • Real-world OpenClaw use: root-SSH Proxmox v6→v8 upgrades, multi-agent "software companies," Tavus video-call mode, and SEO content pipelines.
  • MCP discussions: token costs of prompt-embedded JSON schemas for structured output; proposals to separate text/image/object results and pass context (timezone, etc.) explicitly.
  • Jazz — a terminal-resident agent bundling MCP, git, shell, email, and scheduled tasks; Cloudflare experiments with HTTP endpoints returning Markdown for agents.
  • Infrastructure & Hardware

  • NVIDIA GB300 NVL72: claimed ~50× per-MW performance and 35× lower per-token cost vs. Hopper; the real bottleneck is shifting from GPUs/HBM to datacenter power and distribution. Western Digital's 2026 HDD capacity reportedly booked out, with some AI customers locked through 2027/2028.
  • AccelOpt: self-optimizing LLM agents claim 1.5× faster GQA paged decode and 1.38× faster prefill vs. FlashInfer 0.5.3; code open-sourced. GPU MODE is running a B200 FlashInfer-bench kernel contest.
  • Kernel-tuning gotchas on H100/H200: noisy TFLOPs readings (1400–1500 jitter), Achieved Occupancy excluding idle SMs, and version-mismatch pain across CUTLASS/CuteDSL/Proton on B200.
  • Hesper: WebGPU + BitNet-1.58 2B model hitting ~125 tok/s on M4 Max.
  • Research & Methods

  • CoVe, RLM, Rubric RL: Chain-of-Verification can roughly double accuracy on some tasks; Recursive Language Models (Omar Khattab) advocate recursive code-based reasoning over ever-longer attention; Cameron Wolfe's survey of 15+ rubric-based RL papers replaces fuzzy LLM-judge scoring.
  • Model genealogy: matrix-based weight homology, independence tests reconstructing Llama fine-tune trees from black-box access, and black-box provenance methods — potential anti-"wrapper model" tools.
  • Assistant Axis paper provides measurable evidence that activations drift along a persona axis during long conversations.
  • X-Ware's diffusion-based activation editing and "meta-neurons"; FAR.AI warns deception probes as training targets may teach activation-level camouflage instead of honesty.
  • QED-Nano 4B (Lewis Tunstall): multi-stage distillation + inference caching for IMO-level math proofs on small local models.
  • Products & Industry

  • Perplexity Pro backlash: deep search cut from 200 to 20 queries/month plus upload limits; matching prior usage now costs ~$167/month vs. $20; TrustPilot down to 1.5/5; users migrating to Claude/Opus 4.6 or Kimi.
  • Kimi K2.5 strong on coding/reasoning with a $40/month API tier, but CLI install failures, duplicate billing, quota issues, and scam mirror sites push users toward self-hosting large MoEs (~700GB RAM + 200GB VRAM setups).
  • Practical agentic coding workflows: Claude Cowork for pipeline tasks, planners like Ergo/planbot, executors like Codex/Claude Code/OpenClaw — the workflow (planning + version control + observability) matters more than the model.
  • Security applications: PassLLM (password-guessing LoRA on Qwen3-4B fed millions of real breach pairs) and ATIC (three Claude Opus 4.5 "brains" scoring epistemic uncertainty and flagging to humans).
  • China's "Spring Festival model week": Qwen3.5, GLM-5, MiniMax 2.5, and ByteDance's Seedance 2.0 (with a Jia Zhangke-directed short film) — video generation moving from toy to director-grade workflows.
  • Stripe fee complaints (~8.3% effective take); speculation that Apple is deliberately waiting out the $2T capex wave.
  • Policy, Governance & Safety

  • Pentagon vs. Anthropic: per Axios, the DoD may label Anthropic a "supply-chain risk" for refusing mass surveillance of Americans and fully autonomous weapons — compared by many to the PRISM era.
  • Google reports attackers made 100,000+ prompts against Gemini attempting black-box distillation; community noted the irony given Google's own web-scale training data practices.
  • OpenAI's ChatGPT Lockdown Mode for enterprise restricts tool calls (caching search, weakened web access) to reduce prompt-injection and data-exfiltration risk.
  • Reproducibility controversy over an OpenAI physics paper using GPT-5.2 without disclosing prompts or tooling; journals urged to require conversation logs.
---

📌 Source: Easy AI Daily (#EasyAI)

Tags

#ai-news#qwen3-5#minimax-m2-5#claude-opus-4-6#openclaw#agents#gpu-infrastructure#ai-safety

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169192