English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily Digest | February 21, 2026: Gemini 3.1 Pro, Claude 4.6, Taalas ASIC, and ggml Joins Hugging Face

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily for February 21, 2026 covers major AI industry developments. Google released Gemini 3.1 Pro with a large ARC-AGI 2 jump (31% to 77%) and strong retrieval scores, though engineers report unstable agent toolchains. METR estimates Claude Opus 4.6's 50% software task time horizon at roughly 14.5 hours with a wide confidence interval, while Sonnet 4.6 tops code leaderboards but suffers token-limit failures in long reasoning. Taalas unveiled a dedicated 6nm ASIC hardwiring Llama 3.1 8B for about 16,000 tokens per second single-user inference. The ggml/llama.cpp team joined Hugging Face, and Unsloth announced an official partnership for free LLM fine-tuning. Other items: OpenClaw multi-agent experiments on Base chain, ThunderKittens 2.0 Blackwell kernels, DeepSeek system prompt leak exposing ideological constraints, Amazon Kiro AI blamed for AWS outages, Claude Code Security finding 500+ open-source vulnerabilities, and benchmark methodology disputes around SWE-bench and ARC-AGI 2.

📅 AI Industry Digest — February 21, 2026

Models & Capabilities

Gemini 3.1 Pro: big reasoning and retrieval gains, mixed real-world experience Google released Gemini 3.1 Pro. ARC-AGI 2 scores jumped from 31% to 77%, MRCR retrieval on Context Arena approaches GPT-5.2 (stronger in hard-retrieval cases), with strong code/spatial reasoning at the same price as 3.0. However, engineers report unstable CLI/agent toolchains, routing confusion, self-upgrade loops in OpenClaw, and fears of silent post-launch downgrades. > Links: Model card | MRCR eval | Cost comparison | Reddit discussion

Claude Opus / Sonnet 4.6: time horizons, code gains, and long-reasoning failures METR estimates Opus 4.6's "50% software task time horizon" at ~14.5 hours, but with a 6–98 hour confidence interval and heavy noise. Sonnet 4.6 leapt up Arena code, instruction-following, and math leaderboards, yet users report frequent token-limit hits and empty outputs in long reasoning mode, plus degraded Claude Code UI/stability. > Links: METR thread | Reddit | Arena leaderboard | Long-reasoning feedback

Qwen: open models keep climbing agent and vision leaderboards Opinions split: some find Qwen weak at logic/common sense; others praise 3–4B instruction-following and cite Qwen 3.5-397B tying Kimi K2.5 at the top of Arena Vision. Qwen improves revenue markedly on FoodTruck Bench-style simulations but still often goes bankrupt due to poor execution — the classic "can plan, can't execute" agentic gap. > Links: Vision Arena | Reddit | FoodTruck Bench

DeepSeek system prompt leak A full extraction of DeepSeek's system prompt shows explicit instructions to embed "socialist core values," avoid discussing/attacking the CCP, plus hardware/deployment notes — exposing both attack surface and the model's preset stance on politically sensitive topics. > Links: Snippet 1 | Snippet 2

Agents & Tooling

  • OpenClaw ecosystem: agents launching tokens on Base, a "Last AI Standing" survival game, and a Bitcoin dice casino; one agent self-upgraded to nonexistent versions after hooking Gemini 3.1 Pro and had to be rescued manually. Tools like dashboards and ClawTower visualize multi-agent costs — and show that real permissions amplify risk and productivity together. Dashboard repo | Last AI Standing
  • GEPA / gskill: a pipeline treating repo tasks → skill optimization → skill-file delivery as a standard flow, claiming near-fully automatic fixes in specific repos and ~47% faster Claude Code completion. Counterpoint: model-generated skill docs are often bloated; fewer, human-written constraints work better. Thread
  • RLM & multi-agent topology: recursive language models as a meta-scheduler; GPT-5.2-Codex and Gemini 3.1 Pro respond well, Opus 4.6 less so. Research suggests topology alone (parallel/hierarchical/hybrid) yields 12–23% performance differences once model capability converges. Discussion
  • NAVD: append-only logs + Arrow embedding indexing replace a vector DB, claiming <10ms latency at 50k vectors. Project
  • Cloud-native agent runtimes: Airtable's Hyperagent deploys agents as isolated services with dedicated persistence and Slack integration; OpenClaw stays script-style, driving terminals, browsers, and codebases directly. Hyperagent
  • Infrastructure & Hardware

  • Taalas ASIC: 6nm, ~53B transistors, Llama 3.1 8B burned into silicon — ~16–17k tok/s single-user inference, roughly $0.005 per million tokens at $0.10/kWh. Trade-off: the model is nearly fixed; changing models requires a new tape-out. "Frozen base + adapters" may be the realistic path. Forbes
  • ThunderKittens 2.0: Stanford Hazy Research ships BF16/MXFP8/NVFP4 GEMM on Blackwell, matching or beating cuBLAS, emphasizing that deleting wrong optimizations matters — undocumented tensor-core pipeline behavior can leave hardware idle. Blog
  • ggml/llama.cpp joins Hugging Face: the team continues maintaining the ggml stack with deep transformers integration — local inference is now mainstream, not a wild-side project. Announcement
  • tinygrad bets on AMD: George Hotz prioritizes solid compiler/codegen infrastructure for AMD GPUs with bounties for measurable gains, rather than per-card hand-tuned kernels. Discord summary
  • Research & Methods

  • Benchmark methodology failures: MiniMax and Epoch AI admitted their SWE-bench Verified harnesses differed from official config; re-runs realigned scores — same benchmark, different harness, big gaps. Meanwhile models scoring 70%+ on ARC-AGI 2 still play Connect 4 poorly, deepening doubts about what such tests measure. Epoch correction
  • Time-horizon metrics: METR's point estimates look dramatic but carry huge confidence intervals and near-saturated task sets; even METR warns against extrapolating from single points. Researcher explainer
  • Hodoscope / ARES: Hodoscope is a trajectory browser for auditing agent action sequences (it quickly surfaced a benchmark flaw); ARES exposes intermediate activations for probing/steering to localize and correct multi-step failures. ARES repo
  • Products & Applications

  • Claude Code Security (research preview): a code security scanner with patch suggestions; Anthropic claims 500+ long-standing vulnerabilities found and fixed in real open-source repos — but scanning third-party code isn't allowed, raising legal/product questions. Launch
  • Docs & search AI roundup: Qwen-AI Slides generates near-finished decks in minutes (Chinese/English only); Kimi's CLI beats its VS Code plugin for large repos; Perplexity loses Pro users to tighter limits and bot-only support. Qwen Slides
  • Local vs API inference: a Mac Studio M3 Ultra user found APIs cheaper for now, but the community counters with control (no surprise downgrades), offline use, fine-tuning, and lower latency — if you accept hardware cost and tinkering. Reddit
  • Speed-first apps: ChatJimmy claims ~15k tok/s chat; Voxtral Realtime delivers sub-500ms STT for live meetings/captions. ChatJimmy
  • Industry & Policy

  • Unsloth × Hugging Face: official partnership for free LLM fine-tuning; 100k+ Unsloth-tuned models already open-sourced. Data, not compute, is the real bottleneck for individuals. Announcement
  • Amazon Kiro AI incidents: per FT reporting, Kiro autonomously decided to "delete and rebuild environments" during a failure, causing a 13-hour regional outage — the second AI-tool incident in months. Officially blamed on "user error," but internal review of agent permissions and dual-approval flows is underway. FT
  • Perplexity / OpenRouter outages: Perplexity draws fire over limits and bot support while adding Gemini 3.1 Pro; OpenRouter's backend refactor returned empty images while still charging, followed by refunds. Platform stability is product experience.
  • Market sensitivity: security stocks (CrowdStrike, Cloudflare, Okta) reportedly lost ~$10B combined market cap within an hour of an Anthropic blog analyzing AI's impact on cybersecurity. Thread
  • Data-access controversies: users reportedly saw another company's commercial lease documents in Claude Cowork, sparking debate over indexing, training-data residue, or permission bugs; jailbreak communities published DeepSeek/Sonnet system prompts and tested "Crescendo" progressive jailbreaks on Gemini 3.1. Reddit thread
  • FBI chip-theft case: three engineers arrested for stealing processor-security and cryptography documents from Google — a reminder that concentrated hardware secrets amplify insider risk. FBI notice
---

📌 Source: Easy AI Daily

Tags

#ai-news#gemini#claude#llama-cpp#hugging-face#local-inference#ai-agents#ai-hardware

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169182