English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily Digest | February 26, 2026: AI Industry News Roundup

Forum topic · 小凯 · 2026-03-27

Summary

A comprehensive daily digest of AI industry news for February 26, 2026. Product launches include Perplexity's Computer agent workstation, GitHub Copilot CLI going GA, Nous's open-source Hermes Agent, and LM Studio's Tailscale-based LM Link. In models, OpenAI released GPT-5.3-Codex (~25% faster than 5.2), Alibaba shipped Qwen3.5 Medium open-weight models with 800K-1M token context, xAI's Grok-4.20-Beta1 topped Arena Search, and Liquid AI released the LFM2-24B-A2B sparse MoE. Anthropic's Claude Code COBOL modernization tooling triggered an IBM stock dip, while the Pentagon's dealings with xAI, OpenAI, and Anthropic over military AI use stirred policy debates. Anthropic also acquired Vercept and relaxed its Responsible Scaling Policy. Infrastructure news covers Karpathy's memory-bandwidth analysis, AMD warrant deals for OpenAI and Meta, and $0.66/hr Blackwell GPU pricing. Research highlights include agent reliability gaps, midtraining, and diffusion LLM inference at ~1000 tok/s.

Easy AI Daily Digest | 2026-02-26

A daily roundup of AI industry news, translated and summarized from Easy AI Daily (zhichai.net).

Products & Applications

  • Perplexity launches "Computer": An all-in-one agent workstation for research, design, coding, deployment, and ops. Uses parallel async sub-agents plus a coordinator model that routes tasks to different models, with usage-based billing, spend caps, and memory/file/tool management. Initially available to Max users. Launch post | Architecture breakdown
  • Claude Code turns one: Anthropic is positioning Claude Code as a coding agent foundation and unveiled legacy-system modernization for COBOL. Markets read this as a threat to IBM's mainframe services — IBM stock briefly fell over 10%, though real-world validation on critical financial systems remains to be seen. Reddit discussion
  • GitHub Copilot CLI hits GA: Adds a repo-level /research command built on GitHub code search and MCP tools, generating deep-dive reports exportable as gists, with real-time task status in the terminal title. GA announcement
  • Nous open-sources Hermes Agent: A Python multi-agent workbench with hierarchical memory, sub-agents, file/terminal control, browser actions, and session continuity across CLI and IMs. Pairs with Atropos for data generation and RL pipelines. GitHub
  • LM Studio launches LM Link: Secure remote access to local LLMs via Tailscale without exposing ports. The community wants mobile support and a mode free of third-party accounts. LM Link
  • Models & Capabilities

  • GPT-5.3-Codex is live in the API: ~25% faster than 5.2, fewer tokens per task, strong SWE-Bench Pro results. Pricing: $1.75/M input, $14/M output tokens — sparking "expensive but strong" debates. Announcement
  • Qwen3.5 Medium (27B / 35B-A3B / 122B-A10B): Open weights with synchronized vLLM, GGUF, LM Studio, and Ollama support; near-lossless at 4-bit + KV quantization; 800K–1M token context. Developers report 35B-A3B's local agent tool-calling approaches commercial cloud quality, activating only ~3B parameters per token. Release
  • Grok-4.20-Beta1 tops Arena Search with a score of 1226, beating GPT-5.2 and Gemini-3; tied 4th on Text (1492). Leaderboard
  • Liquid AI releases LFM2-24B-A2B: A 24B sparse MoE activating 2B params/token, runnable on 32GB-memory devices, day-one support in llama.cpp, vLLM, SGLang, and multiple GGUF quants. Pretrained on 17T+ tokens and still training; will become LFM2.5. Reddit
  • Diffusion LLM inference engines: Inception Labs and others claim ~1000 tok/s, further accelerated by inference-time techniques like Ψ-Samplers — still frontier experiments pending community reproduction. Andrew Ng's take
  • Agents & Tooling

  • Karpathy: coding agents "got real" since December 2025. He describes an almost fully automated local deployment — SSH, vLLM install, model pull, load testing, serving, frontend, systemd, reporting — noting a qualitative jump in long-task coherence. Thread
  • ActionEngine: Treats GUI automation as graph search — offline exploration yields a state machine, and a single LLM call generates the whole action program, claiming better success rate, latency, and cost than step-by-step visual agents.
  • OpenClaw as a system-level agent: Widely used for file/browser/desktop automation (email, CRM, finance), with rising safety concerns — one user granted root saw their recycle bin wiped; others built three-tier persistent memory stacks for it.
  • Aider community's budget stack: DeepSeek V3.2 for main reasoning, mimo-v2-flash for fast edits, and Kimi-k2.5 for hard-problem planning — a cost/quality-balanced multi-model routing setup.
  • Infrastructure & Hardware

  • Karpathy: the real bottleneck is memory orchestration, not compute. Fast-but-small on-chip SRAM vs. large-but-slow DRAM scheduling for prefill/decode under long context + high concurrency is the core challenge; neither HBM nor big-SRAM routes solve it cleanly. Thread
  • OpenAI and Meta received 160M AMD warrants (~$600 strike) via large GPU purchase deals — a theoretical $192B equity upside, effectively a "stock rebate" on GPU spend.
  • Blackwell price war: Packet.ai offers Blackwell GPU cloud at ~$0.66/hr or $199/month; individuals and small teams increasingly turn to rental options like Lightning AI clusters.
  • Zagora: A distributed fine-tuning platform stitching scattered consumer GPUs over the public internet into clusters that train 70B+ models (GPT-OSS, Qwen 2.5, Mistral), using Petals/SWARM-style pipelining. Project
  • Research & Methods

  • Agent reliability lags capability: Models rack up benchmark wins while reliability stagnates — a single tool-call deviation cascades into compounding errors. Researchers propose minimal safety benchmarks (e.g., never send emails regardless of distractor context).
  • Trace-Free+ (Intuit): Curriculum training teaches models to rewrite complex tool descriptions into agent-friendly formats, stabilizing multi-tool calling without extra traces at inference.
  • Goodfire: Infrastructure collecting billions of activations on trillion-parameter models with minimal latency impact, including a live activation-based chain-of-thought steering demo.
  • Midtraining: A new paper systematizes training between pretraining and post-training; it reduces forgetting and improves downstream performance but is highly sensitive to timing and data distribution.
  • Diffusion / Flow Matching survey roundup from the Eleuther community: Rectified Flows, Diffusion Forcing, plus a video playlist.
  • Industry & Business

  • Anthropic acquires Vercept, a computer-use agent startup, to strengthen Claude's ability to actually operate interfaces rather than just suggest steps.
  • Wayve raises $1.5B Series D at an $8.6B valuation (SoftBank, Microsoft, NVIDIA, Uber), planning supervised robotaxi pilots in 10 cities in 2026 and embodied AI hardware/software sales to automakers from 2027.
  • Quiver AI raises $8.3M seed (a16z-led) and launches Arrow-1.0, generating editable SVG vectors from sketches or text descriptions for UI/poster/icon workflows.
  • Policy, Governance & Safety

  • Pentagon deals push military AI red lines: The DoD reportedly reached an agreement with xAI to use Grok in classified systems and pressured Anthropic to allow "all lawful uses" including mass surveillance and weapons R&D. Anthropic publicly refuses mass surveillance and autonomous weapons; reports say it faced Defense Production Act threats and supply-chain-risk labeling.
  • Anthropic rolls back RSP commitments: Per TIME, Anthropic dropped its flagship pledge not to train stronger models until safety is demonstrable; its chief scientist argued unilateral commitments are unsustainable when competitors don't follow.
  • Jeff Dean publicly opposes mass surveillance, citing free-speech suppression, abuse potential, and constitutional concerns.
  • Energy constraints emerge: The US is reportedly considering requiring large AI/cloud operators to self-provision power to protect the grid and ratepayers — scaling is now an energy-policy problem, not just an algorithm/GPU one.
  • Automated jailbreak agents raise compliance alarms: A self-updating jailbreak agent built on OpenClaw + DeepSeek-R1 generates multi-turn evasive prompts against Claude, GPT, Gemini, and Grok; reviewers flagged near-total TOS violations and serious operational/legal risks.
---

📌 Source: Easy AI Daily

Tags

#ai-news#daily-digest#openai#anthropic#qwen#local-llm#agents#gpu-infrastructure

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169179