Easy AI Daily Digest | 2026-02-26
A daily roundup of AI industry news, translated and summarized from Easy AI Daily (zhichai.net).
Products & Applications
- Perplexity launches "Computer": An all-in-one agent workstation for research, design, coding, deployment, and ops. Uses parallel async sub-agents plus a coordinator model that routes tasks to different models, with usage-based billing, spend caps, and memory/file/tool management. Initially available to Max users. Launch post | Architecture breakdown
- Claude Code turns one: Anthropic is positioning Claude Code as a coding agent foundation and unveiled legacy-system modernization for COBOL. Markets read this as a threat to IBM's mainframe services — IBM stock briefly fell over 10%, though real-world validation on critical financial systems remains to be seen. Reddit discussion
- GitHub Copilot CLI hits GA: Adds a repo-level
/researchcommand built on GitHub code search and MCP tools, generating deep-dive reports exportable as gists, with real-time task status in the terminal title. GA announcement - Nous open-sources Hermes Agent: A Python multi-agent workbench with hierarchical memory, sub-agents, file/terminal control, browser actions, and session continuity across CLI and IMs. Pairs with Atropos for data generation and RL pipelines. GitHub
- LM Studio launches LM Link: Secure remote access to local LLMs via Tailscale without exposing ports. The community wants mobile support and a mode free of third-party accounts. LM Link
- GPT-5.3-Codex is live in the API: ~25% faster than 5.2, fewer tokens per task, strong SWE-Bench Pro results. Pricing: $1.75/M input, $14/M output tokens — sparking "expensive but strong" debates. Announcement
- Qwen3.5 Medium (27B / 35B-A3B / 122B-A10B): Open weights with synchronized vLLM, GGUF, LM Studio, and Ollama support; near-lossless at 4-bit + KV quantization; 800K–1M token context. Developers report 35B-A3B's local agent tool-calling approaches commercial cloud quality, activating only ~3B parameters per token. Release
- Grok-4.20-Beta1 tops Arena Search with a score of 1226, beating GPT-5.2 and Gemini-3; tied 4th on Text (1492). Leaderboard
- Liquid AI releases LFM2-24B-A2B: A 24B sparse MoE activating 2B params/token, runnable on 32GB-memory devices, day-one support in llama.cpp, vLLM, SGLang, and multiple GGUF quants. Pretrained on 17T+ tokens and still training; will become LFM2.5. Reddit
- Diffusion LLM inference engines: Inception Labs and others claim ~1000 tok/s, further accelerated by inference-time techniques like Ψ-Samplers — still frontier experiments pending community reproduction. Andrew Ng's take
- Karpathy: coding agents "got real" since December 2025. He describes an almost fully automated local deployment — SSH, vLLM install, model pull, load testing, serving, frontend, systemd, reporting — noting a qualitative jump in long-task coherence. Thread
- ActionEngine: Treats GUI automation as graph search — offline exploration yields a state machine, and a single LLM call generates the whole action program, claiming better success rate, latency, and cost than step-by-step visual agents.
- OpenClaw as a system-level agent: Widely used for file/browser/desktop automation (email, CRM, finance), with rising safety concerns — one user granted root saw their recycle bin wiped; others built three-tier persistent memory stacks for it.
- Aider community's budget stack: DeepSeek V3.2 for main reasoning, mimo-v2-flash for fast edits, and Kimi-k2.5 for hard-problem planning — a cost/quality-balanced multi-model routing setup.
- Karpathy: the real bottleneck is memory orchestration, not compute. Fast-but-small on-chip SRAM vs. large-but-slow DRAM scheduling for prefill/decode under long context + high concurrency is the core challenge; neither HBM nor big-SRAM routes solve it cleanly. Thread
- OpenAI and Meta received 160M AMD warrants (~$600 strike) via large GPU purchase deals — a theoretical $192B equity upside, effectively a "stock rebate" on GPU spend.
- Blackwell price war: Packet.ai offers Blackwell GPU cloud at ~$0.66/hr or $199/month; individuals and small teams increasingly turn to rental options like Lightning AI clusters.
- Zagora: A distributed fine-tuning platform stitching scattered consumer GPUs over the public internet into clusters that train 70B+ models (GPT-OSS, Qwen 2.5, Mistral), using Petals/SWARM-style pipelining. Project
- Agent reliability lags capability: Models rack up benchmark wins while reliability stagnates — a single tool-call deviation cascades into compounding errors. Researchers propose minimal safety benchmarks (e.g., never send emails regardless of distractor context).
- Trace-Free+ (Intuit): Curriculum training teaches models to rewrite complex tool descriptions into agent-friendly formats, stabilizing multi-tool calling without extra traces at inference.
- Goodfire: Infrastructure collecting billions of activations on trillion-parameter models with minimal latency impact, including a live activation-based chain-of-thought steering demo.
- Midtraining: A new paper systematizes training between pretraining and post-training; it reduces forgetting and improves downstream performance but is highly sensitive to timing and data distribution.
- Diffusion / Flow Matching survey roundup from the Eleuther community: Rectified Flows, Diffusion Forcing, plus a video playlist.
- Anthropic acquires Vercept, a computer-use agent startup, to strengthen Claude's ability to actually operate interfaces rather than just suggest steps.
- Wayve raises $1.5B Series D at an $8.6B valuation (SoftBank, Microsoft, NVIDIA, Uber), planning supervised robotaxi pilots in 10 cities in 2026 and embodied AI hardware/software sales to automakers from 2027.
- Quiver AI raises $8.3M seed (a16z-led) and launches Arrow-1.0, generating editable SVG vectors from sketches or text descriptions for UI/poster/icon workflows.
- Pentagon deals push military AI red lines: The DoD reportedly reached an agreement with xAI to use Grok in classified systems and pressured Anthropic to allow "all lawful uses" including mass surveillance and weapons R&D. Anthropic publicly refuses mass surveillance and autonomous weapons; reports say it faced Defense Production Act threats and supply-chain-risk labeling.
- Anthropic rolls back RSP commitments: Per TIME, Anthropic dropped its flagship pledge not to train stronger models until safety is demonstrable; its chief scientist argued unilateral commitments are unsustainable when competitors don't follow.
- Jeff Dean publicly opposes mass surveillance, citing free-speech suppression, abuse potential, and constitutional concerns.
- Energy constraints emerge: The US is reportedly considering requiring large AI/cloud operators to self-provision power to protect the grid and ratepayers — scaling is now an energy-policy problem, not just an algorithm/GPU one.
- Automated jailbreak agents raise compliance alarms: A self-updating jailbreak agent built on OpenClaw + DeepSeek-R1 generates multi-turn evasive prompts against Claude, GPT, Gemini, and Grok; reviewers flagged near-total TOS violations and serious operational/legal risks.
Models & Capabilities
Agents & Tooling
Infrastructure & Hardware
Research & Methods
Industry & Business
Policy, Governance & Safety
📌 Source: Easy AI Daily