Easy AI Daily | March 11, 2026 — AI Industry Roundup
A translated digest of the day's AI news, organized by topic.
Agents & Tooling
- Replit Agent 4: Moves beyond an AI-assisted IDE to a collaborative canvas workbench with parallel agents handling apps, websites, slides, and video — the broader trend of coding agents extending into general knowledge work.
- Perplexity Personal Computer: An always-on agent running on a Mac mini with access to local files, apps, and sessions, plus remote control. The enterprise version orchestrates 20 specialized models and 400+ apps, emphasizing unified agent orchestration over a single chat model.
- Base44 Superagents: A batteries-included agent suite for non-technical users, connecting out of the box to Gmail, Slack, Stripe, and CRMs to automate business workflows.
- LangChain Deep Agents context compression: Automatic summarization at task boundaries instead of hard token truncation, improving stability for long multi-step tool-use chains.
- OpenAI computer-use docs: Technical guidance on agent execution loops, filesystem context, network access, and safeguards for letting agents act safely on a computer.
- NVIDIA Nemotron 3 Super: A 120B-parameter (~12B active) hybrid Mamba-Transformer / SSM Latent MoE model with native 1M context, optimized for agents. Weights, data, training recipes, and infrastructure are public; NVIDIA claims up to 2.2x faster FP4 inference than GPT-OSS-120B.
- M5 Max 128GB local LLM benchmarks: Community tests on a 14" M5 Max 128GB using mlx_lm: Qwen3.5-122B-A10B-4bit hit ~1239 t/s prompt throughput at 16K context (73.8GB peak memory); gpt-oss-120b-MXFP4-Q8 reached 2710 t/s (~64.9GB), demonstrating 100B-class models running locally.
- Qwen3.5-35B-A3B Uncensored GGUF: 35B total / ~3B active params, 256 experts (8+1 active), text/image/video input and hybrid attention, near-zero refusals. Community debates capability loss; KL-divergence checks and long-context quality concerns raised.
- Fish Audio S2: TTS supporting 80+ languages with natural-language emotion tags (e.g.,
[whispers sweetly]), multi-speaker dialogue generation, ~100ms first-audio latency. Weights and code released, but commercial use requires a separate license. - Gemini Embedding 2: Multimodal embeddings (text, image, video, audio, PDF) with Matryoshka-style dimensionality reduction. Community analysis finds text pricing relatively expensive; best reserved for multimodal retrieval, and video embeddings can get costly without frame reduction.
- Qwen3.5 multimodal architecture teardown: Gated DeltaNet linear attention + full attention hybrid; a 397B-A17B MoE version and a 27B dense version; native 262k context extensible to ~1M; multi-token prediction in training.
- Reka Edge: A vision-language model for robotics and physical AI — image/video understanding, object detection, tool use — claiming 3x fewer input tokens and 65% higher throughput vs mainstream 8B models.
- PostTrainBench v1.0: Benchmarks whether frontier agents can post-train language models in simplified environments. Notably, on GPT-5.1 Codex Max, medium reasoning effort outperformed high, as extra tokens crowded out context.
- EvoSkill: Executor / proposer / skill-builder trio that distills reusable skill modules from failures; boosted Claude Code + Opus 4.5 exact-match on OfficeQA from 60.6% to 67.9%.
- AgentIR: Encodes agent reasoning traces with queries for reasoning-aware retrieval — 68% accuracy on BrowseComp-Plus vs 52% for larger traditional embedding models and 37% for BM25.
- Karpathy's self-improving swarm: An agentic system proposed ~700 training-pipeline changes, kept 20, cutting GPT-2-level training time from 2.02h to 1.80h (~11% improvement) — a concrete step toward AI-driven AI research.
- Layer-block duplication on Qwen2-72B: Duplicating a 7-layer middle block (without weight changes) topped the Open LLM Leaderboard using 2×4090s; experiments with unconventional layer rewiring suggest Transformer layers are more interchangeable than assumed.
- GPT-5.4 vs Frontier Math: Reports that GPT-5.4 solved an open EpochAI Frontier Math problem previously unsolved by human mathematicians; Epoch researchers preliminarily judge the solution correct, awaiting problem-author confirmation.
- Agent reliability as a safety issue: Princeton's response to NIST argues many agent failures are non-adversarial unreliability, requiring dedicated definitions, measurement, and mitigation.
- Google health AI: An imaging system flags ~25% of interval breast cancers missed by traditional screening; the AMIE conversational clinical-reasoning system was found safe, feasible, and well accepted in real-world pilots.
- r/LocalLLaMA hits 1M members: Growth under a year despite moderation turmoil signals local AI moving from niche hobbyist to mainstream developer interest.
- Replit at $9B valuation (up over 6 months), pivoting into a full productivity suite alongside Claude Cowork and Notion's custom agents.
- The Anthropic Institute: A new public-benefit research institute led by former policy head Jack Clark, spanning ML, economics, and social science.
- AMI Labs: Yann LeCun and Alexandre LeBrun's new venture raised $1.03B to build JEPA-based world models focused on the physical world and common sense; funded by NVIDIA, Samsung, Bezos, and others; code and papers to be open-sourced.
- Anthropic recursive self-improvement reports: Per TIME coverage, 70–90% of code for future models is written by Claude, with iteration cycles compressed from months to weeks; Claude 3.7 Sonnet's 10-day safety delay sparked debate over slowing for safety.
- Claude Code outage as 'intelligence blackout' preview: An auth failure disrupted many developers; Karpathy and others framed it as an infrastructure-level risk when R&D depends heavily on frontier models.
Infrastructure & Hardware
Models & Capabilities
Research & Methods
Products & Applications
Industry & Business
Policy, Governance & Safety
📌 Source: Easy AI Daily