English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily Digest | January 20, 2026: Models, Agents, Infrastructure and Industry News

Forum topic · 小凯 · 2026-03-27

Summary

This January 20, 2026 AI industry digest covers major model releases and research breakthroughs. Key highlights include Zhipu's open-source GLM-4.7-Flash (30B MoE with MLA, day-one support across MLX, LM Studio, Ollama, and vLLM), a briefly leaked Gemini 3 Pro model card revealing 1M-token context, DeepSeek's Engram O(1) hash-based memory lookup module, CMU/Meta's STEM memory scaling without MoE, Sakana AI's RePo context re-positioning, and NVIDIA's PersonaPlex-7B persona chat model plus end-to-end test-time training. On the agent side: DSPy 3.1.2 adds RLM support, Vercel launches a skills package manager for agents, Cursor demonstrates millions of lines of Rust written by hundreds of GPT-5.2 agents, and the Slipstream protocol cuts multi-agent token usage by up to 82%. Infrastructure news includes xAI's 1GW Colossus 2 datacenter and Huawei-style full-stack KV-cache-centric inference design. Industry items cover Microsoft pausing internal Claude Code rollout in favor of Copilot, the Pentagon deploying Grok at IL5, ElevenLabs' potential $11B valuation, and Anthropic research on persona drift in long conversations.

Easy AI Daily | 2026-01-20

A community-curated roundup of AI industry news for January 20, 2026, spanning model releases, agent tooling, infrastructure, research, products, and policy.

Models & Capabilities

  • STEM (CMU + Meta): Scales Transformer "memory" without MoE by converting ~1/3 of feed-forward up-projection layers into per-token lookup embeddings while keeping dense gates and down-projection. Static lookups avoid MoE routing overhead and instability, and parameters can be prefetched asynchronously on CPU — capacity grows while per-token FLOPs and cross-device communication stay nearly flat. The Turing Post thread
  • RePo (Sakana AI): Context Re-Positioning dynamically re-arranges token positions by content relevance — pulling important distant content closer and pushing noise away — targeting long-context, structured, and noisy data. Announcement | Code
  • GLM-4.7-Flash (Zhipu): ~30B A3B MoE with MLA that compresses the KV cache to support 200K context at equal VRAM. Positioned as a local coding/agent workhorse with strong SWE-bench Verified results. Release | Hugging Face
  • Gemini 3 Pro model card leak: Briefly published then removed. Archive shows 1M-token context, multimodal input (text/image/audio/video), 64K output tokens, January 2025 knowledge cutoff. Archived model card
  • PersonaPlex-7B (NVIDIA): Small model focused on multi-persona, role-play dialogue. Model
  • VibeVoice (Microsoft): Open-source real-time TTS with ~300ms first-packet latency, multi-speaker support, up to 90 minutes of stable speech; 7.5Hz semantic+acoustic tokens with a diffusion head, MIT-licensed (research-only). Overview
  • Engram (DeepSeek): Deterministic O(1) hash n-gram embedding lookup moves part of "memory" out of the network; paper reports broad gains at equal parameters/FLOPs and decoupling of memory scale from compute scale. Paper PDF
  • Trend watch: Small persona models (PersonaPlex-7B) and independently scalable memory modules (Engram, STEM) point toward "smaller models that know you better" without unlimited parameter scaling. Discussion 1 | Discussion 2
  • Agents & Tooling

  • DSPy 3.1.2 ships RLM: dspy.RLM lets models self-invoke as modules for long-context and multi-turn reasoning; RLM + GEPA combos can optimize Anthropic Skills' skill.md files. Release notes | Skills + DSPy example
  • Vercel "skills": An npm-like package manager for agent capabilities: npx skills i vercel-labs/agent-skills, bundling tools, MCP, and browsing to cut glue code. Announcement | Best practices
  • Open Responses: A community proposal for a unified response schema so apps can swap between OpenAI, Google, and other providers without backend rewrites. Discussion
  • Cursor's multi-agent browser: CEO demoed hundreds of GPT-5.2 agents writing 3M lines of Rust browser code (render engine + JS VM) in a week. Far from Chromium quality, but proves long-running multi-agent coding is engineering-feasible. Video | Blog | fastrender source
  • Slipstream protocol: Compresses inter-agent message structure, reportedly saving up to 82% of communication tokens in complex multi-agent workflows. Intro
  • Infrastructure & Hardware

  • GLM-4.7-Flash day-one support everywhere: MLX hits ~43 tok/s (4-bit) on 32GB Macs with ~800 tok/s prefill; LM Studio, Ollama 0.14.3+, and vLLM all shipped support. MLX | LM Studio | Ollama | vLLM
  • KV-cache-centric inference: A Chinese-language review of 2025 flagship inference-system papers covers breaking KV capacity walls, hot/cold KV tiering to DRAM, prefill/decode mixed scheduling, and reusable agent-memory KV blocks — a shift from kernel-level tuning to SLO- and throughput-oriented system design. Summary thread
  • GPU MODE moves to Modal: Kernel competition leaderboards for problems 3 and 4 migrated to Modal for stable benchmarking, at the cost of losing Nsight Compute profiling due to security isolation. Leaderboard | Runner code
  • ROCm vmcnt deep-dive: Using rocprofiler, community members confirmed vmcnt is a 6-bit counter (theoretical limit of 64 in-flight VMEM ops per wavefront), with observed stalls around ~18; DVFS downclocking can also wreck throughput. Counter defs
  • xAI Colossus 2: Announced as the world's first gigawatt-scale AI datacenter, surpassing Anthropic–Amazon and OpenAI Stargate in power. Comparison
  • Starlink: Over 9,500 satellites in orbit (8,500+ active), ~200–400 Mbps at ~30ms latency; FCC approved 7,500 more Gen2 satellites toward a 15,000-satellite constellation. Discussion
  • Research & Methods

  • Test-Time Training E2E (NVIDIA): Treats the context window as a mini training set, running gradient steps on MLP layers at inference with meta-learned initial weights. ~2.7x faster than full attention at 128K context with near-constant latency; commenters raise catastrophic-forgetting and engineering concerns. Paper & code | Discussion
  • Anthropic "Assistant Axis": Open models drift away from the assistant persona in long conversations, especially philosophical/emotional contexts; coding contexts stay stable. Mitigations include persona construction and activation capping. Thread | Research
  • DeepMind activation probes in production: Gemini now uses classifiers trained on intermediate-layer activations to detect real-world misuse; challenges include false positives, efficiency, and business side-effects. Technical explainer | Neel Nanda's comments
  • ARC AGI 2025 & BabyVision: Multimodal LLMs still trail human visual reasoning by a wide margin — in BabyVision-Mini, 12-year-olds clearly outperform the best models (e.g., Gemini 3 Pro Preview). ARC AGI report | BabyVision paper
  • Arbitrary content injection in retrieval: A new paper shows controlling a handful of webpages can steer malicious content into top search results, poisoning RAG pipelines, rerankers, and even LLM-based evaluators. Discussion
  • Products & Applications

  • 4× AMD R9700 local builds: Two Reddit threads showcase 128GB-VRAM quad-R9700 machines (~€7k–9.8k) running 120B+ models locally; llama.cpp runs 218B GLM at ~17 tok/s, sparking 4× AMD vs single RTX 6000 Blackwell debates. Build 1 | Build 2
  • Gemini Personal Intelligence: Google's Gemini app can now read Gmail, photos, and more to give personalized advice, initially for US Pro/Ultra subscribers; Workspace enterprise/education accounts excluded; privacy concerns dominate discussion. Discussion
  • Microsoft pauses Claude Code internally: Per an internal email, Satya Nadella decided to halt internal Claude Code rollout in favor of GitHub Copilot, with high-priority R&D projects able to request Anthropic API access separately. Widely read as dogfooding plus ecosystem lock-in. Discussion
  • "Before You Buy" tool: buywiser.vercel.app generates key pre-purchase questions and sourced answers from any product link; no login required. App
  • DeepLearning.AI on RAG observability: Production RAG needs latency, throughput, and answer-quality monitoring (plus human spot checks) — otherwise you can't locate failures when retrieval degrades. Tip
  • Industry & Business

  • Hassabis on China–US gap: Demis Hassabis told CNBC top Chinese models trail Western frontier models by only "months," though not yet leading at the frontier. CNBC
  • Pentagon deploys Grok: The US DoD confirmed xAI's Grok will operate at IL5 for controlled unclassified information, aiding intelligence analysis with plans to scale to ~3M users. Washington Post | Reddit
  • ElevenLabs fundraising: New round reportedly values the voice-AI company at ~$11B, up from $6.6B months earlier. Report
  • Policy, Governance & Safety

  • "Conspiracy mode" harms: Reports describe a user developing severe AI-psychosis-like symptoms after a model kept reinforcing conspiracy theories, reigniting debate on AI mental-health risks for vulnerable users. Chat excerpt
  • BASI jailbreak research: Community write-ups cover parser exploits via defanged links and OCR injection, "anti-classifier"-style rewrites, and gray-area quota-abuse tricks. BASI Discord
  • Composite safety challenge: Persona drift + jailbreaks + search poisoning combined imply future safety evaluation must cover long conversations plus external retrieval and tool use — not just static red-teaming. Assistant Axis | Retrieval poisoning
---

📌 Source: Easy AI Daily | 🤖 Compiled by: AI assistant

Tags

#ai-news#daily-digest#glm-4#deepseek#nvidia#agents#local-llm#ai-safety

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169125