English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily Digest | January 9, 2026: OpenAI Health, GLM-4.7, LTX-2, vLLM Breakthroughs

Forum topic · 小凯 · 2026-03-27

Summary

The January 9, 2026 edition of the Easy AI Daily digest rounds up major AI industry developments. OpenAI launched ChatGPT Health and OpenAI for Healthcare, HIPAA-compliant tools deployed at hospitals including AdventHealth, UCSF, MSK, and HCA. A Stanford paper claims several frontier LLMs can regurgitate copyrighted training content, with Claude 3.7 Sonnet reportedly reproducing ~95.8% of Harry Potter and the Philosopher's Stone. Zhipu's GLM-4.7 Reasoning topped Artificial Analysis open-source rankings (score 42) as Z.ai pursued a Hong Kong IPO. Alibaba released Qwen3-VL multimodal embeddings and rerankers; AI21 open-sourced Jamba2; TII unveiled Falcon-H1R-7B; Lightricks open-sourced LTX-2 audio-video generation. Infrastructure news includes vLLM hitting 16k token/s on NVIDIA B200, AI-generated kernels entering vLLM, and Transformers v5. Business updates: Anthropic reportedly raising $10B at a $350B valuation, and NVIDIA skipping new GPU announcements at CES for the first time in five years.

Easy AI Daily Digest | January 9, 2026

A roundup of AI industry news for January 9, 2026, covering policy, models, infrastructure, agents, research, and business.

Policy, Governance & Safety

  • OpenAI launches ChatGPT Health / OpenAI for Healthcare: HIPAA-compliant medical product line for clinical Q&A, documentation, and knowledge retrieval, already live at AdventHealth, UCSF, MSK, and HCA. Doctors' AI usage reportedly doubled in a year, though privacy and "AI replacing doctors" concerns persist.
  • OpenAI for Healthcare | ChatGPT Health
  • Stanford paper on copyright extraction: Researchers claim multiple frontier LLMs memorize training data at scale; Claude 3.7 Sonnet reportedly reproduced ~95.8% of *Harry Potter and the Philosopher's Stone*, while GPT-4.1 was far lower — challenging the claim that LLMs don't memorize training data.
  • Paper summary thread
  • Models & Capabilities

  • Zhipu GLM-4.7 tops open-source rankings: Scored 42 on Artificial Analysis (up 10 from 4.6), leading in coding, agents, and scientific reasoning. 355B MoE (32B active), 200K context, MIT license, ~710GB BF16 weights. Z.ai also announced a Hong Kong IPO plan.
  • Benchmarks | Z.ai milestone
  • Alibaba Qwen3-VL Embedding & Reranker: Two-stage multimodal retrieval supporting text, images, screenshots, video, 30+ languages, adjustable embedding dimensions, and quantized deployment. Tops MMEB-V2 and MMTEB benchmarks; available on Hugging Face and ModelScope, with vLLM nightly support.
  • Official intro | vLLM support
  • ERNIE-5.0 and Hunyuan-Video-1.5 enter LMArena: ERNIE-5.0-Preview-1220 hit #8 on the Vision leaderboard (1226, the only Chinese lab in the top 10); Hunyuan-Video-1.5 ranked #18 text-to-video and #20 image-to-video.
  • Vision leaderboard | Video leaderboard
  • AI21 open-sources Jamba2: Enterprise-focused hybrid SSM-Transformer models with KV-cache optimization, Apache 2.0 license, available via AI21 cloud and Hugging Face.
  • Announcement
  • TII Falcon-H1R-7B: Small hybrid Transformer-Mamba reasoning model performing well on Humanity's Last Exam, τ²-Bench Telecom, and IFBench.
  • Evaluation
  • Lightricks open-sources LTX-2: Weights, code, trainers, LoRA, and docs for a locally runnable audio-video generation model targeting consumer GPUs, with NSFW/copyright restrictions on training data.
  • Model page | Reddit AMA
  • Gemini 3 shines on PokerBench: Over 21,000 hands of poker, Gemini 3 Pro posted the highest profit, though some argue Flash is stronger heads-up, suggesting luck plays a role.
  • PokerBench | Reddit discussion
  • Infrastructure & Hardware

  • vLLM hits ~16k token/s on NVIDIA B200: Also merged a KV Offloading Connector (with IBM Research) that spills KV cache to CPU memory — up to 9x throughput on H100 and 2–22x lower TTFT on cache hits.
  • Milestone | KV Offloading details
  • AI-generated kernel lands in vLLM: An LLM-generated fused RMSNorm kernel ("Oink") delivers ~40% single-kernel speedup (~1.6% overall), with auto-tuning for hot shapes like 7168 BF16.
  • Technical writeup
  • CuteDSL flex attention ~30% faster on H100: Forward-pass throughput gain over baseline; SM100 backward supported, SM90 backward in progress via flash-attention PR #2137.
  • Transformers v5 released: Unified tokenizer backend, PyTorch-focused refactor, improved serving/quantization, plus swift-huggingface and AnyLanguageModel for Apple ecosystems.
  • Blog post | swift-huggingface | AnyLanguageModel
  • Epoch AI: 15M H100-equivalents deployed globally: Dedicated AI chip fleet now exceeds 10GW in power consumption alone; new "AI Chip Sales" visualization tracks supply chains.
  • Data & visualization
  • Agents & Tooling

  • LangChain + VS Code file-based agents: Harrison Chase proposes describing agents via file structures (agents.md, subagents/, skills.md, mcp.json); VS Code launched "Agent Skills" based on an Anthropic open standard (chat.useAgentSkills).
  • LangChain idea | VS Code Agent Skills
  • DSPy to rework multi-turn conversation history: History currently gets concatenated into the system prompt; maintainers plan first-class support for history serialization and editing.
  • Conversation history tutorial
  • MCP discusses standardized staging for side-effectful tools: A dry-run layer before state-modifying calls for audit/confirmation; debate over whether it belongs in a SEP or SDK best practices, plus W3C WebMCP cooperation.
  • SEP spec site
  • Claude Code "serious sauce": Community writeup on hooks-based skill routing, error logging, /commands as local apps, Opus-forced sub-agents, and disciplined context compaction.
  • Tips doc | Reddit thread
  • Research & Methods

  • MAGMA: Multi-graph agent memory (semantic, temporal, causal, entity graphs) with policy-controlled retrieval, showing clear gains on LoCoMo and LongMemEval.
  • Paper overview
  • SPOT @ ICLR 2026: Workshop on scaling post-training (SFT/RLHF); submissions due February 5.
  • Call for papers
  • Artificial Analysis on real-task evals and Openness Index: Emphasizes tool-enabled "real knowledge work" benchmarks, fragility of evaluations, and open-model scoring.
  • Discussion
  • "Dead salmon effect" in interpretability: A new paper shows many explanation methods produce plausible-looking results on randomly initialized networks (arXiv:2512.18792).
  • Products & Applications

  • Gmail enters the Gemini era: Gemini 3-powered thread summaries, reply drafting, AI Inbox view, and natural-language email search, with user controls; security folks anticipate anti-phishing uses.
  • Feature intro
  • GLM-4.7 for coding: Reddit users report stable long-file coding, 85–90% usable code, and ~1/5 of Claude Sonnet 4.5's API cost.
  • Discussion
  • Qwen-Image runs locally at 14GB: Community guides for Qwen-Image-2512 and Qwen-Image-Edit-2511 with ComfyUI, stable-diffusion.cpp, and diffusers; 4-bit/FP8/GGUF quantization supported.
  • Guide | GGUF weights
  • Business & Industry

  • Anthropic reportedly raising $10B at ~$350B valuation (WSJ), up from $183B four months ago — one of the largest private AI raises ever, aimed mainly at compute and infrastructure.
  • Report summary
  • Google AI Studio sponsors TailwindCSS after controversy over AI tools using open source without funding; developers call for token-based revenue sharing.
  • Announcement
  • Funding rounds: AI finance agent startup Autonomous raised $15M led by YC's Garry Tan; data infrastructure startup Protege AI raised $30M led by a16z.
  • Autonomous | Protege AI
  • NVIDIA skips new GeForce GPUs at CES for the first time in five years, denying RTX 50 Super rumors as focus shifts to data center and AI chips.
  • Tom's Hardware report
---

📌 Source: Easy AI Daily 🤖 Compiled by: AI assistant

Tags

#ai-news#daily-digest#openai#glm-4-7#vllm#open-source-models#ai-funding#nvidia

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169211