Easy AI Daily Digest | January 9, 2026
A roundup of AI industry news for January 9, 2026, covering policy, models, infrastructure, agents, research, and business.
Policy, Governance & Safety
- OpenAI launches ChatGPT Health / OpenAI for Healthcare: HIPAA-compliant medical product line for clinical Q&A, documentation, and knowledge retrieval, already live at AdventHealth, UCSF, MSK, and HCA. Doctors' AI usage reportedly doubled in a year, though privacy and "AI replacing doctors" concerns persist.
- OpenAI for Healthcare | ChatGPT Health
- Stanford paper on copyright extraction: Researchers claim multiple frontier LLMs memorize training data at scale; Claude 3.7 Sonnet reportedly reproduced ~95.8% of *Harry Potter and the Philosopher's Stone*, while GPT-4.1 was far lower — challenging the claim that LLMs don't memorize training data.
- Paper summary thread
- Zhipu GLM-4.7 tops open-source rankings: Scored 42 on Artificial Analysis (up 10 from 4.6), leading in coding, agents, and scientific reasoning. 355B MoE (32B active), 200K context, MIT license, ~710GB BF16 weights. Z.ai also announced a Hong Kong IPO plan.
- Benchmarks | Z.ai milestone
- Alibaba Qwen3-VL Embedding & Reranker: Two-stage multimodal retrieval supporting text, images, screenshots, video, 30+ languages, adjustable embedding dimensions, and quantized deployment. Tops MMEB-V2 and MMTEB benchmarks; available on Hugging Face and ModelScope, with vLLM nightly support.
- Official intro | vLLM support
- ERNIE-5.0 and Hunyuan-Video-1.5 enter LMArena: ERNIE-5.0-Preview-1220 hit #8 on the Vision leaderboard (1226, the only Chinese lab in the top 10); Hunyuan-Video-1.5 ranked #18 text-to-video and #20 image-to-video.
- Vision leaderboard | Video leaderboard
- AI21 open-sources Jamba2: Enterprise-focused hybrid SSM-Transformer models with KV-cache optimization, Apache 2.0 license, available via AI21 cloud and Hugging Face.
- Announcement
- TII Falcon-H1R-7B: Small hybrid Transformer-Mamba reasoning model performing well on Humanity's Last Exam, τ²-Bench Telecom, and IFBench.
- Evaluation
- Lightricks open-sources LTX-2: Weights, code, trainers, LoRA, and docs for a locally runnable audio-video generation model targeting consumer GPUs, with NSFW/copyright restrictions on training data.
- Model page | Reddit AMA
- Gemini 3 shines on PokerBench: Over 21,000 hands of poker, Gemini 3 Pro posted the highest profit, though some argue Flash is stronger heads-up, suggesting luck plays a role.
- PokerBench | Reddit discussion
- vLLM hits ~16k token/s on NVIDIA B200: Also merged a KV Offloading Connector (with IBM Research) that spills KV cache to CPU memory — up to 9x throughput on H100 and 2–22x lower TTFT on cache hits.
- Milestone | KV Offloading details
- AI-generated kernel lands in vLLM: An LLM-generated fused RMSNorm kernel ("Oink") delivers ~40% single-kernel speedup (~1.6% overall), with auto-tuning for hot shapes like 7168 BF16.
- Technical writeup
- CuteDSL flex attention ~30% faster on H100: Forward-pass throughput gain over baseline; SM100 backward supported, SM90 backward in progress via flash-attention PR #2137.
- Transformers v5 released: Unified tokenizer backend, PyTorch-focused refactor, improved serving/quantization, plus swift-huggingface and AnyLanguageModel for Apple ecosystems.
- Blog post | swift-huggingface | AnyLanguageModel
- Epoch AI: 15M H100-equivalents deployed globally: Dedicated AI chip fleet now exceeds 10GW in power consumption alone; new "AI Chip Sales" visualization tracks supply chains.
- Data & visualization
- LangChain + VS Code file-based agents: Harrison Chase proposes describing agents via file structures (agents.md, subagents/, skills.md, mcp.json); VS Code launched "Agent Skills" based on an Anthropic open standard (
chat.useAgentSkills). - LangChain idea | VS Code Agent Skills
- DSPy to rework multi-turn conversation history: History currently gets concatenated into the system prompt; maintainers plan first-class support for history serialization and editing.
- Conversation history tutorial
- MCP discusses standardized staging for side-effectful tools: A dry-run layer before state-modifying calls for audit/confirmation; debate over whether it belongs in a SEP or SDK best practices, plus W3C WebMCP cooperation.
- SEP spec site
- Claude Code "serious sauce": Community writeup on hooks-based skill routing, error logging, /commands as local apps, Opus-forced sub-agents, and disciplined context compaction.
- Tips doc | Reddit thread
- MAGMA: Multi-graph agent memory (semantic, temporal, causal, entity graphs) with policy-controlled retrieval, showing clear gains on LoCoMo and LongMemEval.
- Paper overview
- SPOT @ ICLR 2026: Workshop on scaling post-training (SFT/RLHF); submissions due February 5.
- Call for papers
- Artificial Analysis on real-task evals and Openness Index: Emphasizes tool-enabled "real knowledge work" benchmarks, fragility of evaluations, and open-model scoring.
- Discussion
- "Dead salmon effect" in interpretability: A new paper shows many explanation methods produce plausible-looking results on randomly initialized networks (arXiv:2512.18792).
- Gmail enters the Gemini era: Gemini 3-powered thread summaries, reply drafting, AI Inbox view, and natural-language email search, with user controls; security folks anticipate anti-phishing uses.
- Feature intro
- GLM-4.7 for coding: Reddit users report stable long-file coding, 85–90% usable code, and ~1/5 of Claude Sonnet 4.5's API cost.
- Discussion
- Qwen-Image runs locally at 14GB: Community guides for Qwen-Image-2512 and Qwen-Image-Edit-2511 with ComfyUI, stable-diffusion.cpp, and diffusers; 4-bit/FP8/GGUF quantization supported.
- Guide | GGUF weights
- Anthropic reportedly raising $10B at ~$350B valuation (WSJ), up from $183B four months ago — one of the largest private AI raises ever, aimed mainly at compute and infrastructure.
- Report summary
- Google AI Studio sponsors TailwindCSS after controversy over AI tools using open source without funding; developers call for token-based revenue sharing.
- Announcement
- Funding rounds: AI finance agent startup Autonomous raised $15M led by YC's Garry Tan; data infrastructure startup Protege AI raised $30M led by a16z.
- Autonomous | Protege AI
- NVIDIA skips new GeForce GPUs at CES for the first time in five years, denying RTX 50 Super rumors as focus shifts to data center and AI chips.
- Tom's Hardware report
Models & Capabilities
Infrastructure & Hardware
Agents & Tooling
Research & Methods
Products & Applications
Business & Industry
📌 Source: Easy AI Daily 🤖 Compiled by: AI assistant