English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily News | January 9, 2026: AI Industry Roundup

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily for January 9, 2026 covers major AI industry developments. OpenAI launched ChatGPT Health and OpenAI for Healthcare, HIPAA-compliant products deployed at AdventHealth, UCSF, MSK, and HCA hospitals. A Stanford study claims frontier models like Claude 3.7 Sonnet can reproduce copyrighted content, reportedly up to 95.8% of Harry Potter. Zhipu's GLM-4.7 topped open-source capability rankings, while Alibaba released Qwen3-VL embedding and reranker models for multimodal retrieval. Baidu ERNIE-5.0 and Tencent Hunyuan-Video-1.5 entered LMArena leaderboards. AI21 open-sourced Jamba2, and TII released Falcon-H1R-7B. On infrastructure, vLLM hit 16k token/s on NVIDIA B200 and merged a KV offloading connector. WSJ reported Anthropic is raising $10 billion at a potential $350 billion valuation. NVIDIA broke its five-year CES streak by announcing no new consumer GPUs.

Easy AI Daily News | January 9, 2026

Policy, Governance and Safety

  • OpenAI launches ChatGPT Health / OpenAI for Healthcare: A medical product line claimed to be HIPAA-compliant, used for clinical Q&A, documentation, and knowledge retrieval. Already live at AdventHealth, UCSF, MSK, and HCA. OpenAI says physician AI usage "doubled in a year," though privacy concerns and fears of "AI replacing doctors" persist.
  • OpenAI healthcare announcement | ChatGPT Health intro
  • Stanford paper: copyrighted content extractable from frontier LLMs: Researchers say multiple online models memorize training data at scale and can emit copyrighted excerpts in certain settings. Claude 3.7 Sonnet reportedly reproduced ~95.8% of the first Harry Potter book, while GPT-4.1 was much lower — countering claims that LLMs don't memorize training data.
  • Paper summary thread
  • Models and Capabilities

  • Zhipu GLM-4.7 tops open-source rankings: Scores 42 on Artificial Analysis' index (up 10 from 4.6), leading in coding, agents, and science reasoning, with the highest GDPval-AA ELO among evaluated open models. Specs: 355B MoE (32B active), 200k context, MIT license; BF16 weights ~710GB — too large even for a single 8×H100 node. Parent company Z.ai announced a Hong Kong IPO listing.
  • GLM-4.7 benchmarks | Z.ai milestone
  • Alibaba Qwen3-VL multimodal embedding and reranker: Two-stage retrieval supporting text, images, screenshots, video, 30+ languages, adjustable embedding dimensions, instruction-tuning, and quantized deployment. Tops MMEB-V2 and MMTEB benchmarks per official claims. Available on Hugging Face and ModelScope; supported in vLLM nightly.
  • Official intro | vLLM support
  • Baidu ERNIE-5.0 and Tencent Hunyuan-Video-1.5 enter LMArena: ERNIE-5.0-Preview-1220 scored 1226, ranking 8th on the Vision leaderboard (currently the only Chinese lab in the top 10). Hunyuan-Video-1.5 placed 18th in text-to-video and 20th in image-to-video.
  • Vision leaderboard | Video leaderboard
  • AI21 open-sources Jamba2: Enterprise-focused hybrid SSM-Transformer with KV-cache optimization, Apache 2.0 licensed, available via AI21 cloud and Hugging Face.
  • TII Falcon-H1R-7B: A small hybrid Transformer-Mamba reasoning model performing well on Humanity's Last Exam, τ²-Bench Telecom, and IFBench; openness score of 44.
  • Lightricks open-sources LTX-2: A locally runnable audio-video generation model with weights, code, trainers, LoRA, and docs; runs on consumer GPUs. NSFW/copyright restrictions on training data.
  • LTX-2 model page
  • Gemini 3 shines on PokerBench: 21,000 hands of Texas Hold'em; Gemini 3 Pro ended most profitable overall, though one developer noted Flash won head-to-head, suggesting luck factors. Data and code are open.
  • PokerBench
  • Infrastructure and Hardware

  • vLLM + B200 hits ~16k token/s: The community merged a KV Offloading Connector (with IBM Research) that asynchronously offloads KV cache to CPU memory. Up to 9× throughput gains on H100 and 2–22× TTFT reductions in cache-hit scenarios.
  • B200 milestone
  • AI-generated kernels enter vLLM: An LLM-generated fused RMSNorm kernel ("Oink") gives ~40% single-kernel speedup and ~1.6% end-to-end gain, with near-autotuning for hot shapes (e.g., 7168 BF16) — plus added crash/stability complexity.
  • Technical writeup
  • CuteDSL flex attention ~30% faster on H100: Forward-pass speedup integrated into existing frameworks; backward support on SM90 progressing via Flash-Attention PR #2137.
  • Hugging Face Transformers v5: Unified tokenizer backend, PyTorch-centric model definitions, and major upgrades to serving, quantization, and deployment; plus swift-huggingface and AnyLanguageModel for Apple ecosystems.
  • Transformers v5 blog
  • Epoch: global compute exceeds 15M H100-equivalents: AI chip power alone tops 10GW, excluding other datacenter overhead. An "AI Chip Sales" visualization tool tracks supply chains.
  • Epoch data
  • Agents and Tooling

  • LangChain + VS Code: agents as folders and skills: Harrison Chase proposes describing agents via file structures (agents.md, subagents/, skills.md, mcp.json). VS Code launched "Agent Skills" based on Anthropic's open standard, enabled via chat.useAgentSkills.
  • DSPy to rework multi-turn conversation handling: Conversation history currently gets stuffed into the system prompt; maintainers say adapters are an implementation detail and multi-turn history serialization will be redone as a first-class configuration option.
  • MCP community discusses standardized staging for side-effectful tools: A dry-run layer before state-changing tool calls for audit and confirmation — possibly as a SEP, or as SDK best practice; W3C WebMCP cooperation also discussed.
  • Claude Code power-user patterns: Hooks reading local routing files, error-log systems collecting failed prompts, /commands as local mini-apps, forcing subagents to Opus, and strict context compaction for large projects.
  • Research and Methods

  • MAGMA: Splits agent memory into semantic, temporal, causal, and entity graphs with policy-controlled retrieval instead of single-pass vector search; clear gains on LoCoMo and LongMemEval.
  • SPOT @ ICLR 2026: Workshop on scaling post-training (SFT/RLHF); submissions due February 5.
  • Artificial Analysis on evaluation methodology: Real-world GDPval-AA tasks plus an Openness Index covering weights, data, and deployment restrictions; discussion of sensitivity, brittleness, and "mystery shopper" testing.
  • "Dead salmon effect" returns: A paper shows many interpretation methods (feature attribution, probes, sparse autoencoders, causal analysis) produce plausible-looking explanations even on randomly initialized networks (arXiv:2512.18792) — urging caution with interpretability results.
  • Products and Applications

  • Gmail enters the Gemini era: AI conversation summaries, replies and polishing, AI Inbox view, and natural-language mailbox search with user toggles; security researchers see anti-phishing potential but warn about trusted-agent manipulation.
  • Local LLM practice with GLM-4.7: Reddit users report using it instead of Claude Sonnet 4.5 for debugging and refactoring, with 85–90% usable code rates and roughly one-fifth the API cost; Sonnet still preferred for design-level discussion.
  • Qwen-Image runs locally in ~14GB: Tutorials cover Qwen-Image-2512 and Qwen-Image-Edit-2511 via ComfyUI, stable-diffusion.cpp, and diffusers, with 4bit/FP8/GGUF quantization; GGUF weights received "important-layers-first" quality updates.
  • Industry and Business

  • WSJ: Anthropic raising $10B at a potential $350B valuation: Up from $183B four months ago; among the largest private AI raises ever, with capital seen as going to compute and infrastructure rather than near-term revenue.
  • Google AI Studio sponsors TailwindCSS: After the "AI tools using open source without paying" controversy, developers are calling for token-usage-based or dependency-based revenue sharing for open-source projects.
  • Funding rounds: Autonomous (financial agents) raised $15M led by YC's Garry Tan; Protege AI (data infrastructure) raised $30M led by a16z.
  • NVIDIA skips new GPUs at CES for the first time in 5 years: RTX 50 Super rumors officially quashed; interpreted as a full pivot toward datacenter and AI chips. (Tom's Hardware)
---

📌 Source: Easy AI Daily 🤖 Compiled by: AI assistant

Tags

#ai-news#ai-daily#openai#glm-4.7#qwen3-vl#vllm#anthropic#open-source-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169146