Easy AI Daily News | January 9, 2026
Policy, Governance and Safety
- OpenAI launches ChatGPT Health / OpenAI for Healthcare: A medical product line claimed to be HIPAA-compliant, used for clinical Q&A, documentation, and knowledge retrieval. Already live at AdventHealth, UCSF, MSK, and HCA. OpenAI says physician AI usage "doubled in a year," though privacy concerns and fears of "AI replacing doctors" persist.
- OpenAI healthcare announcement | ChatGPT Health intro
- Stanford paper: copyrighted content extractable from frontier LLMs: Researchers say multiple online models memorize training data at scale and can emit copyrighted excerpts in certain settings. Claude 3.7 Sonnet reportedly reproduced ~95.8% of the first Harry Potter book, while GPT-4.1 was much lower — countering claims that LLMs don't memorize training data.
- Paper summary thread
- Zhipu GLM-4.7 tops open-source rankings: Scores 42 on Artificial Analysis' index (up 10 from 4.6), leading in coding, agents, and science reasoning, with the highest GDPval-AA ELO among evaluated open models. Specs: 355B MoE (32B active), 200k context, MIT license; BF16 weights ~710GB — too large even for a single 8×H100 node. Parent company Z.ai announced a Hong Kong IPO listing.
- GLM-4.7 benchmarks | Z.ai milestone
- Alibaba Qwen3-VL multimodal embedding and reranker: Two-stage retrieval supporting text, images, screenshots, video, 30+ languages, adjustable embedding dimensions, instruction-tuning, and quantized deployment. Tops MMEB-V2 and MMTEB benchmarks per official claims. Available on Hugging Face and ModelScope; supported in vLLM nightly.
- Official intro | vLLM support
- Baidu ERNIE-5.0 and Tencent Hunyuan-Video-1.5 enter LMArena: ERNIE-5.0-Preview-1220 scored 1226, ranking 8th on the Vision leaderboard (currently the only Chinese lab in the top 10). Hunyuan-Video-1.5 placed 18th in text-to-video and 20th in image-to-video.
- Vision leaderboard | Video leaderboard
- AI21 open-sources Jamba2: Enterprise-focused hybrid SSM-Transformer with KV-cache optimization, Apache 2.0 licensed, available via AI21 cloud and Hugging Face.
- TII Falcon-H1R-7B: A small hybrid Transformer-Mamba reasoning model performing well on Humanity's Last Exam, τ²-Bench Telecom, and IFBench; openness score of 44.
- Lightricks open-sources LTX-2: A locally runnable audio-video generation model with weights, code, trainers, LoRA, and docs; runs on consumer GPUs. NSFW/copyright restrictions on training data.
- LTX-2 model page
- Gemini 3 shines on PokerBench: 21,000 hands of Texas Hold'em; Gemini 3 Pro ended most profitable overall, though one developer noted Flash won head-to-head, suggesting luck factors. Data and code are open.
- PokerBench
- vLLM + B200 hits ~16k token/s: The community merged a KV Offloading Connector (with IBM Research) that asynchronously offloads KV cache to CPU memory. Up to 9× throughput gains on H100 and 2–22× TTFT reductions in cache-hit scenarios.
- B200 milestone
- AI-generated kernels enter vLLM: An LLM-generated fused RMSNorm kernel ("Oink") gives ~40% single-kernel speedup and ~1.6% end-to-end gain, with near-autotuning for hot shapes (e.g., 7168 BF16) — plus added crash/stability complexity.
- Technical writeup
- CuteDSL flex attention ~30% faster on H100: Forward-pass speedup integrated into existing frameworks; backward support on SM90 progressing via Flash-Attention PR #2137.
- Hugging Face Transformers v5: Unified tokenizer backend, PyTorch-centric model definitions, and major upgrades to serving, quantization, and deployment; plus swift-huggingface and AnyLanguageModel for Apple ecosystems.
- Transformers v5 blog
- Epoch: global compute exceeds 15M H100-equivalents: AI chip power alone tops 10GW, excluding other datacenter overhead. An "AI Chip Sales" visualization tool tracks supply chains.
- Epoch data
- LangChain + VS Code: agents as folders and skills: Harrison Chase proposes describing agents via file structures (agents.md, subagents/, skills.md, mcp.json). VS Code launched "Agent Skills" based on Anthropic's open standard, enabled via chat.useAgentSkills.
- DSPy to rework multi-turn conversation handling: Conversation history currently gets stuffed into the system prompt; maintainers say adapters are an implementation detail and multi-turn history serialization will be redone as a first-class configuration option.
- MCP community discusses standardized staging for side-effectful tools: A dry-run layer before state-changing tool calls for audit and confirmation — possibly as a SEP, or as SDK best practice; W3C WebMCP cooperation also discussed.
- Claude Code power-user patterns: Hooks reading local routing files, error-log systems collecting failed prompts, /commands as local mini-apps, forcing subagents to Opus, and strict context compaction for large projects.
- MAGMA: Splits agent memory into semantic, temporal, causal, and entity graphs with policy-controlled retrieval instead of single-pass vector search; clear gains on LoCoMo and LongMemEval.
- SPOT @ ICLR 2026: Workshop on scaling post-training (SFT/RLHF); submissions due February 5.
- Artificial Analysis on evaluation methodology: Real-world GDPval-AA tasks plus an Openness Index covering weights, data, and deployment restrictions; discussion of sensitivity, brittleness, and "mystery shopper" testing.
- "Dead salmon effect" returns: A paper shows many interpretation methods (feature attribution, probes, sparse autoencoders, causal analysis) produce plausible-looking explanations even on randomly initialized networks (arXiv:2512.18792) — urging caution with interpretability results.
- Gmail enters the Gemini era: AI conversation summaries, replies and polishing, AI Inbox view, and natural-language mailbox search with user toggles; security researchers see anti-phishing potential but warn about trusted-agent manipulation.
- Local LLM practice with GLM-4.7: Reddit users report using it instead of Claude Sonnet 4.5 for debugging and refactoring, with 85–90% usable code rates and roughly one-fifth the API cost; Sonnet still preferred for design-level discussion.
- Qwen-Image runs locally in ~14GB: Tutorials cover Qwen-Image-2512 and Qwen-Image-Edit-2511 via ComfyUI, stable-diffusion.cpp, and diffusers, with 4bit/FP8/GGUF quantization; GGUF weights received "important-layers-first" quality updates.
- WSJ: Anthropic raising $10B at a potential $350B valuation: Up from $183B four months ago; among the largest private AI raises ever, with capital seen as going to compute and infrastructure rather than near-term revenue.
- Google AI Studio sponsors TailwindCSS: After the "AI tools using open source without paying" controversy, developers are calling for token-usage-based or dependency-based revenue sharing for open-source projects.
- Funding rounds: Autonomous (financial agents) raised $15M led by YC's Garry Tan; Protege AI (data infrastructure) raised $30M led by a16z.
- NVIDIA skips new GPUs at CES for the first time in 5 years: RTX 50 Super rumors officially quashed; interpreted as a full pivot toward datacenter and AI chips. (Tom's Hardware)
Models and Capabilities
Infrastructure and Hardware
Agents and Tooling
Research and Methods
Products and Applications
Industry and Business
📌 Source: Easy AI Daily 🤖 Compiled by: AI assistant