English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily News | January 22, 2026: OpenEvidence's $12B Round, GPU Price Wars, and Agent Benchmark Reality Check

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily for January 22, 2026 covers a fast-moving AI landscape. OpenEvidence, the "ChatGPT for doctors," raised $250M at a $12B valuation with ~40% of US physicians as users. Podium's AI-employee business passed $100M ARR in 21 months, while GPU cloud Runpod hit $120M ARR four years after starting as a Reddit post. Lightning AI merged with Voltage Park amid aggressive GPU price cuts from Voltage and Spheron. Anthropic released Claude's new constitution under a CC0 license. Agent benchmarks showed weak real-world reliability: best models scored only ~24% on Google's APEX-Agents and just over 50% on the prinzbench legal-search benchmark. Tooling news includes Prefect Horizon's enterprise MCP context layer, LangChain Deep Agents, Devin Review for AI code review, and Copilot CLI's new askUserQuestionTool. Research highlights: mixedbread's 17M-parameter multi-vector retrieval beating 8B embedding models, NVIDIA's TTT-E2E constant-time context approach, and community debate over GLM-4.7-Flash inference stack issues.

Easy AI Daily News | January 22, 2026

A structured translation of the Easy AI daily digest covering industry moves, policy, models, agent tooling, infrastructure, products, and research.

Key points

Industry & Company News

  • OpenEvidence raised $250M at a $12B valuation (roughly 12x its October 2025 valuation of $1B). The CEO claims ~40% of US physicians use it; revenue exceeded $100M last year (~120x price-to-sales). CNBC
  • Podium's customer-service/sales AI agent business surpassed $100M ARR in 21 months, with 10,000+ "AI employees" live; burn fell from $95M to zero. Founder thread
  • Runpod reached $120M ARR four years after launching from a /r/LocalLLaMA post. TechCrunch
  • Lightning AI merged with Voltage Park, co-led by William Falcon and Ozan Kaya, seen as consolidating compute plus MLOps to compete with Runpod. Announcement
  • GPU price war: Voltage offers 8×A100 80GB at $6/hr and 2×RTX 5090 at $0.53/hr (claims up to 80% cheaper than AWS/RunPod/Vast.ai); Spheron claims H100/H200/B200 at 40–60% below traditional cloud. Spheron
  • xAI researcher Greg Yang moved to an advisor role due to long-term Lyme disease complications.
  • Policy, Governance & Safety

  • Anthropic released Claude's new "constitution" under a CC0 license, calling it a living document directly used in training; community debates its effectiveness vs. "alignment theater" and the self-referential training loop. Post
  • The BASI Jailbreaking community continued bypassing Gemini (e.g., making it teach Pass-the-hash attacks) and Grok, redirecting findings to Google Bughunters.
  • Systematic adversarial attacks on AI text classifiers demonstrated; Pangram's detection paper holds up at scale, but real-world deployment looks fragile. Attack post
  • Models & Capabilities

  • AirLLM claims running 70B in 4GB and Llama3.1-405B in 8GB VRAM via layer-by-layer paging — technically possible but with terrible latency/throughput; a demo, not production.
  • Gemini 3: education push (SAT practice with The Princeton Review, Writing Coach with Khan Academy) alongside instability of Gemini 3 Pro image/video models on LMArena and OpenRouter.
  • GPT-5.2 Thinking runs 20–30 minute reasoning sessions; leaked GPT-5 mini pricing ~$0.25/M input tokens, positioning it against Haiku 4.5 and Gemini 3 Fast.
  • Agent benchmarks show weak reliability: Google APEX-Agents — Gemini 3 Flash High 24%, GPT-5.2 High 23%, Claude Opus 4.5 18.4%; prinzbench legal search — GPT-5.2 Thinking just over 50%, some Claude versions 0/24 on search. APEX results
  • GLM-4.7-Flash caused cross-framework issues (FlashAttention falling back to CPU, 2.8 tok/s, infinite loops); fixed in llama.cpp PR #18953, model re-uploaded on Hugging Face.
  • Agents & Tooling

  • Prefect Horizon: enterprise "context layer" over MCP with managed deployment, tool registry, gateway, RBAC, and audit logs.
  • LangChain Agent Builder GA plus Deep Agents — agents packaged as organized folders, with sub-agent context isolation.
  • Phil Schmid (Hugging Face): Skills vs. MCP is a false binary — the problem is badly designed MCP servers; design interfaces around outcomes, strongly-typed flat parameters, agent-readable errors.
  • Devin Review (Cognition): AI-powered PR review — re-ranks diffs, flags duplicated code, per-hunk chat.
  • GitHub Copilot CLI added an askUserQuestionTool so the assistant asks clarifying questions before acting — a shift toward conversational CLI agents.
  • New open-source Coderrr positions itself as a free Claude Code alternative (GitHub); the community worries Aider is stalling while the Aider-CE fork adds MCP/agent features.
  • Infrastructure & Hardware

  • Anthropic's VLIW kernel takehome became a leaderboard: hand-written CUDA/Triton reached 2200 cycles; Claude Opus 4.5 in Claude Code achieved ~1790 cycles, near top human results. Task repo
  • PyTorch maintainers flooded by low-quality AI-generated PRs; proposal to auto-filter with Claude/Pangram, then Cursor Bugbot + GPT-5 Pro triage before human review.
  • AMD AI Bundle in Adrenalin drivers ships one-click Windows installs of PyTorch, ComfyUI, Ollama, LM Studio, and Amuse.
  • Used high-end GPU prices keep climbing (3090 ~€850 on eBay; 5090 resale up to £2,659) — local AI hardware is becoming an "appreciating asset."
  • GPU MODE discussions on Blackwell warp/TMA utilization, NCCL all-reduce pipelining (issue), nvshmem, and HBM supply point to memory and interconnect as the real bottlenecks, not FLOPs.
  • Products & Applications

  • Google × Khan Academy Writing Coach guides students through drafting and revising rather than writing for them.
  • Runway Gen-4.5 image-to-video emphasizes character consistency, camera motion, and narrative continuity; evaluation is shifting from single-clip quality to multi-shot storytelling.
  • LMArena's Text Arena passed 5M votes; Video Arena opened web access with 3 generations/day, battle mode only. Video Arena
  • OpenRouter-based desktop frontends: Inforno (multi-model chat, .rno history files, GitHub) and Soulbotix (virtual-avatar Windows client with local Whisper ASR, site).
  • Discussion of AI-generated virtual "OnlyFans models" undercutting human creators on cost and availability.
  • Research & Methods

  • DSPy RLM: treat large contexts as Python variables manipulated via function calls, turning context management into a code problem.
  • Mixedbread's 17M-parameter multi-vector (ColBERT-style) model beats 8B single-vector embeddings on long-document benchmarks (LongEmbed), with p50 < 50ms serving 1B+ documents; TurboPuffer announces 100B-vector ANN indexing. Multi-vector late-interaction is winning recall — if you have heavy retrieval infrastructure.
  • NVIDIA TTT-E2E: treat context as online training data so inference time scales with steps, not context length — at the cost of weaker needle-in-haystack recall.
  • Cute/CUTLASS layout algebra: a categorical foundations article argues kernel writers need layout-algebra intuition for correct memory access under complex tiling.
  • Community Gripes

  • LM Studio users report GLM-4.7-Flash crashes/slowness; sentiment that post-Qwen3 models lack a GPT-4-style leap.
  • Manus.im complaints: of 38 built modules only ~20 still work, Manus 1.6 regression, and a $42 upgrade's promised 8,000 credits not delivered — risky signals for an "engineering agent" product.
---

📌 Source: Easy AI Daily | 🤖 Compiled by: AI assistant

Tags

#ai-news#daily-digest#openevidence#gpu-cloud#agents#mcp#anthropic#llm-benchmarks

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169124