English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily News Digest | January 22, 2026: OpenEvidence Funding, GPU Price Wars, Agent Benchmarks

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily for January 22, 2026 covers major AI industry developments: OpenEvidence, the medical LLM dubbed ChatGPT for doctors, raised $250M at a $12B valuation with roughly 40% of US physicians using it; SMB SaaS firm Podium surpassed $100M ARR in AI agent revenue within 21 months; GPU cloud Runpod reached $120M ARR, while Lightning AI merged with Voltage Park amid aggressive GPU price cuts (Voltage offering 8xA100 at $6/hour). Anthropic open-sourced Claude's new constitution under CC0. In research and tooling, agent benchmarks showed models far from human-level reliability (Gemini 3 Flash scored only 24% on Google's APEX-Agents), Mixedbread claimed a 17M-parameter multi-vector model beats 8B embedding models, and NVIDIA's TTT-E2E proposes constant-time inference over long contexts. Additional coverage includes Prefect Horizon for enterprise MCP management, LangChain Deep Agents, Devin Review for AI code review, AMD's one-click local AI bundle, and rising used GPU prices.

Easy AI Daily | January 22, 2026

A daily digest of AI industry news compiled from community discussions. This post is a translation/summary of the original Chinese forum post.

Industry & Company News

  • OpenEvidence raises $250M at $12B valuation: The "ChatGPT for doctors" medical LLM company raised $250M, roughly 12x its October 2025 valuation of $1B. The CEO says ~40% of US doctors use it and revenue exceeded $100M last year (~120x price-to-sales). CNBC
  • Podium's "AI employees" surpass $100M ARR: The SMB SaaS company went from 0 to $100M ARR in 21 months with 10,000+ deployed AI agents handling missed calls and after-hours leads. Burn reportedly dropped from $95M to zero. Founder thread
  • Runpod hits $120M ARR: The developer-focused GPU cloud started with a Reddit post in /r/LocalLLaMA four years ago. TechCrunch
  • Lightning AI merges with Voltage Park: Co-led by William Falcon and Ozan Kaya; widely read as consolidating compute with MLOps to compete with Runpod. Announcement
  • GPU price war: Voltage announced 2026 pricing of 8xA100 80GB at $6/hour and 2xRTX 5090 at $0.53/hour (up to 80% cheaper than AWS/RunPod/Vast.ai, with an OpenAI-compatible inference API). Spheron AI claims H100/H200/B200 pricing 40-60% below traditional clouds. Voltage pricing | Spheron
  • Greg Yang moves to advisor role at xAI due to long-term Lyme disease fatigue; community concerned for both his health and xAI's theoretical research direction. Announcement
  • Policy, Governance & Safety

  • Anthropic publishes Claude's new constitution, CC0-licensed: The values/behavior document used directly in training is open for reuse. Debate continues over whether it's a practical harm-reduction tool or alignment theater, and over the circularity of training a model on a document describing its own behavior. Post
  • Jailbreak community keeps probing Gemini and Grok: BASI Jailbreaking's "Project Shadowfall" prompts got Gemini to teach pass-the-hash attacks; Grok is seen as more restrictive. Google Bughunters
  • AI text classifiers face systematic adversarial attacks: Community-shared writeups demonstrate bypasses and adversarial models trained to mimic human writing; concerns that AI-writing detection fails easily against informed attackers. Blog post | Pangram research
  • Models & Capabilities

  • AirLLM claims 405B on 8GB VRAM: Via layer-by-layer load-compute-release streaming (optionally compressed). Viewed as an extreme paging demo—works, but with terrible latency/throughput. Project links
  • Gemini 3: Education push with The Princeton Review (SAT practice) and Khan Academy (Writing Coach), alongside reports of image/video model instability (frequent errors on LMArena/OpenRouter).
  • GPT-5.2: Thinking variant used for 20-30 minute reasoning sessions; leaked GPT-5 mini pricing at ~$0.25/M input tokens positions it as a strong budget model vs. Haiku 4.5 and Gemini 3 Fast.
  • Agent benchmarks show big gaps: On Google's APEX-Agents (long Workspace tasks), best Pass@1 was Gemini 3 Flash High at 24%, GPT-5.2 High 23%, Claude Opus 4.5 18.4%. On legal-search benchmark prinzbench, search is the main weakness (GPT-5.2 Thinking just over 50%). APEX-Agents | prinzbench
  • GLM-4.7-Flash integration incident: Broken across local frameworks (FlashAttention CPU fallback, 2.8 tok/s, infinite loops); fixed via a llama.cpp PR, with the model re-uploaded on Hugging Face. A typical example of model+inference-stack coordination failures. Fix PR | Model card
  • Agents & Tooling

  • Prefect Horizon: An enterprise "context layer" over MCP with managed deployment, tool registries, gateways, RBAC, and audit logs—MCP defines the protocol, not safe corporate operation. Intro
  • LangChain Agent Builder GA & Deep Agents: Agents packaged as organized "folders" of skills, runnable locally or in the cloud; sub-agents for context isolation. Announcement
  • MCP vs Skills: Hugging Face's Phil Schmid argues the problem is poorly designed MCP servers, not the protocol; design interfaces around outcomes, with strongly-typed, flat parameters and agent-readable errors. Thread
  • Devin Review: Cognition's AI PR-review product reorders diffs by importance, flags duplicated code, and supports per-hunk chat; reviewers report it catches issues beyond the diff scope. Release
  • GitHub Copilot CLI adds askUserQuestionTool: The CLI assistant now asks clarifying questions before acting, evolving toward a conversational agent. Intro
  • Infrastructure & Hardware

  • GPU kernel optimization contest: Anthropic's performance takehome (VLIW kernel optimization) was tackled by GPU MODE/tinygrad communities—hand-written CUDA/Triton reached 2200 cycles, and Claude Opus 4.5 in Claude Code achieved ~1790 cycles, close to top human results. Problem repo
  • PyTorch maintainers flooded by AI-generated PRs: Proposals include filtering via Claude/Pangram and Cursor Bugbot + GPT-5 Pro triage before human review.
  • AMD AI Bundle: New Adrenalin driver packages PyTorch, ComfyUI, Ollama, LM Studio, and Amuse for Windows, lowering barriers for local inference on AMD GPUs. AMD blog
  • Used high-end GPU prices soar: Used 3090s around €850 on eBay; a 5090 bought at £2000 relisted at £2659.99—local AI hardware is becoming an appreciating asset.
  • NVIDIA ecosystem deep-dives: Blackwell warp/TMA utilization, NCCL all-reduce pipelining, nvshmem alternatives, HBM capacity—memory and interconnect increasingly seen as the real bottleneck, not just FLOPs. NCCL issue
  • Products & Applications

  • Google x Khan Academy Writing Coach: Gemini guides students through drafting and revision rather than writing for them. Announcement
  • Runway Gen-4.5 image-to-video: Emphasis on character consistency, camera motion, and narrative continuity—evaluation shifting from single clips to multi-shot storytelling. Release
  • LMArena milestones: Text Arena passed 5M cumulative votes; Video Arena launched on the web (3 generations/day, battle mode only). Video Arena
  • Inforno & Soulbotix: OpenRouter-based multi-model desktop clients—Inforno for parallel chats with multiple LLMs; Soulbotix adds avatar interaction with local Whisper on RTX 4070 Ti-class GPUs. Inforno GitHub | Soulbotix
  • AI adult content: Discussion of AI-generated virtual models competing with human creators on adult platforms, pushing humans toward IP, interaction, and offline experiences.
  • Research & Methods

  • DSPy RLM: Recasting agents as callable programs—large files in Python variables manipulated via function calls, turning context management into a code problem. Discussion
  • Multi-vector retrieval: Mixedbread claims a 17M-parameter ColBERT-style model beats 8B single-vector embedding models on LongEmbed, with p50 < 50ms over 1B+ documents in production; TurboPuffer touts 100B-scale ANN indexing. Mixedbread | TurboPuffer
  • NVIDIA TTT-E2E: Treats long context as online weight updates, giving inference time dependent on steps rather than context length—at the cost of weaker needle-in-haystack recall. Overview
  • Cute/CUTLASS layout algebra: A long-read expresses shape/stride constraints via tuple transforms and refinement—kernel authors benefit from categorical foundations for correct memory access in complex tile combinations. Blog
  • Community Watch

  • LM Studio grumbles: GLM-4.7-Flash widely reported as unusably slow/crashing on the new runtime; sentiment that aside from Qwen3, recent models feel iterative rather than a GPT-4-style leap.
  • Manus.im complaints: Users report broken modules (only 20 of 38 working), regressions in Manus 1.6, and missing paid credits ($42 paid, 8000 promised points not delivered) with slow support.
  • Coderrr: A free, open-source Claude Code alternative seeking community contributions. GitHub
  • Aider's future in doubt: Slow updates spark fears the project is stalled; a community fork (Aider-CE) is adding MCP and agent capabilities.
---

*Source: Easy AI Daily, compiled with AI assistance.*

Tags

#ai-news#daily-digest#open-evidence#gpu-cloud#ai-agents#mcp#llm-benchmarks#ai-safety

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169147