Easy AI Daily News | January 22, 2026
A structured translation of the Easy AI daily digest covering industry moves, policy, models, agent tooling, infrastructure, products, and research.
Key points
Industry & Company News
- OpenEvidence raised $250M at a $12B valuation (roughly 12x its October 2025 valuation of $1B). The CEO claims ~40% of US physicians use it; revenue exceeded $100M last year (~120x price-to-sales). CNBC
- Podium's customer-service/sales AI agent business surpassed $100M ARR in 21 months, with 10,000+ "AI employees" live; burn fell from $95M to zero. Founder thread
- Runpod reached $120M ARR four years after launching from a /r/LocalLLaMA post. TechCrunch
- Lightning AI merged with Voltage Park, co-led by William Falcon and Ozan Kaya, seen as consolidating compute plus MLOps to compete with Runpod. Announcement
- GPU price war: Voltage offers 8×A100 80GB at $6/hr and 2×RTX 5090 at $0.53/hr (claims up to 80% cheaper than AWS/RunPod/Vast.ai); Spheron claims H100/H200/B200 at 40–60% below traditional cloud. Spheron
- xAI researcher Greg Yang moved to an advisor role due to long-term Lyme disease complications.
- Anthropic released Claude's new "constitution" under a CC0 license, calling it a living document directly used in training; community debates its effectiveness vs. "alignment theater" and the self-referential training loop. Post
- The BASI Jailbreaking community continued bypassing Gemini (e.g., making it teach Pass-the-hash attacks) and Grok, redirecting findings to Google Bughunters.
- Systematic adversarial attacks on AI text classifiers demonstrated; Pangram's detection paper holds up at scale, but real-world deployment looks fragile. Attack post
- AirLLM claims running 70B in 4GB and Llama3.1-405B in 8GB VRAM via layer-by-layer paging — technically possible but with terrible latency/throughput; a demo, not production.
- Gemini 3: education push (SAT practice with The Princeton Review, Writing Coach with Khan Academy) alongside instability of Gemini 3 Pro image/video models on LMArena and OpenRouter.
- GPT-5.2 Thinking runs 20–30 minute reasoning sessions; leaked GPT-5 mini pricing ~$0.25/M input tokens, positioning it against Haiku 4.5 and Gemini 3 Fast.
- Agent benchmarks show weak reliability: Google APEX-Agents — Gemini 3 Flash High 24%, GPT-5.2 High 23%, Claude Opus 4.5 18.4%; prinzbench legal search — GPT-5.2 Thinking just over 50%, some Claude versions 0/24 on search. APEX results
- GLM-4.7-Flash caused cross-framework issues (FlashAttention falling back to CPU, 2.8 tok/s, infinite loops); fixed in llama.cpp PR #18953, model re-uploaded on Hugging Face.
- Prefect Horizon: enterprise "context layer" over MCP with managed deployment, tool registry, gateway, RBAC, and audit logs.
- LangChain Agent Builder GA plus Deep Agents — agents packaged as organized folders, with sub-agent context isolation.
- Phil Schmid (Hugging Face): Skills vs. MCP is a false binary — the problem is badly designed MCP servers; design interfaces around outcomes, strongly-typed flat parameters, agent-readable errors.
- Devin Review (Cognition): AI-powered PR review — re-ranks diffs, flags duplicated code, per-hunk chat.
- GitHub Copilot CLI added an
askUserQuestionToolso the assistant asks clarifying questions before acting — a shift toward conversational CLI agents. - New open-source Coderrr positions itself as a free Claude Code alternative (GitHub); the community worries Aider is stalling while the Aider-CE fork adds MCP/agent features.
- Anthropic's VLIW kernel takehome became a leaderboard: hand-written CUDA/Triton reached 2200 cycles; Claude Opus 4.5 in Claude Code achieved ~1790 cycles, near top human results. Task repo
- PyTorch maintainers flooded by low-quality AI-generated PRs; proposal to auto-filter with Claude/Pangram, then Cursor Bugbot + GPT-5 Pro triage before human review.
- AMD AI Bundle in Adrenalin drivers ships one-click Windows installs of PyTorch, ComfyUI, Ollama, LM Studio, and Amuse.
- Used high-end GPU prices keep climbing (3090 ~€850 on eBay; 5090 resale up to £2,659) — local AI hardware is becoming an "appreciating asset."
- GPU MODE discussions on Blackwell warp/TMA utilization, NCCL all-reduce pipelining (issue), nvshmem, and HBM supply point to memory and interconnect as the real bottlenecks, not FLOPs.
- Google × Khan Academy Writing Coach guides students through drafting and revising rather than writing for them.
- Runway Gen-4.5 image-to-video emphasizes character consistency, camera motion, and narrative continuity; evaluation is shifting from single-clip quality to multi-shot storytelling.
- LMArena's Text Arena passed 5M votes; Video Arena opened web access with 3 generations/day, battle mode only. Video Arena
- OpenRouter-based desktop frontends: Inforno (multi-model chat, .rno history files, GitHub) and Soulbotix (virtual-avatar Windows client with local Whisper ASR, site).
- Discussion of AI-generated virtual "OnlyFans models" undercutting human creators on cost and availability.
- DSPy RLM: treat large contexts as Python variables manipulated via function calls, turning context management into a code problem.
- Mixedbread's 17M-parameter multi-vector (ColBERT-style) model beats 8B single-vector embeddings on long-document benchmarks (LongEmbed), with p50 < 50ms serving 1B+ documents; TurboPuffer announces 100B-vector ANN indexing. Multi-vector late-interaction is winning recall — if you have heavy retrieval infrastructure.
- NVIDIA TTT-E2E: treat context as online training data so inference time scales with steps, not context length — at the cost of weaker needle-in-haystack recall.
- Cute/CUTLASS layout algebra: a categorical foundations article argues kernel writers need layout-algebra intuition for correct memory access under complex tiling.
- LM Studio users report GLM-4.7-Flash crashes/slowness; sentiment that post-Qwen3 models lack a GPT-4-style leap.
- Manus.im complaints: of 38 built modules only ~20 still work, Manus 1.6 regression, and a $42 upgrade's promised 8,000 credits not delivered — risky signals for an "engineering agent" product.
Policy, Governance & Safety
Models & Capabilities
Agents & Tooling
Infrastructure & Hardware
Products & Applications
Research & Methods
Community Gripes
📌 Source: Easy AI Daily | 🤖 Compiled by: AI assistant