Easy AI Daily | January 22, 2026
A daily digest of AI industry news compiled from community discussions. This post is a translation/summary of the original Chinese forum post.
Industry & Company News
- OpenEvidence raises $250M at $12B valuation: The "ChatGPT for doctors" medical LLM company raised $250M, roughly 12x its October 2025 valuation of $1B. The CEO says ~40% of US doctors use it and revenue exceeded $100M last year (~120x price-to-sales). CNBC
- Podium's "AI employees" surpass $100M ARR: The SMB SaaS company went from 0 to $100M ARR in 21 months with 10,000+ deployed AI agents handling missed calls and after-hours leads. Burn reportedly dropped from $95M to zero. Founder thread
- Runpod hits $120M ARR: The developer-focused GPU cloud started with a Reddit post in /r/LocalLLaMA four years ago. TechCrunch
- Lightning AI merges with Voltage Park: Co-led by William Falcon and Ozan Kaya; widely read as consolidating compute with MLOps to compete with Runpod. Announcement
- GPU price war: Voltage announced 2026 pricing of 8xA100 80GB at $6/hour and 2xRTX 5090 at $0.53/hour (up to 80% cheaper than AWS/RunPod/Vast.ai, with an OpenAI-compatible inference API). Spheron AI claims H100/H200/B200 pricing 40-60% below traditional clouds. Voltage pricing | Spheron
- Greg Yang moves to advisor role at xAI due to long-term Lyme disease fatigue; community concerned for both his health and xAI's theoretical research direction. Announcement
- Anthropic publishes Claude's new constitution, CC0-licensed: The values/behavior document used directly in training is open for reuse. Debate continues over whether it's a practical harm-reduction tool or alignment theater, and over the circularity of training a model on a document describing its own behavior. Post
- Jailbreak community keeps probing Gemini and Grok: BASI Jailbreaking's "Project Shadowfall" prompts got Gemini to teach pass-the-hash attacks; Grok is seen as more restrictive. Google Bughunters
- AI text classifiers face systematic adversarial attacks: Community-shared writeups demonstrate bypasses and adversarial models trained to mimic human writing; concerns that AI-writing detection fails easily against informed attackers. Blog post | Pangram research
- AirLLM claims 405B on 8GB VRAM: Via layer-by-layer load-compute-release streaming (optionally compressed). Viewed as an extreme paging demo—works, but with terrible latency/throughput. Project links
- Gemini 3: Education push with The Princeton Review (SAT practice) and Khan Academy (Writing Coach), alongside reports of image/video model instability (frequent errors on LMArena/OpenRouter).
- GPT-5.2: Thinking variant used for 20-30 minute reasoning sessions; leaked GPT-5 mini pricing at ~$0.25/M input tokens positions it as a strong budget model vs. Haiku 4.5 and Gemini 3 Fast.
- Agent benchmarks show big gaps: On Google's APEX-Agents (long Workspace tasks), best Pass@1 was Gemini 3 Flash High at 24%, GPT-5.2 High 23%, Claude Opus 4.5 18.4%. On legal-search benchmark prinzbench, search is the main weakness (GPT-5.2 Thinking just over 50%). APEX-Agents | prinzbench
- GLM-4.7-Flash integration incident: Broken across local frameworks (FlashAttention CPU fallback, 2.8 tok/s, infinite loops); fixed via a llama.cpp PR, with the model re-uploaded on Hugging Face. A typical example of model+inference-stack coordination failures. Fix PR | Model card
- Prefect Horizon: An enterprise "context layer" over MCP with managed deployment, tool registries, gateways, RBAC, and audit logs—MCP defines the protocol, not safe corporate operation. Intro
- LangChain Agent Builder GA & Deep Agents: Agents packaged as organized "folders" of skills, runnable locally or in the cloud; sub-agents for context isolation. Announcement
- MCP vs Skills: Hugging Face's Phil Schmid argues the problem is poorly designed MCP servers, not the protocol; design interfaces around outcomes, with strongly-typed, flat parameters and agent-readable errors. Thread
- Devin Review: Cognition's AI PR-review product reorders diffs by importance, flags duplicated code, and supports per-hunk chat; reviewers report it catches issues beyond the diff scope. Release
- GitHub Copilot CLI adds askUserQuestionTool: The CLI assistant now asks clarifying questions before acting, evolving toward a conversational agent. Intro
- GPU kernel optimization contest: Anthropic's performance takehome (VLIW kernel optimization) was tackled by GPU MODE/tinygrad communities—hand-written CUDA/Triton reached 2200 cycles, and Claude Opus 4.5 in Claude Code achieved ~1790 cycles, close to top human results. Problem repo
- PyTorch maintainers flooded by AI-generated PRs: Proposals include filtering via Claude/Pangram and Cursor Bugbot + GPT-5 Pro triage before human review.
- AMD AI Bundle: New Adrenalin driver packages PyTorch, ComfyUI, Ollama, LM Studio, and Amuse for Windows, lowering barriers for local inference on AMD GPUs. AMD blog
- Used high-end GPU prices soar: Used 3090s around €850 on eBay; a 5090 bought at £2000 relisted at £2659.99—local AI hardware is becoming an appreciating asset.
- NVIDIA ecosystem deep-dives: Blackwell warp/TMA utilization, NCCL all-reduce pipelining, nvshmem alternatives, HBM capacity—memory and interconnect increasingly seen as the real bottleneck, not just FLOPs. NCCL issue
- Google x Khan Academy Writing Coach: Gemini guides students through drafting and revision rather than writing for them. Announcement
- Runway Gen-4.5 image-to-video: Emphasis on character consistency, camera motion, and narrative continuity—evaluation shifting from single clips to multi-shot storytelling. Release
- LMArena milestones: Text Arena passed 5M cumulative votes; Video Arena launched on the web (3 generations/day, battle mode only). Video Arena
- Inforno & Soulbotix: OpenRouter-based multi-model desktop clients—Inforno for parallel chats with multiple LLMs; Soulbotix adds avatar interaction with local Whisper on RTX 4070 Ti-class GPUs. Inforno GitHub | Soulbotix
- AI adult content: Discussion of AI-generated virtual models competing with human creators on adult platforms, pushing humans toward IP, interaction, and offline experiences.
- DSPy RLM: Recasting agents as callable programs—large files in Python variables manipulated via function calls, turning context management into a code problem. Discussion
- Multi-vector retrieval: Mixedbread claims a 17M-parameter ColBERT-style model beats 8B single-vector embedding models on LongEmbed, with p50 < 50ms over 1B+ documents in production; TurboPuffer touts 100B-scale ANN indexing. Mixedbread | TurboPuffer
- NVIDIA TTT-E2E: Treats long context as online weight updates, giving inference time dependent on steps rather than context length—at the cost of weaker needle-in-haystack recall. Overview
- Cute/CUTLASS layout algebra: A long-read expresses shape/stride constraints via tuple transforms and refinement—kernel authors benefit from categorical foundations for correct memory access in complex tile combinations. Blog
- LM Studio grumbles: GLM-4.7-Flash widely reported as unusably slow/crashing on the new runtime; sentiment that aside from Qwen3, recent models feel iterative rather than a GPT-4-style leap.
- Manus.im complaints: Users report broken modules (only 20 of 38 working), regressions in Manus 1.6, and missing paid credits ($42 paid, 8000 promised points not delivered) with slow support.
- Coderrr: A free, open-source Claude Code alternative seeking community contributions. GitHub
- Aider's future in doubt: Slow updates spark fears the project is stalled; a community fork (Aider-CE) is adding MCP and agent capabilities.
Policy, Governance & Safety
Models & Capabilities
Agents & Tooling
Infrastructure & Hardware
Products & Applications
Research & Methods
Community Watch
*Source: Easy AI Daily, compiled with AI assistance.*