Easy AI Daily Digest | February 28, 2026
A roundup of AI industry news, model releases, infrastructure updates, research, and policy developments from the zhichai.net community.
Industry & Company News
OpenAI Completes Record $110B Round at ~$840B Valuation
OpenAI announced a $110 billion financing round (pre-money ~$730B, post-money ~$840B). SoftBank invested $30B; NVIDIA invested $30B plus 3GW inference and 2GW training compute; Amazon committed $50B total with deepened cloud cooperation, and AWS will be the exclusive third-party cloud for OpenAI Frontier. Microsoft did not join this round but maintains a streamlined partnership.- OpenAI funding announcement | Partnership details | Amazon partnership | Microsoft joint statement | EpochAI analysis
- Anthropic statement | Axios report | NPR report | Legal/policy discussion
- Reuters
- Artificial Analysis | Qwen blog
- Official blog
- Experiments on Hugging Face
- Paper
- CuTeDSL examples | cuTile docs
- Notes
- Blog
- Model
- Paper
- Design write-up
- MCP ping spec
- Onyx leaderboard
- SWE-fficiency Benchmark
Anthropic Confronts the Pentagon, Threatened with "Supply Chain Risk" Label
Anthropic refused to provide Defense Department versions of Claude usable for large-scale domestic surveillance and fully autonomous weapons, rejected an ultimatum, and said it will go to court if labeled a supply chain risk. Reports say the DoD is considering requiring contractor audits and discontinuation of Anthropic services, sparking industry concern about government overreach. Many users posted subscriptions in support of Claude.Sam Altman: OpenAI Shares Anthropic's "Red Lines"
Per Axios, Altman said OpenAI aligns with Anthropic on opposing mass surveillance and fully autonomous lethal weapons. OpenAI is negotiating with the DoD using cloud hosting and technical guardrails to control military usage.DeepSeek V4 Prioritizes Domestic Hardware; NVIDIA/AMD Denied Early Access
DeepSeek reportedly gave Huawei and other Chinese chipmakers early access to V4 for adaptation, while NVIDIA and AMD have not received the same treatment. Analysts note DeepSeek still relies heavily on NVIDIA for training; this is seen as catching up on non-NVIDIA hardware support rather than a boycott.Burger King Pilots Voice AI "Patty" That Scores Employee Politeness
Burger King is testing BK Assistant in 500 US stores: an OpenAI-based voice bot on employee headsets that answers recipe questions and tracks whether staff say greetings and polite phrases, generating a store "friendliness" score. The community worries it is workplace surveillance disguised as training.Models & Capabilities
OpenAI Discloses Metrics: 900M Weekly ChatGPT Users, 9M Paying Businesses
First systematic disclosure: Codex weekly actives hit 1.6M (3x since start of year), ChatGPT surpassed 900M weekly actives, 50M+ individual subscriptions, and 9M+ paying business customers, up sharply from 3M in mid-2025.Qwen3.5 Series Expands with Strong Open-Weight Models
New additions include a 27B dense model, 122B A10B MoE, and 35B A3B MoE under Apache 2.0, with 260K context (extendable to 1M). Artificial Analysis scored up to 42 on its Intelligence Index, approaching closed-source flagships on some metrics.LMArena Leaderboards: GLM-5, Qwen3.5, Kimi-K2.5 Lead Open-Source
Text top-3: GLM-5, Qwen3.5-397B A17B, Kimi-K2.5 Thinking. Code: GLM-5 first; Kimi-K2.5 tied with MiniMax-M2.5 second. Arena also open-sourced its Arena-Rank tool.Google Nano Banana 2 (Gemini 3.1 Flash Image) Released
Community testing shows notable improvements in spatial layout and proportions, though text hallucination persists. Pricing: $0.50 input / $3.00 output — about half the Pro version.Qwen3.5-35B-A3B GGUF Quantization Deep-Dive
Unsloth ran hundreds of GGUF quantization experiments with full PPL/KL metrics and 9TB of model files. Findings: KV q8_0 is nearly a "free lunch" for throughput; the 35B MoE runs ~10x faster than the 27B dense on a single card, hitting 60+ tok/s on a local 4070S.Doc-to-LoRA & Text-to-LoRA: Sakana "Compiles" LoRAs in One Pass
Sakana AI's hypernetwork generates LoRA weights directly from natural language descriptions or long documents, effectively one-forward-pass "fine-tuning." Claims include encoding document knowledge into adapters beyond raw context windows and cross-modal knowledge transfer from VLMs to text models.Local/Cloud Inference Speeds: Qwen3.5-35B and GPT-OSS 20B Real TPS
Qwen3.5-35B MoE Q4_K_M hits ~62 tok/s on a 4070 Super and ~25 tok/s on 7900XT 16GB; GPT-OSS 20B runs ~100 tok/s locally on MacBooks (1M tokens in ~3 hours), fueling interest in hybrid local+API setups.Infrastructure & Hardware
vLLM on AMD ROCm: Up to 4.4x Decode Throughput Gains
Seven new attention backends plus KV cache layout and batching optimizations deliver up to 4.4x decode speedups on MI300X-class GPUs viaVLLM_ROCM_USE_AITER=1. Combined with MLA KV compression, ~8K dimensions compress to 576.DeepSeek DualPath: Splitting KV Cache I/O Bottlenecks with RDMA
A PKU/Tsinghua/DeepSeek paper proposes an inference system where RDMA lets prefill and decode nodes cooperatively use idle storage and network bandwidth, claiming ~1.9x speedups on models like DeepSeek 660B for agentic workloads.Google Colab Adds RTX PRO 6000 at ~$0.81/hour
Roughly an order of magnitude cheaper than the old A100 high-memory tier (~$7.5/h in credits), making Colab viable again for individual pretraining/fine-tuning.GPU MODE: PTX Memory Models, cuTile, and CuTeDSL
Developers debated PTX acquire-release memory semantics vs. distributed consistency models, and explored cuTile/CuTeDSL with multimem-based reduce-scatter in CUTLASS examples toward fused compute+communication training kernels.Research & Methods
Logit Fusion Gains Momentum
Discussions around fusing logits from multiple models/checkpoints during training — effectively ensembling + curriculum in the training loop with no inference overhead — with calls to make it a first-class training method like LoRA.NNsight 0.6 Released
Interpretability intervention tracing is 2.4–3.9x faster, with vLLM multi-GPU/multi-node support, vision-language and diffusion model support on Hugging Face, plus AI-friendly docs for automated probe writing.CoDA-GQA-L: Compressing KV Cache with Two Triton Kernels
An open-source attention design using "landmark banks + grouped queries" significantly cuts KV cache memory (Mistral-7B demo available). Eleuther analysis notes accuracy costs, however: replacing 32 attention layers while fine-tuning 18.6% of parameters raised PPL from 4.81 to 5.75.World Models Salon: From Sora to V-JEPA, "Mirror vs. Map"
Chipro MLOps is hosting two paper clinics on "Understanding World or Predicting Future?" covering JEPA/V-JEPA, Dreamer, Genie, Sora, and World Labs, discussing generative vs. representational approaches, spatial intelligence, and causal modeling.Benchmark Debates: Does Chain-of-Thought Count as "Template Bias"?
Eleuther community debates whether multi-turn CoT few-shot prompting is a form of template gaming, and whether benchmarks should simulate messy real user input or controlled comparisons.Agents & Tooling
LLM Connection Strings: Model Config as a URL
Dan Levy's proposal formats provider, model, and parameters as URIs likellm://provider/model?..., letting scripts and agents switch models/routing with a single argument.
MCP ping Semantics Pitfall: Must You initialize First?
The Python SDK requires initialization before ping, but the spec's "still available" wording implies existing connections. Bedrock AgentCore works around it by creating a temporary session to ping, highlighting spec ambiguity affecting production.OpenClaw / Cursor / Claude Code: Real-World Multi-Agent Coding
One engineer shipped 118 commits/day over 72 days using 5–10 agents, but reports context loss and file-change conflicts. Cursor community warns "vibe coding" doesn't suit serious projects; many find a hand-written orchestrator with cheaper models more cost-effective than Claude Code's built-in orchestration.Products & Applications
Local LLM Selection Tools: LLmFit and Onyx Launch
LLmFit recommends models for your hardware in one command (users report tok/s estimates often mismatch reality); Onyx offers a self-hosted LLM leaderboard scoring coding, math, reasoning, and efficiency.Claude Code at Scale: 76K Lines, 118 Functions Slower Than They Should Be
Codeflash built two large features (76K lines) with Claude Code, then found 118 functions optimizable by up to 446x — typical issues being naive algorithms, repeated computation, and poor data structures. Lesson: LLMs write correct code first but rarely self-profile; benchmarks and code review must be mandatory in the workflow.Nano Banana 2 in Production: Better Spatial Understanding, Still Hallucinates Text
Reddit comparisons show big gains in spatial sense and complex scenes (e.g., interior redesign), but text tasks still hallucinate. Banana 2's lower price suits high-volume generation; Banana Pro remains preferred for high-fidelity, character-consistent work.Policy, Governance & Safety
DoD vs. Anthropic: Supply Chain Labels and a Chilling Effect
Beyond the contract dispute, reports suggest the DoD may designate Anthropic a "national security supply chain risk" and require contractors to assess Claude usage. Legal observers note the DoD can restrict contractor usage on military projects but can't easily ban commercial use. The case is seen as a landmark precedent for who draws the line on "acceptable use" in AI governance.Jailbreaks and Safety Filters: Ongoing Escalation
BASI community continues probing Gemini Pro 3/3.1, Grok, Claude, and ChatGPT with new jailbreak prompts. Some say models are getting much harder to jailbreak; others claim new short-prompt bypasses — indicating both stronger alignment and persistent attacks.---
📌 Source: Easy AI Daily (zhichai.net)