Key points
Models and Capabilities
- Google Gemini 3.1 Flash-Lite (preview): Positioned as the lowest-latency, highest-throughput multimodal model in the Gemini 3 family. 1M context, ~360+ token/s measured, ~5.1s average response. Jeff Dean cited pricing of ~$0.25/M input and $1.50/M output tokens — roughly 2.5–3.75x pricier than 2.5 Flash-Lite, sparking value debates. LMArena Elo: 1432. Release thread, Arena leaderboard, Cost analysis
- OpenAI GPT-5.3 Instant: Rolled out to all ChatGPT users; claims more natural tone, fewer unnecessary refusals, hallucination reductions of 26.8% (with search) and 19.7% (without).
gpt-5.3-chat-latestappeared in the API. OpenAI teased "GPT-5.4 sooner than you think," which some read as deflection from the DoD/NSA contract controversy. Announcement, GPT-5.4 teaser - Alibaba Qwen 3.5: The 0.8B model with a vision encoder runs in-browser via WebGPU and on phones (~12 token/s on older hardware); 27B/35B variants perform near larger models with linear attention efficiency gains. Community still warns of notable hallucination risks. WebGPU demo discussion
- Apple M5 Pro / M5 Max: M5 Pro supports 64GB unified memory at 307GB/s; M5 Max 128GB at 614GB/s. Apple claims up to 4x faster LLM prompt processing vs M4 Pro/Max; SSD throughput up to 14.5GB/s; integrated N1 chip adds Wi-Fi 7. LocalLLaMA discussion
- Together AI: Context Parallel + sequence parallelism trains 5M-context 8B models on a single 8×H100 node, cutting attention memory up to 87%.
- Databricks FlashOptim (open-sourced): Reduces AdamW memory from ~16 to 7 bytes/param; 8B fine-tuning peak memory drops from 175GiB to 113GiB.
- SkyPilot Job Groups: Splits RL training across high-end GPUs, cheap GPUs, and large-memory CPUs.
- NVIDIA Blackwell split: Data center (CC 10.0) vs consumer (CC 12.0) lines with family-specific features only on sm_100a/sm_100f — a forward-compatibility headache for CUDA developers. NVIDIA blog
- ByteDance CUDA Agent: Auto-generates CUDA kernels claimed 2x faster than torch.compile on small/medium kernels (paper, site).
- Inference silicon: Taalas HC1 hard-wires Llama-3.1-8B at ~17k token/s (paper); Apple M4 ANE shows up to 80x energy-efficiency advantage over A100 for Llama2 110M (benchmarks).
- MCP ecosystem grows despite security concerns: Notion shipped MCP/API integration with Meeting Notes; Cursor launched MCP Apps with interactive in-chat UI. A security writeup documented 5 easily exploited attack patterns (prompt-injected tool descriptions, filesystem/network overreach, third-party abuse).
- ShadowClaw: A single-file C agent that calls local LLMs (e.g., Ollama) via curl, with shell execution, file I/O, HTTP, and persistent state. Repo
- RLM (Recursive Language Modeling): Ongoing debate — giving models a REPL to write their own code vs traditional ReAct tool calling; supporters cite flexibility, critics cite controllability and safety.
- Perplexity Computer and Cursor cloud agents: Both run agents in isolated VMs/sandboxes to operate browsers, terminals, and IDEs end-to-end, trading transparency and cost for convenience.
- Benchmark realism: New work notes agent benchmarks over-index on math/coding vs real labor distribution; Arena launched Document Arena for PDF-based reasoning (Claude Opus 4.6 leads).
- Multi-agent consensus: Byzantine experiments show LLM multi-agent systems struggle to agree even without malicious nodes, with timeouts and stalemates worsening with scale. Theory-of-Mind modules help only depending on base model capability.
- Spectral norm scaling and muP: Scaling weight spectral norms/updates by √(fan-out/fan-in) explains when networks shift from kernel-method-like behavior to feature learning (paper, Modula).
- SAEs on text-to-image diffusion: Image composition is largely predictable early in reverse diffusion; style settles mid-way; final stages only refine texture; targeted interventions at each stage demonstrated (paper).
- Anthropic: Claude and Claude Code traffic far exceeded expectations, prompting emergency capacity expansion; reported gains in US enterprise share; Claude Code voice mode in gray release (hold spacebar to dictate, no extra charge).
- Cursor ecosystem: IDE revamp with Zen mode, cloud agents (including Android WebApp), and "AI coworker" Viktor built on Cursor for Slack — marketing audits, campaign management, 3000+ SaaS integrations.
- Local LLM tooling: LM Studio fixed LM Link multi-device discovery; easytranscriber from KBLab offers WhisperX-like ASR with timestamps, measured 35%–102% faster (post).
- Qwen team departures: Tech lead Justin Lin and other key members left Alibaba; concerns about future open-source/licensing strategy even as Qwen 3.5 training scripts and quantization guides continue shipping.
- OpenAI talent loss: RLHF/post-training VP Max Schwarzer left for Anthropic to do RL research.
- DoD/NSA contract fallout: Reports of the US DoD threatening to label Anthropic a "supply chain risk"; OpenAI's contract sparked privacy concerns and a claimed 295% spike in ChatGPT uninstalls (base-size caveats apply); Sam Altman said terms were revised to bar domestic surveillance of US citizens, though independent legal review was still demanded.
Infrastructure and Hardware
Agents and Tooling
Research
Products and Industry
Source: Easy AI Daily (zhichai.net)