Key points
Models & Capabilities
- Google Gemini 3.1 Flash-Lite (preview): Positioned as the lowest-latency, highest-throughput multimodal model in the Gemini 3 family. 1M context, measured 360+ tokens/s, ~5.1s average response. Jeff Dean quoted ~$0.25/M input and $1.5/M output tokens; LMArena Elo of 1432 — but pricing is 2.5–3.75x higher than 2.5 Flash-Lite, sparking value debates. (DeepMind thread, API notes, Jeff Dean pricing)
- OpenAI GPT-5.3 Instant: Rolled out to all ChatGPT users with more natural responses, fewer refusals, and hallucination rate reductions of 26.8% (with search) and 19.7% (without).
gpt-5.3-chat-latestappeared in the API; OpenAI teased "GPT-5.4 sooner than you think." (Announcement) - Alibaba Qwen 3.5: Hot on Reddit. The 0.8B model includes a vision encoder, runs in-browser via WebGPU and on old phones (~12 tokens/s); 27B/35B versions show linear-attention efficiency and near-frontier reasoning. Hallucination risks remain.
- Apple M5 Pro / M5 Max: Up to 4x faster LLM prompt processing vs M4 Pro/Max; 64GB/128GB unified memory at 307/614GB/s bandwidth, 14.5GB/s SSD, N1 chip with Wi-Fi 7. (LocalLLaMA thread)
- Together: Context Parallel + sequence parallelism trains 8B models at 5M context on 8×H100, cutting attention memory up to 87%.
- Databricks FlashOptim (open-sourced): Reduces AdamW optimizer memory from ~16 to 7 bytes/param; 8B fine-tuning peak memory 175GiB → 113GiB.
- SkyPilot Job Groups: Schedules RL training across high-end GPUs, cheap GPUs, and big-memory CPUs.
- NVIDIA Blackwell split: Data-center (CC 10.0) vs consumer (CC 12.0) lines with family-specific features (sm_100a/sm_100f), complicating forward compatibility for CUDA/kernel developers. (NVIDIA blog)
- ByteDance CUDA Agent: Automatically generates high-performance CUDA kernels, reportedly 2x faster than torch.compile on small/mid kernels. Related RL-CUDA research: arXiv:2602.24286, project site cuda-agent.github.io.
- Inference chips: Taalas HC1 hard-wires Llama-3.1-8B at ~17k tokens/s (arXiv:2412.18511); Apple ANE shows up to 80x better energy efficiency than A100 on Llama2 110M (ANE benchmarks).
- MCP ecosystem expands despite "MCP is dead" takes: Notion MCP integration, Cursor MCP Apps with interactive UI rendering; a security analysis catalogs 5 easily exploited attack patterns.
- ShadowClaw: Single-file C agent calling local LLMs via curl with shell/file/HTTP tools and persistent state. (Repo)
- RLM (Recursive Language Modeling): DSPy community debates REPL-style agents vs traditional ReAct/tool-calling.
- Perplexity Computer & Cursor cloud agents: Sandboxed VMs operating browsers/terminals/IDEs to produce PRs and documents.
- New work argues agent benchmarks skew heavily toward math/coding vs real job distributions; Arena launched Document Arena (PDF-based), where Claude Opus 4.6 currently leads.
- Byzantine consensus experiments: LLM multi-agent systems struggle to reach consensus even without malicious nodes; failures grow with participant count. ToM/BDI + formal verification gains depend heavily on the base model.
- Spectral norm scaling & muP: Scaling spectral norms by √(fan-out/fan-in) explains when networks perform true feature learning (arXiv:2310.17813, Modula).
- SAE analysis of text-to-image diffusion: Image composition is largely determined early in reverse diffusion; style mid-way; texture last. Targeted interventions demonstrated (arXiv:2504.15473).
- Claude / Claude Code traffic surged beyond Anthropic's expectations; voice mode (hold-space-to-talk) rolling out to Claude Code.
- Cursor updated its IDE with Zen mode and cloud agents; "AI coworker" Viktor in Slack integrates 3000+ SaaS tools with persistent memory.
- LM Studio, OpenClaw, Manus: local inference tooling and community updates.
- easytranscriber (KBLab): ASR with accurate timestamps, 35%–102% faster than WhisperX depending on hardware. (Blog post)
- Qwen leadership exodus: Tech lead Justin Lin and other key members left Alibaba; community worries about future open-source/licensing strategy, though Qwen 3.5 releases continue.
- OpenAI → Anthropic talent move: Max Schwarzer, VP of RLHF/post-training, joined Anthropic as an RL researcher.
- OpenAI DoD/NSA contract controversy: Privacy concerns, demands for contract disclosure; Sam Altman says terms now bar domestic surveillance of US citizens. Reported 295% spike in ChatGPT app uninstalls — though analysts caution the absolute numbers may be small; Claude downloads reportedly rose in parallel.
- MCP security criticized as "a mess": prompt injection via tool descriptions, privilege escalation, and third-party abuse among 5 attack patterns; sandboxing and policy controls recommended.