Models & Capabilities
- Anthropic launches 1M-context Opus 4.6 as default: Anthropic quietly made the 1M-token context version of Opus 4.6 the default for Max/Team/Enterprise plans, removed long-context surcharges and beta headers, and raised per-request image/PDF limits to roughly 600 pages. It scored 78.3% on MRCR v2 1M tokens, seen by many as the new high-water mark for long context. (Latent Space coverage)
- OmniCoder-9B: Tesslate released OmniCoder-9B, fine-tuned from Qwen3.5-9B for code-agent scenarios using 425k+ agentic coding trajectories (including data generated by Claude Opus 4.6 and GPT-5.4). Native 262k context, expandable to 1M+, strong error recovery, fully open under Apache 2.0. (Reddit discussion)
- Qwen3.5-9B hailed as a small-but-strong model: Local LLM users report it runs on a single 12GB RTX 3060 with agentic coding performance approaching much larger models, praised for cost-effectiveness on limited hardware. (Reddit discussion)
- Qwen 3.5 fine-tunes called 'notably stronger': A community post highlighted 33 fine-tuned variants of Qwen 3.5, with the 40B dense and a Claude-Opus-style model singled out for stronger reasoning and customization appeal. (Reddit thread)
- MCP debate: demand is real, usability is the problem: Engineers agree MCP isn't dead but has high onboarding friction. LlamaIndex's take: MCP suits scenarios needing stable APIs and real-time data; local skills are lighter but more fragile. (LlamaIndex, Pamela Fox)
- Chrome adds Web MCP support (v146): A demo shows a LangChain Deep Agent continuously browsing X and auto-generating daily reports, pushing MCP toward 'browser as agent host.' (Discussion)
- Hermes Agent: A self-hosted agent with long-term memory and self-improvement, storing user preferences and skills over time; frequently cited as a representative self-hosted agent. (Discussions)
- AI coding workflows evolve from assistant to 'software factory': Engineers describe multi-agent pipelines — five agents handling review, testing, security, and performance, plus two merging PRs and running regressions — closer to fully automated CI. (Share, swyx)
- Automated research heats up: Karpathy's autoresearch sparked the topic, though veterans note continuity with DSPy, GEPA, and Bayesian optimization pipelines. Together AI open-sourced Open Deep Research v2's app, eval set, and code. (Karpathy, Together AI)
- 'Context Drought': Latent Space argues 1M context has been available since 2024 but has grown less than an order of magnitude since, bottlenecked by HBM/DRAM supply; the podcast predicts future 'context rationing.' (Article, Podcast)
- IndexCache: Reusing sparse-attention indices across layers in DeepSeek Sparse Attention yields ~1.2x end-to-end speedup on GLM-5 744B, and 1.82x prefill / 1.48x decode on a 30B-class model at 200k context, cutting index computation ~75% at equal quality. (Thread)
- Klein KV extends KV-cache optimization to image generation: Black Forest Labs injects reference-image KV caches into subsequent DiT denoising steps, speeding multi-reference editing up to ~2.5x. (Intro)
- Microsoft first to validate NVIDIA Vera Rubin NVL72: Nadella says Azure is the first cloud to validate the system; Lambda advocates bare-metal over virtualized deployments for the Rubin era. (Satya, Lambda)
- tinygrad's exabox vision: tinygrad claims its end goal is a Python-driven 2027 'exabox' exposed as one giant GPU, hiding all distributed complexity. (Post)
- RandOpt / Neural Thickets: MIT-led work adds Gaussian noise to pretrained weights and ensembles, approaching or exceeding GRPO/PPO on reasoning, coding, writing, chemistry, and VLM tasks — suggesting pretrained models are surrounded by 'task experts' and late-stage tuning is simpler than assumed. (Overview)
- Universal data replay (Stanford): Replaying general data during training yields ~1.87x gains in fine-tuning and ~2.06x in mid-training, including +4.5 points on a web-navigation agent and ~2% on Basque QA. (Summary)
- Multi-agent memory as a computer-architecture problem: A paper models shared agent memory as cache/memory hierarchies, addressing consistency and permissions rather than just 'bigger context.' (Summary)
- BrokenArXiv: With subtly tampered math claims from recent papers, GPT-5.4 rejects only ~40% of false propositions — 'pseudo-rigor detection' remains unsolved, though GPT-5.4 slightly outperforms Claude on this style of task. (Project, Paul's comparison)
- Personal agent UX goes always-on and cross-device: Perplexity Computer launched on iOS with phone/desktop sync; Claude Code demoed starting desktop coding sessions from a phone; Genspark's Claw is marketed as a cloud-resident 'AI employee.' Common thread: remote execution + persistent sessions + multi-model orchestration. (Perplexity, Claude Code, Genspark Claw)
- Gemini task automation: The Verge tried Gemini automating real tasks like hailing an Uber and ordering from a menu — acting rather than just advising. (The Verge)
- Gemini UI/UX 2.0 and $250/month Ultra tier: The redesign emphasizes personalization while pushing the Google AI Ultra subscription, criticized as priced for enterprises rather than consumers. (Discussion)
- Nano Banana Pro quality complaints: Users report pixelation and blur after March 10, suspecting model or safety-policy changes; sentiment shifting from impressed to disappointed. (Discussion)
- Claude's interactive charts UI: Users shared Claude's interactive chart interface for manipulating data within chat, widely shared as a promising direction for in-conversation data analysis. (Screenshot post)
- xAI restarts hiring: Musk said xAI is reviewing past interviews and re-contacting strong candidates previously rejected, effectively admitting screening flaws. (Post)
- Altman: 'selling intelligence' like utilities: Sam Altman framed the future as metered intelligence, like electricity or water — OpenAI as a global 'intelligence utility.' (Discussion)
- Carmack backs permissive training-data stance: John Carmack argued open-source code is a gift, and AI training amplifies rather than steals its value, resonating with part of the open-source community. (Tweet)
- AINews joins Latent Space: AINews is now integrated into the Latent Space site with searchable archives; the Discord channel won't reopen in its original form. (Announcement)
- Palantir CEO's political comments: Alex Karp claimed AI will reduce the influence of highly educated, Democratic-leaning voters while increasing the power of technically skilled working-class men — seen as dragging AI into US political polarization. (Discussion)
- Sanders bill to ban new AI data centers: Bernie Sanders formally proposed legislation banning new AI data centers, citing existential risk; the blanket approach sparked strong controversy in policy circles. (Discussion)
- OpenFold3 Preview 2: Mo AlQuraishi announced a release claiming major progress closing the gap with AlphaFold3 across modalities, with weights, training datasets, and configs public — billed as the only AF3-family model fully reproducible from scratch. (Announcement)
- WAXAL speech dataset: 2,400+ hours of speech covering 27 Sub-Saharan languages — TTS for 17, ASR for 19 — serving 100M+ speakers; a significant step for low-resource speech models. (Intro, Google Research)
Agents & Tooling
Infrastructure & Hardware
Research & Methods
Products & Applications
Industry & Companies
Policy, Governance & Safety
More Research
📌 Source: Easy AI Daily