Industry & Company News
- Replit valued at $9B, pivoting to a full AI productivity suite: Replit's valuation tripled in six months as it expands from an online IDE into a complete AI productivity platform — canvas, apps, websites, slides, and video — targeting knowledge work beyond coding. (Latent Space analysis)
- Anthropic launches The Anthropic Institute: Led by former policy head Jack Clark, now "Head of Public Benefit," spanning ML, economics, and social science to study how strong models affect society. (Official announcement)
- Yann LeCun founds AMI Labs with $1.03B: Advanced Machine Intelligence Labs raised funding from NVIDIA, Samsung, and Bezos Expeditions to build JEPA-based world models, with open-source code and papers planned; no near-term product or revenue expectations. (TechCrunch coverage)
- r/LocalLLaMA passes 1 million members: The local-model community, created in March 2023, has grown into a mainstream hobbyist movement reminiscent of the early Linux enthusiast scene. (Celebration post)
- NVIDIA Nemotron 3 Super: 120B parameters (~12B active), 1M context, hybrid Mamba-Transformer + Latent MoE, native multi-token prediction, more KV-cache-efficient than Qwen3.5-122B. NVIDIA claims up to 2.2x faster inference than GPT-OSS-120B on Blackwell; support already in vLLM, llama.cpp, and Ollama. (Technical thread)
- Google Gemini Embedding 2: Fully multimodal (text, image, video, audio, PDF) with Matryoshka embeddings; community notes text pricing is relatively expensive — best for multimodal retrieval, and downsample video frames to control cost. (Feature overview)
- Qwen3.5 multimodal architecture breakdown: Hybrid Gated DeltaNet linear attention + global attention, a 397B-A17B MoE variant and a 27B dense variant, native 262k context extendable to ~1M, multi-token prediction used in training. (Detailed thread)
- Fish Audio S2 TTS: 80+ languages, natural-language emotion tags like
[whispers sweetly], multi-speaker dialogue in one pass, ~100ms first-frame latency; claimed to beat Google/OpenAI TTS on several benchmarks, but non-commercial license only. (Hugging Face model) - Qwen3.5-35B-A3B "Aggressive" GGUF: Community uncensored release claiming 0/435 refusals, 256-expert MoE with 8+1 active per token, image/video input; some users urge KL-divergence verification of capability claims. (Hugging Face page)
- Apple M5 Max 128GB local LLM benchmarks: Qwen3.5-122B, Qwen3 Coder, and gpt-oss-120b ran at 16–32k context with prompt throughput up to 2700+ tok/s, but memory usage hit 60–90GB. (Benchmark post)
- Perplexity "Personal Computer": A Mac mini as an always-on hybrid local+cloud agent server with file/app/browser access and remote control; an enterprise edition orchestrates 400+ SaaS apps with 20 dedicated models. (Announcement)
- Replit Agent 4: A multi-agent collaborative canvas that builds apps, websites, and slides in parallel rather than chat-based single-file coding. (Release)
- Base44 Superagents: One-stop workflow agents for non-technical users, pre-connected to Gmail, Slack, Stripe, CRM, etc. (Product launch)
- LangChain Deep Agents auto context compression: Agents summarize history at task boundaries rather than hard token cutoffs, balancing long-task memory and token cost. (Announcement)
- OpenAI computer-use technical brief: Documentation of the agent execution loop, filesystem context, networking, and guardrails for safe computer operation. (Technical brief)
- PostTrainBench v1.0: Tests whether frontier agents can perform post-training in simplified environments; notably, medium reasoning length outperformed very long reasoning on GPT-5.1 Codex Max due to context compression overhead. (Benchmark thread)
- EvoSkill: Extracts reusable skills from agent failures via executor/proposer/skill-builder roles, lifting Claude Code + Opus 4.5 accuracy on OfficeQA from 60.6% to 67.9%. (Project page)
- AgentIR: Encodes agent reasoning traces with queries for retrieval, reaching 68% accuracy on BrowseComp-Plus vs. 52% for larger standard embedding models and 37% for BM25. (Method intro)
- Layer-block copying pushes Qwen2-72B up the leaderboard: Duplicating 7 middle layers without weight changes improved Open LLM Leaderboard scores, suggesting functional circuit blocks in pretrained layer stacks; experiments even explored cyclic layer reuse for early-exit inference. (Research blog)
- Karpathy's self-improving agent swarm: ~700 autonomous edits, 20 effective ones, cut GPT-2-level training time from 2.02 to 1.80 hours (−11%). (GitHub repo)
- GPT-5.4 reportedly solves an open EpochAI Frontier Math problem: Initial review by Epoch researchers found the solution plausible; awaiting problem-author confirmation. (Discussion thread)
- 70–90% of Anthropic's R&D code reportedly written by Claude: TIME reporting suggests iteration cycles shrank from months to weeks, with some researchers predicting fully automated AI research within a year — fueling early recursive self-improvement concerns. (Discussion)
- Many agent failures are reliability issues, not attacks: Princeton-led NIST feedback argues non-adversarial failures lack definitions, metrics, and mitigations — evaluation and monitoring are becoming safety problems. (Random Walker thread)
- Claude Code login outage as an "intelligence brownout": An OAuth failure disrupted many developers for a day; Karpathy's autoresearch lab stalled too, highlighting dependence on single cloud models as an infrastructure risk. (Karpathy tweet)
- Google medical AI in practice: An AI system detected 25% of interval breast cancers missed by routine screening, and the AMIE conversational clinical-reasoning system passed real-world pilots on safety, feasibility, and patient acceptance. (Screening system)
- Reka Edge: A vision-language model for robotics/physical AI, claiming 3x fewer input tokens and 65% higher throughput vs. mainstream 8B-class models. (Release)
- Faceless YouTube channels with Claude: One creator earns ~$70k over 9 months using Claude scripts, ElevenLabs voice, Magic Hour visuals, and CapCut editing — drawing both skepticism and curiosity in the comments. (Workflow post)
- Claude rewrites a technical email — and a city changed its traffic lights: A user's congestion complaint, translated by Claude into engineering language, led to retimed signals admitting 2–3 more cars per cycle. (Original post)
- Four AI models stock-picking for 9 weeks: With $1,000 each via Alpaca API, ChatGPT led at +21.1%, Perplexity +1.1%, Gemini −6.6%, Claude −11.5%. A single-trail, entertainment-grade experiment — not empirical evidence. (Experiment summary)
- Anthropic launches free Claude Academy: Free online courses covering Claude on Amazon Bedrock, GCP Vertex, and education/nonprofit use cases — a no-cost alternative to expensive AI bootcamps. (Course intro)
Models & Capabilities
Agents & Tooling
Research & Methods
Policy, Governance & Safety
Products & Applications
📌 Source: Easy AI Daily