Horizon AI Daily Digest - May 28, 2026 (covering May 27). From 41 items, 30 highlights were selected.
Top Stories
1. The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence ⭐️ 10.0/10 — The MiniMax-M2 series releases a 229.9B-parameter model with only 9.8B activations, designed end-to-end for agentic intelligence and self-evolution. (arXiv)
2. YouTube to automatically label AI-generated videos ⭐️ 8.0/10 — YouTube will automatically label AI-generated videos, sparking debate on content quality and detection methods. (HN discussion)
3. I think Anthropic and OpenAI have found product-market fit ⭐️ 8.0/10 — Simon Willison argues Anthropic and OpenAI have achieved product-market fit; discussion centers on profitability and market impact. (HN discussion)
4. What Apple and Google are doing to your push notifications ⭐️ 8.0/10 — How Apple and Google intercept push notifications, raising privacy and UX concerns. (HN discussion)
5. DuckDuckGo saw 28% more visits after Google said people love AI mode ⭐️ 8.0/10 — User backlash against Google's aggressive AI search drove a ~28% traffic increase for DuckDuckGo. (HN discussion)
Research Highlights
- Can LLMs Introspect? A Reality Check — Behavioral evidence is insufficient to prove LLM introspection; introspection must be distinguished from pattern matching.
- Is Agent Memory a Database? — Long-term agent memory should be treated as a state-trajectory-driven data management workload, not traditional storage.
- Your Agents Are Aging Too — Introduces agent lifespan engineering and the AgingBench benchmark, revealing post-deployment agent degradation.
- Experiments in Agentic AI for Science — Two autonomous frameworks (DeepTS/DeepCollector and DeepScribe) automate scientific data curation using hybrid local-remote architectures.
- Anchor — Formal constrained generation pipelines mitigate artifact drift in agent benchmark generation.
- OmniToM — Theory-of-mind benchmark with explicit belief modeling for LLMs.
- JobBench — A benchmark aligning agent work with human will; current models show limited performance.
- ScientistOne — Chain-of-Evidence verifiability framework exposes citation fabrication in existing autonomous research outputs.
- Automatic Layer Selection for Hallucination Detection — FEPoID criterion auto-selects LLM intermediate layers to improve hallucination detection.
- Reusable Skills in Offline Hierarchical RL — Exploits local dynamics regularity for offline skill reuse.
- Creative Physical Intelligence in LMMs — A benchmark for creative tool use in physical scenarios.
- Calibrated Interactive RL for Multi-turn Dialogue — Mitigates distribution shift in multi-turn dialogue with an aligned simulator.
- Trustworthy Legal AI (LexGuard) — Relevance-sensitive evaluation and solver-grounded formal reasoning for legal AI.
- Last.fm is now independent ⭐️ 7.0/10 — Last.fm splits from CBS; the community welcomes the move and praises API stability. (HN discussion)
- Tech CEOs suffering from AI psychosis ⭐️ 7.0/10 — CEO cognitive biases around AI mirror past tech hype cycles. (HN discussion)
- Alternative Internets Beyond HTTPS ⭐️ 7.0/10 — A tour of Gemini, Gopher, and Finger protocols. (HN discussion)
- BrickAnything — Geometry-conditioned generation of physically buildable brick structures with structure-aware tokenization.
- POLAR — Personalizing embodied multimodal LLM agents via multimodal knowledge graphs and episodic memory.
- Constraint acquisition needs better benchmarks — The MPMMine benchmark suite for constraint acquisition research.
- LLM Procedural Knowledge for Virtual Labs — Managing uncertainty in LLM-generated procedures for virtual laboratory planning.
- Reasoning, Code, or Both? — Chain-of-thought is more robust than code execution for LLM math reasoning under question variations.
- PolyFusionAgent — Multimodal foundation model and agent for polymer property prediction and inverse design.
- anthropics/claude-code v2.1.152 — Improves code review and skill flexibility.
- SimCity 3k in 4k — Revisiting SimCity 3000 in 4K. (HN discussion)
- Facebook launches a 'Plus' subscription ⭐️ 6.0/10 — Meta rolls out paid Plus subscriptions and tests AI subscription tiers.