Horizon AI Daily Digest - May 27, 2026
> 36 items tracked, 24 selected as the best.
Key points
- LLM introspection under scrutiny: A new paper argues existing research fails to distinguish genuine introspection from surface pattern matching (arXiv:2605.26242) ⭐️ 9.0
- JobBench: Evaluates AI agents by human-delegated needs rather than economic value, covering 35 professions and 130 tasks; the strongest model reaches only 45.9% (arXiv:2605.26329) ⭐️ 9.0
- ScientistOne: Introduces a Chain-of-Evidence framework and end-to-end autonomous research system to address verifiability of AI-generated research outputs (arXiv:2605.26340) ⭐️ 9.0
- MiniMax-M2 series: MoE LLMs with extremely small activation sizes delivering strong real-world intelligence, designed for agentic scenarios with self-evolution capabilities (arXiv:2605.26494) ⭐️ 9.0
- Regulation: Spain blocks prediction markets Polymarket and Kalshi over missing gambling licences, sparking global regulatory and ethical debate (Reuters, HN discussion) ⭐️ 8.0
- Agent memory as a database: Proposes treating long-term agent memory as a new data-management workload; introduces the GEM model for state-trajectory correctness (arXiv:2605.26252)
- POLAR: Personalizes embodied multimodal LLM agents over long-term user interactions using multimodal knowledge graphs (arXiv:2605.26256)
- AgingBench: Benchmarks and diagnoses reliability degradation of long-running deployed AI agents (arXiv:2605.26302)
- Anchor: Mitigates artifact drift in agent benchmark generation via constrained optimization (arXiv:2605.26321)
- OmniToM: Benchmarks Theory of Mind in LLMs via explicit belief modeling, going beyond endpoint Q&A (arXiv:2605.26322)
- Hallucination detection: Automatic middle-layer selection based on intrinsic-dimension first effective peaks (arXiv:2605.26366)
- CARL: Contrastive alignment of local dynamics and action sequences for reusable skills in offline hierarchical RL (arXiv:2605.26371)
- MM-CreativityBench: Evaluates creative physical intelligence in large multimodal models (arXiv:2605.26396)
- Multi-turn dialogue RL: A calibrated interactive RL framework mitigating cumulative distribution shift with an aligned simulator (arXiv:2605.26403)
- LexGuard: Relevance-sensitive evaluation and solver-grounded reasoning for trustworthy legal AI (arXiv:2605.26530)
- Claude Code v2.1.152 released with improved code-review auto-fixes and skill management (GitHub release)
- Cloudflare Flagship: New feature-management service with global flagging and evaluation (docs, HN discussion)
- Minicor (YC P26): Windows desktop automation at scale for AI companies, enabling integration with API-less systems (minicor.com, HN discussion)
- BrickAnything: Geometry-conditioned, structure-aware autoregressive generation of buildable brick structures from 3D shapes (arXiv:2605.26182)
- MPMMine: Benchmark suite addressing the lack of constraint-acquisition benchmarks (arXiv:2605.26279)
- Agentic AI for science: DeepTS/DeepCollector and DeepScribe automate scientific workflows with hybrid local/cloud architectures (arXiv:2605.26305)
- Virtual lab planning: Framework for managing uncertainty in LLM-generated procedural knowledge (arXiv:2605.26333)
- Math reasoning robustness: Chain-of-thought proves more robust than code execution on question variants (arXiv:2605.26414)
- PolyFusionAgent: Multimodal foundation model and autonomous assistant for polymer property prediction and inverse design (arXiv:2605.26543)