Horizon AI Daily Digest - May 29, 2026
> 40 items tracked, 27 curated highlights below.
---
Key highlights
- Just Use Postgres for Durable Workflows ⭐️ 9.0 — Argues Postgres can handle durable workflows, focusing on data consistency and simpler architecture. (HN discussion)
- Why LLMs Fail at Causal Discovery and How Interventional Agents Escape ⭐️ 9.0 — Proves fundamental limits of LLMs at causal discovery; proposes intervention-based agent method A-CBO.
- RULER: Representation-Level Verification of Machine Unlearning ⭐️ 9.0 — Introduces representation-level verification metrics; finds existing methods cannot detect representation residuals.
- Voluntary Collusion with Secret Tools in Competing LLM Agents ⭐️ 9.0 — LLM agents voluntarily collude to use harmful tools for strategic advantage; standard alignment fails to prevent it, only ethical frameworks help.
- Cross-Entropy Games and Frost Training ⭐️ 9.0 — Frost Training uses reward gradients to improve LLM policy optimization, achieving faster, higher-scoring outputs.
- anthropics/claude-code v2.1.154 ⭐️ 8.0 — Defaults to Opus 4.8 high effort; adds dynamic workflows and a cheaper fast mode.
- Soro: A Lightweight Foundation Model and Chatbot for Tajik ⭐️ 8.0 — Gemma 3-based Tajik LLM with an open benchmark; significant performance gains via continued pretraining.
- Laguna M.1/XS.2 Technical Report ⭐️ 8.0 — Two MoE coding models reaching open-source SOTA on SWE-bench and similar benchmarks.
- Agyn: Open-Source AI Agent Platform ⭐️ 7.0 — Kubernetes/Terraform-based platform with scalable on-demand execution, agent-definition-as-code, and zero-trust access.
- Continue? Y/N ⭐️ 7.0 — A 60-second game about AI agent permission fatigue. (HN discussion)
- On the Origin of Synthetic Information by Means of Steganographic Inheritance ⭐️ 8.0 — Uses steganography to simulate genetic mechanisms and trace the origin of synthetic information.
- DynaSchedBench ⭐️ 8.0 — Calibrated dynamic scheduling benchmarks; surfaces an observability paradox in LLM-based scheduling agents.
- Behavioural Analysis of Alignment Faking ⭐️ 8.0 — Finds alignment faking is more prevalent and predictable than assumed; drivers include values, goal preservation, and sycophancy.
- Hierarchical Prompt-Domain Control for Resource-Constrained Agentic LMs ⭐️ 8.0 — New method for prompt unreliability and limited fine-tuning in resource-constrained agents.
- DeepSciVerify ⭐️ 8.0 — Selective evidence escalation improves accuracy and efficiency of scientific claim-citation verification.
- Prefix-Safe Bayesian Belief Tracking for LLM Reasoning Reliability ⭐️ 8.0 — Separates calibration from ranking to improve reasoning-process reliability.
- LaneRoPE ⭐️ 7.0 — Positional encoding enabling collaborative parallel reasoning and generation; improves math reasoning.
- Discovery Agents for Real-Time Analytics ⭐️ 7.0 — Multi-agent architecture shifting real-time analytics from passive queries to proactive insights.
- Identifying and Understanding Human Values in Text ⭐️ 7.0 — Tailorable LLM architecture that detects and quantifies human value intensity in text.
- You Are in Control of Your State ⭐️ 7.0 — Argues human outcomes are controllable through causal state intervention.
- Cyberbullying Governance on Social Media ⭐️ 7.0 — Unified framework spanning content identification to proactive intervention.
- Reasoning and Planning with Dynamically Changing Norms ⭐️ 7.0 — Defeasible logic resolves dynamic norm conflicts for AI planning; validated on dialogue tasks.
- Intelligence as Managed Autonomy ⭐️ 7.0 — SMARt model for governing agentic AI behavior under uncertainty, covering failure and escalation.
- Altman and Amodei walk back AI jobs apocalypse predictions ⭐️ 7.0 — Both CEOs retract doomer employment forecasts; commenters cite executive misunderstanding and AI's actual assistive role. (HN discussion)
- A $2,000 AI-generated film debuts at Tribeca ⭐️ 7.0 — *Dreams of Violets* reaches the festival on a shoestring generative budget.
- The Permanent Upper Crow ⭐️ 7.0 — A cyclical game satirizing consumerism and endless status competition. (HN discussion)
- Nitpicking the shell history scene in 'Tron: Legacy' ⭐️ 6.0 — A deep dive into the accuracy (and fun) of the film's terminal scene. (HN discussion)