> An in-depth technical review of AGEL-Comp: A Neuro-Symbolic Framework for Compositional Generalization in Interactive Agents (arXiv:2604.26522, IntelliSys 2026, Mahnoor Shahid & Hannes Rothe, University of Duisburg-Essen), plus its companion paper *Grounding vs. Compositionality* (arXiv:2604.26521, AAAI MAKE 2026).
Feynman-style one-liner
Think of a law firm: the LLM is a brilliant but reckless junior lawyer full of ideas; the NTP (neural theorem prover) is a stern auditor that checks every proposal against the legal code (world model) and rejects anything unprovable; the CPG (causal program graph) is a legal code that grows by itself, where each Horn clause is a statute; and ILP (inductive logic programming) is a legislator who, after each lost case (prediction error), does a minimal contrastive analysis, finds the single decisive factor, and writes it into the code. Winning depends not on the lawyer's talent but on the code's completeness and the auditor's strictness — that is why a 7B model can reach 100%: it is not the one doing the gatekeeping. The framework's own admitted weakness: learning only triggers when two cases differ by exactly one literal — rarely true in the real world.
What it is
- Neuro-symbolic agent architecture: LLM (divergent proposal) + NTP (convergent verification) + CPG (growable world model) + ILP (experience-driven rule induction), targeting compositional generalization in interactive environments.
- Authors: Mahnoor Shahid (PhD student, DFKI background) and Hannes Rothe (IS professor), University of Duisburg-Essen.
- Environment: Retro Quest, a self-built 2D top-down RPG (Unity + ML-Agents), 10 quests, 5 difficulty levels, 3 seeds.
- Models tested: GPT-4o, Gemini Pro 2.5, LLaVA-1.6 (Mistral-7B), DeepSeek-VL-7B.
- Open science: fully open-sourced (github.com/Place-Beyond-Bytes/AGEL-Comp) with OSF preregistration (osf.io/a6j4c).
- Stage 1 (MCS): minimal contrasting-case causal attribution — find a successful memory with the same action; if the state difference ΔS is a single literal, that literal is deemed the cause (deliberately avoiding LLM prior confounds).
- Stage 2 (MIL): meta-interpretive induction of a Horn clause from the positive example against the CPG as background knowledge, e.g.
causes_damage(X) :- is_harmful(X).
6. Fire example: the agent approaches fire to grab a coin, loses 50 HP, attributes causation via a single-literal contrast, legislates the rule, and on later encounters the NTP pre-emptively vetoes approaching fire. - Planner–verifier separation: let the LLM diverge, let an independent mechanism gatekeep.
- Prediction-error + minimal-contrast rule induction: more decidable than natural-language reflection (Reflexion-style), more composable than code skill libraries (Voyager-style).
- The companion paper's lesson: compositional reasoning does not emerge for free from grounding — "the model will get bigger and figure it out" is an assumption worth challenging.
- Main paper: arXiv:2604.26522 (full-text review, Table 3 verified cell by cell)
- Companion: arXiv:2604.26521
- Code: github.com/Place-Beyond-Bytes/AGEL-Comp; Preregistration: osf.io/a6j4c
- Crisis literature: Lake & Baroni 2018 (SCAN); Kim & Linzen 2020 (COGS); Keysers 2020 (CFQ); Dziri 2023 (Faith and Fate); CompWoB (arXiv:2311.18751)
- Foundations: Rocktäschel & Riedel 2017 (NTP); Muggleton 2014/2015 (MIL/Metagol); Cropper & Dumančić 2021/2022; Evans & Grefenstette 2018
- Related agents: Voyager (arXiv:2305.16291); ExpeL (ICLR 2024); CLIN (ACL 2024); Reflexion (2023)
Architecture
1. Perception: environment state → structured ground literals.
2. LLM planner: generates a pool of "plausible but unverified" candidate subgoals.
3. NTP verifier: differentiable backward-chaining proof over the CPG yields a soft unification score σ; plans with σ≈0 are rejected before execution, forcing the LLM to re-propose.
4. CPG world model: a causal directed hypergraph where each Horn clause h ← b₁,…,bₙ is a hyperedge; grows via W_{t+1} ← W_t ∪ ΔW.
5. Two-stage grounding on failure:
Ablations confirm the loop: w/o NTP, first-try drops 60% → 22.5%; w/o ILP, the agent repeats "logically valid but factually wrong" plans.
Verified numbers (Table 3, 30 attempts per cell)
| Model | Config | Task success % | First-try % | Time (s) | | --- | --- | --- | --- | --- | | GPT-4o | Baseline | 86.67±5.77 | 6.67 | 347.88 | | GPT-4o | AGEL full | 100±0 | 66.67 | 36.40 | | Gemini-2.5-Pro | Baseline | 83.33±5.77 | 3.33 | 302.02 | | Gemini-2.5-Pro | AGEL full | 100±0 | 66.67 | 28.06 | | DeepSeek-VL-7B | Baseline | 66.67±5.77 | 3.33 | 739.81 | | DeepSeek-VL-7B | AGEL full | 100±0 | 56.67 | 74.36 | | LLaVA-1.6-7B | Baseline | 63.33±5.77 | 0.00 | 1028.72 | | LLaVA-1.6-7B | AGEL full | 100±0 | 50.00 | 96.81 |
Aggregates: baseline 75.0% / 3.33% first-try → full system 100% / 60.0% first-try (≈18×, +56.7pp); sample efficiency 6.8×; baselines are 5–10× slower.
Statistical caution
±5.77 = 10/√3 means the three seeds differ by only one quest's outcome. No significance tests, no confidence intervals anywhere in the paper; Fisher exact test on GPT-4o's main comparison gives p≈0.11. There is no independent train/test split — compositional generalization is proxied by novel quests in the latter curriculum.
Most counterintuitive ablation
"Learns but cannot verify" (w/o NTP: 91.7% / 22.5%) beats "verifies but cannot learn" (w/o ILP: 76.7% / 8.3%, nearly identical to baseline). The symbolic learner contributes more than the verifier; verification without learning is worthless.
Fact corrections vs. popular claims
1. "3.3% baseline success" — true but it is the aggregate first-try mean, not overall success; GPT-4o's baseline task success was already 86.67%. 2. "Small local model hits 100%" — true only for retry-allowed task success; LLaVA's zero-retry first-try rate is just 50%. 3. "7B closes the gap with GPT-4o" — overstated: gaps persist (first-try 50% vs 66.67%; 96.81s vs 36.40s; 17.22 vs 11.38 iterations). The 100% is a ceiling effect. 4. "Agent knows fire is dangerous but fails on fire-breathing dragons" — fabricated: no dragons in the paper, and the agent initially does *not* know fire is harmful; it learns from interaction. 5. "Billions of parameters, locally deployable" — accurate (both 7B open models). 6. "Improvement of 60%" — misstated; 60% is the improved absolute first-try rate, the gain is +56.7pp (~18×).
Devil's advocate: three sharpest critiques
1. Examiner, exam, and grader are the same authors — all evidence comes from a 10-task self-built environment; MCS's single-literal requirement is itself a confession of environmental simplicity; no third-party benchmark comparison (TextWorld/ALFWorld/BabyAI). 2. Perception is fed for free — the environment API hands over clean structured literals; the learning loop may never trigger in noisy open-world settings where it would matter most. The NTP/embedding pre-training also injects hand-crafted inductive bias absent from baselines. 3. 100% is a ceiling effect — gains shrink as the base model strengthens (GPT-4o +13.3pp, DeepSeek +33.3pp, LLaVA +36.7pp): the framework patches weakness, not strength. Fairness note: OSF preregistration, full open-sourcing, and honest multi-metric reporting indicate solid procedure — the weakness is statistical rigor, not integrity.
Verdict
A mechanistically clear, procedurally rigorous, but narrowly validated workshop-level paper — not a breakthrough. It demonstrates that in small, closed, symbol-friendly worlds, "LLM proposes + NTP verifies + prediction-error-driven ILP legislates" strongly outperforms raw LLMs. Three genuinely transferable ideas:
Read the architecture sections (§3–§4), play with the open code, but cite the numbers with their measurement caveats — the lawsuit was won by the auditor and the legal code, not because the junior lawyer got smarter.