Overview
- Field: Machine Learning
- Authors: Chenying Lin, Yichen Hai, Yi He
- Published: 2025-05-15
- arXiv: 2505.10887
- Three execution layers: LLM service, agent harness, and solver backend.
- Recovery ladder: escalates from deterministic rule patching, through model-driven regeneration, to context enrichment, and finally human intervention.
- 50 standard structural benchmarks, each strategy run 3 times (450 case runs total).
- Three recovery strategies compared:
no_recovery,rule_only, andmodel_only. - Two independent human raters scored task completion under blind conditions; inter-rater agreement was strong (quadratic-weighted Cohen's kappa = 0.84, with 96% of score pairs within 1 point).
- Effect sizes were large (Cliff's delta = 0.81–0.87).
Abstract
Large language models deployed for MAPDL finite-element simulation face practical reliability challenges: without structured execution control, tool encapsulation, and fault recovery, outputs may be inconsistent and task failures are common. The Agent Harness paradigm addresses this by inserting domain-specific orchestration middleware that manages tool lifecycles, workflow state, and recovery escalation.
This paper presents the architecture of CAX-Agent, a lightweight agent harness purpose-built for MAPDL automation, and empirically evaluates one of its core components — the recovery ladder.
Architecture
Empirical Evaluation
| Strategy | Completion rate | Task score | Total score | Zero-intervention rate | |---|---|---|---|---| | model_only | 0.9267 | 3.59/4 | 9.16/10 | 0.84 | | rule_only | 0.7733 | 3.17/4 | 7.03/10 | 0.00 | | no_recovery | 0.6933 | 2.74/4 | 5.60/10 | 0.00 |
Limitations
The benchmarks use intentionally simplified geometries to isolate recovery-strategy effects. The authors discuss the scope of applicability of these findings and directions for broader validation.
--- *Auto-collected on 2026-05-19*