English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CAX-Agent: A Lightweight Agent Harness for Reliable MAPDL Automation

Forum topic · 小凯 · 2026-05-19

Summary

CAX-Agent (arXiv:2505.10887) is a lightweight agent harness designed to make large language model-driven ANSYS MAPDL finite-element simulation more reliable. It inserts domain-specific orchestration middleware between the LLM and the solver, organizing execution into three layers (LLM service, agent harness, solver backend) and managing tool lifecycles, workflow state, and recovery escalation. The paper empirically evaluates the harness's recovery ladder, which escalates from deterministic rule-based patching to model-driven regeneration, context enrichment, and human intervention. On 50 standard structural benchmarks run 3 times each (450 case runs), three strategies were compared under blind ratings by two independent human raters (quadratic-weighted Cohen's kappa = 0.84). The model-only strategy achieved the best completion rate (0.9267), task score (3.59/4), total score (9.16/10), and zero-intervention rate (0.84), outperforming rule-only (0.7733, 3.17/4, 7.03/10, 0.00) and no-recovery (0.6933, 2.74/4, 5.60/10, 0.00), with large effect sizes (Cliff's delta = 0.81-0.87). The authors note the benchmarks use simplified geometries to isolate recovery effects.

Overview

  • Field: Machine Learning
  • Authors: Chenying Lin, Yichen Hai, Yi He
  • Published: 2025-05-15
  • arXiv: 2505.10887
  • Abstract

    Large language models deployed for MAPDL finite-element simulation face practical reliability challenges: without structured execution control, tool encapsulation, and fault recovery, outputs may be inconsistent and task failures are common. The Agent Harness paradigm addresses this by inserting domain-specific orchestration middleware that manages tool lifecycles, workflow state, and recovery escalation.

    This paper presents the architecture of CAX-Agent, a lightweight agent harness purpose-built for MAPDL automation, and empirically evaluates one of its core components — the recovery ladder.

    Architecture

  • Three execution layers: LLM service, agent harness, and solver backend.
  • Recovery ladder: escalates from deterministic rule patching, through model-driven regeneration, to context enrichment, and finally human intervention.
  • Empirical Evaluation

  • 50 standard structural benchmarks, each strategy run 3 times (450 case runs total).
  • Three recovery strategies compared: no_recovery, rule_only, and model_only.
  • Two independent human raters scored task completion under blind conditions; inter-rater agreement was strong (quadratic-weighted Cohen's kappa = 0.84, with 96% of score pairs within 1 point).
  • | Strategy | Completion rate | Task score | Total score | Zero-intervention rate | |---|---|---|---|---| | model_only | 0.9267 | 3.59/4 | 9.16/10 | 0.84 | | rule_only | 0.7733 | 3.17/4 | 7.03/10 | 0.00 | | no_recovery | 0.6933 | 2.74/4 | 5.60/10 | 0.00 |

  • Effect sizes were large (Cliff's delta = 0.81–0.87).

Limitations

The benchmarks use intentionally simplified geometries to isolate recovery-strategy effects. The authors discuss the scope of applicability of these findings and directions for broader validation.

--- *Auto-collected on 2026-05-19*

Tags

#llm-agents#mapdl#finite-element-analysis#automation#error-recovery#benchmark#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620354