Overview
Field: NLP Authors: Depeng Su, Yuyu Luo, Guobiao Hu Published: 2026-07-20 arXiv: 2607.18181 Categories: cs.CL, cs.SE
Abstract (translated)
Batteryless IoT requires iterative design of vibration energy harvesters (VEH) under coupled physical constraints, and LLMs are becoming an interface layer for engineering workflows. However, existing engineering benchmarks primarily assess the effectiveness of final artifacts, providing limited insight into LLM behavior at different stages of coupled physical design.
The authors introduce VEHBench, an engineering-native diagnostic benchmark for LLM-assisted VEH design. It contains 763 literature-grounded tasks, scored by analytic physics oracles.
VEHBench evaluates four design roles:
- Specification classification
- Verifier-guided search
- Corrupted-state recovery
- Policy-conditioned selection
- Engineering benchmarks that only measure end artifacts hide where LLMs actually fail; VEHBench isolates performance by workflow stage.
- Physics-oracle scoring makes the benchmark objective and grounded in real VEH design constraints.
- No single LLM dominates all stages, motivating stage-aware model routing in engineering pipelines.
Experimental results show that LLM capability is strongly stage-dependent: no single model consistently dominates across the entire workflow, and response-control profiles expose distinct behavioral patterns across design roles. VEHBench thus provides a stage-aware foundation for evaluating, selecting, routing, and improving verifier-grounded engineering LLMs.
Key Takeaways
*Auto-collected on 2026-07-22*