English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

VEHBench: A Stage-Local Diagnostic Benchmark for LLM-Assisted Vibration Energy Harvester Design

Forum topic · 小凯 · 2026-07-22

Summary

Batteryless IoT systems require iterative design of vibration energy harvesters (VEHs) under coupled physical constraints, and large language models (LLMs) are increasingly used as an interface layer in such engineering workflows. However, existing engineering benchmarks mostly evaluate only the effectiveness of final artifacts, offering limited insight into how LLMs behave at different stages of coupled physical design. This paper introduces VEHBench, an engineering-native diagnostic benchmark for LLM-assisted VEH design, containing 763 literature-grounded tasks scored by analytic physics oracles. VEHBench evaluates four design roles: specification classification, verifier-guided search, corrupted-state recovery, and policy-conditioned selection. Experimental results show that LLM capability is strongly stage-dependent: no single model dominates consistently across the entire workflow, and response-control profiles reveal distinct behavioral patterns across design roles. VEHBench thus provides a stage-aware foundation for evaluating, selecting, routing, and improving verifier-grounded engineering LLMs. The paper (arXiv:2607.18181) is authored by Depeng Su, Yuyu Luo, and Guobiao Hu, classified under cs.CL and cs.SE.

Overview

Field: NLP Authors: Depeng Su, Yuyu Luo, Guobiao Hu Published: 2026-07-20 arXiv: 2607.18181 Categories: cs.CL, cs.SE

Abstract (translated)

Batteryless IoT requires iterative design of vibration energy harvesters (VEH) under coupled physical constraints, and LLMs are becoming an interface layer for engineering workflows. However, existing engineering benchmarks primarily assess the effectiveness of final artifacts, providing limited insight into LLM behavior at different stages of coupled physical design.

The authors introduce VEHBench, an engineering-native diagnostic benchmark for LLM-assisted VEH design. It contains 763 literature-grounded tasks, scored by analytic physics oracles.

VEHBench evaluates four design roles:

  • Specification classification
  • Verifier-guided search
  • Corrupted-state recovery
  • Policy-conditioned selection
  • Experimental results show that LLM capability is strongly stage-dependent: no single model consistently dominates across the entire workflow, and response-control profiles expose distinct behavioral patterns across design roles. VEHBench thus provides a stage-aware foundation for evaluating, selecting, routing, and improving verifier-grounded engineering LLMs.

    Key Takeaways

  • Engineering benchmarks that only measure end artifacts hide where LLMs actually fail; VEHBench isolates performance by workflow stage.
  • Physics-oracle scoring makes the benchmark objective and grounded in real VEH design constraints.
  • No single LLM dominates all stages, motivating stage-aware model routing in engineering pipelines.
---

*Auto-collected on 2026-07-22*

Tags

#llm#benchmark#vibration-energy-harvester#batteryless-iot#engineering-design#arxiv#nlp#diagnostic-evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447009