Paper Overview
Research Area: Chip/Hardware Verification Author: FNU Aditi Published: 2026-09-22 arXiv: 2609.26751
Summary
Large language models are increasingly used to generate SystemVerilog Assertions from natural-language specifications and register-transfer-level designs. Existing datasets and benchmarks support important goals such as large-scale training, formal evaluation, specification-to-assertion generation, and mutation-based testing. A complementary need is to study whether a generated assertion captures externally observable behavior or depends on incidental details of one RTL implementation.
EquivSVA addresses this gap as a formally verified dataset organized around behavior families. Each family contains:
- Four structurally distinct RTL implementations of the same externally observable behavior
- Shared interface-level gold properties
- Three controlled mutants
- Formal-validation evidence
- Of 293 interface-only generated properties, 93 are formally sound
- The number of sound properties varies across equivalent implementations for 14 of 24 test families
- Paper: https://arxiv.org/abs/2609.26751
- Code & Data: https://github.com/aditigupta96/EquivSVA
The dataset comprises 120 behavior families across 12 categories, 480 reference RTL implementations, 914 gold properties, and 360 mutants. Every final family passes a fixed 17-job validation suite covering RTL equivalence, gold-property proofs, property reachability, mutant distinguishability, and gold-property checks on mutants. Fixed family-safe train, development, and test splits are also provided.
Case Study
As a small demonstration of the analyses enabled by the dataset, the publicly released, Apache-2.0-licensed Qwen2.5-Coder-7B-Instruct model was evaluated on the held-out test split:
These results illustrate how behavior-family organization can support controlled studies of assertion-generation robustness without requiring changes in intended functionality.