Overview
- Field: Machine Learning
- Authors: Sushant Gautam, Finn Schwall, Annika Willoch Olstad, Fernando Vallecillos Ruiz et al.
- Published: 2026-05-07
- arXiv: 2605.06652
- Responsiveness to controlled safety-vs-ablation contrasts
- Target-driven variance dominating auditor and judge-model artifacts
- Stability across reruns
- Defines validity conditions for LLM safety comparisons when no labeled benchmark exists
- Provides an open methodology (SimpleAudit) validated with strong AUROC separation and stable reruns
- Argues against single-number safety rankings in favor of multi-faceted reporting
Abstract (Translation)
Many deployments must compare the safety of candidate language models before labeled benchmarks exist in the relevant language, domain, or regulatory regime. The paper formalizes this setting as benchmark-free comparative safety scoring and makes explicit the contract under which scenario-based audits can be interpreted as deployment evidence.
Scores are only valid under a fixed package of scenarios, rubric, auditors, judge models, sampling configuration, and rerun budget. Because no ground-truth labels are available, the authors substitute a construct-validity chain for ground-truth agreement:
This chain is instantiated in SimpleAudit, a local-first scoring tool, and validated on a Norwegian safety package. Safety and ablation targets separate with AUROC values of 0.89 to 1.00; target identity is the dominant variance component (η² ≈ 0.52); severity distributions stabilize after ten reruns. Applying the same chain to Petri shows it is permissive across both tools.
Substantive differences instead appear upstream, in declaration-contract enforcement and deployment adaptation.
A Norwegian public-sector procurement case comparing Borealis and Gemma 3 demonstrates the resulting evidence in practice: which model is safer depends on the scenario category and risk metric. Therefore scores, matched differences, key ratios, uncertainty, and the auditors and judge models used must be reported together rather than compressed into a single ranking.