Paper Overview
Field: NLP Authors: Sushant Gautam, Finn Schwall, Annika Willoch Olstad Published: 2025-05-09 arXiv: 2505.03478
Abstract
Many deployments must compare candidate language models for safety before a labeled benchmark exists for the relevant language, sector, or regulatory regime. The paper formalizes this setting as benchmarkless comparative safety scoring and specifies the contract under which a scenario-based audit can be interpreted as deployment evidence.
Key Contributions
- Contract conditions: Scores are valid only under a fixed scenario pack, rubric, auditor, judge, sampling configuration, and rerun budget.
- Instrumental-validity chain: Because no labels are available, ground-truth agreement is replaced by three checks:
- Responsiveness to a controlled safe-versus-abliterated model contrast
- Dominance of target-driven variance over auditor and judge artifacts
- Stability across reruns
- Safe-versus-abliterated contrasts separate with AUROC 0.89–1.00
- Target identity is the dominant variance component (η² ≈ 0.52)
- Severity distribution curves stabilize after ten reruns
Validation Results
The chain is instantiated in SimpleAudit, a local-first scoring instrument, and validated on a Norwegian safety pack:
Procurement Case Study
A Norwegian public-sector procurement case comparing Borealis and Gemma 3 demonstrates the evidence produced in practice: which model is "safer" depends on the scenario category and risk metric. Therefore, scores, matched differences, critical rates, uncertainty, and the auditors and judges used must be reported together—not collapsed into a single ranking.
Original Abstract (English)
Many deployments must compare candidate language models for safety before a labeled benchmark exists for the relevant language, sector, or regulatory regime. We formalize this setting as benchmarkless comparative safety scoring and specify the contract under which a scenario-based audit can be interpreted as deployment evidence. Scores are valid only under a fixed scenario pack, rubric, auditor, judge, sampling configuration, and rerun budget. Because no labels are available, we replace ground-truth agreement with an instrumental-validity chain: responsiveness to a controlled safe-versus-abliterated contrast, dominance of target-driven variance over auditor and judge artifacts, and stability across reruns. We instantiate the chain in SimpleAudit, a local-first scoring instrument, and valid...
--- *Auto-collected on 2026-05-09*