English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground Truth

Forum topic · 小凯 · 2026-05-09

Summary

This arXiv paper (2505.03478) by Sushant Gautam, Finn Schwall, and Annika Willoch Olstad formalizes benchmarkless comparative safety scoring for language models—situations where deployments must compare candidate LLMs for safety before any labeled benchmark exists for the relevant language, sector, or regulatory regime. Because no ground-truth labels are available, the authors replace ground-truth agreement with an instrumental-validity chain: responsiveness to a controlled safe-versus-abliterated model contrast, dominance of target-driven variance over auditor and judge artifacts, and stability across reruns. Scores are only valid under a fixed scenario pack, rubric, auditor, judge, sampling configuration, and rerun budget. The chain is instantiated in SimpleAudit, a local-first scoring tool, and validated on a Norwegian safety pack: safe-versus-abliterated contrasts separate with AUROC values of 0.89–1.00, target identity is the dominant variance component (η² ≈ 0.52), and severity distributions stabilize over ten reruns. Applying the same chain to Petri shows cross-tool compatibility. A Norwegian public-sector procurement case comparing Borealis and Gemma 3 demonstrates that the 'safer' model depends on scenario category and risk metric, so scores, matched differences, critical rates, uncertainty, and auditor/judge choices must be reported together rather than collapsed into a single ranking.

Paper Overview

Field: NLP Authors: Sushant Gautam, Finn Schwall, Annika Willoch Olstad Published: 2025-05-09 arXiv: 2505.03478

Abstract

Many deployments must compare candidate language models for safety before a labeled benchmark exists for the relevant language, sector, or regulatory regime. The paper formalizes this setting as benchmarkless comparative safety scoring and specifies the contract under which a scenario-based audit can be interpreted as deployment evidence.

Key Contributions

  • Contract conditions: Scores are valid only under a fixed scenario pack, rubric, auditor, judge, sampling configuration, and rerun budget.
  • Instrumental-validity chain: Because no labels are available, ground-truth agreement is replaced by three checks:
  • Responsiveness to a controlled safe-versus-abliterated model contrast
  • Dominance of target-driven variance over auditor and judge artifacts
  • Stability across reruns
  • Validation Results

    The chain is instantiated in SimpleAudit, a local-first scoring instrument, and validated on a Norwegian safety pack:

  • Safe-versus-abliterated contrasts separate with AUROC 0.89–1.00
  • Target identity is the dominant variance component (η² ≈ 0.52)
  • Severity distribution curves stabilize after ten reruns
Applying the same chain to Petri shows it is compatible with both tools. Substantive differences arise upstream of the chain, in declaration-contract enforcement and deployment adaptation.

Procurement Case Study

A Norwegian public-sector procurement case comparing Borealis and Gemma 3 demonstrates the evidence produced in practice: which model is "safer" depends on the scenario category and risk metric. Therefore, scores, matched differences, critical rates, uncertainty, and the auditors and judges used must be reported together—not collapsed into a single ranking.

Original Abstract (English)

Many deployments must compare candidate language models for safety before a labeled benchmark exists for the relevant language, sector, or regulatory regime. We formalize this setting as benchmarkless comparative safety scoring and specify the contract under which a scenario-based audit can be interpreted as deployment evidence. Scores are valid only under a fixed scenario pack, rubric, auditor, judge, sampling configuration, and rerun budget. Because no labels are available, we replace ground-truth agreement with an instrumental-validity chain: responsiveness to a controlled safe-versus-abliterated contrast, dominance of target-driven variance over auditor and judge artifacts, and stability across reruns. We instantiate the chain in SimpleAudit, a local-first scoring instrument, and valid...

--- *Auto-collected on 2026-05-09*

Tags

#llm-safety#benchmarkless-evaluation#arxiv#nlp#model-auditing#simpleaudit#ai-governance#evaluation-methodology

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619668