English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

When No Benchmark Exists: Validating Comparative LLM Safety Scoring with Toolchains

Forum topic · 小凯 · 2026-05-10

Summary

A new arXiv paper (2605.06652) formalizes benchmark-free comparative LLM safety scoring for deployments where labeled benchmarks are unavailable in the relevant language, domain, or regulatory regime. Instead of ground-truth agreement, the authors propose a construct-validity chain as deployment evidence: responsiveness on controlled safety-vs-ablation contrasts, target-driven variance dominating auditor and judge-model artifacts, and stability across reruns. They instantiate this chain in SimpleAudit, a local-first scoring tool, and validate it on a Norwegian safety package. Safety and ablation targets separate with AUROC values of 0.89-1.00, target identity is the dominant variance component (eta-squared ~0.52), and severity distributions stabilize over ten reruns. Applying the same chain to Petri shows the chain is tool-agnostic. Substantive differences arise upstream in declaration-contract enforcement and deployment adaptation. A Norwegian public-sector procurement case comparing Borealis and Gemma 3 shows that the safer model depends on scenario category and risk metric, so scores, matched differences, key ratios, uncertainty, and the auditors and judge models used must be reported together rather than compressed into a single ranking.

Overview

  • Field: Machine Learning
  • Authors: Sushant Gautam, Finn Schwall, Annika Willoch Olstad, Fernando Vallecillos Ruiz et al.
  • Published: 2026-05-07
  • arXiv: 2605.06652
  • Abstract (Translation)

    Many deployments must compare the safety of candidate language models before labeled benchmarks exist in the relevant language, domain, or regulatory regime. The paper formalizes this setting as benchmark-free comparative safety scoring and makes explicit the contract under which scenario-based audits can be interpreted as deployment evidence.

    Scores are only valid under a fixed package of scenarios, rubric, auditors, judge models, sampling configuration, and rerun budget. Because no ground-truth labels are available, the authors substitute a construct-validity chain for ground-truth agreement:

  • Responsiveness to controlled safety-vs-ablation contrasts
  • Target-driven variance dominating auditor and judge-model artifacts
  • Stability across reruns
  • This chain is instantiated in SimpleAudit, a local-first scoring tool, and validated on a Norwegian safety package. Safety and ablation targets separate with AUROC values of 0.89 to 1.00; target identity is the dominant variance component (η² ≈ 0.52); severity distributions stabilize after ten reruns. Applying the same chain to Petri shows it is permissive across both tools.

    Substantive differences instead appear upstream, in declaration-contract enforcement and deployment adaptation.

    A Norwegian public-sector procurement case comparing Borealis and Gemma 3 demonstrates the resulting evidence in practice: which model is safer depends on the scenario category and risk metric. Therefore scores, matched differences, key ratios, uncertainty, and the auditors and judge models used must be reported together rather than compressed into a single ranking.

    Key Takeaways

  • Defines validity conditions for LLM safety comparisons when no labeled benchmark exists
  • Provides an open methodology (SimpleAudit) validated with strong AUROC separation and stable reruns
  • Argues against single-number safety rankings in favor of multi-faceted reporting

Tags

#llm-safety#evaluation#benchmarks#arxiv#machine-learning#model-auditing#simpleaudit

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619695