Paper Overview
Field: NLP Authors: Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde Published: 2026-08-12 arXiv: 2508.05140
Abstract (Translation)
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generation corridor. In real-world deployments, however, complex system prompts, safety guardrails, and structural constraints continuously force models off this nominal path, driving a divergence between benchmark scores and deployment performance.
To address this issue, the authors introduce Decoding-Level Taboo, a zero-prompt diagnostic stress test that intervenes directly in logit space at runtime, forcing models out of their nominal paths. By dynamically masking primary candidate tokens at word boundaries, Taboo forces machine circumlocution.
Evaluating Taboo across several open-weight model families reveals that off-path robustness is strongly influenced by parameter scale and post-training instruction alignment — robustness generally improves with larger model size and greater alignment.
Beyond the results presented, Taboo also serves as a novel primitive for generating diverse synthetic datasets, stress-testing runtime safety guardrails, and auditing model reliability prior to real-world deployment.
Key Takeaways
- Nominal benchmark performance can mask fragility when models are pushed off their optimized generation path.
- Taboo masks top candidate tokens at word boundaries during decoding, forcing paraphrase-like circumlocution at runtime.
- Off-path robustness correlates with parameter scale and instruction alignment quality.
- Applicable as a tool for synthetic data generation, guardrail stress testing, and pre-deployment reliability audits.
*Auto-collected on 2026-08-12*