Imagine watching a film without seeing the picture. A voice tells you: "Mark walks into the room, wearing a suit, expression serious." "It's raining outside the window." "He puts a cup of coffee on the table."
This is Audio Description (AD)—a service that narrates key visual information for visually impaired audiences. Doing it well is far harder than it sounds:
- Descriptions must fit into gaps between dialogue without interrupting it
- You must choose what is worth describing—expressions, actions, scene changes, on-screen text—each with different priority
- Wording must be concise; a sentence that runs too long will overlap the next line of dialogue
- Candidate generation: an LLM generates multiple candidate descriptions per scene (varying length and detail)
- Scoring: each candidate is scored on content relevance, temporal fit, and linguistic quality
- Constrained solving: MILP selects the optimal combination under global constraints:
- Descriptions must not overlap dialogue
- Each description must fit entirely within one gap
- Minimum spacing between adjacent descriptions
- Overall description density must not be too high (avoiding information overload)
- Output: the selected descriptions and their timing
- CIDEr/METEOR: n-gram overlap with reference descriptions
- SODA-M/SODA-T: content coverage and timing quality
- QEval/QEval-T: question-answering-based evaluation—can questions about the film be answered from the descriptions, and is the timing correct
- MILP optimizer wins across the board. Gemini MILP reaches QEval=45.9 and QEval-T=25.5, significantly better than all other automatic systems (p<0.01).
- The largest gain is in timing. Gemini MILP achieves SODA-T of 55.7 vs only 36.0 for pure Gemini prompting. Timing improved far more than content quality—the bottleneck is not "what to say" but "when to say it".
- CIDEr actually drops. MILP's CIDEr (13.6) is lower than pure LLM prompting (19.0). But the authors argue this is a feature, not a bug: CIDEr rewards wording overlap with references, while MILP produces more concise, more fragmented descriptions conveying the same content differently. Low lexical overlap does not mean low quality.
- LLM schedulers underperform MILP. LLMs lack global vision for timing—they greedily fill each gap without considering overall pacing.
- Human experts still substantially outperform all automatic systems. Expert QEval-T is 61.2 vs 25.5 for the best MILP.
What to describe, when to describe it, and how to phrase it—these three decisions are coupled and cannot be made independently. In *Audio Description as Constrained Global Optimization*, Sterner and Lapata (University of Glasgow) formally model this as a constrained optimization problem.
Three Coupled Decisions
Traditional automatic AD pipelines separate the decisions: one model picks the content, another decides placement, a third generates text. This staged approach ignores the coupling between constraints.
Example: you decide to describe "Mark is wearing a suit"—but the next dialogue gap is only 2 seconds, and the description takes 3 seconds to speak. You must either shorten it ("suit"), postpone it to a longer gap, or drop it. This trade-off spans content, timing, and wording simultaneously.
The authors solve this with Mixed-Integer Linear Programming (MILP):
The REFRAMED Benchmark
The authors built REFRAMED, an evaluation benchmark with audio description data from multiple films. Evaluation covers both text quality and timing:
Key Findings: Optimization Beats Prompting
Experiments compared several systems: pure LLM prompting (Qwen 3.5, Gemini 3.1), an LLM scheduler (LLM decides timing itself), and the MILP optimizer (LLM generates candidates, MILP optimizes globally).
Why Global Optimization Beats End-to-End
The result seems counterintuitive: if LLMs can do everything, why does splitting the task into "generate candidates + optimize selection" work better?
The answer is constraint satisfaction. LLMs are strong at generating natural language but weak at satisfying hard constraints. Ask an LLM to "fit a description into a 2.3-second gap without exceeding it" and it will often violate the limit. MILP is natively built for constraint satisfaction.
This mirrors the idea of separating judgment from gating: the LLM handles generation (judgment), MILP handles selection and scheduling (gating). The LLM's creativity is unconstrained; MILP's constraint satisfaction doesn't depend on the LLM's "self-discipline." Each does what it does best.
Commentary
What's most admirable about this paper is that it resists the "end-to-end LLM solves everything" narrative. Decomposing a problem into "what LLMs are good at" and "what classical algorithms are good at" requires clear-eyed problem analysis.
The finding that CIDEr drops while QEval rises is worth remembering: the choice of evaluation metric changes the conclusion. If you only look at CIDEr, MILP looks worse than pure prompting; with QEval, the conclusion reverses. Metrics measure what they define—not necessarily what you care about.
The finding that timing improved far more than content quality also carries meaning: in multimodal generation, the bottleneck often lies in coordination, not content generation. LLMs already produce decent descriptive text; they just don't know when to say it.
Human expert QEval-T (61.2) is more than double the best MILP (25.5)—a sobering gap. But lifting the baseline from 19.0 to 25.5 is real, tangible progress.
---
Paper: Sterner, A., & Lapata, M. (2026). Audio Description as Constrained Global Optimization. arXiv:2609.30121 Link: https://arxiv.org/abs/2609.30121