> Paper: https://arxiv.org/abs/2608.07460 > Code: https://github.com/ananya-sahu/CreativeInstruct
A Counterintuitive Fact
Ask GPT-4 to write three different stories and you get three versions. Read them closely: the character names change, the settings change, but the narrative skeleton is nearly identical.
This is not a coincidence. Post-training systematically kills output diversity while improving quality. The model becomes "better," but also more "like itself." This paper quantifies the cost: across multiple semantic diversity metrics, SFT/RLHF-trained models lose on average 20-40% diversity compared to their base models.
The problem is that many tasks cannot rely on "good" alone:
- Story generation: requires genuine narrative branching, not reskinned rewrites
- Reinforcement learning: policy exploration needs enough entropy in the action distribution, otherwise the agent falls into local optima
- Data augmentation: if LLM-generated training data all looks the same, augmentation accomplishes nothing
- Lexical: n-gram, distinct-n, self-BLEU — only detect literal repetition
- Semantic: embedding distance, sentence similarity — detect meaning repetition
- Diversity: 70.3% higher on LLM-GED than the SFT baseline, and 28% higher than a multi-model ensemble baseline
- Quality: human evaluation shows parity with the post-trained model, no sacrifice
- Efficiency: only one model needed at inference, no ensembling
- Creativity mode depends on the base model's existence. Without the base model, you cannot construct the contrastive data — a limitation for real deployments.
- LLM-GED is LLM-evaluated, risking circularity. The authors align it with human evaluation, but the metric's reliability needs more validation.
- The RL environments are fairly simple. Whether diversity gains translate to reward gains in more complex RL settings (multi-agent, long horizon) remains to be seen.
CreativeInstruct's core question is simple: can a model retain post-training quality while recovering the base model's creativity?
Method: A Learnable "Creativity Switch"
The idea is strikingly simple. During instruction fine-tuning, the authors insert a special token, [StartCreativity], telling the model: "from here on, generate in the base model's style."
The approach has three steps:
Step 1: Collect contrastive data. For the same instruction, generate outputs with the base model (high diversity, lower quality) and the post-trained model (high quality, lower diversity). This forms the creativity-quality spectrum.
Step 2: Build labeled training data. Base model outputs are labeled "creativity mode," post-trained outputs "quality mode," and the model is trained to switch to creativity mode when it sees [StartCreativity]. Essentially, the model learns a conditional distribution: given the creativity signal, output like the base model; without it, output like the post-trained model.
Step 3: Switch on demand at inference. Want creativity? Insert [StartCreativity]. Want stability? Leave it out. One model, two modes.
This is fundamentally different from "temperature tuning." Temperature merely resamples an existing distribution, changing how aggressively you sample; CreativeInstruct changes the distribution itself, biasing the model toward the base model's wider one. To use an analogy: temperature is changing your dive pose in the same pool; CreativeInstruct is switching to a different pool.
LLM-GED: A Metric That Actually Measures Narrative Diversity
The paper's second contribution is LLM-GED (Graph Edit Distance). Existing diversity metrics fall into two camps:
But story diversity is not just lexical or sentence-level. Two stories can use completely different words yet tell the same "hero's journey" narrative structure. LLM-GED uses an LLM to parse stories into "abstract narrative graphs" (nodes are narrative units, edges are relations), then compares graphs via edit distance. Larger graph differences mean higher narrative-level diversity.
This resembles AST comparison in software engineering: don't compare code literally, compare syntax trees — except here the syntax tree becomes a narrative graph.
Results: Quality Unchanged, Diversity Way Up
On narrative generation, CreativeInstruct delivers:
The RL experiments are especially interesting. Using CreativeInstruct as policy initialization for RL training, the authors find exploration efficiency improves markedly — final reward is 29% higher than standard SFT initialization across several RL environments. This confirms an intuition: the exploration bottleneck in RL is fundamentally a lack of policy diversity. If the initial policies all cluster in the same region, no exploration algorithm can escape.
Why This Paper Is Worth Reading
1. It puts "post-training kills diversity" on the table. Many people know SFT narrows models, but few have systematically quantified the cost. CreativeInstruct gives concrete numbers: a 20-40% diversity loss.
2. The [StartCreativity] design is remarkably engineering-friendly. No architecture changes, no extra inference cost — just one token. This "minimally invasive" design philosophy is worth learning from.
3. LLM-GED has independent value. Evaluating LLM output diversity has long been difficult; the graph-edit-distance approach is closer to "narrative-level difference" than purely lexical or semantic metrics.
4. The RL experiments reveal a deep connection: creativity is not just "writing prettier stories" — it is a foundational resource for exploration. A policy without creativity is, in RL terms, "spinning in a local optimum."
An Honest Assessment
The paper is not without problems:
The answer appears to be: yes — and it can be recovered.
---
Paper: Sahu, Bansal, Stengel-Eskin. *CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity*. arXiv:2608.07460, 2026.
Code: https://github.com/ananya-sahu/CreativeInstruct