English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Paper: Cross-Entropy Games and Frost Training

Forum topic · 小凯 · 2026-05-29

Summary

This arXiv paper (2605.27701) by Arthur Renard, Franck Gabriel, Valentin Hartmann, et al. introduces Frost Training, a method that improves Monte Carlo-based policy optimization for a broad family of LLM-as-a-judge tasks the authors call Cross-Entropy Games. The core idea is to exploit the gradient of the reward function in embedding space—a signal previously used in the Greedy Coordinate Gradient (GCG) jailbreaking technique. The authors show for the first time that this signal can also boost model training. They validate the approach with GRPO training on maximum-likelihood infilling tasks. Results show Frost Training improves the model's ability to generate high-scoring outputs, achieves higher maximum scores in best-of-k settings, and trains faster than baseline methods. The work was shared on zhichai.net on 2026-05-29.

Paper Overview

  • Research area: AI
  • Authors: Arthur Renard, Franck Gabriel, Valentin Hartmann, et al.
  • Published: 2026-05-28
  • arXiv: 2605.27701
  • Key Contributions

  • Presents Frost Training, a method for improving Monte Carlo-based policy optimization for a large family of LLM-as-a-judge tasks called Cross-Entropy Games.
  • Key idea: exploit the gradient of the reward function in embedding space.
  • This gradient signal is already used in the Greedy Coordinate Gradient (GCG) jailbreaking technique; the paper demonstrates for the first time that it can also be used to boost model training.
  • Validation is performed using GRPO training for maximum-likelihood infilling.
  • Results

  • Improved ability of the model to generate high-scoring outputs.
  • Higher maximum scores in a best-of-k setting.
  • Increased training speed.

Original Abstract

> We present Frost Training, a method for improving Monte Carlo-based policy optimization for a large family of LLM-as-a-judge tasks called Cross-Entropy Games. The key idea is to exploit the gradient of the reward function in embedding space. This signal is used in the Greedy Coordinate Gradient (GCG) jailbreaking technique; we demonstrate for the first time that it can also be used to boost model training. We validate our method using GRPO training for maximum-likelihood infilling. Frost Training improves the model's ability to generate high-scoring outputs, reaching higher maximum scores in a best-of-k setting, and does so at an increased speed.

---

*Auto-collected on 2026-05-29.*

Tags

#ai#llm-as-a-judge#reinforcement-learning#frost-training#grpo#arxiv-paper#policy-optimization#embedding-gradients

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980517