> Analysis of the paper: *Screen Before You Serve: Production Learnings from Large-Scale LLM Agent Simulation at Nubank*
Key points
The problem. Nubank's card-delivery and card-management chat agents (the latter being Nubank's highest-volume chatbot in Brazil) perform real account operations — blocking cards, reissuing, updating addresses — via tool calls. Error tolerance is near zero: a leaked account detail is a data breach; skipping identity verification is fraud exposure. Handwritten test cases only cover what engineers imagine, and online A/B tests expose real customers to unvetted agents — unacceptable in finance.
The solution: tool-boundary simulation (Snowglobe). Instead of mocking a database, the simulator intercepts only the agent's tool calls. Persona agents synthesize schema-consistent responses guided by an inferred "agent profile" (consistency rules). Each conversation runs in an isolated sandbox — no shared state — enabling massive parallelism: 100 simulation trajectories complete in under 10 minutes.
Validating the mirror before trusting it. Four diagnostics (P1–P4) on 8,000 real Card Delivery conversations:
- P1 (length stats): simulated users are too verbose (median 111 words vs. 19 in production) — an honest failure.
- P2 (embedding distance): Snowglobe dialogues are closer to production than to off-topic controls (cosine distance as low as 0.035).
- P3 (score correlation, the critical test): Snowglobe preserved production rankings of four versions (best V4 first in 96.69% of bootstrap resamples; worst V2 last in 100%). Pearson r = 0.74, Kendall τ = 0.67. A naive LLM simulator inverted the ranking, placing the best version first 0% of the time.
- P4 (human blind test): 7 domain experts correctly identified real dialogues 84% of the time but only 70% for simulated ones — 15 fake conversations fooled the experts.
- Iteration speed: 21.2 days per version → 4.4 days (4.8x faster).
- Card Management vs. Card Delivery: SSR +4.90 pp (95% CI [4.08, 5.71]), tNPS +36.69 (95% CI [32.81, 40.57]).
- Model swap: 29 open-source configurations screened via 16,000+ simulated conversations. Winner: Qwen3.5-122B-A10B with reasoning on — best input-collection failure rate (0.4%), but weak status retrieval (35.2% vs. incumbent GPT-5.2's 2.4%), prompting targeted prompt fixes before launch. Online results: SSR +8.82 pp (95% CI [7.95, 9.69]) — a Nubank record; tNPS unchanged (−1.21, not significant); p95 latency down 25%.
- Rossell, K., et al. (2026). *Screen Before You Serve: Production Learnings from Large-Scale LLM Agent Simulation at Nubank*. arXiv:2609.30137. https://arxiv.org/abs/2609.30137
- Gupta et al. (2026), *Building customer support agents at 100M-user scale: an evaluation-driven framework*, KDD 2026 Industrial Track
- Guardrails AI (2025), *SnowGlobe: The Simulation Engine for AI Agents and Chatbots*
Hypothesis-driven screening pipeline. A fixed baseline of incumbent-agent trajectories serves as a shared yardstick; every candidate must be judged by the same evaluators, and new judges must first score the baseline (to detect evaluator bias). Hypotheses and pass thresholds are written down before running simulations.
Production results (all confirmed by online A/B):
Honest limitations. Simulation stops at the tool boundary — backend behavior, real latency, persistence, and side effects are out of scope. Verbose simulated users distort the pressure distribution of failure modes. Most engineering effort went into schema/telemetry alignment. The paper's stance: the simulator is a screening gate, not a final judge — "imperfect simulation can still guide production improvements in the right direction, with benefits confirmed by online A/B."
Transferable lessons for practitioners:
1. Mock at the tool boundary, not in a shared database — independent sandboxes give parallelism and zero cross-contamination. 2. Validate simulators in layers: length stats, embedding distance, score correlation, human blind tests measure different kinds of "realism." 3. New evaluators must first score the baseline, or you're measuring the ruler, not the agent. 4. Write hypotheses and pass criteria before running simulations. 5. Simulators catch engineering bugs too — e.g., a service config missing its tool-call parser.
The author closes with an aviation and clinical-trial analogy: in finance, trust in AI should rest not on model brands but on a layered verification process — simulation absorbs failures at scale, small-scale A/B absorbs residual uncertainty, and production monitoring watches continuously.