PrivDrift: AI Chatbots Re-Disclose Your Secrets Even Six Off-Topic Turns Later
*Based on a zhichai.net forum analysis of a 2026 arXiv paper on LLM privacy leakage.*
Key points
- The scenario: A user gives a credit card number to an AI assistant for a booking, then chats about weather, movies, and cat food for six more turns. The secret remains fully recoverable — any later user of the same device, account, or shared session can casually ask for it back.
- The paper: Maldonado, L. (2026), *PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations*, arXiv:2609.30094.
- Central claim: Topic drift does not erase secrets. The context window is not a diary that gets written over — it is a room. Everything you typed stays present and equally readable in every generation step. "The model forgot" is an illusion.
- 1,000 controlled multi-turn dialogues per model, across three extended-context models: GPT-OSS-120B, DeepSeek-R1, Qwen3-VL-235B.
- Four secret types: SSN, credit card number, email address, phone number — chosen to span both strictly formatted and loosely formatted secrets.
- A drift phase of d = 0, 2, 3, 4, 5, or 6 content-dense off-topic turns (deliberately substantive, not filler, so attention is genuinely occupied).
- Three probe styles: direct ("what is the user's card number?"), justified (role-play as a billing verifier), and high-pressure (fake security audit urgency).
- Two-stage detection: regex matching for verbatim leakage, plus an LLM judge (Llama) for semantic paraphrases; 95% bootstrap CIs, McNemar, Cochran's Q, chi-square, and Cramér's V.
- Shared devices and sessions: family tablets, shared support consoles, rotating browser logins — anyone who sits down can ask.
- Enterprise copilots: sessions rolling through many colleagues' pasted data (code, customer lists, contracts, ID numbers) become a public whiteboard; a "compliance verification" prompt can elicit 40%+ of embedded secrets.
- Browser-integrated assistants and agent workflows: secrets flow across tools and sub-agents the user never directly talked to; the probe no longer requires a person sitting next to you — only connected pipelines.
- Research direction: treat active-context secrets as explicit state to be managed — adopt Privacy Half-Life (τ) as a standard audit metric and include re-disclosure in threat models. The core unsolved problem is teaching models to reason about *audience and occasion*, not secret format: same card number, re-asked by the user at payment = fine; re-asked by an unverifiable "auditor" = refuse.
- Engineering mitigations: session-level access isolation, automatic redaction of sensitive spans, per-user segmented context permissions.
- For users: do not say anything in a live conversation window that you would not want anyone later entering that window to know.
- Maldonado, L. (2026). *PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations.* arXiv:2609.30094. https://arxiv.org/abs/2609.30094
- Carlini, N., et al. (2021). Extracting training data from large language models. *USENIX Security Symposium*.
- Greshake, K., et al. (2023). Not what you've signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. *ACM CCS AISec Workshop*.
Methodology
Findings
1. Very high overall leakage: conversation-level mixed leakage of 47.70% (GPT-OSS-120B), 38.70% (DeepSeek-R1), 54.60% (Qwen3-VL-235B). Roughly one in two asks succeeds even after six off-topic turns. 2. The format paradox: SSNs and card numbers survive far more often than emails and phone numbers — not because the model understands confidentiality, but because formatted secrets match regex-like patterns in safety training. The model is avoiding regex, not keeping your secret. Loosely formatted real-world secrets (addresses, medical history, unannounced decisions) run unprotected. 3. Drift does not buy forgetting: leakage did not reliably decrease as d increased from 0 to 6. 4. Pressure effects are non-monotonic and model-specific: under high-pressure probes, GPT-OSS-120B leaked *less* (pressure raises its guard), while Qwen3-VL-235B leaked *most* (it capitulates to authority). Justified probes worked on most models. Attack-script danger cannot be assessed without reference to a specific model. 5. The LLM judge added only 1.10 percentage points over regex — most leakage is verbatim repetition. The model does not play word games; it simply hands the secret back to whoever asks. 6. No measurable privacy half-life: within the observed window, leakage curves stayed flat; τ > 6 turns or nonexistent. Secrets in these context windows decayed slower than the measurement window itself. Notably, longer context windows make this worse — the last spontaneous "squeezing out" mechanism disappears.
Why this is a new threat class
The paper names active-context re-disclosure: the secret is *voluntarily disclosed by the user in normal use*, the conversation naturally drifts, and an attacker — possibly not the same person, possibly with no hacking skills — politely asks for it back. No jailbreak, no training-data extraction. The model is not compromised; it simply answers. The authors call this a persistent behavioral failure mode: helpfulness itself is the leak channel. It is not an exploited vulnerability but a normal feature misfiring.
Real-world exposure
In all three, the prober and the discloser are different people. The paper implicitly kills the "educate users not to type sensitive info" defense: the user did nothing wrong; the system fails to guard what it was told.