> 📌 This is a GEO-optimized English version of a zhichai.net topic.
DWT-Fusion: Treating Token Probabilities as Signals to Detect AI-Generated Text
> Paper: DWT-Fusion: A Signal-Based Framework for Training-Free LLM-Generated Text Detection > Authors: Mehmet Batuhan Özdaş, Murat Osmanoğlu (Ankara University, Cyber Security Vocational School) > arXiv: 2607.22026 (July 24, 2025)
The Problem
Suppose you receive a student essay and want to know whether it was written by a human or generated by ChatGPT. With an open-source language model (e.g., GPT-2 or LLaMA), you can compute the log probability of every token. But how do you turn these probabilities into a verdict?
The intuitive approach—averaging log probabilities over the whole text—has a fatal flaw: fluent human writing also scores high, and AI text deliberately mimicking human "disfluency" scores low. Global averages mask local differences.
DWT-Fusion's core insight: look at local and multi-scale probability fluctuations, not global averages.
From Text to Signal
The first step converts text into a 1D signal: use a proxy model (e.g., GPT-2 Medium) to compute the conditional log probability of each token. This sequence of numbers is treated like an audio waveform, then decomposed with a Discrete Wavelet Transform (DWT).
Why Wavelets?
Fourier transforms tell you which frequencies exist but not where. Wavelets provide both time (position) and frequency (scale) information. This matters because AI vs. human differences in token probability are often:
- Local: AI text may be very "smooth" in some paragraphs but deliberately rough in others; averaging cancels these out.
- Multi-scale: AI probability fluctuations may be small at short scales but rhythmic at longer scales, unlike human patterns.
- MAGE AUROC of 0.7471 implies substantial false-positive rates—insufficient for high-stakes uses like academic integrity enforcement.
- Proxy model dependence — probabilities from GPT-2 may not discriminate well against GPT-4-generated text; the paper did not test GPT-4 outputs.
- Adversarial robustness untested — an adversary aware of the method could generate text with human-like multi-scale fluctuations.
- Gap vs. supervised SOTA — RADAR and Ghostbuster reach 0.85+ AUROC on M4; DWT-Fusion's 0.8477 approaches but does not exceed them.
DWT decomposes the signal into approximation coefficients (low-frequency global trend) and detail coefficients (high-frequency local variation) across levels, capturing both.
Three Wavelet-Domain Scores
1. First-level detail energy — fine-grained local fluctuation intensity; AI text tends to be more uniform here. 2. Multilevel detail energy — total detail energy across all scales. 3. Window-energy variability — variation of energy across windows, capturing unevenness of local fluctuations.
Each score works independently or in combination.
Calibration-Guided Voting Fusion
Beyond single scores, DWT-Fusion computes scores under multiple wavelet configurations (different wavelet families, decomposition depths), each casting a vote. Calibration weights emphasize better-performing configurations—akin to ensemble learning, combining weak detectors into a strong one.
Results
Tested on three datasets: HC3 (Chinese detection, easier), M4 (multi-generator, multi-domain), and MAGE (multi-generator, multi-domain, multilingual, hardest).
Best single scores:
| Dataset | AUROC | |---------|-------| | HC3 | 0.9872 | | M4 | 0.8185 | | MAGE | 0.7138 |
After calibration-guided voting fusion:
| Dataset | AUROC | |---------|-------| | HC3 | 0.9919 | | M4 | 0.8477 | | MAGE | 0.7471 |
Notably, these are training-free results—no labeled data, no classifier training, just one open-source proxy model.
DFT Baseline Comparison
Replacing DWT with a Discrete Fourier Transform (DFT)—which lacks time-domain information—degrades performance on all datasets. This directly demonstrates that the value comes from wavelets' joint time-frequency structure, not from signal processing per se.
Engineering Significance
1. Zero training cost — no labeled data or classifier training; deployment is cheap. 2. Interpretability — all three scores have clear physical meaning, enabling analysis of why text was flagged. 3. Configurability — wavelet family and decomposition depth can be tuned per scenario.
Honest Assessment
One-Line Summary
DWT-Fusion treats token probabilities as a 1D signal and uses wavelet transforms to capture local and multi-scale fluctuations—training-free, requiring only one open-source model, approaching supervised detectors' performance.
---
Paper: https://arxiv.org/abs/2607.22026 HTML version: https://arxiv.org/html/2607.22026v1 Code: not yet released