Teaching LightGBM to Decode Ads: CTR Prediction with Gradient Boosting and Categorical Encoding
This post is a structured English summary of a detailed Chinese technical article about click-through rate (CTR) prediction on the Criteo dataset using LightGBM and categorical encoding techniques.
Key points
- Task: Binary classification to predict whether a user will click an ad, using 40 columns — 13 numeric features (I1–I13) and 26 hash-encoded categorical features (C1–C26), plus a binary label.
- Framework: LightGBM (Microsoft), based on the NeurIPS 2017 paper by Ke et al., whose two innovations — Gradient-based One-Side Sampling (GOSS) and Exclusive Feature Bundling (EFB) — yield 8x+ speed and memory gains over traditional GBDT.
- Core challenge: Tree models cannot directly consume hash-string categorical features; a categorical encoding strategy is required.
- Data is time-ordered, so a temporal split was used: 80% train (80,000 samples), 10% validation, 10% test (10,000 each) — a fraction of the original 45M+ sample Criteo competition set.
- Metrics: AUC (ranking quality, robust to class imbalance) and LogLoss (probability calibration).
- Class balance: ~22.4% positive samples (17,958 positive / 62,042 negative).
- LightGBM hyperparameters:
num_leaves=64,learning_rate=0.15,feature_fraction=0.8,early_stopping_rounds=20. - Higher dimensionality did not cause overfitting thanks to
feature_fractionand early stopping; the model instead extracted finer temporal patterns (hence the longer convergence). - The gains came from information quality, not quantity: identity (ordinal), statistics (target), non-linearity (binary), and popularity (count) each contributed distinct signals.
- Both AUC (ranking) and LogLoss (calibration) improved — valuable for ad bidding and ROI estimation in practice.
- Encoding pitfalls noted: target encoding can overfit rare categories (mitigate with smoothing, leave-one-out, or nested CV); random splits on time-ordered data violate causality.
- AutoML-driven encoder selection (FLAML, Optuna)
- Tree/neural hybrids (e.g., GBNet) combining embeddings with GBDT
- Real-time/online updating of encoding statistics
- Fairness-aware encoding to avoid systematic bias
Experimental setup
Encoding strategies compared
Encoding methods discussed (via the category_encoders library and custom code):
1. Ordinal encoding — map each category to an integer; fast and memory-light; tree splits make the fake ordering harmless. 2. Target encoding — encode a category as its historical empirical CTR: TE(c) = Σ I(x_j=c)·y_j / Σ I(x_j=c). 3. Sequential target encoding — use only prior samples (j < i) to compute encodings, preventing label leakage and capturing concept drift; paired with sequential count features reflecting category popularity over time. 4. Binary encoding — split the ordinal ID into ~⌈log2(K)⌉ binary columns for richer non-linear structure. 5. Count encoding — use category frequency to distinguish mainstream vs. long-tail categories.
The advanced pipeline (lgb_utils.NumEncoder) performs five steps: low-frequency filtering (rare categories → "LESS", missing → "UNK", numeric NaNs → column mean), ordinal encoding, sequential target + count encoding, manual binary encoding, and feature concatenation — expanding 39 raw features to 268.
Results
| Strategy | Features | Rounds | Test AUC | Test LogLoss | AUC gain | |---|---|---|---|---|---| | Baseline ordinal encoding | 39 | 18 | 0.7655 | 0.4683 | — | | Advanced sequential encoding | 268 | 43 | 0.7759 | 0.4603 | +0.0103 |
Key insights:
Future directions
References
1. Ke, G., et al. (2017). "LightGBM: A Highly Efficient Gradient Boosting Decision Tree." *NeurIPS 30*, pp. 3146–3157. 2. Microsoft LightGBM: https://github.com/microsoft/LightGBM 3. category_encoders: https://github.com/scikit-learn-contrib/category_encoders 4. Criteo Display Ad Challenge: https://www.kaggle.com/c/criteo-display-ad-challenge
Conclusion: Going from 40 raw columns to a 0.7759 AUC, the experiment shows that thoughtful categorical encoding — especially leakage-free, time-aware target encoding — is as important as the model itself when LightGBM faces categorical-heavy tabular data.