English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Teaching LightGBM to Decode Ads: CTR Prediction with Gradient Boosting and Categorical Encoding on the Criteo Dataset

Forum topic · ✨步子哥 · 2025-11-27

Summary

This Chinese-language forum post walks through a complete click-through rate (CTR) prediction experiment on the Criteo advertising dataset using Microsoft's LightGBM framework. It explains the theory behind gradient boosting decision trees and LightGBM's two key innovations from its NeurIPS 2017 paper: Gradient-based One-Side Sampling (GOSS) and Exclusive Feature Bundling (EFB), which deliver roughly 8x speed and memory improvements. The core challenge addressed is encoding 26 hash-based categorical features into numeric form. The author compares two pipelines: a baseline using ordinal encoding (39 features, test AUC 0.7655, LogLoss 0.4683) and an advanced pipeline using sequential target encoding, count features, and binary encoding (268 features, test AUC 0.7759, LogLoss 0.4603), a +0.0103 AUC gain. Because the data is time-ordered, temporal train/validation/test splits are used, and target encoding only leverages prior samples to avoid label leakage. The post includes model hyperparameters (num_leaves=64, learning_rate=0.15, feature_fraction=0.8, early stopping), evaluation discussion of AUC and LogLoss under class imbalance, and future directions such as AutoML encoder selection and tree-neural hybrid models.

Teaching LightGBM to Decode Ads: CTR Prediction with Gradient Boosting and Categorical Encoding

This post is a structured English summary of a detailed Chinese technical article about click-through rate (CTR) prediction on the Criteo dataset using LightGBM and categorical encoding techniques.

Key points

  • Task: Binary classification to predict whether a user will click an ad, using 40 columns — 13 numeric features (I1–I13) and 26 hash-encoded categorical features (C1–C26), plus a binary label.
  • Framework: LightGBM (Microsoft), based on the NeurIPS 2017 paper by Ke et al., whose two innovations — Gradient-based One-Side Sampling (GOSS) and Exclusive Feature Bundling (EFB) — yield 8x+ speed and memory gains over traditional GBDT.
  • Core challenge: Tree models cannot directly consume hash-string categorical features; a categorical encoding strategy is required.
  • Experimental setup

  • Data is time-ordered, so a temporal split was used: 80% train (80,000 samples), 10% validation, 10% test (10,000 each) — a fraction of the original 45M+ sample Criteo competition set.
  • Metrics: AUC (ranking quality, robust to class imbalance) and LogLoss (probability calibration).
  • Class balance: ~22.4% positive samples (17,958 positive / 62,042 negative).
  • LightGBM hyperparameters: num_leaves=64, learning_rate=0.15, feature_fraction=0.8, early_stopping_rounds=20.
  • Encoding strategies compared

    Encoding methods discussed (via the category_encoders library and custom code):

    1. Ordinal encoding — map each category to an integer; fast and memory-light; tree splits make the fake ordering harmless. 2. Target encoding — encode a category as its historical empirical CTR: TE(c) = Σ I(x_j=c)·y_j / Σ I(x_j=c). 3. Sequential target encoding — use only prior samples (j < i) to compute encodings, preventing label leakage and capturing concept drift; paired with sequential count features reflecting category popularity over time. 4. Binary encoding — split the ordinal ID into ~⌈log2(K)⌉ binary columns for richer non-linear structure. 5. Count encoding — use category frequency to distinguish mainstream vs. long-tail categories.

    The advanced pipeline (lgb_utils.NumEncoder) performs five steps: low-frequency filtering (rare categories → "LESS", missing → "UNK", numeric NaNs → column mean), ordinal encoding, sequential target + count encoding, manual binary encoding, and feature concatenation — expanding 39 raw features to 268.

    Results

    | Strategy | Features | Rounds | Test AUC | Test LogLoss | AUC gain | |---|---|---|---|---|---| | Baseline ordinal encoding | 39 | 18 | 0.7655 | 0.4683 | — | | Advanced sequential encoding | 268 | 43 | 0.7759 | 0.4603 | +0.0103 |

    Key insights:

  • Higher dimensionality did not cause overfitting thanks to feature_fraction and early stopping; the model instead extracted finer temporal patterns (hence the longer convergence).
  • The gains came from information quality, not quantity: identity (ordinal), statistics (target), non-linearity (binary), and popularity (count) each contributed distinct signals.
  • Both AUC (ranking) and LogLoss (calibration) improved — valuable for ad bidding and ROI estimation in practice.
  • Encoding pitfalls noted: target encoding can overfit rare categories (mitigate with smoothing, leave-one-out, or nested CV); random splits on time-ordered data violate causality.
  • Future directions

  • AutoML-driven encoder selection (FLAML, Optuna)
  • Tree/neural hybrids (e.g., GBNet) combining embeddings with GBDT
  • Real-time/online updating of encoding statistics
  • Fairness-aware encoding to avoid systematic bias

References

1. Ke, G., et al. (2017). "LightGBM: A Highly Efficient Gradient Boosting Decision Tree." *NeurIPS 30*, pp. 3146–3157. 2. Microsoft LightGBM: https://github.com/microsoft/LightGBM 3. category_encoders: https://github.com/scikit-learn-contrib/category_encoders 4. Criteo Display Ad Challenge: https://www.kaggle.com/c/criteo-display-ad-challenge

Conclusion: Going from 40 raw columns to a 0.7759 AUC, the experiment shows that thoughtful categorical encoding — especially leakage-free, time-aware target encoding — is as important as the model itself when LightGBM faces categorical-heavy tabular data.

Tags

#lightgbm#ctr-prediction#criteo#categorical-encoding#gradient-boosting#target-encoding#machine-learning#kaggle

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/176415025