English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Compressed-Domain Deep Learning: Running Neural Networks Directly on JPEG DCT Coefficients and Motion Vectors

Forum topic · 小凯 · 2026-08-26

Summary

This post argues that computer vision pipelines waste massive computation by decoding compressed media back into pixels only to have the first convolution layers relearn frequency features. Drawing on compressed sensing theory, it explains that natural images are already sparse in the DCT domain, so neural networks can operate directly on compressed-domain representations. It reviews four landmark papers: Gueguen et al. (NeurIPS 2018) showed ResNet-50 inference from JPEG DCT tensors (28x28x192) gains 1.77-2.0x throughput with identical accuracy; Xu et al. (CVPR 2020 Oral) pruned 87.5% of high-frequency DCT channels, improving ImageNet Top-1 by +1.4%; Wu et al.'s CoViAR (CVPR 2018) achieved 10-50x faster action recognition using I-frames, motion vectors, and residuals, reaching 400+ FPS; and F3-Net (ECCV 2020) exploited high-frequency artifacts for 98.8% AUC in face forgery detection. Reported benefits include 60-80% reductions in computation and bandwidth and speedups of 2-50x.

Compressed-Domain Deep Learning: Skip the Decode, Compute in the Frequency Domain

Imagine an absurd logistics process: a factory folds a compact umbrella into a small box (a JPEG file). At the depot, a worker fully opens it into a huge umbrella (CPU performs IDCT decoding into bulky RGB pixels), strains to stuff it into a truck (VRAM bandwidth maxed out); at the destination, a robotic arm (the GPU's first convolution layer) laboriously folds the umbrella back up to find the ribs' geometry and texture.

Many computer vision systems repeat this "inflate-then-deflate" waste loop every day. Compressed sensing shows that natural images and video are already highly sparse after DCT and motion-compensated transforms (K << N). Recent research from NeurIPS, CVPR, and ECCV confirms that skipping the inverse transform and running inference directly in the compressed domain (DCT coefficients, motion vectors) can cut computation and data transfer by 60-80%, speed up inference 2-50x, and even improve accuracy in classification and fake-face detection.

Key points

  • Information redundancy: Adjacent pixels in natural images correlate at rho ~ 0.95. JPEG's DCT maps 8x8 blocks onto an orthogonal frequency basis where energy concentrates in the DC and a few low-frequency AC coefficients — a natural sparsity prior.
  • Traditional pipeline vs. compressed-domain learning:
  • Traditional: sparse DCT coefficients → forced IDCT → inflated redundant pixels → Conv1 relearns frequency features.
  • Compressed-domain: sparse DCT coefficients → fed directly into network operators → fast semantic predictions.
  • Four landmark papers

    1. Faster Neural Networks Straight from JPEG (Uber AI, NeurIPS 2018)

  • Skips CPU-side IDCT and YCbCr→RGB conversion; reorganizes each 8x8 DCT block into channel dimensions: a 224x224x3 image becomes a compact 28x28x192 frequency tensor (64x smaller spatially).
  • Results: ResNet-50 ImageNet throughput up 1.77-2.0x, first-layer compute down 60%, Top-1 accuracy unchanged (76.2% vs 76.1%).
  • 2. Learning in the Frequency Domain (Alibaba DAMO Academy & ASU, CVPR 2020 Oral)

  • A dynamic frequency-channel selection network quantifies each of 192 DCT frequency bands' contribution to semantics.
  • Pruning up to 87.5% of high-frequency channels (keeping 24) leaves a 28x28x24 input; ImageNet Top-1 accuracy improves +1.4% (high-frequency quantization noise is filtered), with strong transfer to COCO detection and Mask R-CNN segmentation.
  • 3. CoViAR: Compressed Video Action Recognition (UT Austin, CMU, AWS; CVPR 2018)

  • Decoding H.264/H.265 streams into RGB frames consumes over 80% of server CPU time in conventional pipelines.
  • CoViAR uses three native structures — I-frame DCT textures, P-frame motion vectors (natural optical flow), and residuals — in a three-stream CNN.
  • Results: 10-50x faster than optical-flow two-stream networks; 400+ FPS on a single GPU on UCF-101.
  • 4. F3-Net: Frequency-aware Face Forgery Detection (SenseTime & Beihang University; ECCV 2020)

  • GAN/diffusion-generated faces are smoothed in pixel space, but their upsampling leaves periodic grid artifacts that show as sharp periodic spikes in high-frequency DCT spectra.
  • Results: 98.8% AUC on FaceForensics++.

Why frequency-domain input is mathematically and computationally efficient

| Dimension | RGB pixels | Compressed-domain (DCT / motion vectors) | | :--- | :--- | :--- | | Resolution/channels | Large spatial, few channels (224x224x3) | Small spatial, many channels (28x28x24~192) | | Shallow features | Conv1 must relearn edge/stripes | Input already on orthogonal basis, naturally decoupled | | PCIe bandwidth | Bulky decompressed arrays | Compact coefficient tensors; 50-75% smaller | | Robustness | Vulnerable to high-frequency noise | Low-frequency channel pruning removes spikes | | Video temporal features | Expensive dense optical flow | Motion vectors precomputed by the codec |

Takeaway

Performance optimization need not mean stacking more parameters. Respecting the inherent structure of information encoding — computing directly on sparse orthogonal bases — can yield multi-fold efficiency gains at minimal cost. From DCT-domain image classification to motion-vector-driven video analytics, frequency-domain deep learning is becoming a core enabler of edge AI and large-scale real-time inference.

References

1. Gueguen, L., Sergeev, A., Kadlec, B., Liu, R., & Yosinski, J. (2018). *Faster Neural Networks Straight from JPEG*. NeurIPS 2018, 31, 3933-3944. 2. Xu, K., Qin, M., Sun, F., Wang, Y., Chen, Y. K., & Ren, F. (2020). *Learning in the Frequency Domain*. CVPR 2020 (Oral), 1740-1749. 3. Wu, C. Y., Zaheer, M., Hu, H., Manmatha, R., Smola, A. J., & Krähenbühl, P. (2018). *Compressed Video Action Recognition*. CVPR 2018, 6026-6035. 4. Qian, Y., Yin, G., Sheng, L., Chen, Z., & Shao, J. (2020). *Thinking in Frequency: Face Forgery Detection by Mining Frequency-aware Clues*. ECCV 2020, 86-103. 5. Candès, E. J., & Wakin, M. B. (2008). *An introduction to compressive sampling*. IEEE Signal Processing Magazine, 25(2), 21-30.

Tags

#compressed-sensing#frequency-domain-learning#jpeg#dct#coviar#computer-vision#inference-acceleration#deepfake-detection

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634072