English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PEPS: Treating Positional Encoding as Sampling Points on Lissajous Trajectories — an AMD Paper at ACM CCGT 2026

Forum topic · ✨步子哥 · 2026-07-20

Summary

A Chinese tech forum post reviews "PEPS: Positional Encoding Projected Sampling" by Guillaume Perez, Janarbek Matai, and Takahiro Harada (AMD), published in Proc. ACM Comput. Graph. Interact. Tech., Vol. 9, No. 1 (May 2026), DOI 10.1145/3806062. The paper's core insight is geometric: each pair of (sin, cos) values produced by positional encoding traces a point on a Lissajous curve, and Proposition 2.1 proves these trajectories are unique per input coordinate. PEPS treats these projected points as legitimate query locations against a shared learned grid encoder, yielding new high-frequency representation power for implicit neural representations (INR). Pink-PEPS further allocates latent dimensions per frequency following the 1/f power law of natural images. Experiments show state-of-the-art results on Kodak image representation (PSNR 48.07 with NTC_PinkPEPS vs ~45.3 baselines), 4K texture compression (14x compression, SOTA across 8 texture types, only +12.5% inference overhead on RDNA4), and SDF representation (global IoU 0.816; the hard Pitted Stonefish instance improves from 0.39 to 0.47). Notably, PEPS matches SOTA with 25% fewer parameters and boosts hash-grid IoU from 0.547 to 0.731, suggesting it is a natural extension to InstantNGP-style encoders rather than a replacement.

Overview

This forum post reviews PEPS: Positional Encoding Projected Sampling (Perez, Matai, Harada — AMD), published in *Proc. ACM Comput. Graph. Interact. Tech.*, Vol. 9, No. 1, Article 8, May 2026. DOI: 10.1145/3806062.

Key points

  • The problem: Plain MLPs suffer from spectral bias — fitting a 1024×1024 image with a 3-layer, 64-unit MLP yields only ~18 dB PSNR because high frequencies (edges, texture) are not learned. Positional encoding (Tancik et al. 2020; NeRF) fixes this by mapping coordinates through sinusoidal features, raising PSNR to ~45 dB, and has been standard in NeRF-style methods for six years.
  • The geometric insight: Each (sin(x·φ_i), cos(x·φ_i)) pair is a point on a circle; for 2D coordinates, stacking these across frequencies traces a Lissajous curve. Proposition 2.1: two distinct input points can never produce identical Lissajous trajectories — PE gives every coordinate a unique curve.
  • The method: PEPS treats these projected points as *sampling locations*. For input x with L frequencies, it builds a query sequence P_x = (x, S_1…S_L, C_1…C_L) where S_i(x) = (1+sin(x·φ_i))/2, C_i(x) = (1+cos(x·φ_i))/2, then queries a single shared grid encoder at every point, aggregating the 2L+1 latent vectors before the MLP. The frequency-wise queries act as independent "evidence windows" at different observation scales, straightening the parameter-to-capacity curve of grid encoders.
  • Pink-PEPS: Since natural images follow a 1/f^α power spectrum (α ≈ 1–2), high frequencies carry little energy. Pink-PEPS allocates a_n = max(1, ⌊d/f_n⌋) latent dimensions to frequency n, with a shifted slicing offset to ensure gradient flow through the whole latent vector. This compiles domain knowledge (the 1/f law) directly into the encoder shape, cutting compute with minimal quality loss.
  • Results

  • Kodak image representation (24 images, 2.7x compression): Grid-PEPS 47.72 dB / Grid-PinkPEPS 47.83 dB / NTC_PinkPEPS 48.07 dB, vs. BI Grid 45.30, LPE 45.06, NTC_N 44.87. With 25% fewer parameters, Grid-PinkPEPS still reaches 44.89 dB — matching SOTA.
  • 4K neural texture compression (18 texture sets from polyhaven.com, 14x compression): NTC_PinkPEPS leads at 41.89 dB mean PSNR, beating LPE/NTC_N (~40.2) and BI Grid (41.25) across all 8 texture types. On RDNA4 (Radeon RX 9070 XT), inference is 4.86 ms for Grid-PinkPEPS vs 4.32 ms for a plain BI grid — only +12.5% overhead (Grid-PEPS: +26.6%).
  • SDF representation (512³ voxels, 227x compression): Grid-PEPS global IoU 0.816 vs Grid 0.799, LPE 0.793, PE 0.770. On the hard, high-detail Pitted Stonefish instance, Grid-PEPS reaches 0.470 IoU vs 0.390 for a plain grid and 0.242 for hash grids under L1 — matching an 8x-larger grid's quality (0.456) with far fewer parameters.
  • vs. hash grids (InstantNGP): Hash 0.547 global IoU vs Hash-PEPS 0.731. The frequency-diverse query points land in different hash buckets, mitigating hash collisions — PEPS is a natural extension of hash grids, not a replacement.

Notable design detail

All frequency queries share one grid. Per Appendix A, the reason is gradient flow: independent per-frequency grids would starve high-frequency tables of gradient signal on sparse training coordinates; a shared grid lets gradients flow through all frequencies.

Open questions raised in the post

1. Pink-PEPS assumes a 1/f^α spectrum — unclear whether SDF data satisfies this; α=2 ("Brownian-PEPS") is mentioned but not deeply tested. 2. Experiments use L = 3–4 frequencies; scaling further increases query counts and table-lookup overhead (Pink-PEPS at L=4 is slower: 5.0 ms vs 4.86 ms). 3. A rigorous theoretical account of *why* combining per-frequency queries beats single queries is still open; Proposition 2.1 provides uniqueness, not a full spectral analysis.

Takeaway

The post argues PEPS stands out among incremental INR work because it pairs strong empirical results with an actual explanation: a long-used tool (Transformer-style positional encoding) carried an overlooked geometric structure — unique Lissajous trajectories — that turns PE from a dimension-lifting trick into a principled sampler. Code availability was not yet confirmed at the time of writing.

References

1. Perez, G., Matai, J., Harada, T. (2026). *PEPS: Positional Encoding Projected Sampling*. Proc. ACM Comput. Graph. Interact. Tech. 9(1), Article 8. DOI: 10.1145/3806062. 2. Mildenhall et al. (2021). *NeRF*. Communications of the ACM 65(1). 3. Müller et al. (2022). *Instant Neural Graphics Primitives*. ACM TOG 41(4). 4. Rahaman et al. (2019). *On the Spectral Bias of Neural Networks*. ICML. 5. Tancik et al. (2020). *Fourier Features Let Networks Learn High Frequency Functions*. NeurIPS. 6. Vaidyanathan et al. (2023). *Random-Access Neural Compression of Material Textures*. ACM TOG 42. 7. Fujieda, Yoshimura, Harada (2023). *Local Positional Encoding for MLPs*. Pacific Graphics. 8. Vaswani et al. (2017). *Attention Is All You Need*. NeurIPS. 9. Lissajous, J. A. (1857). *Mémoire sur l'étude optique des mouvements vibratoires*. 10. Field, D. J. (1987). *Relations between the statistics of natural images and the response properties of cortical cells*. JOSA A 4(12).

Tags

#positional-encoding#implicit-neural-representation#lissajous-curves#neural-texture-compression#nerf#hash-grid#sdf#amd

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178446950