Abstract
Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces. We identify this as a major, directly correctable decoder bottleneck and introduce CRISP: a pixel-space diffusion decoder with a backbone-agnostic latent adapter, DiT-based denoiser, and support mask predictor. CRISP replaces video-VAE and LiDAR-native decoders alike while keeping the encoder and latent generator fixed. Across KITTI-360, SemanticKITTI, and nuScenes, replacing only the decoder reduces FSVD/FPVD by 50.5% on average across frozen backbones; for generic video VAEs, the reductions reach 71%/74%. On the LiDAR-native LiDM backbone, FRID drops by 71%, with the largest gains at depth discontinuities. In a pretrained LiDM world model, the same zero-shot replacement improves FSVD by 15.5%, narrowing the sim-to-real gap.
The problem: flying pixels
The consequence is a concrete and recurring artifact: flying pixels. Convolutional VAE decoders smooth across the sharp depth contours that define LiDAR geometry, emitting interpolated depth at object edges. Even LiDM's LiDAR-native decoder, whose circular padding closes the azimuthal seam, still convolves across radial depth discontinuities. When the blurred range map is back-projected to 3D, those boundary pixels become floating points anchored to neither surface.
Method
CRISP is composed of three components:
- Latent adapter: maps the frozen encoder's latent, from any backbone, to a common sequence of conditioning tokens for the denoiser.
- DiT pixel-diffusion decoder: denoises a Gaussian noise sample directly in range-map space into a dense depth map, with the latent tokens injected early and again at the network midpoint.
- Support mask predictor: looks at the latent and the dense depth to decide which pixels hold a valid LiDAR return, so the final output is sparse like a real scan.
Results
Not every metric improves: FPVD on LiDM, FSVD/FPVD on SVD ×1, and JSD/EMD on Wan ×1 regress, and in the world model the edge F-score drops 20%. CRISP (≈516M parameters) is faster end-to-end on Wan ×3 (57→86 fps) but slower on LiDM (369→41 fps). Full quantitative results, ablations and limitations are in the paper.
Qualitative results
BibTeX
@inproceedings{ceron2026crisp,
title = {CRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion Decoding},
author = {Ceron, Andrea and Schmidt, Michael and Marcos-Ramiro, Alvaro and Schmidt, Sebastian and Busam, Benjamin},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2026}
}