CRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion Decoding

Andrea Ceron1,2, Michael Schmidt2, Alvaro Marcos-Ramiro2, Sebastian Schmidt1,2, Benjamin Busam1

1TU MΓΌnchen   2BMW AG

NeurIPS 2026 · Poster, Main Track

Case 1, scene-scene discontinuities: ground truth, baseline decoder (SVD) with the void lost, and CRISP with the void retained. Case 2, object-scene discontinuities: ground truth, baseline decoder (Wan) with the edge degraded, and CRISP with the edge preserved. Case 3, object-object discontinuities: ground truth, baseline decoder (LiDM) with the gap lost, and CRISP with the gap retained.
Qualitative comparison across three frozen latent backbones: blue denotes the ground-truth point cloud, green the inherited VAE decoder, and red our method (CRISP). Across all cases, CRISP suppresses flying pixels and boundary-bridging artifacts that persist in the baseline reconstructions, yielding cleaner geometry and sharper depth contours.

Abstract

Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces. We identify this as a major, directly correctable decoder bottleneck and introduce CRISP: a pixel-space diffusion decoder with a backbone-agnostic latent adapter, DiT-based denoiser, and support mask predictor. CRISP replaces video-VAE and LiDAR-native decoders alike while keeping the encoder and latent generator fixed. Across KITTI-360, SemanticKITTI, and nuScenes, replacing only the decoder reduces FSVD/FPVD by 50.5% on average across frozen backbones; for generic video VAEs, the reductions reach 71%/74%. On the LiDAR-native LiDM backbone, FRID drops by 71%, with the largest gains at depth discontinuities. In a pretrained LiDM world model, the same zero-shot replacement improves FSVD by 15.5%, narrowing the sim-to-real gap.

The problem: flying pixels

The consequence is a concrete and recurring artifact: flying pixels. Convolutional VAE decoders smooth across the sharp depth contours that define LiDAR geometry, emitting interpolated depth at object edges. Even LiDM's LiDAR-native decoder, whose circular padding closes the azimuthal seam, still convolves across radial depth discontinuities. When the blurred range map is back-projected to 3D, those boundary pixels become floating points anchored to neither surface.

Method

Latent Adapter𝓐𝝍FusionLN & ModAttentionLN & ModMLPFusionLN & ModAttentionLN & ModMLPCoarse DiT StageRefinement DiT StagePredicted Depthπ’šΰ·π’…Noisy Depth𝒙𝒕(training)Gaussian Noise𝒙𝑻Input LiDAR Range Mapπ’šPredicted Support Maskπ’Žΰ·Reconstructed LiDAR RMπ’šΰ·ProjectionComposePre-trainedEncoderπ“”πŽMaskPredic.π“œπ“Tokenization3.2CascadeDiTDepth Decoderπ““πœ½3.33.43.63.5𝒛𝝉𝝉
Architecture of CRISP. Left: the frozen encoder EΟ‰ maps the input range map y to latent z. Top: the latent adapter Aψ projects and tokenizes z into conditioning tokens Ο„; subsequently, the support mask predictor MΟ† takes z and Ε·d to produce the binary mask mΜ‚. Centre: the cascade DiT decoder DΞΈ denoises xt on Ο„ at two fusion points (red dashed arrows), yielding the dense depth Ε·d. Right: mΜ‚ and Ε·d are composed as Ε· = mΜ‚ βŠ™ Ε·d βˆ’ (1βˆ’mΜ‚) to produce the final LiDAR range map.

CRISP is composed of three components:

Results

−50.5% avg. FSVD/FPVD across frozen backbones
71% / 74% FSVD/FPVD reduction on generic video VAEs
−71% FRID on LiDM
−15.5% world-model FSVD, zero-shot decoder swap
Case 1, narrow objects lost: ground truth, baseline decoder (SVD Γ—1) with the edge degraded, and CRISP with the edge retained. Case 2, void regions incorrectly filled: ground truth, baseline decoder (Wan Γ—1) with the void lost, and CRISP with the void retained. Case 3, foreground-background flying pixels: ground truth, baseline decoder (LiDM) with the gap lost, and CRISP with the gap retained.
Qualitative LiDAR reconstruction results (colors as in the teaser above). Each column, from left to right, highlights a distinct failure mode addressed by CRISP: narrow objects lost; void regions incorrectly filled; foreground-background flying pixels.

Not every metric improves: FPVD on LiDM, FSVD/FPVD on SVD ×1, and JSD/EMD on Wan ×1 regress, and in the world model the edge F-score drops 20%. CRISP (≈516M parameters) is faster end-to-end on Wan ×3 (57→86 fps) but slower on LiDM (369→41 fps). Full quantitative results, ablations and limitations are in the paper.

Qualitative results

Bird's-eye view over time. Left: ground-truth LiDAR point cloud (blue). Middle: inherited baseline VAE decoder (green). Right: CRISP (red).
Baseline Decoder (SVD Γ—3) Our Method (CRISP) Baseline Decoder (SVD Γ—3) Our Method (CRISP)
Boundary degradedBoundary preserved
Compare CRISP with
Extended qualitative comparison, SVD Γ—3 (frozen encoder, nuScenes). CRISP restores continuous scan-line geometry and suppresses the boundary-bridging artifacts highlighted by the inset box.

BibTeX

@inproceedings{ceron2026crisp,
  title     = {CRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion Decoding},
  author    = {Ceron, Andrea and Schmidt, Michael and Marcos-Ramiro, Alvaro and Schmidt, Sebastian and Busam, Benjamin},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}