Latent Dynamics · 2026-09-30
Self-Reconstruction Dynamics for Autoencoder Reconstruction Refinement
Hitoshi Iyatomi
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
The authors study Self-Reconstruction Dynamics (SRD), produced by repeatedly applying a frozen autoencoder to its own reconstruction, which forms transient image/latent trajectories. Although repeated self-reconstruction degrades fidelity, SRD encodes sample-specific correction information. They propose SRD-guided Reconstruction Refinement (SRD-RR), predicting a latent correction from a short SRD with the AE frozen and no per-sample test-time optimization. Across six datasets, SRD-RR recovers 38.6% (one) to 45.3% (two) of the empirically recoverable MSE gap.
Why it matters
This work reframes iterative self-application of an autoencoder: instead of using later iterates as improved outputs, it treats the transient SRD trajectory as conditioning signal for a separate, decoder-compatible latent correction. The authors also show that sample correspondence matters (shuffling trajectories causes severe degradation) and that the refinement objective affects the trade-off between pixel fidelity and perceptual quality, including on a pretrained DINOv2-based representation autoencoder.
Method
- Generate Self-Reconstruction Dynamics (SRD) by repeatedly applying a frozen AE to its own reconstruction to obtain observable image- and latent-space trajectories.
- Train SRD-RR to predict a decoder-compatible latent correction from a short SRD, keeping encoder/decoder frozen and using no per-sample test-time optimization.
- Evaluate improvements via an MSE recovery ratio (MSE-recov) relative to an empirical decoder-optimized reference obtained by per-sample latent optimization.
Limitation
Removing trajectory information substantially reduces the gain, while cross-sample trajectory assignment causes severe degradation, showing that the useful SRD information is strongly sample-specific.
Abstract (from arXiv)
Standard autoencoder (AE) inference uses a single encoder-decoder pass, though the latent may not be optimal for each sample under a fixed decoder. We ask whether a trained AE can reveal information for improving its own reconstruction. Repeated application of a frozen AE to its reconstruction produces transient image- and latent-space trajectories, termed Self-Reconstruction Dynamics (SRD). Although this degrades fidelity in the AEs studied here, SRD contains sample-specific information for correcting the reconstruction. We propose SRD-guided Reconstruction Refinement (SRD-RR), which predicts a latent correction from a short SRD with the AE frozen and no per-sample test-time optimization. We also introduce MSE-recov, an MSE recovery ratio relative to an empirical decoder-optimized reference. Across six datasets, SRD-RR recovers 38.6% of the empirically recoverable MSE gap with one transition and 45.3% with two. A two-transition variant trained without direct access to original images, using an SRD-derived pseudo-target, achieves 40.7% recovery and a 1.74 dB average PSNR gain. Removing trajectory information reduces the gain, while cross-sample trajectory assignment causes severe degradation, confirming strong sample specificity. Nonlinear SRD-conditioned refinement consistently outperforms fixed and trained linear latent correction. On a pretrained DINOv2-based representation autoencoder (RAE) with substantially different latent dynamics, SRD conditioning again improves a matched trajectory-free predictor. However, pixel-MSE latent refinement reveals a strong mismatch between pixel fidelity and perceptual quality, while the SRD-derived pseudo-target mitigates this degradation. Overall, SRD is a useful sample-specific refinement signal, while the objective determines how it translates into pixel and perceptual quality.
Related papers
- STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
- JEPA Learns What the Mask Leaves Unrecoverable
- Stochastic World Models for Verifying Vision-Based Neural Feedback Systems
- Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning
- LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models