University of Geneva · Meta
PoolDINO
Pooling Representation Autoencoders
for Efficient Diffusion
1 University of Geneva · 2 Meta
Fewer tokens. Faster generation.
A simple learned pooling layer is all it takes.
Selected ImageNet-256 samples · 2 × 2 pooling · 180 training epochs · 100 sampling steps
The idea
Keep the features.
Pool the tokens.
Representation autoencoders generate images from pretrained visual features. But dense feature grids make the generative Transformer expensive.
PoolDINO learns to merge neighboring encoder tokens into a single, high-dimensional token. We train this pooling layer jointly with the RGB decoder, then train a generator directly on the smaller grid. No separate feature autoencoder. No extra training stage.
Quality & efficiency
A smaller grid goes a long way.
1.09 FID · 3.7× faster sampling
1.44 FID · 9.0× faster sampling
Learn pooling with RGB reconstruction

| Spatial pooling | Tokens | FID ↓ | Latent samples/s ↑ |
|---|---|---|---|
| RAEv2 reference | 256 | 1.08 | 6.52 |
| Learned 2 × 2 | 64 | 1.09 | 24.16 |
| Learned 2 × 4 | 32 | 1.19 | 40.10 |
| Learned 4 × 4 | 16 | 1.44 | 58.46 |
FID measures generation quality; lower is better. Speedups are relative to the unpooled RAEv2 reference under these sampling and hardware settings.
Method
Simple pooling. Same two stages.
Learn to reconstruct
Freeze the vision encoder. Train a shared affine pooling operator and an RGB decoder together. Repeat the pooled tokens before decoding to keep the decoder architecture unchanged.

Learn to generate
Freeze the encoder, pooling operator, and image decoder. Train a flow-matching Transformer on the compressed latents, with auxiliary heads for internal guidance and encoder reconstruction.

Across compression levels
From 256 tokens to 16.
Selected generations from the four spatial grids. All generators train for 80 epochs.
Examples are selected independently across columns—not matched samples. All use 100 Euler steps and the selected IG-only scale. The random galleries below provide a broader view.
Extended training
Spend less per update.
Train for longer.
Extending training from 80 to 180 epochs for 2 × 2 pooling, and to 300 epochs for 4 × 4 pooling, improves FID at the reported guidance settings.
These are extended-training runs, not compute-matched runs. Further training to approximately match the baseline’s compute budget did not improve the best observed FID; the paper reports the complete sweeps.
What is retained?
Good generation is not
the whole story.
Learned pooling does not simply recover averaging or PCA. It retains a different subspace, and its alignment with these baselines decreases as compression increases.
Comparable guided generation quality at 4× compression coexists with lower classification accuracy and weaker dense-task transfer. Pooling trained for RGB reconstruction does not consistently outperform averaging on segmentation and depth estimation.
Classification and dense prediction results
| Pooling | Window | Linear top-1 ↑ | ADE20K mIoU ↑ | NYUv2 AbsRel ↓ |
|---|---|---|---|---|
| Unpooled | 1 × 1 | 85.31% | 46.86% | 0.0817 |
| Learned | 2 × 2 | 83.47% | 39.38% | 0.0895 |
| Average | 2 × 2 | 85.33% | 42.59% | 0.0869 |
| Learned | 4 × 4 | 80.80% | 38.61% | 0.1022 |
| Average | 4 × 4 | 85.33% | 36.76% | 0.0997 |
mIoU: mean intersection over union (segmentation). AbsRel: absolute relative error (depth). Higher mIoU and lower AbsRel are better.
Sample gallery
Look a little closer.
48 randomly selected images per model, without quality filtering. Corresponding positions share the same class condition across galleries.

A closer look
More measurements.
Additional figures from the paper. Open any plot as a vector PDF for a closer look.
Inference & training throughput


Sampling in fewer steps
At 50 rather than 100 steps, latent-sampling throughput approximately doubles. The 180-epoch 2 × 2 model reaches 1.0629 FID, and the 300-epoch 4 × 4 model reaches 1.2981 FID.


Guidance and Inception Score


Small change. Shorter sequences.
Explore the implementation. The paper will be linked here when it is available on arXiv.