PoolDINO

University of Geneva · Meta

PoolDINO

Pooling Representation Autoencoders
for Efficient Diffusion

Ramón Calvo-González1Youssef Saied1François Fleuret1,2

1 University of Geneva · 2 Meta

Fewer tokens. Faster generation.
A simple learned pooling layer is all it takes.

Generated portrait of a dog Generated motorcycle Generated shark underwater Generated leather holster

Selected ImageNet-256 samples · 2 × 2 pooling · 180 training epochs · 100 sampling steps

The idea

Keep the features.
Pool the tokens.

Representation autoencoders generate images from pretrained visual features. But dense feature grids make the generative Transformer expensive.

PoolDINO learns to merge neighboring encoder tokens into a single, high-dimensional token. We train this pooling layer jointly with the RGB decoder, then train a generator directly on the smaller grid. No separate feature autoencoder. No extra training stage.

Quality & efficiency

A smaller grid goes a long way.

4×fewer tokens

1.09 FID · 3.7× faster sampling

16×fewer tokens

1.44 FID · 9.0× faster sampling

2 stagessame training pipeline

Learn pooling with RGB reconstruction

FID versus latent samples per second. PoolDINO's compressed models sample faster; extended training improves FID at the same throughput.
ImageNet-256 generation with internal guidance (IG), using 100 Euler steps. Throughput is measured on an H100 NVL at batch size 128 and excludes RGB decoding. Red: 80 epochs. Gold: extended training. Literature points pair published FIDs with measured throughput.
Generation with internal guidance · 80 training epochs · 100 Euler steps
Spatial poolingTokensFID ↓Latent samples/s ↑
RAEv2 reference2561.086.52
Learned 2 × 2641.0924.16
Learned 2 × 4321.1940.10
Learned 4 × 4161.4458.46

FID measures generation quality; lower is better. Speedups are relative to the unpooled RAEv2 reference under these sampling and hardware settings.

Method

Simple pooling. Same two stages.

01

Learn to reconstruct

Freeze the vision encoder. Train a shared affine pooling operator and an RGB decoder together. Repeat the pooled tokens before decoding to keep the decoder architecture unchanged.

Image passes through frozen DINOv3, a learned pooling operator, token repetition and a trainable ViT-XL decoder to reconstruct the image.
Each local window becomes one token. The small grids illustrate the operation; the encoder produces a 16 × 16 grid.
02

Learn to generate

Freeze the encoder, pooling operator, and image decoder. Train a flow-matching Transformer on the compressed latents, with auxiliary heads for internal guidance and encoder reconstruction.

Pooled encoder features are noised and passed to the generator. The generator learns flow matching and auxiliary internal-guidance and encoder-reconstruction objectives.
The generator models the smaller grid directly. Internal guidance uses an intermediate prediction to guide sampling.

Across compression levels

From 256 tokens to 16.

Selected generations from the four spatial grids. All generators train for 80 epochs.

Examples are selected independently across columns—not matched samples. All use 100 Euler steps and the selected IG-only scale. The random galleries below provide a broader view.

Extended training

Spend less per update.
Train for longer.

Extending training from 80 to 180 epochs for 2 × 2 pooling, and to 300 epochs for 4 × 4 pooling, improves FID at the reported guidance settings.

These are extended-training runs, not compute-matched runs. Further training to approximately match the baseline’s compute budget did not improve the best observed FID; the paper reports the complete sweeps.

What is retained?

Good generation is not
the whole story.

Learned pooling does not simply recover averaging or PCA. It retains a different subspace, and its alignment with these baselines decreases as compression increases.

Comparable guided generation quality at 4× compression coexists with lower classification accuracy and weaker dense-task transfer. Pooling trained for RGB reconstruction does not consistently outperform averaging on segmentation and depth estimation.

Classification and dense prediction results
Frozen representations; dense-task models trained from scratch
PoolingWindowLinear top-1 ↑ADE20K mIoU ↑NYUv2 AbsRel ↓
Unpooled1 × 185.31%46.86%0.0817
Learned2 × 283.47%39.38%0.0895
Average2 × 285.33%42.59%0.0869
Learned4 × 480.80%38.61%0.1022
Average4 × 485.33%36.76%0.0997

mIoU: mean intersection over union (segmentation). AbsRel: absolute relative error (depth). Higher mIoU and lower AbsRel are better.

A closer look

More measurements.

Additional figures from the paper. Open any plot as a vector PDF for a closer look.

Inference & training throughput
Forward-pass throughput versus batch size on H100 and A100 GPUs for different pooling geometries
Forward-pass throughput across batch sizes, on H100 and A100 GPUs.
Forward-backward throughput versus batch size on H100 and A100 GPUs
Forward–backward throughput. This benchmark excludes optimizer updates and auxiliary objectives; it is not full training-step throughput.
Sampling in fewer steps

At 50 rather than 100 steps, latent-sampling throughput approximately doubles. The 180-epoch 2 × 2 model reaches 1.0629 FID, and the 300-epoch 4 × 4 model reaches 1.2981 FID.

Quality-throughput comparison with PoolDINO at 50 steps and literature baselines
PoolDINO uses 50 steps here; the RAEv2 reference uses 100. See the paper for each literature baseline’s protocol.
FID and Inception Score across 10 to 100 Euler steps for the 300-epoch 4 by 4 model
4 × 4 pooling, 300 epochs, fixed IG scale 2.75. Gold: FID (left). Red: Inception Score (right).
Guidance and Inception Score
Generation quality and throughput with internal guidance alone and with classifier-free guidance
Internal guidance alone and combined with classifier-free guidance. The extra unconditional evaluations change throughput.
Inception Score versus sampling throughput for PoolDINO and literature baselines
Inception Score provides a complementary view of the quality–throughput trade-off.

Small change. Shorter sequences.

Explore the implementation. The paper will be linked here when it is available on arXiv.