Your Transformer Is Secretly Elastic
Walter Simoncini (opens in a new tab)1, Dheeraj Varghese (opens in a new tab)1, Gertjan J. Burghouts (opens in a new tab)2, Cees G.M. Snoek (opens in a new tab)1
1University of Amsterdam2TNO, Intelligent Imaging
● MLP Neuron■ Attention Head
Sorted by Score: High → LowWidth Proportional to FLOPs
Rank every MLP neuron and attention head on a common scale, once, with frozen weights. Any pre-trained transformer then becomes a single model you can prune to any budget at inference time, without retraining. Rankings are optimized in minutes.
Abstract
Scaling up transformers has produced models with broad capabilities, deployed across heterogeneous hardware and fluctuating workloads. Yet, model families offer only a few fixed sizes (e.g., S, B, L), forcing users to commit to a model based on hardware and latency constraints. This rigidity is often suboptimal under varying query loads. We introduce a retraining-free method that elastifies pre-trained transformers into a single model that adapts on-the-fly across a continuum of computational budgets. Our key contribution is a differentiable relaxation of the functional unit-ranking problem: unlike prior work, which optimizes a single importance weight per MLP block or attention head, we rank every individual MLP neuron and attention head jointly, on a common scale, by minimizing the model's loss under soft pruning across many sparsity targets, all at once. Across 27 vision and text transformer encoders ranging from 22M to 8B parameters, our method yields smooth degradation up to 60% sparsity and outperforms one-shot baselines under deep pruning. For example, pruning DINOv3 ViT-H+/16 to 50% sparsity removes 427M parameters while reducing ImageNet-1k linear-probing accuracy by only 0.8%, without the need for weight correction or retraining. At 60% sparsity, a pruned AugReg ViT-B/16 model achieves up to 1.80× higher throughput and 1.26× lower single-query latency on an A100 GPU. Elastic models are produced in minutes, with construction time scaling sub-linearly with parameter count.
- 27Vision & Text encoders22M to 8B parameters
- 427MPruned params, for −0.8% in ImageNet-1k linear probingDINOv3 ViT-H+/16, 50% sparse
- 1.80×Throughput. 1.26× lower single-query latencyAugReg ViT-B/16, 60% sparse
- 0.46msTo switch budgetsMean. Excluding compilation
Results Explorer
- Ours
- 56.4
- down 26.1vs dense 82.5
- Best Baseline
- 16.1SnapViT
- up 40.3Ours ahead
- Compute Kept
- 51%
- 23.5 of 46.4 GFLOPs
- Params Kept
- 49%
- 42M of 86M params
| Method | Dense | 10% | 20% | 30% | 40% | 50% | 60% |
|---|---|---|---|---|---|---|---|
| Ours | 82.5 | 81.9 | 80.2 | 76.4 | 69.4 | 56.4 | 38.8 |
| SnapViT | 82.5 | 80.8 | 75.8 | 60.1 | 25.2 | 16.1 | 10.4 |
| SparseGPT | 82.5 | 40.8 | 17.3 | 12.5 | 10.3 | 8.3 | 6.9 |
| VBP | 82.5 | 78.7 | 73.4 | 42.8 | 28.7 | 8.8 | — |
Scaling
Pruned Large Models Outperform Small Dense Ones
Across the six-model DINO WebSSL family, the pruned 3B and 7B models pareto-dominate the dense 2B and 5B in ImageNet-1k k-NN. Each curve is one model, from dense to 60% sparsity. The 3B curve passes above the dense 2B, and the 7B curve above the dense 5B.
- DINO WebSSL 300M
- DINO WebSSL 1B
- DINO WebSSL 2B
- DINO WebSSL 3B
- DINO WebSSL 5B
- DINO WebSSL 7B
- Dense
Deployment
One Dense Index, Many Budgets
In deployment, the index can be built once offline with the dense encoder, while incoming queries are encoded elastically depending on the current workload. Elastic queries thus must remain compatible with a fixed dense index, allowing query budgets to change without rebuilding the index. Our method makes this possible in both text and image retrieval.
See it in the explorer (Dense Index):
Speed
- 1.80×
- Throughput
- Batch 128
- 1.26×
- Lower Latency
- Single query
- 0.46 ms
- Budget Switch
- Mean cost
AugReg ViT-B/16 at 60% sparsity vs dense, A100, fp16, with compilation and GPU-friendly pruning. Switching cost excludes the initial compilation.
Where Sparsity Goes
Kept fraction per transformer block for the initial importance scores versus our optimized ranking.
Citation
@article{simoncini2026secretly,
title = {Your Transformer Is Secretly Elastic},
author = {Simoncini, Walter and Varghese, Dheeraj and Burghouts, Gertjan J. and Snoek, Cees G. M.},
journal = {arXiv preprint},
year = {2026}
}This research has received funding from the NWO perspectiefprogramma Foundation models for Industry (FIND) under grant number P23.016. Cees G. M. Snoek is (partially) funded by the Horizon Europe project ELLIOT (GA No. 101214398). We thank SURF for the support in using the National Supercomputer Snellius, and DAS-6 for providing the compute to run the experiments presented in this paper.