Your Transformer Is Secretly Elastic

Walter Simoncini (opens in a new tab)1, Dheeraj Varghese (opens in a new tab)1, Gertjan J. Burghouts (opens in a new tab)2, Cees G.M. Snoek (opens in a new tab)1

1University of Amsterdam2TNO, Intelligent Imaging

b110% sparseb230% sparseb350% sparseb470% sparse
Initial ranking. Estimate initial importance scores and rank every MLP neuron ● and attention head ■ on a common scale. Bars represent unit scores, and width represents the unit cost in FLOPs. Watch the white neuron.

● MLP Neuron■ Attention Head

Sorted by Score: High → LowWidth Proportional to FLOPs

Rank every MLP neuron and attention head on a common scale, once, with frozen weights. Any pre-trained transformer then becomes a single model you can prune to any budget at inference time, without retraining. Rankings are optimized in minutes.

Abstract

Scaling up transformers has produced models with broad capabilities, deployed across heterogeneous hardware and fluctuating workloads. Yet, model families offer only a few fixed sizes (e.g., S, B, L), forcing users to commit to a model based on hardware and latency constraints. This rigidity is often suboptimal under varying query loads. We introduce a retraining-free method that elastifies pre-trained transformers into a single model that adapts on-the-fly across a continuum of computational budgets. Our key contribution is a differentiable relaxation of the functional unit-ranking problem: unlike prior work, which optimizes a single importance weight per MLP block or attention head, we rank every individual MLP neuron and attention head jointly, on a common scale, by minimizing the model's loss under soft pruning across many sparsity targets, all at once. Across 27 vision and text transformer encoders ranging from 22M to 8B parameters, our method yields smooth degradation up to 60% sparsity and outperforms one-shot baselines under deep pruning. For example, pruning DINOv3 ViT-H+/16 to 50% sparsity removes 427M parameters while reducing ImageNet-1k linear-probing accuracy by only 0.8%, without the need for weight correction or retraining. At 60% sparsity, a pruned AugReg ViT-B/16 model achieves up to 1.80× higher throughput and 1.26× lower single-query latency on an A100 GPU. Elastic models are produced in minutes, with construction time scaling sub-linearly with parameter count.

Results Explorer

DINOv3 ViT-B/16: IN1K k-NN, same model. Ours versus baselines across compute budgets.0204060804640332619GFLOPsTop-1 k-NN Accuracy (%)Dense 82.520%40%60%
Ours
Sparsity50% sparsity
Ours
56.4
down 26.1vs dense 82.5
Best Baseline
16.1SnapViT
up 40.3Ours ahead
Compute Kept
51%
23.5 of 46.4 GFLOPs
Params Kept
49%
42M of 86M params
DINOv3 ViT-B/16: IN1K k-NN, same model. Ours versus baselines across compute budgets.
MethodDense10%20%30%40%50%60%
Ours82.581.980.276.469.456.438.8
SnapViT82.580.875.860.125.216.110.4
SparseGPT82.540.817.312.510.38.36.9
VBP82.578.773.442.828.78.8—

Scaling

DINO WebSSL: Top-1 k-NN Accuracy (%) versus GFLOPs, each model from dense to 60% sparsity.606570758010020050010002000GFLOPs (log scale)Top-1 k-NN Accuracy (%)DINO WebSSL 300M, dense: 75.3 at 162.3 GFLOPsDINO WebSSL 300M, 10% sparsity: 75.0 at 146.3 GFLOPsDINO WebSSL 300M, 20% sparsity: 74.5 at 129.9 GFLOPsDINO WebSSL 300M, 30% sparsity: 73.4 at 113.7 GFLOPsDINO WebSSL 300M, 40% sparsity: 71.9 at 97.7 GFLOPsDINO WebSSL 300M, 50% sparsity: 68.9 at 81.2 GFLOPsDINO WebSSL 300M, 60% sparsity: 62.9 at 65.2 GFLOPsDINO WebSSL 1B, dense: 78.0 at 599 GFLOPsDINO WebSSL 1B, 10% sparsity: 77.9 at 539.3 GFLOPsDINO WebSSL 1B, 20% sparsity: 77.8 at 479.3 GFLOPsDINO WebSSL 1B, 30% sparsity: 77.4 at 419.4 GFLOPsDINO WebSSL 1B, 40% sparsity: 76.7 at 359.8 GFLOPsDINO WebSSL 1B, 50% sparsity: 75.3 at 299.5 GFLOPsDINO WebSSL 1B, 60% sparsity: 71.5 at 239.8 GFLOPsDINO WebSSL 2B, dense: 76.9 at 1087.5 GFLOPsDINO WebSSL 2B, 10% sparsity: 76.8 at 978.9 GFLOPsDINO WebSSL 2B, 20% sparsity: 76.5 at 870.2 GFLOPsDINO WebSSL 2B, 30% sparsity: 76.2 at 761.5 GFLOPsDINO WebSSL 2B, 40% sparsity: 75.5 at 652.8 GFLOPsDINO WebSSL 2B, 50% sparsity: 74.3 at 544.9 GFLOPsDINO WebSSL 2B, 60% sparsity: 70.5 at 435.5 GFLOPsDINO WebSSL 3B, dense: 78.1 at 1535.6 GFLOPsDINO WebSSL 3B, 10% sparsity: 78.1 at 1383 GFLOPsDINO WebSSL 3B, 20% sparsity: 78.0 at 1228.6 GFLOPsDINO WebSSL 3B, 30% sparsity: 77.8 at 1075.2 GFLOPsDINO WebSSL 3B, 40% sparsity: 77.3 at 922.5 GFLOPsDINO WebSSL 3B, 50% sparsity: 76.5 at 767.6 GFLOPsDINO WebSSL 3B, 60% sparsity: 73.9 at 614.2 GFLOPsDINO WebSSL 5B, dense: 77.0 at 2567.3 GFLOPsDINO WebSSL 5B, 10% sparsity: 76.9 at 2310.7 GFLOPsDINO WebSSL 5B, 20% sparsity: 77.0 at 2054 GFLOPsDINO WebSSL 5B, 30% sparsity: 76.8 at 1797.5 GFLOPsDINO WebSSL 5B, 40% sparsity: 76.5 at 1540.8 GFLOPsDINO WebSSL 5B, 50% sparsity: 75.8 at 1284.2 GFLOPsDINO WebSSL 5B, 60% sparsity: 73.9 at 1027.6 GFLOPsDINO WebSSL 7B, dense: 76.8 at 3348.6 GFLOPsDINO WebSSL 7B, 10% sparsity: 76.9 at 3015 GFLOPsDINO WebSSL 7B, 20% sparsity: 76.9 at 2679 GFLOPsDINO WebSSL 7B, 30% sparsity: 77.1 at 2343.2 GFLOPsDINO WebSSL 7B, 40% sparsity: 77.3 at 2009.7 GFLOPsDINO WebSSL 7B, 50% sparsity: 77.1 at 1674.9 GFLOPsDINO WebSSL 7B, 60% sparsity: 75.9 at 1340.1 GFLOPs300M1B2B3B5B7B

Pruned Large Models Outperform Small Dense Ones

Across the six-model DINO WebSSL family, the pruned 3B and 7B models pareto-dominate the dense 2B and 5B in ImageNet-1k k-NN. Each curve is one model, from dense to 60% sparsity. The 3B curve passes above the dense 2B, and the 7B curve above the dense 5B.

  • DINO WebSSL 300M
  • DINO WebSSL 1B
  • DINO WebSSL 2B
  • DINO WebSSL 3B
  • DINO WebSSL 5B
  • DINO WebSSL 7B
  • Dense

Deployment

One Dense Index, Many Budgets

In deployment, the index can be built once offline with the dense encoder, while incoming queries are encoded elastically depending on the current workload. Elastic queries thus must remain compatible with a fixed dense index, allowing query budgets to change without rebuilding the index. Our method makes this possible in both text and image retrieval.

See it in the explorer (Dense Index):

Speed

1.80×
Throughput
Batch 128
1.26×
Lower Latency
Single query
0.46 ms
Budget Switch
Mean cost

AugReg ViT-B/16 at 60% sparsity vs dense, A100, fp16, with compilation and GPU-friendly pruning. Switching cost excludes the initial compilation.

Throughput vs Dense, Batch 128EagerCompiled
1.0×1.2×1.4×1.6×1.8×1.08×10%31.7 GFLOPs1.17×20%28.2 GFLOPs1.28×30%24.8 GFLOPs1.42×40%21.2 GFLOPs1.59×50%17.8 GFLOPs1.80×60%14.2 GFLOPs

Where Sparsity Goes

Kept fraction per transformer block for the initial importance scores versus our optimized ranking.

50% sparsity
MLP NeuronsAttention Heads123456789101112
Initial ScoresOptimized Ranking (Ours)Block 1 → 12

Citation

@article{simoncini2026secretly,
  title   = {Your Transformer Is Secretly Elastic},
  author  = {Simoncini, Walter and Varghese, Dheeraj and Burghouts, Gertjan J. and Snoek, Cees G. M.},
  journal = {arXiv preprint},
  year    = {2026}
}

This research has received funding from the NWO perspectiefprogramma Foundation models for Industry (FIND) under grant number P23.016. Cees G. M. Snoek is (partially) funded by the Horizon Europe project ELLIOT (GA No. 101214398). We thank SURF for the support in using the National Supercomputer Snellius, and DAS-6 for providing the compute to run the experiments presented in this paper.