We probed six frozen vision towers across stages. No tower wins everywhere, and spatial tasks peak before the last layer. Projectors preserve spatial info, with the size of the gain varying by dataset.
Adding vision to a base LLM means choosing a tower, layer, and projector. These are choices full-model benchmarks like MMMU never isolate. We probe frozen features across stages, before any multimodal training, and publish Capability Profiles.
Can a probe identify the object, with plenty of labels and with almost none?
Can it recover depth and surface orientation from a single image?
Can it find the same physical point again in a second photograph?
How much recognition survives when half the image is hidden?
Each column measures a different capability, so the results are not combined into an overall score. Some apparent differences are too small to distinguish on the current test set.
| Tower | Few-label1% top-1 โ | Full-label100% top-1 โ | Depthd1 โ | Normalsdeg โ | Cross-viewNAVI@2cm โ | Occlusion50% kept โ |
|---|---|---|---|---|---|---|
| DINOv2 | 0.369 | 0.908 | 0.702 | 18.85 | 0.539 | 52% |
| SigLIP2 | 0.656 | 0.917 | 0.670 | 23.90 | 0.401 | 49% |
| Muse Glimmer | 0.433 | 0.921 | 0.659 | 24.30 | 0.214 | 52% |
| Kimi K2.6 | 0.345 | 0.886 | 0.660 | 25.51 | 0.365 | 28% |
| Qwen3.8 | 0.245 | 0.879 | 0.674 | 24.24 | 0.378 | 30% |
| Kimi K3 | 0.219 | 0.837 | 0.655 | 27.03 | 0.333 | 20% |
Towers stay frozen. We tap eight Relative Depth points, plus merged and projected when present. Features cache once; only the readout trains. Recognition is a 1.6M attention pool on ImageNet-100, geometry is Probe3D decoders on DIODE, and correspondence is training-free matching.
Heads are capacity matched with frozen PCA so wider stages don't get bigger heads; the unmatched arm runs as a check. Learning rate is picked per cell on validation, three seeds, with paired bootstrap intervals on headlines.
DINOv2 sweeps depth, normals and cross-view but sits third at recognition; Muse Glimmer leads recognition and sits last on correspondence. On the attention readout, SigLIP2 leads by 0.22 at 1% labels.
Geometry falls before the last layer in 23 of 24 arms while recognition rises to it in 19 of 24. Only Qwen3.8 declines into the last layer on recognition, and only in unmatched arms.
On DIODE, all 16 projected-against-tower depth and normal comparisons favor the Projector, so it preserves spatial information. However, on KITTI driving scenes none resolve in its favor and two reverse. Qwen3.8 also reverses on ScanNet while gaining on NAVI.
The paper sets out every limitation in full, along with the controls that failed.
Qualitative PCA of frozen patch tokens. Read structure, not color.
The eagle separates as one solid region. First on all three correspondence sets, third on recognition.
Vision Tower Bench v1 released, covering six Towers on geometry, semantic and correspondence.
Vision Towers for Kimi K2.6, Kimi K3, Muse Glimmer and Qwen3.8 published on Hugging Face.