Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
ECCV 2026 Long Oral Presentation
A plain Transformer architecture for point clouds can achieve state-of-the-art performance for 3D scene understanding, when properly trained and scaled.