header

Welcome


bdrp


Welcome to the Computer Vision Group at RWTH Aachen University!

The Computer Vision group has been established at RWTH Aachen University in context with the Cluster of Excellence "UMIC - Ultra High-Speed Mobile Information and Communication" and is associated with the Chair Computer Sciences 8 - Computer Graphics, Computer Vision, and Multimedia. The group focuses on computer vision applications for mobile devices and robotic or automotive platforms. Our main research areas are visual object recognition, tracking, self-localization, 3D reconstruction, and in particular combinations between those topics.

We offer lectures and seminars about computer vision and machine learning.

You can browse through all our publications and the projects we are working on.

News

RLC'26

Our paper Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Dynamics Models was accepted at the 2026 Reinforcement Learning Conference (RLC)!

Sept. 2, 2026

ECCV'26

We have two papers accepted at the European Conference on Computer Vision (ECCV) 2026!

Sept. 1, 2026

SPAICE'26

Our paper "AstroPIE-53: Towards Understanding Failure Modes of Human Pose Estimation in Space" will be presented at the European Space Agency (AI for Space Applications Conference (SPAICE)) in Noordwijk, Netherlands!

June 24, 2026

WACV'26

Our paper "We Still See Broken Limbs: Towards Anatomical Realism in GenAI via Human Preference Learning" was accepted at the 5th Workshop on Image/Video/Audio Quality Assessment in Computer Vision, VLM and Diffusion Models at IEEE/CVF Winter Conference on Applications of Computer Vision 2026. See you in Tucson, Arizona!

Jan. 23, 2026

ICCV'25

Our paper DONUT: A Decoder-Only Model for Trajectory Prediction was accepted at the 2025 International Conference on Computer Vision (ICCV)!

Our project: Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference achieves 3rd Place of LSVOS Workshop, MeViS Track.

Oct. 1, 2025

RO-MAN'25

Our paper How do Foundation Models Compare to Skeleton-Based Approaches for Gesture Recognition in Human-Robot Interaction? has been accepted!

June 12, 2025

Recent Publications

pubimg
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding

European Conference on Computer Vision (ECCV)

Transformers have become a common foundation across deep learning, yet 3D scene understanding still relies on specialized backbones with strong domain priors. This isolates the field from the broader Transformer ecosystem, limiting the transfer of research advances from other domains and the benefits of increasingly optimized software and hardware stacks. To bridge this gap, we propose the Volume Transformer (Volt), which adapts the vanilla Transformer encoder to 3D scenes with minimal modifications. Specifically, Volt partitions 3D scenes into volumetric patch tokens, processes them with full global self-attention, and injects positional information via 3D rotary positional embeddings (RoPE). Our initial experiments reveal that naively training Volt on standard 3D benchmarks leads to poor generalization, highlighting the limited scale of current 3D supervision. To overcome this, we introduce a data-efficient training recipe based on strong 3D augmentations, regularization, and distillation from a convolutional teacher, making Volt competitive with state-of-the-art methods. We then scale supervision through joint training on multiple datasets and show that Volt benefits more from increased scale than domain-specific 3D backbones, achieving state-of-the-art results on several indoor and outdoor semantic segmentation benchmarks. Finally, as a drop-in backbone in a standard 3D instance segmentation pipeline, Volt also sets a new state of the art, highlighting its potential as a simple, scalable, and general-purpose backbone for 3D scene understanding.

fadeout
 
pubimg
Towards Metric-Agnostic Trajectory Forecasting

European Conference on Computer Vision (ECCV) 2026

Accurate trajectory forecasting of surrounding traffic participants is a core capability for autonomous driving, enabling vehicles to anticipate behavior and plan safe maneuvers. We observe that current state-of-the-art forecasting models on Argoverse 2 and the Waymo Open Motion Dataset tailor their training objectives to the different benchmark metrics. Because these metrics encourage conflicting behavior, we propose a paradigm change for trajectory forecasting: training models with metric-agnostic probabilistic objectives and treating metric optimization as a downstream task applied to the predictive distribution. Concretely, we introduce Trajectory Distribution Evaluation (TraDiE) policies, metric-specific policies that map a predictive distribution to the set of K trajectories and confidences required by trajectory forecasting metrics. We evaluate this framework by introducing DONUT-NLL, which adapts the training objective of the state-of-the-art trajectory forecasting model DONUT to directly optimize the predictive distribution. Using our policies, DONUT-NLL achieves state-of-the-art results on all metrics of the Waymo motion prediction benchmark.

fadeout
 
pubimg
Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers

IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2026

Efficient and accurate feed-forward multi-view reconstruction has long been an important task in computer vision. Recent transformer-based models like VGGT, $\pi^3$ and MapAnything have demonstrated remarkable performance with relatively simple architectures. However, their scalability is fundamentally constrained by the quadratic complexity of global attention, which imposes a significant runtime bottleneck when processing large image sets. In this work, we empirically analyze the global attention matrix of these models and observe that the probability mass concentrates on a small subset of patch-patch interactions corresponding to cross-view geometric correspondences. Building on this insight and inspired by recent advances in large language models, we propose a training-free, block-sparse replacement for dense global attention, implemented with highly optimized kernels. Our method accelerates inference by more than 3x while maintaining comparable task performance. Evaluations on a comprehensive suite of multi-view benchmarks demonstrate that our approach seamlessly integrates into existing global attention-based architectures such as VGGT, $\pi^3$, and MapAnything, while substantiallyimproving scalability to large image collections.

fadeout
Datenschutzerklärung/Privacy Policy Home Visual Computing institute RWTH Aachen University