Log In

2461 SW Campus Way, Corvallis, OR 97331

View map

Advancing Visual Perception and Spatial Understanding in Dynamic Environments

Understanding dynamic scenes through visual inputs such as videos is crucial in various domains, including autonomous driving, robotics, augmented reality, and human-computer interaction. This research aims to advance our comprehension of dynamic environments by exploring methods that integrate inputs from diverse domains. Our first work proposes a visual memory architecture tailored for learning and inference in the spatio-temporal domain. We maintain a fixed set of memory slots managed by a data-driven memory module based on Gumbel-Softmax. This approach allows the memory module to adapt its strategy for each specific target task, efficiently constraining memory size—a critical aspect for modeling prolonged video sequences. Next, we address the challenge of long-term point tracking to enhance our understanding of non-rigid motion in the physical world. While prior methods primarily focus on 2D inputs, modeling target motion in this domain is challenging due to its entanglement with camera motion. Consequently, we introduce a deep learning framework designed for long-term point tracking in 3D, capable of generalizing to novel points and videos without the need for fine-tuning during testing. Our framework includes a cost volume fusion module that integrates multiple past appearances and motion information via a transformer architecture, significantly enhancing tracking performance. Finally, we explore extending this work. Low-level information like long-term point tracks is crucial for tasks such as non-rigid structure from motion and dynamic SLAM, while higher-level information, such as objectness, is vital for modeling physical interactions and autonomous exploration. In unknown environments with unseen objects, motion is essential for characterizing different objects. Therefore, we propose to utilize long-term point tracking to enable novel object discovery and interaction modeling. Specifically, given RGB-D input and noisy camera poses, we aim to construct a framework that simultaneously discovers objects based on motion grouping and models interactions among them. In more detail, the framework will predict the future state of the environment as a series of 3D point clouds based on the modeled interactions. Overall, this work aims to advance real-world understanding by harnessing inputs from both 2D and 3D data domains.

MAJOR ADVISOR: Fuxin Li
COMMITTEE: Alan Fern
COMMITTEE: Stefan Lee
COMMITTEE: Sinisa Todorovic
GCR: Harold Bae