Efficient Transformers
Efficient transformers for vision
Designing propagation and attention modules that make high-resolution vision models faster and more memory-efficient.
I am a Principal Research Scientist and Tech Lead at NVIDIA Research, where I work with the LPR team led by Jan Kautz. My research spans efficient transformers, spatial vision-language models, and robotics. I develop efficient transformers with GSPN, grounded spatial VLMs such as GR3D, and agents that learn from robot execution, including ASENA and Recova. I also contribute to VLM and VLA foundation model efforts across Cosmos and Isaac GR00T.
Before joining NVIDIA, I earned my Ph.D. at the VLLAB at UC Merced, advised by Ming-Hsuan Yang. I have been fortunate to receive the Baidu Graduate Fellowship, the NVIDIA Pioneering Research Award, and the Rising Star EECS recognition.
Research Themes
Efficient Transformers
Designing propagation and attention modules that make high-resolution vision models faster and more memory-efficient.
Spatial VLMs
Connecting language, images, and 3D geometry for localization, spatial reasoning, and embodied understanding.
Robotics
Building agents that use perception, planning, and feedback to navigate, manipulate, and recover on real robots.
ASENA brings coding agents into the physical world through a real robot harness for perception, navigation, whole-body control, and execution feedback. Agents write, test, and refine robot programs, turning experience into reusable skills.
Recova uses an agent to build digital twins, discover recovery strategies in simulation, and recover from manipulation failures on real robots. Task and recovery policies improve from successful rollouts and human-assisted corrections, supporting longer autonomous operation with less human intervention.
GSPN is a fast vision attention module that accelerates generic vision foundation models for high-resolution input images.
Cosmos 3 is NVIDIA's omnimodal world foundation model that unifies understanding, generation, simulation, and action across text, image, video, audio, and robot actions for Physical AI. I serve as a core contributor on its spatial and embodied capabilities.
LoHo-Manip is a modular framework that scales short-horizon vision-language-action policies to long-horizon manipulation, using a task-management VLM that predicts subtask sequences and 2D visual traces to guide the executor with implicit progress tracking, replanning, and recovery.
GR3D is a spatial vision-language model that unifies explicit 2D, implicit 2D, and monocular 3D grounding in a single framework, decomposing spatial understanding into grounded 2D perception followed by 3D inference.
Compact GSPN (C-GSPN) is a ViT block with compressed spatial propagation and fused CUDA kernels that cuts propagation latency by nearly 10x, using a two-stage distillation scheme to scale subquadratic spatial propagation networks to vision foundation models.
SR-3D unifies single-view 2D and multi-view 3D representations for flexible region prompting and grounded spatial reasoning.
DAM generates detailed localized captions for user-specified regions in images and videos, preserving both local detail and global context.
TEVA improves high-resolution image understanding by dynamically selecting detail-rich regions while keeping token usage efficient.
NaVILA is a two-level framework that combines VLAs with locomotion skills for navigation. It generates high-level language-based commands, while a real-time locomotion policy ensures obstacle avoidance.
Efficient frontier VLM models with efficient training and inference.
SpatialRGPT is a grounded spatial reasoning model that can reason about spatial relationships in images.
3D Gaussian Splatting without COLMAP computation.
We present ODISE: Open-vocabulary DIffusion-based panoptic SEgmentation, which unifies pre-trained text-image diffusion and discriminative models to perform open-vocabulary panoptic segmentation.