Current and past directions are below. See the publications page for the full list, including recent works.

Current directions

Robust Finetuning of Foundation Models

Robust Finetuning of Foundation Models

Adapting large pretrained models to a target task (including for decision-makingin VLAs) without destroying the out-of-distribution generalization. We develop projection- and geometry-based constraints on how far finetuning may move the weights, and benchmarks to measure progress.

  • MAPS C. Huang et al. · CVPR 2026 arXiv Project Code
  • The Geometry of Robustness S. Chopra et al. · CVPR 2026 arXiv ★ Highlight
  • Grounding Descriptions in Images (GRAIN) S. Halbe et al. · WACV 2026 arXiv
  • Directional Gradient Projection (DiGraP) C. Huang et al. · ICLR 2025 arXiv
  • FRAMES-VQA C. Huang et al. · CVPR 2025 arXiv Code
  • Rethinking Weight Decay J. Tian et al. · NeurIPS 2024 arXiv
  • Fast Trainable Projection (FTP) J. Tian et al. · NeurIPS 2023 arXiv Code
  • Trainable Projected Gradient (TPGM) J. Tian et al. · CVPR 2023 arXiv Code
Vision-Language-Action Models and Generalist Agents

Vision-Language-Action Models & Generalist Agents

Turning multi-modal foundation models into policies that act, through supervised finetuning, reinforcement learning, and post-training — and understanding what transfers from web-scale pretraining to embodied control.

  • Generalist Embodied Agents (GEA) A. Szot et al. · CVPR 2025 arXiv ★ Oral presentation
  • Hierarchical RL for Quadruped Locomotion J. Coholich et al. · ACC 2025
  • Grounding MLLMs in Actions A. Szot et al. · NeurIPS 2024 arXiv
  • ReLIC A. Elawady et al. · arXiv 2024 arXiv
  • RL via Auxiliary Task Distillation A.N. Harish et al. · ECCV 2024 arXiv
  • Foundation Models for Robotics (survey) Y. Hu et al. · Preprint 2023 arXiv Project
  • BC-IRL A. Szot et al. · ICLR 2023 arXiv OpenReview
Memory, Verification, and Reasoning in Agents

Memory, Verification, and Reasoning in Agents

We build methods and benchmarks that isolate memory from perception, memory architectures that stay tractable over tens of thousands of steps, and verifiers that resist agreement bias.

  • FindingDory K. Yadav et al. · ECCV 2026 arXiv Project Code
  • Self-Grounded Verification M. Andrade et al. · ICLR 2026 arXiv
  • Reasoning Errors in VLM Verifiers J. Cha et al. · ICLR Workshop 2026 OpenReview
  • EVE Y. Ali et al. · ICRA Workshop 2026 arXiv PDF Talk PDF
  • Memo G. Gupta et al. · NeurIPS 2025 arXiv
Simulation, Real-to-Sim-to-Real, and Benchmarks

Simulation, Real-to-Sim-to-Real, and Benchmarks

Simulators and benchmarks for embodied AI, and the pipelines that carry policies between simulation and the real world. This includes personalized real-to-sim capture and viewpoint-robust transfer from fixed-camera data.

  • Sim2real Image Translation J. Coholich et al. · ICRA 2026 arXiv Project Code
  • EmbodiedSplat G. Chhablani et al. · ICCV 2025 arXiv Video
  • Habitat 3.0 X. Puig et al. · ICLR 2024 arXiv Project
  • GOAT-Bench M. Khanna et al. · CVPR 2024 arXiv Project Code
  • Seeing the Unseen R. Ramrakhya et al. · CVPR 2024 arXiv Code
  • HomeRobot S. Yenamandra et al. · CoRL 2023 arXiv Code
  • Social Embodied Rearrangement A. Szot et al. · ICML 2023 arXiv
3D Representations and Neural Fields

3D Representations and Neural Fields

Self-supervised pretraining and sparse-view synthesis for neural fields, and single-shot 3D shape, appearance, and pose estimation.

  • EscherNet++ X. Zhang et al. · CVPR Findings 2026 arXiv
  • EmbodiedSplat G. Chhablani et al. · ICCV 2025 arXiv Video
  • NeRF-MAE M.Z. Irshad et al. · ECCV 2024 arXiv Project Code
  • Neural Fields in Robotics (survey) M.Z. Irshad et al. · arXiv 2024 arXiv
  • FSD M. Lunayach et al. · ICRA 2024 arXiv Project
  • ICE-G V. Jaganathan et al. · CVPR Workshop 2024 arXiv
  • NeO 360 M.Z. Irshad et al. · ICCV 2023 arXiv Project Code
  • ShAPO M.Z. Irshad et al. · ECCV 2022 arXiv projec Code Video
  • CenterSnap M.Z. Irshad et al. · ICRA 2022 arXiv Project Code
Continual, Lifelong, and Open-World Learning

Continual, Lifelong, and Open-World Learning

Learning across a stream of tasks without rehearsing old data or forgetting past tasks, from prompt-based continual learning to continual customization of diffusion models, plus discovering categories in an unsupervised manner.

  • Domain Generalization meets GCD V. Rathore et al. · CVPR 2025 arXiv
  • Continual Diffusion (C-LoRA) J.S. Smith et al. · TMLR 2024 arXiv
  • STAMINA J.S. Smith et al. · CVPR Workshop 2024 arXiv
  • Adaptive Memory Replay J.S. Smith et al. · CVPR Workshop 2024 arXiv
  • Continual Federated Learning (HePCo) S. Halbe et al. · TMLR 2024 arXiv Code
  • CODA-Prompt J. Smith et al. · CVPR 2023 arXiv Code
  • ConStruct-VL J. Smith et al. · CVPR 2023 arXiv Code
  • Rehearsal-Free Continual Learning J. Smith et al. · CVPR Workshop 2023 arXiv
  • CLIP-GCD R. Ouldnoughi et al. · arXiv 2023 arXiv
  • Lifelong Learning Machines D. Kudithipudi et al. · Nature Mach. Intell. 2022 Nature

Earlier projects

Beyond Supervised Learning

Beyond Supervised Learning

A range of machine learning problems that move beyond needing large amounts of explicit human annotation: continual learning, cross-task learning (unlabeled datasets with entirely new categories), semi-supervised learning, one/few-shot learning, and domain adaptation.

  • Unbiased Teacher v2 Y.C. Liu et al. · CVPR 2022 arXiv Project Code
  • Open-Set Semi-Supervised Detection Y.C. Liu et al. · ECCV 2022 arXiv Project ★ Oral presentation
  • Joint Self-Supervised Temporal Domain Adaptation M.H. Chen et al. · CVPR 2020 arXiv Code
  • Manifold Graph with Learned Prototypes C.W. Kuo et al. · arXiv 2019 arXiv Project
  • A Closer Look at Few-shot Classification W. Chen et al. · ICLR 2019 PDF
  • Multi-Class Classification without Multi-Class Labels Y.C. Hsu et al. · ICLR 2019 PDF
  • Re-evaluating Continual Learning Scenarios Y.C. Hsu et al. · NeurIPS Continual Learning Workshop 2018 arXiv
Distributed Perception

Distributed Perception

Integrating information from multiple sensor modalities and robots in a principled way, including when a heterogeneous set of sensors sits on different platforms.

  • When2com Y.C. Liu et al. · CVPR 2020
  • Who2com Y.C. Liu et al. · ICRA 2020 arXiv
  • UNO J. Tian et al. · ICRA 2020 arXiv
Cross-Task Learning, Clustering, and Object Discovery

Cross-Task Learning, Clustering, and Object Discovery

Methods for automatically discovering object categories in unlabeled data, using cross-task learning and a novel deep learning-based clustering loss (part of the National Robotics Initiative project).

  • Learning to Cluster Across Domains and Tasks Y.C. Hsu et al. · ICLR 2018 arXiv
  • Probabilistic Constrained Clustering Y.C. Hsu et al. · CVPR Deep-Vision Workshop 2018 arXiv Code
  • Learning to Cluster for Instance Segmentation Y.C. Hsu et al. · IJCNN 2018 arXiv
  • Neural Network-Based Clustering Y.C. Hsu et al. · ICLR-W 2016 PDF arXiv (extended) Code
Goal-Driven Perception

Goal-Driven Perception

Incorporating decision-making — specifically top-down cues and goal reasoning — to dynamically change how perceptual inputs are processed.

  • The Regretful Agent C.Y. Ma et al. · CVPR 2019 arXiv Code Project
  • Self-Monitoring Navigation Agent C.Y. Ma et al. · ICLR 2019 PDF Code
Scene Flow

Scene Flow

Estimating scene flow (dense 3D motion fields) using factor graphs with continuous optimization, and more recently deep learning-based methods. Joint work with Frank Dellaert’s group.

  • Continuous Optimization for Scene Flow Z. Lv et al. · ECCV 2016 Project
Multi-Modal / Multi-Cue Fusion

Multi-Modal / Multi-Cue Fusion

Combining multiple modalities (e.g. LIDAR and images), including cueing and fusion at multiple levels — with improved performance from mid-level fusion and stable training at a small parameter cost.

  • Fusing LIDAR and Images for Pedestrian Detection J. Schlosser et al. · ICRA 2016
  • Long-Range Pedestrian Detection Z. Kira et al. · IROS 2012 PDF
Fine-Grained Video Analysis

Fine-Grained Video Analysis

Recurrent and convolutional networks that better exploit spatio-temporal structure in videos, modeling activities in a fine-grained manner. Joint work with Prof. AlRegib’s lab and NEC Labs.

  • Attend and Interact C.Y. Ma et al. · CVPR 2018 arXiv
  • TS-LSTM and Temporal-Inception C.Y. Ma et al. arXiv Code
Game Theory for Implicit Generative Models

Game Theory for Implicit Generative Models

Viewing machine learning problems through the lens of game theory, starting with implicit generative modeling (GANs) — with connections to online learning that inspired the DRAGAN regularizer for stable training.

  • How to Train Your DRAGAN N. Kodali et al. arXiv Code
Knowledge Transfer Across Heterogeneous Robots

Knowledge Transfer Across Heterogeneous Robots

Mid-level representations for learning object models and transferring them across heterogeneous robots with differing sensors — principles that remain relevant in the age of deep learning.

  • Inter-Robot Transfer Learning Z. Kira · AAMAS 2010 PDF
  • Grounded Symbolic Knowledge Across Robots (Ph.D.) Z. Kira · Ph.D. Dissertation, Georgia Tech 2010 PDF