ResearchMy work connects robotics, language, and vision with an emphasis on reliability in real-world and long-horizon settings. Across DRIC, NTU, and Columbia, a recurring theme has been how multimodal systems gather evidence, revise beliefs, and remain useful when observations are incomplete, ambiguous, or physically constrained. Research Interests
Real-World Robot Learning and Embodied AIMy current work focuses on robot-learning infrastructure, tactile and visual policy learning, VLA-style manipulation, world-model-based reasoning, rollout evaluation, and failure-driven improvement for real-world manipulation. A recurring question in this line of work is how embodied systems can monitor uncertainty, gather missing evidence, and recover when contact-rich execution or navigation plans start to fail. Representative work: CARe, AED, Affordance-Guided Base Placement, VLN-NF. Trustworthy LLM and MLLM ReasoningI study how language and multimodal models reason under incomplete, misleading, or physically constrained evidence, with an emphasis on false premises, underspecified affordances, abstention, and failure analysis. This direction centers on evidence-seeking behavior and constraint-aware reasoning in embodied and multimodal settings. Long-Horizon Multimodal Understanding and EvaluationAcross movies and long-form multimodal narratives, I build datasets, benchmarks, and evaluation setups for reasoning that unfolds over extended time horizons. This direction focuses on temporal coherence, narrative structure, evidence tracking, and evaluation protocols for long-context multimodal understanding. Earlier Representative WorkEarlier projects that still connect to my current research directions include the following themes. Multimodal Question Generation / QA: CaKE , GPN , SRCMSA . 3D Vision: MonoDTR , S^3 , OCID-Ref . Robotics and Embodied Grounding: GDN , SCAN , OCID-Ref . Data-Efficient Learning: SCAN , Class-agnostic Few-shot Object Counting , ReDAL . |