Research
I'm interested in building systems that learn structured, reusable representations of actions and scenes, with a focus on post-training vision-language models for long-video understanding and temporal reasoning. I'm also interested in world models for embodied reasoning and self-improving agents.
|
|
CrossView: Can Vision-Language Models Reason Across Cameras?
Sahil Shah*,
S. P. Sharan*,
Harsh Goel*,
Manvik Pasula,
Adithya Hebbalae,
Minkyu Choi,
Sandeep Chinchali
ECCV, 2026
project page
/
arXiv
/
code
A multi-camera video QA benchmark spanning driving, surveillance, ego/exo, and robotics, where today's best models still fall short of reasoning across simultaneous views.
|
|
Incentivizing Vision Language Models to Search for Long Video Question Answering
Harsh Goel*,
S. P. Sharan*,
Sahil Shah,
Minkyu Choi,
Joungbin An,
Kristen Grauman,
Sandeep Chinchali
ECCV, 2026   (Spotlight Presentation)
project page
/
arXiv
/
code
Turning long-video QA into a multi-turn search, with temporal-logic-derived verifiable rewards for RL post-training, lifts Pass@1 by 8% and Pass@4 by 15%.
|
|
ObjectAlign: Neuro-symbolic Object Consistency Verification and Correction
Mustafa Munir,
Harsh Goel,
Xiwen Wei,
Minkyu Choi,
Sahil Shah,
Kartikeya Bhardwaj,
Paul Whatmough,
Sandeep Chinchali,
Radu Marculescu
CVPRW, 2026
arXiv
Verifying object consistency across frames catches objects that drift, morph, or vanish during video editing, then corrects them in place.
|
|
RT-NeuS: Towards Real-Time Neuro-Symbolic Video Understanding via Adaptive Temporal Verification
Shawn Liang,
Sahil Shah,
Chengwei Zhou,
S. P. Sharan,
Harsh Goel,
Arnab Sanyal,
Sandeep Chinchali,
Gourav Datta
NeuS, 2026
arXiv
Adaptive sampling and keyframe selection bring neuro-symbolic temporal verification into the real-time regime without giving up multi-step reasoning.
|
|
NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic Reasoning
Sahil Shah,
S. P. Sharan*,
Harsh Goel*,
Minkyu Choi,
Mustafa Munir,
Manvik Pasula,
Radu Marculescu,
Sandeep Chinchali
AAAI, 2026
project page
/
paper
/
code
Training-free pipeline that identifies logical event sequences in video, boosting VQA accuracy by over 10% on causal and multi-step reasoning tasks.
|
|
COFFEE: a High-Performance Approach to Convex Optimization for Thermodynamic Equilibrium Computations
Fu-Yao Yu*,
Sahil Shah*,
Yash Mittal,
Paul Bessler,
Aamir Mohsin,
Jeffrey Geng,
Arnav Vats,
David Soloveichik
SIEDS, 2025   (Best Paper)
project page
/
paper
/
code
A trust-region convex solver for molecular equilibrium runs 2× faster and 10⁷× more accurately than prior tools, scaling to large biochemistry datasets.
|
|
A Challenge to Build Neuro-Symbolic Video Agents
Sahil Shah,
Harsh Goel,
Sai Shankar Narasimhan,
Minkyu Choi,
S. P. Sharan,
Oguzhan Akcin,
Sandeep Chinchali
NeuS, 2025
arXiv
/
code
Combining neuro-symbolic reasoning with video perception enables agents that can interpret, predict, and act on temporal events, not just recognize them.
|
|
We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback
Minkyu Choi*,
S. P. Sharan*,
Harsh Goel,
Sahil Shah,
Sandeep Chinchali
arXiv preprint, 2025
arXiv
Neuro-symbolic feedback enables zero-shot refinement of generated videos, boosting temporal and semantic alignment by nearly 40% without retraining.
|
|
Neuro-Symbolic Evaluation of Text-to-Video Models using Formal Verification
S. P. Sharan*,
Minkyu Choi*,
Sahil Shah,
Harsh Goel,
Mohammad Omama,
Sandeep Chinchali
CVPR, 2025
project page
/
paper
/
code
Formally verifying a video against a temporal logic specification yields a text-to-video metric that aligns 5× better with human judgment than existing scores.
|
|
Real-Time Privacy Preservation for Robot Visual Perception
Minkyu Choi*,
Yunhao Yang*,
Neel P Bhatt*,
Kushagra Gupta,
Sahil Shah,
Aditya Rai,
David Fridovich-Keil,
Ufuk Topcu,
Sandeep Chinchali
TMLR, 2025
arXiv
Blurring objects based on logical specifications enables real-time video privacy with 95%+ compliance, provable guarantees, and seamless robot deployment.
|
|
Towards Neuro-Symbolic Video Understanding
Minkyu Choi,
Harsh Goel*,
Mohammad Omama*,
Yunhao Yang,
Sahil Shah,
Sandeep Chinchali
ECCV, 2024   (Oral Presentation)
project page
/
paper
/
code
Decoupling perception and temporal reasoning with TL-based state machines boosts long-range event identification by up to 15% on self-driving datasets.
|
|