Research

Understanding, or just correlation?

My work asks whether a model's apparent competence reflects real understanding of the structure in its data, or a correlation that happened to hold where it was tested. I study this through robustness, multimodal reasoning benchmarks, and physics-grounded generative models.

01

Robustness

When inputs shift in several ways at once, what breaks, and can a model learn to withstand all of it together?

02

Multimodal reasoning

Are vision-language models reading the image, or leaning on priors and guesses?

03

Scientific imaging

How do we make generative models respect the physics of the instrument, instead of hallucinating detail?

Publications

Google Scholar →
Blurry elongated microscopy volume transforming into a crisp isotropic 3D structure
CVPR 2026Microscopy · Generative models

MicroFM: Physics-guided Flow Matching for Isotropic Microscopy Reconstruction

Xingzu Zhan, Runmin Jiang, Vatsal Gupta, Tanush Swaminathan, Yanwen Wang, Genpei Zhang, Haili Wang, Min Xu

Microscopes see far less detail along the depth axis. Can a generative model fill it in without making things up? MicroFM trains on data simulated with realistic optics and starts its flow from the actual observation rather than noise, which keeps reconstructions faithful and makes sampling fast.

Workflow diagram: language model perturbations, analysis and robustness strategies
EMNLP 2024Robustness · NLP

Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets

Vatsal Gupta*, Pranshu Pandya*, Tushar Kataria, Vivek Gupta, Dan Roth

Robustness is usually tested one perturbation at a time, but real inputs are messy in several ways at once. We study how language models behave under multiple simultaneous perturbations, and show a mixed-training strategy can recover robustness to all of them without hurting clean accuracy.

* and † denote equal contribution. For a complete, up-to-date list see Google Scholar.