Machine Learning Research at Georgia Tech
I study artificial intelligence, machine learning, and language models, with a focus on social reasoning, evaluation, and interpretability in finance, wargaming, law, and other high-stakes domains.

Each project asks the same question from a different angle: what do models infer about people, and how do we test whether that understanding is real?
COLM 2026
Mapped which slices of a model’s pretraining data produce social reasoning versus STEM skill, then validated the link by unlearning those slices and measuring the drop.
ACL 2025 Findings
Ran 23 foundation models across 20 core financial tasks to find where domain fluency is real and where it collapses under evaluation pressure.
COLM 2025
Tested whether models feel narrative tension where readers do, and found where their sense of suspense lines up with people and where it doesn’t.
AAAI 2026 workshop
Generated financial evaluation data on demand so benchmarks stay usable when privacy, rarity, or cost make real datasets too thin.
“Capability Provenance in Language Models” accepted to COLM 2026.
FinForge presented at the AAAI 2026 Agentic AI in Financial Services workshop.
FIFE presented at the NeurIPS 2025 GenAI Finance workshop, evaluating 53 models.
Suspense study accepted to COLM 2025, in the top 3% of submissions.
FLaME published in Findings of ACL 2025.
I work where models interact with people, institutions, and incentives rather than clean toy tasks.
Language models meet regulation, incentives, and real decision costs here. That makes evaluation grounded rather than abstract.
Open-ended planning exposes whether a model can track beliefs, goals, and uncertainty over time instead of pattern-matching one step at a time.
Law, healthcare, and other institutional settings are where social misunderstanding turns into brittle automation, bad advice, or manipulation risk.
If you’re working on model behavior, interpretability, or high-stakes evaluation, I’d love to compare notes.