Research
How do language models come to understand people, and how do we know when that understanding is real rather than convincing imitation? I work on the social reasoning, evaluation, and interpretability of these systems: where the capability comes from in training, how it takes shape inside the model, and how to test it in domains where a wrong answer carries real cost.
Core Agenda
Data provenance Tracing social capability back to concrete training data choices.
Representations Studying how social roles and personas appear in activation space.
Evaluation Building benchmarks where institutions, incentives, and instructions interact.
Safety Understanding how useful social reasoning and misuse risk rise together.
Most of my work sits in one of four lanes, but the questions are connected: what in the data teaches models about people, what structure that creates inside the model, and what evaluation is strong enough to show whether the behavior generalizes.
Provenance
I study where social capability comes from in the training corpus. My COLM 2026 paper pairs training-data attribution with targeted unlearning to tie specific pretraining regions to social reasoning, so a shift in behavior maps to a data decision instead of a black box.
Representations
I look at how professions, personas, and other social categories are organized in activation space, and when that structure is compositional or steerable instead of merely descriptive.
Measurement
I build evaluation tools such as FLaME, FIFE, and FinForge to test model behavior where institutions, incentives, and instructions interact in ways that are difficult to fake with shallow pattern matching.
Implications
I treat safety as part of the same research program. Better models of people can create utility, but they also make misuse and manipulation easier, so the explanatory work has to keep pace with the capability work.
These projects are the clearest examples of the broader agenda: measurement where the stakes are concrete, and interpretability that ties model behavior back to a mechanism instead of a vague story.
COLM 2026
Provenance: training-data attribution (TrackStar) plus targeted unlearning locate the pretraining regions behind social versus STEM reasoning in OLMo3-7B, turning a correlation into a causal test.
ACL 2025 Findings
Measurement: a benchmark of 20 core finance tasks across 23 models, built to test whether apparent domain competence holds up when institutions and incentives are in play.
COLM 2025
Representations: a study of whether models register narrative suspense where readers do, probing how well their social and affective judgment matches ours.
NeurIPS 2025 workshop
Measurement: a high-difficulty benchmark of chainable, verifiable constraints for financial instructions; none of the 53 models evaluated passes it cleanly.
AAAI 2026 workshop
Measurement: a semi-synthetic pipeline that manufactures finance evaluation data when privacy, rarity, or collection cost make real corpora too narrow to trust.
EMNLP 2025 workshop
Implications: open-ended wargames test whether a model can hold beliefs, goals, and uncertainty across many turns instead of reacting one prompt at a time.