Publications
This is the full publication archive. For the broader research agenda and representative projects, visit the Research page.
Google Scholar Semantic Scholar
Training-data attribution maps which regions of a model’s pretraining corpus support social versus STEM reasoning, and the two draw on qualitatively distinct regions.
BELLA (Budget-Efficient LLM Selection via Automated skill-profiling): interpretable, skill-based model selection that makes cost-performance trade-offs explicit instead of black-box.
A scalable semi-synthetic pipeline for financial evaluation benchmarks, producing FinForge-5k: 5,000+ human-validated QA pairs across 11 finance subdomains from a 100k-document corpus.
A high-difficulty benchmark of 88 human-authored prompts with chainable, verifiable constraints, evaluating 53 models on complex financial instruction following.
A position paper and scoping review of 100 studies on AI in wargames, with a novel ontology of open-endedness, deployment recommendations, and open research challenges.
Replicating four seminal psychology studies with language models: LMs can tell whether a text is suspenseful, but not how suspenseful, nor how suspense rises and falls.
The first holistic benchmarking suite for financial NLP: 23 foundation language models evaluated across 20 core finance tasks, with open-source framework, data, and results.
A cost-aware and uncertainty-based framework for dynamic 2D prediction in multi-stage classification systems.