Glenn Matlin
Chandreyi Chakraborty
Saehee Eom
Mika Okamoto
Rayan Castilla
Louis Jaburi
Alvin Deng
Taywon Min
Lucia Quirke
Stella Biderman
Mark Riedl
July 21, 2026
Publication
We use training-data attribution to map which regions of a model’s pretraining corpus support social reasoning versus STEM reasoning in OLMo3-7B, and find the two draw on qualitatively distinct parts of the corpus.
Published
July 21, 2026
Authors
Glenn Matlin, Chandreyi Chakraborty, Saehee Eom, Mika Okamoto, Rayan Castilla, Louis Jaburi, Alvin Deng, Taywon Min, Lucia Quirke, Stella Biderman, Mark Riedl
Venue
Conference on Language Models (COLM) 2026

We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social-reasoning versus STEM-reasoning in OLMo3-7B. Training-data attribution measures how strongly each training document influences a model’s predictions on a benchmark, but document-level scores are too noisy to identify which corpus regions support which capabilities, and prior work has emphasized factual knowledge rather than reasoning. We compute gradient-based attribution (TrackStar via Bergson) over a working set drawn from the de-duplicated Dolma3 mix, aggregate influence across WebOrganizer’s 24-format × 24-topic taxonomy (576 bins), and contrast benchmark pairs in a 2×2 design that varies domain (social vs. STEM) and capability type (reasoning vs. knowledge): SocialIQA and MMLU Social Sciences against ARC-Challenge and MMLU STEM. Social and STEM reasoning draw on qualitatively distinct corpus regions, and the contrast is sharper at the reasoning level than at the knowledge level. Targeted machine unlearning provides partial causal validation: forgetting high-attribution topic bins (e.g., Literature for SocialIQA) degrades the aligned benchmark more than within-bin random baselines. Re-running the pipeline on two further open-data models, Comma v0.1 (Common Pile) and DCLM-7B, keeps it causally selective while the recovered provenance map stays ecosystem-specific.
@inproceedings{matlin2026capability,
title = {Capability Provenance in Language Models: A Case Study in Social Reasoning},
author = {Matlin, Glenn and Chakraborty, Chandreyi and Eom, Saehee and Okamoto, Mika and Castilla, Rayan and Jaburi, Louis and Deng, Alvin and Min, Taywon and Quirke, Lucia and Biderman, Stella and Riedl, Mark},
booktitle = {Conference on Language Models (COLM)},
year = {2026}
}Continue exploring
Return to the publication archive or step back to the broader research agenda.