Glenn Matlin
  • Home
  • About
  • Research
  • Work With Me
  • Reading Lists
  • Publications
  • Guides
  • Blog
  • CV

Collaboration

Reading Lists

These are the papers I want collaborators to have read before we talk. Each area below has a short description of what I care about and a short stack of papers to start with. Find the area you care about, read its list, and let it shape what you pitch me: this is the reading half of Work With Me, so read the list for your area before you fill out the interest form.

Fill Out the Interest Form Back to Work With Me

Capability Provenance & Training Data Attribution

I care about where language models learn to conceptualize humanity. If a model can infer motives, roles, norms, or patterns of social behavior, I want to know what data put that there. That is why I care about attribution and unlearning: can we trace a behavior back to specific parts of training data, and can we show that those parts were actually causal?

This area is for people who like attribution, unlearning, and hard questions about where model behavior actually came from.

5 papers to read first
  • Akyürek et al. — Towards Tracing Knowledge in Language Models Back to the Training Data
  • Chang et al. — Scalable Influence and Fact Tracing for Large Language Model Pretraining
  • Cheng et al. — Training Data Attribution (TDA): Examining Its Adoption & Use Cases
  • Ruis et al. — Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
  • Grosse et al. — Studying Large Language Model Generalization with Influence Functions

Persona Vectors & Internal Representations

I care about giving people tools to understand and use their language models effectively. For me, that means getting past black-box vibes and asking what the model is actually encoding about people internally. Persona vectors are one way into that question: are there stable directions for roles, perspectives, or styles of reasoning, and can we steer them without fooling ourselves?

This area is for people who like activation spaces, steering, and making model internals less mystical.

5 papers to read first
  • Zou et al. — Representation Engineering: A Top-Down Approach to AI Transparency
  • Rimsky et al. — Steering Llama 2 via Contrastive Activation Addition
  • Lee et al. — Do LLMs Have Distinct and Consistent Personality? TRAIT
  • Serapio-García et al. — A Psychometric Framework for Evaluating and Shaping Personality Traits in Large Language Models
  • Chen et al. — Persona Vectors: Monitoring and Controlling Character Traits in Language Models

Social Reasoning & Theory of Mind

I care about whether language models can actually reason about people: who knows what, who believes what, what someone intends, and how those mental states change over time. A lot of “social intelligence” claims are fake because the benchmark is weak. This area is about building and stress-testing evaluations for belief tracking, perspective-taking, common ground, deception, and social inference — then using those tests to separate real social reasoning from shallow pattern matching.

This area is for people who like social cognition, benchmark design, and careful evaluation of mental-state reasoning.

6 papers to read first
  • Chen et al. — ToMBench: Benchmarking Theory of Mind in Large Language Models
  • Kim et al. — FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions
  • Xu et al. — OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
  • He et al. — HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models
  • Xiao et al. — Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
  • Gandhi et al. — Understanding Social Reasoning in Language Models with Language Models

Safety & Misuse of Social Modeling

If a model can reason about beliefs, intentions, and vulnerability, it can do more than help people. It can flatter them, manipulate them, mislead evaluators, or target the users most likely to be persuaded. I care about measuring those failure modes, understanding their mechanisms, and building ways to detect or constrain them.

This area is for people who can hold both ideas at once: the capability matters, and the failure mode does too.

5 papers to read first
  • Sharma et al. — Towards Understanding Sycophancy in Language Models
  • Wen et al. — Language Models Learn to Mislead Humans via RLHF
  • Williams et al. — On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
  • Liu et al. — LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models
  • Scheurer et al. — Large Language Models Can Strategically Deceive Their Users When Put Under Pressure

Wargaming & Strategic Decision-Making

Wargames force models out of canned benchmark land. They have to plan, bluff, negotiate, interpret uncertainty, and deal with other agents pushing back. I care most about open-ended wargames, where language matters and the stakes are not fake just because the environment is simulated.

Wargaming is not strictly about violent military conflict. Wargames are a tool for reasoning about any situation with opposing forces and high stakes — climate change, pandemics, rogue AI, market crises. The point is structured reasoning under pressure, not the battlefield.

This area is for people who like multi-agent reasoning, planning, and open-ended environments where the benchmark does not do the thinking for you.

5 papers to read first
  • Matlin et al. — Shall We Play a Game? Language Models for Open-ended Wargames
  • Bakhtin et al. — Human-level Play in the Game of Diplomacy by Combining Language Models with Strategic Reasoning
  • Kramár et al. — Negotiation and Honesty in Artificial Intelligence Methods for the Board Game of Diplomacy
  • Lamparth et al. — Human vs. Machine: Behavioral Differences Between Expert Humans and Language Models in Wargame Simulations
  • Rivera et al. — Escalation Risks from Language Models in Military and Diplomatic Decision-Making

Negotiation & Cooperation

I do not just care about bargaining in the narrow sense. I care about deals, coalitions, cooperation under conflict, and what happens when models have to manage incentives over multiple turns with other agents who want different things.

This area is for people who like multi-turn social interaction, incentives, coalitions, and strategic communication.

5 papers to read first
  • Bianchi et al. — How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis
  • Abdelnabi et al. — Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
  • Vaccaro et al. — Advancing AI Negotiations: New Theory and Evidence from a Large-Scale Autonomous Negotiations Competition
  • Mukobi et al. — Welfare Diplomacy: Benchmarking Language Model Cooperation
  • Akata et al. — Playing Repeated Games with Large Language Models

Finance & Financial NLP

Finance is where plausible-sounding bullshit gets expensive. I care about whether models can actually reason over filings, calculations, instructions, and evidence — not just sound finance-coded.

This area is for people who care about domain-specific reasoning, evaluation, and high-stakes use cases where “sounds right” is useless.

5 papers to read first
  • Chen et al. — FinQA: A Dataset of Numerical Reasoning over Financial Data
  • Chen et al. — ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering
  • Islam et al. — FinanceBench: A New Benchmark for Financial Question Answering
  • Matlin et al. — Financial Language Model Evaluation (FLaME)
  • Xie et al. — FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning

Read the list for your area, then head back to Work With Me and fill out the interest form. Tell me what you read, what you think, and what you want to do.

© 2025-2026 Glenn Matlin