Collaboration
These are the papers I want collaborators to have read before we talk. Each area below has a short description of what I care about and a short stack of papers to start with. Find the area you care about, read its list, and let it shape what you pitch me: this is the reading half of Work With Me, so read the list for your area before you fill out the interest form.
I care about where language models learn to conceptualize humanity. If a model can infer motives, roles, norms, or patterns of social behavior, I want to know what data put that there. That is why I care about attribution and unlearning: can we trace a behavior back to specific parts of training data, and can we show that those parts were actually causal?
This area is for people who like attribution, unlearning, and hard questions about where model behavior actually came from.
I care about giving people tools to understand and use their language models effectively. For me, that means getting past black-box vibes and asking what the model is actually encoding about people internally. Persona vectors are one way into that question: are there stable directions for roles, perspectives, or styles of reasoning, and can we steer them without fooling ourselves?
This area is for people who like activation spaces, steering, and making model internals less mystical.
If a model can reason about beliefs, intentions, and vulnerability, it can do more than help people. It can flatter them, manipulate them, mislead evaluators, or target the users most likely to be persuaded. I care about measuring those failure modes, understanding their mechanisms, and building ways to detect or constrain them.
This area is for people who can hold both ideas at once: the capability matters, and the failure mode does too.
Wargames force models out of canned benchmark land. They have to plan, bluff, negotiate, interpret uncertainty, and deal with other agents pushing back. I care most about open-ended wargames, where language matters and the stakes are not fake just because the environment is simulated.
Wargaming is not strictly about violent military conflict. Wargames are a tool for reasoning about any situation with opposing forces and high stakes — climate change, pandemics, rogue AI, market crises. The point is structured reasoning under pressure, not the battlefield.
This area is for people who like multi-agent reasoning, planning, and open-ended environments where the benchmark does not do the thinking for you.
I do not just care about bargaining in the narrow sense. I care about deals, coalitions, cooperation under conflict, and what happens when models have to manage incentives over multiple turns with other agents who want different things.
This area is for people who like multi-turn social interaction, incentives, coalitions, and strategic communication.
Finance is where plausible-sounding bullshit gets expensive. I care about whether models can actually reason over filings, calculations, instructions, and evidence — not just sound finance-coded.
This area is for people who care about domain-specific reasoning, evaluation, and high-stakes use cases where “sounds right” is useless.
Read the list for your area, then head back to Work With Me and fill out the interest form. Tell me what you read, what you think, and what you want to do.
Social Reasoning & Theory of Mind
I care about whether language models can actually reason about people: who knows what, who believes what, what someone intends, and how those mental states change over time. A lot of “social intelligence” claims are fake because the benchmark is weak. This area is about building and stress-testing evaluations for belief tracking, perspective-taking, common ground, deception, and social inference — then using those tests to separate real social reasoning from shallow pattern matching.
This area is for people who like social cognition, benchmark design, and careful evaluation of mental-state reasoning.
6 papers to read first