Glenn Matlin
  • Home
  • About
  • Research
  • Work With Me
  • Reading Lists
  • Publications
  • Guides
  • Blog
  • CV

Financial Instruction Following Evaluation (FIFE)

NLP
Finance
Evaluation
A high-difficulty benchmark of 88 human-authored prompts with chainable, verifiable constraints, evaluating 53 models on complex financial instruction following.
Authors

Glenn Matlin

Siddharth

Anirudh JM

Aditya Shukla

Yahya Hassan

Sudheer Chava

Published

December 1, 2025

Publication

Financial Instruction Following Evaluation (FIFE)

A high-difficulty benchmark of 88 human-authored prompts with chainable, verifiable constraints, evaluating 53 models on complex financial instruction following.

Published

December 1, 2025

Authors

Glenn Matlin, Siddharth, Anirudh JM, Aditya Shukla, Yahya Hassan, Sudheer Chava

Venue

GenAI Finance Workshop at NeurIPS 2025

Read on arXiv Code & Data All Publications

Abstract

Language Models (LMs) struggle with complex, interdependent instructions, particularly in high-stakes domains like finance where precision is critical. We introduce FIFE, a novel, high-difficulty benchmark designed to assess LM instruction-following capabilities for financial analysis tasks. FIFE comprises 88 human-authored prompts and employs a verification system with chainable, verifiable constraints for fine-grained reward signals. We evaluate 53 models (proprietary, open-weight, open-source) in a zero-shot setting. Our key findings reveal a clear performance hierarchy: the top open-weight model (76.1 strict / 79.5 loose) surpasses the leading proprietary system (65.9 strict / 70.5 loose), while the best open-source models lag significantly (45.5 strict / 48.9 loose). However, even top-performing models struggle with FIFE’s complex requirements, failing to achieve perfect compliance. We release our dataset and code as an open-source resource to promote research in Reinforcement Learning for the financial domain.

At a Glance

  • 88 human-authored prompts with chainable, verifiable constraints for fine-grained reward signals
  • 53 models evaluated zero-shot across proprietary, open-weight, and open-source families
  • Clear capability hierarchy: top open-weight (76.1 strict / 79.5 loose) > leading proprietary (65.9 / 70.5) > best open-source (45.5 / 48.9) — and no model achieves perfect compliance
  • Dataset and code released to support RL for the financial domain

Cite This Paper

BibTeX
@article{matlin2025fife,
  title   = {Financial Instruction Following Evaluation (FIFE)},
  author  = {Matlin, Glenn and Siddharth and JM, Anirudh and Shukla, Aditya and Hassan, Yahya and Chava, Sudheer},
  year    = {2025},
  journal = {arXiv preprint arXiv:2512.08965},
  note    = {GenAI Finance Workshop at NeurIPS 2025}
}

Continue exploring

Return to the publication archive or step back to the broader research agenda.

Publications Research

© 2025-2026 Glenn Matlin