Benchmarking Large Language Model Rationality Using Measurement Axioms

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

While Large language models (LLMs) are increasingly deployed as decision-makers, current evaluation practice emphasizes benchmark accuracy rather than the structural properties that make decisions coherent and interpretable. We introduce a measurement-theoretic framework for evaluating AI decision-making using formal axioms derived from utility theory, focusing on transitivity of preference. When transitivity is violated over a choice set, alternatives can no longer be coherently ordered as globally better or worse, undermining the interpretability of model preferences and complicating alignment with human values. We apply our framework across 20 distinct LLMs, systematically varying question format, generation temperature, and contextual memory, using classic choice paradigms previously shown to induce intransitive preferences in humans. We find substantial and systematic transitivity violations in many models, although some patterns partially mirror known features of human decision behavior. Measurement axioms thus offer a principled, complementary foundation for evaluating the rationality and coherence of AI systems beyond benchmark accuracy.

Article activity feed