How should we validate LLMs used as human proxies?
A new Nature Computational Science Review distinguishes four roles for LLM human proxies and explains why each role requires its own validity criteria.
Nature Computational Science has published a Review by Trust Lab researcher Nikita Karetnikov, Iyad Rahwan, and Trust Lab lead Davor Svetinovic. The paper examines how large language models are used to stand in for people across scientific and applied settings.
Its starting point is that “human-like” is too broad to serve as an evaluation standard. A convincing character, a capable task agent, an experimental subject, and a model intended to predict a population make different claims. Treating them as equivalent can produce evidence that does not support the conclusion being drawn.
For trusted agentic systems, the practical consequence is direct: evaluation should begin with the role and the claim, then trace both to the relevant human reference and test. Believability does not establish task competence, and task competence does not show that a model represents a person or population.
Role-specific validity
Four roles, four different claims
Believable agents
Assess whether people accept a coherent character or persona in its intended context.
Task agents
Evaluate whether an agent performs work correctly, robustly, and efficiently against a suitable human baseline.
Experimental subjects
Study the model’s own behaviour with controls for construct validity, reproducibility, and model sensitivity.
Silicon samples
Test predictions about people or populations against human ground truth, calibration, and subgroup validity.
Contributions and support
How the Review was developed
Nikita Karetnikov surveyed the literature and wrote the paper. Iyad Rahwan and Davor Svetinovic contributed to its framing and structure, and Davor Svetinovic initiated and supervised the work.
The research acknowledges support from the Khalifa University Research Center for Advanced Intelligent Systems.
Publication record
Large language models as human proxies
Nikita Karetnikov, Iyad Rahwan, Davor Svetinovic
Nature Computational Science
