top of page

A Reporting Checklist for Large Language Models in Behavioural Science

Feuerriegel, S., Christopher Barrie, M. J. Crockett… [72 authors] … Robb Willer, Dirk U. Wulff, Renwen Zhang, Simone Zhang, Steve Rathje, Manoel Horta Ribeiro.

Nature Human Behaviour.

Large language models (LLMs) are deep neural network architectures (typically transformers) trained on a large body of textual data that can generate human-like text. Many researchers across the behavioural and social sciences are enthusiastic about the potential of LLMs to open new avenues for studying human behaviour1. For example, researchers have proposed using LLMs to conduct in silico experiments that simulate human judgments and decisions in response to interventions, and to facilitate large-scale data annotation and analysis2. Other works have deployed LLMs as interventions to foster creativity, persuade, teach, or reduce misinformation beliefs3,4. However, the rapidly evolving role of LLMs in shaping empirical evidence, theoretical frameworks, policy decisions, and public discourse poses challenges for research rigour.

The alternative text for this image may have been generated using AI.
Here, we present a consensus-based checklist (‘GUIDE-LLM’) for research that involves LLMs in behavioural and social science, with the aim of strengthening transparency, reproducibility and ethical accountability. GUIDE-LLM stands for ‘guidelines for the use of LLMs in behavioural and social science’. The GUIDE-LLM checklist provides researchers with concrete guidance on improving transparency, reproducibility and ethical use, to strengthen rigour and documentation throughout the entire research workflow (including when LLMs function as research tools, as well as studies in which LLMs themselves are the object of empirical investigation).

bottom of page