Mapping claimed AI constructs, measurement instruments, and related behaviors/tasks
NeurIPS 2026 Evaluations and Datasets track (to appear).
Olawale Salaudeen1,†, Florian Dorner2,3∗, Tom Sühr2∗, Sang Truong4∗, Haoran Zhang1∗, Patrick Bissett4, James Fox5, Marzyeh Ghassemi1, Tom Hartvigsen6, Peter Hase4,5, Augustin Kelava7, Suhas Mahesh5, Margaret Mitchell8, Jamelle Watson-Daniels9, Sanmi Koyejo4, Angelina Wang10
1Massachusetts Institute of Technology
2Max Planck Institute for Intelligent Systems
3ETH Zürich
4Stanford University
5Schmidt Sciences AI Center
6University of Virginia
7University of Tübingen
8Hugging Face
9Meta AI (FAIR)
10Cornell Tech
†Correspondence: aiconstructlexis@gmail.com
∗ denotes equal contribution
Read more here.
Funded by the Schmidt Sciences AI Institute Fellow in Residence Program
AI Construct Lexis is a collaborative effort to map and define AI constructs, measurement instruments, and related behaviors/tasks—building a shared framework for understanding and evaluating AI systems. We ultimately aim to establish a nomological network for AI [1, 2, 3]: a connected system linking theoretical constructs, the ways we measure them, and the behaviors they predict in the real world (see Conjointly’s overview). We take inspiration from the Cognitive Atlas [4] in cognitive science and aim to represent the current state of thought in AI.
[1] Salaudeen, O., et al. (2025). Measurement to Meaning: A Validity-Centered Framework for AI Evaluation. arXiv arXiv:2505.10573. https://arxiv.org/abs/2505.10573
[2] MacCorquodale, K., & Meehl, P. E. (1948). On a distinction between hypothetical constructs and intervening variables. Psychological Review, 55(2), 95–107. https://doi.org/10.1037/h0056029
[3] Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302. https://doi.org/10.1037/h0040957
[4] Poldrack, R. A., et al. (2011). The cognitive atlas: toward a knowledge foundation for cognitive neuroscience. Frontiers in Neuroinformatics, 5, 17. https://doi.org/10.3389/fninf.2011.00017
Defined theoretical concepts that characterize AI systems, including capabilities (e.g., reasoning and software engineering ability) and risks (e.g., safety and bias). Each construct anchors how we interpret behavior and design measurements.
Operational definitions and measurements that translate constructs into observable quantities (e.g., benchmarks and user surveys). Each measurement is evaluated for its validity—whether it captures the intended construct and supports sound inferences about model behavior.
Observed behavioral patterns that reveal how constructs manifest, interact, and generalize across models, tasks, and environments (e.g., goal misalignment during tool use and the emergence of social bias in generated outputs).
For questions or collaboration, email aiconstructlexis@gmail.com.
Subscribe to our Substack for project updates and insights: @aiconstructlexis