Publication type
Survey Futures Working Paper Series
Series Number
22
Series
Survey Futures Working Paper Series
Authors
Publication date
July 17, 2026
Summary:
The Question Appraisal System (QAS-99) is a widely used protocol for the systematic assessment of survey questions, requiring evaluators to make 27 binary coding decisions across eight appraisal steps. QAS coding is slow and resource-intensive, creating the possibility that large language models (LLMs) can replace or supplement human evaluators. We compared LLM evaluations of 118 draft survey questions with those of a QDET expert and novice coders. Agreement was further assessed using additional human evaluators and a second LLM on a validation subset. We find that LLMs produced actionable QAS evaluations at scale, with agreement levels broadly comparable to those observed between novice and expert human evaluators. LLMs identified fewer problems than novices, who in turn identified fewer problems than experts, suggesting that expertise currently remains important for questionnaire review. Agreement varied substantially across QAS decisions, indicating that some question characteristics are inherently more difficult to evaluate than others. Compared with the expert coder, the LLM was notably less likely to identify certain data quality risks, particularly those related to social desirability bias. Nevertheless, the findings suggest that LLMs can support QAS-based survey appraisal at scale and large-scale research on questionnaire design and survey measurement error.
Subjects
Paper download#589116