Automated survey analysis with AI for the World Bank

1. Context and challenge

The World Bank’s communications and business intelligence team runs perception surveys across dozens of countries. Each survey includes open-ended questions, such as what the Bank could do to increase its effectiveness in a given country. Coding thousands of free-text answers by hand took analysts many hours per survey, and consistency between coders was difficult to guarantee.

2. Our role

The World Bank commissioned Research for Purpose to automate the analysis of open-ended answers across a few pilot country surveys. We were responsible for the platform architecture, the optimization of the AI prompts, and the quality control framework.

3. What we did

We built a platform in which analysts set up each survey question, load the answers and define the categories to be analyzed. Next, we tested the prompts the World Bank’s team had been using and improved them substantially. For each category, the AI returns a classification together with a confidence score and a written explanation of its decision. AI agents check the outputs, and any low-confidence result is flagged for review.

Refining the prompts produced several practical rules. Categories work best when they are defined first by what they include and then by what they exclude. Worked examples inside a prompt were found to bias results, so we removed them. Each prompt was tested on the first 50 to 100 rows before being applied to a full survey, and further reviewed. We documented these rules in prompting guidelines so the team can write new category prompts on its own, and we supported them with AI training for the Bank’s global communications and external affairs team.

4. Results and learnings

AI categorization exceeded 90-95% accuracy, which is more accurate than the manual analysis it replaced. Analysts can segment answers by any filter available in the data and generate an AI summary of each segment. The platform saves the team tens of hours of manual work per survey.

The remaining 5-10% of AI data categorization was not inaccurate as there was no consensus among analysts what categorization is correct.

Facebook
Twitter
LinkedIn