Examples and research
These are real, attributed examples of development and public-sector organizations using AI in their work, plus a few research-backed ones. They are here so the rest of this resource is grounded in things that actually happened, not in promises.
Read them honestly. This is an early, fast-moving, uneven space. Most of these are pilots, single studies, or documented experiments, not proof that a method is settled or that it will work the same way for you. Several of them are valuable precisely because they document where AI fell short, not just where it helped. We have grouped them by what each one shows. Every example is attributed; where we could not stand behind a claim, we left it out.
Synthesis: reading and mapping across a large body of evidence
A large children's agency used text mining and search to speed up extraction across 631 evaluation reports for an evidence synthesis, instead of reading every report by hand.
Source: UNICEF evaluation evidence synthesis across 631 reports, write-up on medRxiv.
A UN-system evaluation office built an AI tool to map and summarize evaluations across many agencies, so findings held in separate archives can be searched as one body.
Source: UN System-Wide Evaluation Office (SWEO), AI evaluation-mapping tool.
A national government set up an AI-assisted platform to make its evaluation evidence easier to find and reuse across the public sector.
Source: Finland "OpenEval" national AI-assisted evaluation evidence platform.
A development bank's independent evaluation group tested machine learning to assist evaluative synthesis in a private-sector evaluation.
Source: Bravo et al. (2023), World Bank Group Independent Evaluation Group (IEG), ML in evaluative synthesis.
A large evidence intermediary maintains a portal of 13,000-plus impact evaluations and evidence gap maps, which is the kind of structured corpus that AI-assisted synthesis can sit on top of.
Source: 3ie Evidence Gap Maps and impact evaluation portal.
Analysis: extracting, coding, and gap-checking with a framework
A government department, working with an M&E-tech group, used a general AI chat tool to find gaps in evaluation reports against an equity framework, using a defined four-step method with a human in the loop. It worked when the framework was supplied; it missed disability and accessibility gaps when the framework did not name them. This is one of the best-documented end-to-end pilots precisely because it reports its failure modes.
Source: Shared Services Canada x MERL Tech, GBA+ gap-finding pilot.
A research study used generative AI to help code semi-structured interviews in a maternal-health study, testing how well AI-assisted thematic analysis holds up against human coding.
Source: GenAI for thematic analysis, maternal-health interview study, medRxiv.
A pair of qualitative-analysis case studies documented one project where AI-assisted coding succeeded and one where it disappointed, side by side, which is unusually honest about when the method does and does not deliver.
Source: Friese, "Success and Disenchantment," paired qualitative-analysis case studies.
Reporting and service delivery: turning data into something people use
A data-and-delivery partnership built a COVID-19 data system for a major city that produced daily customized reports for city leadership, turning a fast-moving data stream into decision-ready briefings.
Source: IDinsight x Dimagi, Delhi COVID-19 data system.
A UN agency's independent evaluation office publicly described standing up a generative-AI-supported evaluation function, an early institutional commitment rather than a one-off experiment.
Source: UNFPA Independent Evaluation Office, GenAI-powered evaluation function (2024); referenced in UNEG ethical-principles documentation.
What the adoption picture actually looks like
These are not success stories; they are the honest backdrop. They explain why this resource is cautious.
A large humanitarian-sector survey (over 2,500 workers across 144 countries, with most respondents in the Global South) found that almost everyone has tried AI and most use it weekly, but only a small share of organizations have integrated it widely, few have a formal AI policy, and most offer little or no training. People are using consumer AI tools on their own, ahead of any guidance.
Source: Humanitarian Leadership Academy + Data Friendly Space survey (2025), the "humanitarian AI paradox" / shadow AI.
A philanthropy-sector technology survey of more than 350 foundations found a large majority reporting some AI use but only around a third having an AI policy, with concrete uses clustering on triage and synthesis rather than final decisions.
Source: TAG State of Philanthropy Tech (2024).
Sector research has documented that AI assistance can lower the English-language barrier for local and Global-South organizations and, at the same time, that AI-writing detectors over-flag non-native-English writers. Both are real, and they pull in opposite directions.
Source: Detector false-positive findings from Turnitin and GPTZero studies, as summarized in our research for this resource.
How to read all of this
The pattern across the credible examples is consistent. AI helped most when it was pointed at a defined body of material, given a clear framework or question, asked to extract rather than to conclude, and kept under human review. It disappointed when it was trusted to author judgments, when no framework was supplied, or when its outputs were taken at face value. That is the same throughline as the rest of this resource: use AI for the reading, sorting, and first-draft work, and keep the thinking, the context, and the final call with a person.