Where AI fits across your program
A practical map of where AI helps in development work, where it does not, and how to adopt it without getting burned. M&E is the anchor and the lens: we start from the function you already run, then widen out to the rest of the program.
This is a peer's view, not a principles lecture and not a sales page: what AI does well at each stage, where a human still has to hold the pen, and what to do when an output looks confident and is wrong. Everything here is free.
Start where you are
- I want to orient.You are not sure where AI even belongs in your work. Start with the map below.The landscape map
- I have a task to do.You know the job (clean a dataset, code open-ends, draft a report) and want the method. Go straight to the task guides.The guides
- I want to adopt this responsibly.You are curious but cautious, or someone on your team is. Start with readiness, governance, quality assurance, and data protection.The adopt-responsibly guides
How to read this map
AI is not one decision. It is a different decision at every stage of your work, and the honest answer is “it depends on the stage.” Two ideas make the rest of this page easier to read.
AI-native vs AI-skinned
There is a real difference between applying AI inside the work (at data collection, at analysis, where it changes what is possible) and bolting an AI summary onto the end of a process that has not changed at all. The first can shift what your team can do. The second is mostly a coat of paint. When you evaluate any AI use, ask which one it is.
The augmentation gradient
“Does the AI do it, or do I?” is the wrong question. Real uses sit on a gradient:
- Assist. AI drafts or suggests; you do the work and decide everything.
- Automate with review. AI does a first pass on every item; a human reviews every item before it counts.
- Automate with spot-check. AI runs at scale; a human validates a sample and the items the system flags as uncertain.
The right point depends on the stakes and on how well you can check the output. Nothing in M&E that informs a decision should run with no human in the loop. More on that in the quality-assurance guide.
Where AI fits across the M&E lifecycle
Suitability is how much AI can safely take on at that stage. Every level is named in words, not just color.
| Stage | AI suitability | What AI does well | Where a human is mandatory | Guides |
|---|---|---|---|---|
| DesignTheory of change, indicators, evaluation questions, instruments | Medium to high | Draft items, menus, and variants; check coverage gaps against a framework you supply. | Context and political sensitivity, and the soundness of the underlying theory. |
|
| Data collectionFieldwork, instruments, consent | Low to medium | Translation, transcription, formatting instruments. | Consent, sampling integrity, field ethics. AI does not make these decisions. | |
| Data cleaningCleaning and harmonization | High | De-duplication, categorization, record linkage, normalizing inconsistent entries. | Silent error propagation. Every transformation needs validation before you trust it. | |
| Quantitative analysisPatterns, summaries, analysis code | Medium | Pattern detection, descriptive summaries, drafting analysis code. | Causal inference and interpreting significance stay with you. | |
| Qualitative codingThematic analysis, open-ended responses | High, with a human in the loop | Coding open-ends against a codebook at scale, surfacing candidate themes. | Codebook validity and intercoder reliability. This is both the most-cited win and the most-cited disappointment. | |
| Evidence synthesisLiterature review, text-mining across reports | High | Text-mining across hundreds of reports, surfacing and highlighting relevant passages. | Bias toward written, English-language evidence. What is not in the documents stays invisible. | No how-to guide yet. Background: Evidence synthesis, Systematic review. |
| ReportingDrafting, gap-checks, visualization | High | Drafting from structured evidence, gap-checks, first-pass visualization. | Strategic recommendations, framing, and accountability are yours. Always disclose and cite. | |
| LearningAdaptive management | Medium | Near-real-time insight in place of periodic reporting. | Insight without sense-making is not learning. A faster dashboard is not a decision. | No how-to guide yet. Background: Adaptive management, Learning agenda. |
Design
Theory of change, indicators, evaluation questions, instruments
AI suitability: Medium to high
Does well: Draft items, menus, and variants; check coverage gaps against a framework you supply.
Human mandatory: Context and political sensitivity, and the soundness of the underlying theory.
- How to Build a Theory of Change with AI
- How to Use AI to Design an M&E Framework
- How to Use AI for Indicator Development
- How to Write a Logframe with AI
- How to Use AI to Write a MEL Plan
- How to Build Better Surveys with AI
- Build a Theory of Change with AI (playbook)
- Build a Results Framework with AI (playbook)
- Develop Indicators with AI (playbook)
- Build a MEL Plan with AI (playbook)
- Build an Evaluation Plan with AI (playbook)
Data collection
Fieldwork, instruments, consent
AI suitability: Low to medium
Does well: Translation, transcription, formatting instruments.
Human mandatory: Consent, sampling integrity, field ethics. AI does not make these decisions.
Data cleaning
Cleaning and harmonization
AI suitability: High
Does well: De-duplication, categorization, record linkage, normalizing inconsistent entries.
Human mandatory: Silent error propagation. Every transformation needs validation before you trust it.
Quantitative analysis
Patterns, summaries, analysis code
AI suitability: Medium
Does well: Pattern detection, descriptive summaries, drafting analysis code.
Human mandatory: Causal inference and interpreting significance stay with you.
Qualitative coding
Thematic analysis, open-ended responses
AI suitability: High, with a human in the loop
Does well: Coding open-ends against a codebook at scale, surfacing candidate themes.
Human mandatory: Codebook validity and intercoder reliability. This is both the most-cited win and the most-cited disappointment.
Evidence synthesis
Literature review, text-mining across reports
AI suitability: High
Does well: Text-mining across hundreds of reports, surfacing and highlighting relevant passages.
Human mandatory: Bias toward written, English-language evidence. What is not in the documents stays invisible.
No how-to guide yet. Background: Evidence synthesis, Systematic review.
Reporting
Drafting, gap-checks, visualization
AI suitability: High
Does well: Drafting from structured evidence, gap-checks, first-pass visualization.
Human mandatory: Strategic recommendations, framing, and accountability are yours. Always disclose and cite.
Learning
Adaptive management
AI suitability: Medium
Does well: Near-real-time insight in place of periodic reporting.
Human mandatory: Insight without sense-making is not learning. A faster dashboard is not a decision.
No how-to guide yet. Background: Adaptive management, Learning agenda.
Beyond M&E: AI across the wider program
The same questions show up outside the M&E function, in service delivery and beneficiary-facing tools, and the answers get more cautious because real people are on the other end.
Service delivery and beneficiary-facing tools
Chatbots, triage, translation in the field
AI suitability: Use with care
Does well: Translation, drafting information materials, routing and triage support behind a human.
Human mandatory: Anything that gives advice, makes an eligibility call, or touches a vulnerable person directly. Consent, data protection, and a clear human fallback are not optional.
In-project analysis and operational reporting
Operational data, internal updates
AI suitability: Medium
Does well: Summarizing operational data, drafting internal updates, flagging anomalies for review.
Human mandatory: The same silent-error and oversight cautions as the M&E stages. Keep a person between the output and any decision.
The point of widening out is not to chase every shiny use. It is to see the real approaches teams are taking across their programs and bring it back to the question that anchors this whole pillar: what does this mean for how you measure, evaluate, and learn?
Go further
- The nine playbooksMulti-step prompt workflows that chain small, checked steps into a finished deliverable: a theory of change, a MEL plan, a donor report.
- The toolkitFilter the guides and playbooks by what you are trying to do, the stage you are at, and how sensitive your data is.
- The prompt libraryTested prompts for the M&E tasks you already do, by topic.
- Language for disclosing AI useCopyable statements for declaring AI use and describing a responsible, human-led approach. Written for proposals; most travel well.
Do a specific task
If you already know the task, the 21 how-to guides are the method, grouped the same way the map is.
Foundations
Design
Analyze
Report
The teams getting the most out of AI are not the ones moving fastest. They are the ones who know where it fits, where it does not, and what their own checks are. That is the point of this map: not to push you to adopt more, but to help you adopt the right things, in the right places, with your eyes open.