Where AI fits across your program

A practical map of where AI helps in development work, where it does not, and how to adopt it without getting burned. M&E is the anchor and the lens: we start from the function you already run, then widen out to the rest of the program.

This is a peer's view, not a principles lecture and not a sales page: what AI does well at each stage, where a human still has to hold the pen, and what to do when an output looks confident and is wrong. Everything here is free.

How to read this map

AI is not one decision. It is a different decision at every stage of your work, and the honest answer is “it depends on the stage.” Two ideas make the rest of this page easier to read.

AI-native vs AI-skinned

There is a real difference between applying AI inside the work (at data collection, at analysis, where it changes what is possible) and bolting an AI summary onto the end of a process that has not changed at all. The first can shift what your team can do. The second is mostly a coat of paint. When you evaluate any AI use, ask which one it is.

The augmentation gradient

“Does the AI do it, or do I?” is the wrong question. Real uses sit on a gradient:

  • Assist. AI drafts or suggests; you do the work and decide everything.
  • Automate with review. AI does a first pass on every item; a human reviews every item before it counts.
  • Automate with spot-check. AI runs at scale; a human validates a sample and the items the system flags as uncertain.

The right point depends on the stakes and on how well you can check the output. Nothing in M&E that informs a decision should run with no human in the loop. More on that in the quality-assurance guide.

Where AI fits across the M&E lifecycle

Suitability is how much AI can safely take on at that stage. Every level is named in words, not just color.

  1. Design

    Theory of change, indicators, evaluation questions, instruments

    AI suitability: Medium to high

    Does well: Draft items, menus, and variants; check coverage gaps against a framework you supply.

    Human mandatory: Context and political sensitivity, and the soundness of the underlying theory.

  2. Data collection

    Fieldwork, instruments, consent

    AI suitability: Low to medium

    Does well: Translation, transcription, formatting instruments.

    Human mandatory: Consent, sampling integrity, field ethics. AI does not make these decisions.

  3. Data cleaning

    Cleaning and harmonization

    AI suitability: High

    Does well: De-duplication, categorization, record linkage, normalizing inconsistent entries.

    Human mandatory: Silent error propagation. Every transformation needs validation before you trust it.

  4. Quantitative analysis

    Patterns, summaries, analysis code

    AI suitability: Medium

    Does well: Pattern detection, descriptive summaries, drafting analysis code.

    Human mandatory: Causal inference and interpreting significance stay with you.

  5. Qualitative coding

    Thematic analysis, open-ended responses

    AI suitability: High, with a human in the loop

    Does well: Coding open-ends against a codebook at scale, surfacing candidate themes.

    Human mandatory: Codebook validity and intercoder reliability. This is both the most-cited win and the most-cited disappointment.

  6. Evidence synthesis

    Literature review, text-mining across reports

    AI suitability: High

    Does well: Text-mining across hundreds of reports, surfacing and highlighting relevant passages.

    Human mandatory: Bias toward written, English-language evidence. What is not in the documents stays invisible.

    No how-to guide yet. Background: Evidence synthesis, Systematic review.

  7. Reporting

    Drafting, gap-checks, visualization

    AI suitability: High

    Does well: Drafting from structured evidence, gap-checks, first-pass visualization.

    Human mandatory: Strategic recommendations, framing, and accountability are yours. Always disclose and cite.

  8. Learning

    Adaptive management

    AI suitability: Medium

    Does well: Near-real-time insight in place of periodic reporting.

    Human mandatory: Insight without sense-making is not learning. A faster dashboard is not a decision.

    No how-to guide yet. Background: Adaptive management, Learning agenda.

Beyond M&E: AI across the wider program

The same questions show up outside the M&E function, in service delivery and beneficiary-facing tools, and the answers get more cautious because real people are on the other end.

Service delivery and beneficiary-facing tools

Chatbots, triage, translation in the field

AI suitability: Use with care

Does well: Translation, drafting information materials, routing and triage support behind a human.

Human mandatory: Anything that gives advice, makes an eligibility call, or touches a vulnerable person directly. Consent, data protection, and a clear human fallback are not optional.

In-project analysis and operational reporting

Operational data, internal updates

AI suitability: Medium

Does well: Summarizing operational data, drafting internal updates, flagging anomalies for review.

Human mandatory: The same silent-error and oversight cautions as the M&E stages. Keep a person between the output and any decision.

The point of widening out is not to chase every shiny use. It is to see the real approaches teams are taking across their programs and bring it back to the question that anchors this whole pillar: what does this mean for how you measure, evaluate, and learn?

The teams getting the most out of AI are not the ones moving fastest. They are the ones who know where it fits, where it does not, and what their own checks are. That is the point of this map: not to push you to adopt more, but to help you adopt the right things, in the right places, with your eyes open.