How to bring AI in, and the safeguards

A practical, low-drama way to bring AI into your team's work, and the things to weigh before you use it on anything that matters. This is not a transformation program. It is a sensible way to start small, learn what AI is good for, and grow into it without getting burned.

The teams that get into trouble with AI are rarely the ones who were too cautious. They are the ones who used it confidently in the wrong place, or in the right place with no check. So this page has two halves. The first is how to start, in a way that is small enough to be safe and real enough to be useful. The second is the set of things to weigh before AI touches anything sensitive, important, or public. Read both before you scale anything.

How to bring AI in

There is a natural path most teams move along, from a couple of curious people trying things, to AI being a normal, governed part of how the team works. You cannot skip to the end, and you should not try. The point is to know roughly where you are now and what the next concrete step is. There is nothing to score and nothing to pass.

  • Further along is not automatically better. The right place to be depends on your data, your team, and what you are accountable for. A small team using AI carefully on a couple of low-risk tasks is in good shape. The trap is being further along than your checks are.
  • You will be in different places for different tasks. It is normal to be quite advanced on drafting and barely started on anything sensitive. Locate yourself by where most of the work sits.
  • Growing means adding checks, not removing them. Every step toward more use is also a step toward more validation, never less. That is the whole discipline.
  1. Start small, on something low-risk

    Pick one task where AI is genuinely good (drafting a first version of something, summarizing a long document, cleaning up messy notes) and where a mistake is cheap and easy to catch. Avoid anything sensitive, public, or high-stakes for now. Run that one task deliberately, and have a person check every output before it goes anywhere. The goal here is not productivity. It is learning where this thing helps and where it slips, on work where getting it wrong costs you nothing.

  2. Get one thing genuinely working

    After a while, you will find AI is reliably useful for at least one specific job, and you will know how to get a good result there. That is real progress, even if it lives in one person's head and there is no written rule yet. The signal is concrete: something real went out the door with AI in its production, and a person checked it first, and you could explain exactly how AI helped and where you corrected it.

    The risk at this point is over-trusting the one thing that works, stretching it to tasks it is not suited for, or scaling it before anyone has written down how it is done. The next move is small: write down the method, including the checking step, so a second person can run it the same way. Give the AI your own materials to work from (your template, your data, your framework) so its output is anchored to your reality rather than invented. That single habit, grounding it in your own inputs, prevents most of what goes wrong.

  3. Weave it into a workflow, with checks at each step

    The next stage up is when AI is part of how a whole piece of work runs, not just one task, and the work is broken into small, checked steps. More than one person works this way, and there are written norms: which tools are allowed, what data must never go in, who checks the output, and how AI use is disclosed. The work is divided into "AI does this, a person checks that" at each point, rather than one big handoff.

    The risk here is subtle and worth naming. A workflow can run smoothly and produce a lot of output while the actual human thinking quietly thins out. Fast output gets mistaken for understanding. The move is to shift your attention from doing the work to checking the system: spot-check samples, look hard at whatever the workflow flags as uncertain, and re-check your own assumptions on a schedule rather than once at the start.

  4. Make it a normal, governed part of the work

    The far end of the path is when AI is a routine, governed part of how the team operates, across more than one workflow, with a clear owner, a data policy, a disclosure standard, and a regular review of where it is helping and where it is not. New people are brought into the practice rather than reinventing it. The team's energy goes into judgment and interpretation, and into the cases the system flags, rather than into the mechanical work.

    The risk at this point is complacency: governance written once and never revisited, errors that are harder to spot precisely because the system is trusted and large. So the discipline at this stage is to treat governance as a living thing. Schedule a periodic look at what is working, retire the uses that have stopped earning their place, and pair the fact that you use AI widely with visible evidence of how you check it. This is maintenance, not a finish line.

The honest summary of the whole path: start on something where a mistake is cheap, keep a person responsible at every step, learn what AI is actually good at on your work, write down what works so it is not trapped in one head, and add checks as you add use. That is the entire method. Everything else is detail.

The safeguards

Before AI touches anything sensitive, important, or public, there are a handful of things to weigh. None of these is a reason not to use AI. Together they are what makes the times you do use it trustworthy.

Data sensitivity and privacy

This is the one to get right first, because the harm is real and hard to undo. Development work routinely handles data about people whose safety can depend on it staying private: program participants, vulnerable groups, anyone covered by a consent agreement.

When you put text into a typical cloud AI tool, that text travels to a company's servers to be processed, and depending on the tool's terms and tier, it may be retained or used to improve the service. Pasting a spreadsheet of names, a set of case notes, or a partner's confidential document into a free chat tool means that information now sits in someone else's system, possibly under terms your consent forms never anticipated.

The practical rules are simple. Remove identifying details before any AI step. Never put sensitive raw data into an ungoverned tool. Make sure any consent you collected actually covers this kind of re-use, and if it does not, do not assume it does.

For genuinely sensitive work, there is a strong option: a local model, an AI model that runs on a machine you control (a laptop, a workstation, an office server) instead of a cloud service. The data never leaves the device. Local models have caught up enough that smaller ones now run on ordinary office hardware and handle a lot of routine text work well, though they are generally weaker than the largest cloud models on the hardest reasoning. They also work offline, which matters for field settings with no reliable connection. You do not have to go local for everything. Most teams use cloud tools for non-sensitive work and keep a local option for the sensitive and offline jobs where it earns its keep.

Accuracy and made-up information

AI produces confident, fluent answers, and some of them are wrong. The dangerous ones are the plausible ones: a fabricated statistic, an invented citation, a subtly mis-stated fact, all reading exactly like the real thing. Research on AI-generated citations has found large shares of them fabricated or erroneous. Fluent and well-formatted is not the same as correct, and the polish actively hides the errors.

The guardrail is to ground and verify. Do not ask AI for facts out of thin air; give it your own source material and ask it to work from that, then check what it produces against the source. A useful instruction to add to your prompts: tell it not to invent statistics, sources, or specifics, and to say "no support found" rather than guess. Treat anything factual it gives you as a draft claim to verify, never as a finished fact.

Consumer tools versus enterprise tools

Most people doing quiet, individual AI use are on the wrong tool tier for sensitive work, and most do not realize there is a distinction. As a general pattern: free and basic consumer tiers may use what you type to improve the underlying model, while enterprise, team, and paid API tiers typically do not, and add data controls and agreements. The exact terms vary by product and change over time, so check the tool you actually use rather than assuming.

The practical upshot: a free consumer tool is fine for non-sensitive drafting and brainstorming. It is the wrong place for anything confidential, personal, or competitive. Before your team standardizes on a tool, someone should read its terms with the question "what happens to what we type in," and you should match the tier to the sensitivity of the work.

Cost: a shift, not just a saving

AI genuinely saves time on a specific set of tasks: drafting, cleaning up records, summarizing and searching long documents, translation. Those savings are real and show up quickly. But it adds cost on a different set, and the optimistic pitch usually leaves that half out.

The biggest hidden cost is verification: everything AI produces has to be checked, and on work you have to defend, that checking is real labor that has to be staffed. There is also the cost of the errors that slip through, the price of capable tools across a team, and the time it takes people to learn what AI is good at and how to use it well, during which the team is often slower before it is faster.

The honest way to think about it: AI tends to move effort rather than remove it. The reading you save becomes reviewing the AI's reading. The drafting you save becomes editing and verification. The volume it makes easy to produce becomes volume someone has to process. This is not an argument against AI. It is an argument against promising a "cut" you cannot deliver. The defensible claim is that AI changes the shape of the work and frees skilled people to spend their time on judgment instead of mechanics, not that it will halve anyone's workload. If you are making the case to leadership, name the specific task, count both sides of the ledger, and budget the checking time and the learning curve from the start.

Governance and disclosure

Two simple practices cover most of what "governance" needs to mean for a normal team. The first is to agree internally on the basics: which tools are allowed, what data must never go into them, and who is responsible for checking AI-assisted work before it goes out. You do not need a long policy document to start; you need a shared, written answer to those three questions.

The second is disclosure. Being open about where you used AI is the right default, and it can also make a skeptical reader discount the work if it arrives bare. The answer is not to hide it; it is to disclose with rigor. Pair any statement that you used AI with what you did to check it: the source material you grounded it in, the part a person reviewed, the finding you confirmed another way. Disclosure plus visible rigor builds trust. Disclosure on its own can spend it. Where an external body (a funder, a client, a partner) has its own rules about AI use and disclosure, follow theirs; the landscape is still forming and varies a lot, so when in doubt, a sensible default is to keep a record of which parts used AI, which tool and version, and the prompts.

When not to use AI

Knowing when to keep AI out of a task is not caution for its own sake. It is what makes the times you do use it trustworthy. Do not reach for AI when:

  • The data is sensitive or personal and you cannot keep it controlled. If you cannot de-identify it, and you do not have a tool or a local model that keeps it private, do not put it in. The convenience is not worth the breach.
  • You need a fact, figure, or source you cannot independently verify. If there is no way to check what it tells you against a real source, treat its answer as unusable.
  • The specific, local, or marginal context is the whole point. AI flattens toward the average and quietly underweights the small voice, the local nuance, the group that is a minority in the data. Where catching exactly that is the job, a person has to do it. At minimum, ask the AI directly who or what might be under-represented before you trust its summary.
  • Real judgment, strategy, or accountability is at stake. Deciding what matters, what is fair, what to commit to, and what to put your name on stays with a person. AI can inform those calls; it cannot make them or own them.
  • No one is going to check the output. If a result will go straight into a decision or a deliverable with no human review, do not use AI to produce it. The check is not optional; it is the thing that makes the use safe.
  • The work runs at a volume nobody is sampling. Once AI is processing far more than a person reviews, a small consistent error can scale undetected. If you cannot spot-check a meaningful sample on a schedule, you are running unmonitored, and that is when silent errors do their damage.

The throughline across all of it is the same handful of moves: ground AI in your own materials rather than asking it to invent; keep a person in the loop before any decision; check at least one important result the way you would have without AI; protect sensitive data and consent; and be open about your AI use in a way that shows the rigor behind it. Get those right and you can use AI confidently. Skip them and the tool that was supposed to help becomes the thing that burns you.

Back to what AI can do for your work