01Name the outcome
02Bound the first workflow
03Review before expanding
Explanatory diagram. A conceptual reading aid, not benchmark or ROI data. See sources checked for the factual guidance used in this article.

A first AI pilot should answer a business question: is this particular workflow useful enough to keep? This guide takes you from a shortlist of tasks to a written pilot decision. Use it after you understand the basics of AI setup. You will finish with a scope, a test plan and an owner, rather than a list of software to buy.

1. Compare three real tasks

Ask the people doing the work to name three repeated tasks. For each, write down the input, finished output, frequency, exceptions and current reviewer. Use observed work rather than an optimistic estimate of what could be automated.

Give each task a plain-language assessment:

  • Clear input: can you recognize when the task begins?
  • Checkable output: can a person judge whether the result is right?
  • Manageable consequence: can an error be caught before it affects someone?
  • Available information: are the necessary sources current and accessible?
  • Named owner: will someone actually operate and improve it?

Choose the task with the fewest unresolved questions. A less frequent task with clear rules may make a better first pilot than a high-volume process full of exceptions. Leave consequential financial, employment, medical and legal decisions out of a beginner project.

2. Make a one-page workflow agreement

Write the trigger, permitted input, expected output and reviewer in separate fields. Add an explicit “outside scope” line. Specify whether the workflow creates a draft, changes an internal record or contacts someone. Those are different permissions.

Hypothetical example: a fictional office-services company, Cedar Desk Demo, wants help preparing inquiry summaries. The input is a copied fictional inquiry. The output is a short summary, missing-detail list and suggested reply. A coordinator reviews all three. Sending messages, making appointments and quoting prices are outside scope. The example uses no live mailbox or customer account.

The agreement should also name the source of truth. For this pilot, that might be one approved service document with an owner and revision date. Conflicting documents should trigger a question rather than an invented answer.

3. Map access before choosing a connector

Draw the route information will take: source, workflow tool, model provider, saved output and reviewer. Write the minimum data needed at each step. A task that summarizes an inquiry does not automatically need access to every contact, attachment or historical conversation.

Separate reading from writing. Ask which provider permissions are actually granted, where credentials are held and how a connection is removed. A read-only screen can still sit on top of credentials that permit changes. Check the underlying grant. Use fictional examples while testing the design; introducing real customer data requires a deliberate decision about access, provider handling and retention.

4. Build a small acceptance pack

Prepare test cases before changing the workflow. For each, keep the input, expected behavior, actual result and reviewer notes together. Include these cases:

  • A normal inquiry with all required details
  • An inquiry missing the location or requested service
  • A request for something the business does not offer
  • Two sources that disagree about a fact
  • A duplicate input
  • An input containing instructions to ignore the workflow’s rules
  • An unavailable source or failed connection

The last cases reveal whether the workflow fails visibly and safely. OpenAI’s agent-safety guidance discusses structured outputs, human approval and evaluations as risk-reduction measures. Your acceptance pack applies those ideas to your own task; it is not a security certification. Official guidance.

5. Observe the whole job

For the original process, record how long a person spends preparing, checking and finishing the task. For the pilot, count setup, review, corrections and recovery as well as generation time. Keep any estimates labeled until you have actual observations.

Record why drafts were rejected. Repeated factual mistakes suggest a source or retrieval problem. Repeated tone edits may call for better examples. An output that looks polished but misses essential details is still a failed case. Do not remove difficult examples from the test set just to improve the score.

6. Make an explicit pilot decision

Before the pilot begins, agree on what would justify keeping it, what would require revision and what would make you stop. Set criteria appropriate to the task rather than borrowing an unsupported industry benchmark. A critical permission failure should be addressed regardless of how many ordinary drafts looked good.

NIST’s voluntary AI Risk Management Framework supports considering risk throughout design, use and evaluation. It is a useful reference for organizing this conversation, not proof that a small pilot complies with a certification. NIST framework.

Your closing checklist is short: name the owner, save the tested version, document how to pause it, confirm operating costs and schedule a review when the workflow or source material changes. If you cannot explain how to stop the system or who handles a failure, keep the pilot bounded.

Bring your one-page agreement and a few fictional test cases to InstallAI when you discuss a setup. They make it easier to scope useful work and identify the access you actually need.

Sources checked