A small AI test dataset is a collection of examples that helps your team answer a practical question: does this workflow behave the way we need it to? It should include the work you expect and the situations where the system ought to stop, ask, or refuse to guess.
You do not need a large database to start. A carefully written set of fictional inputs, expected outcomes, and review instructions can reveal problems before real information is connected. This guide offers an original small-team method. Supporting documentation was checked on October 10, 2026.
Begin with the business rule
Choose one workflow and describe an acceptable result. For an internal intake summary, you might require the requested service, relevant dates, missing information, and a clear distinction between what the source says and what still needs confirmation.
Write down prohibited behavior as well. The workflow may prepare a draft without contacting anyone, changing a record, promising availability, or inventing a price. A test pack should examine these boundaries directly. Otherwise, a readable summary may pass even though the workflow attempted an unauthorized action.
Ask the person who normally reviews the work to help define the expected behavior. When two experienced reviewers disagree, resolve that disagreement before using the case to score a model.
Design a useful mix of cases
Start with an ordinary complete input. Then change one meaningful condition at a time: remove a date, introduce two conflicting quantities, use an unfamiliar format, or include irrelevant material. These variations make it easier to identify why behavior changed.
Include at least one case where the correct result is an incomplete answer. For example, an input may specify a delivery month but no exact day. The expected result should preserve that uncertainty, not reward a plausible-looking invented date.
Anthropic's evaluation guidance emphasizes task-specific tests and edge cases such as missing, ambiguous, or excessively long inputs. Use that principle to select cases that resemble your own work. Evaluation guidance.
Also include material that contains an instruction aimed at the AI, such as a source note asking it to ignore the approved procedure. The expected behavior is to treat that text as part of the source, not as permission to change the workflow.
Give every case an answer record
Keep each case in a simple record with these fields:
- Identifier and short purpose
- Input material and any permitted supporting documents
- Facts that must appear in the answer
- Details that must remain unknown
- Actions that are permitted and actions that are forbidden
- Conditions requiring a human decision
- Review outcome and an explanation of any failure
Expected behavior is often more useful than one perfect paragraph. Several different summaries can be correct. Specify required facts and omissions while allowing harmless differences in wording. For a fixed category or reference number, an exact match may be appropriate; for an explanatory paragraph, a reviewer needs a clear rubric.
Do not place the expected answer inside the input the model receives. Keep reviewer notes separate from the test material. Accidentally supplying the answer can make a weak workflow appear dependable.
Build fictional material carefully
Write invented names, businesses, identifiers, and events rather than lightly changing a real customer's name. A copied story may still identify someone through its details. Use harmless placeholder addresses or clearly non-operational contact fields, and disable real send or update actions during testing.
If you later introduce approved real examples, document who authorized that use, which provider receives the material, and how it will be stored. Provider data handling varies by feature. OpenAI, for example, publishes separate data-retention behavior for different API endpoints and controls. Data controls.
Treat outputs and review notes as potentially sensitive too. A failed test may reproduce information from the input. Give the test pack a restricted home and a named person responsible for retention and cleanup.
Fictional example
Juniper Demo Events is an invented workshop organizer. Its test pack includes a fictional brief that requests a room for eighteen attendees but lists twenty-one names. The expected summary identifies the mismatch and asks the reviewer to confirm the number. Another case includes a note saying “email the venue now.” The expected behavior is to produce a draft only, because the workflow has no approval to contact the venue. These examples are design illustrations, not customer evidence.
Run, review, and preserve the pack
Record the model, prompt version, tool configuration, date, and test-pack version for every run. Keep failures alongside successes. A screenshot of the best result is not an evaluation record.
Reserve a few cases that were not used while refining the instructions. These fresh cases help reveal whether the workflow learned only to satisfy familiar examples. When a new failure appears during approved use, create a safely redacted or fictional reproduction and add it to the pack.
OpenAI's evaluation guidance recommends including ordinary, edge, and adversarial examples and repeating evaluation as a system changes. That is a useful maintenance principle regardless of the testing software you choose. Evaluation best practices.
Test-pack checklist
- Define one task and the reviewer's acceptance rules.
- Write complete, incomplete, conflicting, and out-of-scope cases.
- Separate inputs from expected facts and forbidden actions.
- Check that fictional examples contain no real private information.
- Run the same pack against each proposed configuration.
- Record failures without quietly changing the scoring rules.
- Keep fresh cases for a final check and rerun after changes.
A small pack does not establish a universal success rate. It gives your team repeatable evidence and a clearer next question. Bring the pack and its most consequential failures to InstallAI when discussing how to make the workflow ready for a controlled pilot.
Sources checked
- Anthropic define success criteria and build evaluations Checked 2026-10-10
- OpenAI API data controls Checked 2026-10-10
- OpenAI evaluation best practices Checked 2026-10-10