Most AI agent pilots fail before they begin. Not because the technology was wrong, but because nobody wrote down what success looked like. This checklist covers every definition a business leader needs to establish before an agent touches a live workflow.
Definitions
The following terms are used throughout this checklist. Agreeing on their meaning with your vendor before a pilot starts prevents most end-of-pilot disputes.
A pilot is a time-bounded, scoped deployment of an AI agent against a specific workflow, with pre-defined success and failure criteria.
A workflow brief is a written document describing the specific process the agent will run, its trigger, steps, systems, outputs, and volume.
A success threshold is the minimum measured performance level at which you agree to expand the deployment. It is a number, not a description.
A failure threshold is a defined condition that triggers an immediate pause of the pilot, regardless of overall performance.
A hard stop is a specific event that requires the pilot to pause immediately. Hard stops are defined before the pilot begins, not during it.
An escalation is any workflow instance the agent routes to a human reviewer rather than completing autonomously. Escalation rate is a performance metric, not a failure indicator. What constitutes an acceptable rate depends on the workflow.
An edge case is any input that falls outside the standard scenario a vendor demo was built to handle. Real operations consist largely of edge cases.
An exit outcome is one of three defined decisions at the end of a pilot: expand, extend, or exit. Exit outcomes are defined before the pilot starts.
Why define this upfront
Pilots without pre-defined criteria produce ambiguity. The vendor reads results optimistically. The business reads them cautiously. Both interpretations are defensible because success was never specified.
Pre-definition creates three concrete benefits:
Anchors the evaluation. Vendors frequently expand a pilot's claimed value as it progresses. A written success brief limits what counts.
Creates a decision, not a negotiation. When the pilot ends, results either met the threshold or they did not. The outcome is a finding, not a conversation.
Protects against silent failure. Without defined thresholds, poor performance accumulates unnoticed until it becomes expensive to reverse.
The pre-pilot checklist
Work through each section in order. Every item should be written and shared with the vendor before the agent runs on any live workflow.
Step 1: Write the workflow brief
A workflow brief is the foundation of the entire pilot. Without one, the vendor will test their workflow, not yours.
Component | What to specify |
Trigger | The event that starts the workflow — inbound message, form submission, calendar event, data update |
Steps | The sequence of actions from trigger to completion, in order |
Systems touched | Every tool, database, or platform the agent reads from or writes to |
Output | The specific end state that constitutes a completed workflow |
Volume | How many instances of this workflow run per day and per week |
Exceptions | Known scenarios where the standard steps do not apply |
Owner | The person on your team responsible for reviewing agent outputs during the pilot |
A workflow brief should be one to two pages. If it cannot be written in that length, the workflow is not yet scoped tightly enough to pilot.
Step 2: Define the success threshold
The success threshold is a single sentence with three variables. Fill it in before the pilot starts.
The pilot succeeds if accuracy exceeds [X]%, escalation rate stays below [Y]%, and average completion time is under [Z] minutes, measured over [N] live workflow runs. This is important to be ruthlessly clear in your checklist.
Guidance on setting each variable:
Variable | How to set it | Common mistake |
Accuracy [X] | Match the accuracy standard you hold your current human process to | Setting it higher than your current process, which creates an unfair bar |
Escalation [Y] | Based on team capacity to handle escalated reviews during the pilot period | Setting it at zero, which penalises appropriate caution |
Completion time [Z] | Based on the time requirement your customers or operations actually impose | Using average demo speed rather than your operational requirement |
Sample size [N] | Minimum 200 instances for most business workflows | Ending the pilot before enough volume has run to produce reliable results |
Step 3: Define the failure threshold
Failure thresholds are separate from success thresholds. They are conditions that trigger an immediate pause, not performance metrics to trend over time.
Hard stops. Define each of the following before the pilot begins:
The agent produces a customer-facing error that cannot be reversed
The agent accesses data or systems outside its documented scope
The agent produces a compliance or policy violation
Accuracy drops more than 50%
in a single week with no explanation from the vendor
The vendor makes a change to the agent without advance notice
Share the hard-stop list with the vendor in writing before the pilot starts. A vendor who objects to any item on this list is providing useful information about their confidence in the product.
Step 4: Build the test dataset
Before the agent runs on live inputs, assemble a test dataset with three categories:
Category | Volume | Description |
Standard inputs | 20 to 30 | Typical workflow instances — the cases that make up the majority of daily volume |
Edge cases | 5 to 10 | Unusual but real inputs your team encounters regularly that a vendor demo never includes |
Known-bad inputs | 3 to 5 | Inputs the agent should escalate or reject rather than process |
Run the agent against the full test dataset in a non-live environment before it touches real operations. Any failure on a known-bad input at this stage is a finding to address before the live pilot begins, not a surprise to manage after.
Step 5: Define the review process
Every pilot needs a structured review cadence. Without one, issues accumulate silently.
Review type | Frequency | Owner | Purpose |
Output sample review | Daily, first two weeks | Operations or team lead | Identify systematic errors before they scale |
Edge case audit | Weekly | Department head | Confirm unusual inputs are handled correctly |
Vendor sync | Weekly | Business lead and vendor CSM | Surface issues before they compound |
Performance snapshot | Weekly | Both parties | Track accuracy, escalation rate, and completion time against threshold |
End-of-pilot report | Final day of pilot | Both parties | Compare measured results against the pre-defined success threshold |
The review cadence does not need to be heavy. A fifteen-minute daily sample check in week one catches most systematic issues before they become significant.
Step 6: Define the exit outcomes
At the end of the pilot, one of three outcomes applies. Define what each one means before the pilot starts.
Expand: Results met or exceeded the success threshold across the defined sample size. The agent moves to full deployment with a defined monitoring process in place.
Extend: Results were mixed. Specific underperforming areas are documented. The vendor has a written remediation plan. A second pilot phase is scoped with a revised success threshold and a shorter timeline.
Exit: Results fell below the success threshold, or a hard stop was triggered. The deployment ends. A post-mortem documents what failed and why before any alternative vendor is evaluated.
Pre-defined exit outcomes convert the end of a pilot from a negotiation into a decision.
Pilot timeline
Phase | Duration | Primary activity |
Pre-pilot setup | 1 week | Workflow brief, test dataset, success threshold, failure threshold, and exit outcomes — all documented and shared with vendor |
Non-live validation | 2 to 3 days | Agent runs against test dataset; findings addressed before go-live |
Live pilot | 3 to 4 weeks | Agent runs on real or parallel workflows with daily and weekly review cadence |
Assessment | 3 to 5 days | Results measured against pre-defined success threshold |
Decision | 1 week | Expand, extend, or exit based on written criteria |
A pilot shorter than three weeks does not generate enough volume to produce reliable results on most business workflows. A pilot longer than six weeks without a structured review process becomes an informal continuation rather than an active evaluation.
Common failure patterns
The following are the most frequent causes of inconclusive or failed pilots. Each is preventable at the pre-pilot stage.
Failure pattern | Root cause | Prevention |
"It mostly worked" | Success threshold was never defined | Complete Step 2 before go-live |
Pilot drifts past deadline | Exit outcomes were not specified | Complete Step 6 before go-live |
Vendor disputes the results | Metrics were not agreed on in writing | Share the full checklist with the vendor before the pilot starts |
Hard stop triggered with no plan | Failure threshold was not defined | Complete Step 3 before go-live |
Edge cases discovered during live pilot | Test dataset was not built | Complete Step 4 in non-live environment |
Issues found too late to fix | No review cadence in place | Complete Step 5 before go-live |
Sample size too small to conclude | Volume threshold was not defined | Specify [N] in the success threshold before the pilot starts |
What Signal does with this
When a business submits a workflow to Signal, this checklist is the starting point.
Signal works through the workflow brief with you, defines realistic success thresholds based on evidence from comparable deployments, builds the test dataset including edge cases and known-bad inputs, and runs the agent evaluation before you commit to a live pilot.
The trust profile Signal produces at the end maps directly to the criteria above, accuracy, escalation rate, failure modes, and edge case handling documented against your specific workflow.
If you are preparing for a pilot and want a structured evaluation before you commit, submit your workflow here.