PROJECT_ID: SIGNAL_NOT_NOISE

[How-to Guide] What to Define Before You Start Your AI Journey

May 15, 2026

Article

Most AI agent pilots fail before they begin. Not because the technology was wrong, but because nobody wrote down what success looked like. This checklist covers every definition a business leader needs to establish before an agent touches a live workflow.


Definitions

The following terms are used throughout this checklist. Agreeing on their meaning with your vendor before a pilot starts prevents most end-of-pilot disputes.

A pilot is a time-bounded, scoped deployment of an AI agent against a specific workflow, with pre-defined success and failure criteria.

A workflow brief is a written document describing the specific process the agent will run, its trigger, steps, systems, outputs, and volume.

A success threshold is the minimum measured performance level at which you agree to expand the deployment. It is a number, not a description.

A failure threshold is a defined condition that triggers an immediate pause of the pilot, regardless of overall performance.

A hard stop is a specific event that requires the pilot to pause immediately. Hard stops are defined before the pilot begins, not during it.

An escalation is any workflow instance the agent routes to a human reviewer rather than completing autonomously. Escalation rate is a performance metric, not a failure indicator. What constitutes an acceptable rate depends on the workflow.

An edge case is any input that falls outside the standard scenario a vendor demo was built to handle. Real operations consist largely of edge cases.

An exit outcome is one of three defined decisions at the end of a pilot: expand, extend, or exit. Exit outcomes are defined before the pilot starts.


Why define this upfront

Pilots without pre-defined criteria produce ambiguity. The vendor reads results optimistically. The business reads them cautiously. Both interpretations are defensible because success was never specified.

Pre-definition creates three concrete benefits:

Anchors the evaluation. Vendors frequently expand a pilot's claimed value as it progresses. A written success brief limits what counts.

Creates a decision, not a negotiation. When the pilot ends, results either met the threshold or they did not. The outcome is a finding, not a conversation.

Protects against silent failure. Without defined thresholds, poor performance accumulates unnoticed until it becomes expensive to reverse.


The pre-pilot checklist

Work through each section in order. Every item should be written and shared with the vendor before the agent runs on any live workflow.


Step 1: Write the workflow brief

A workflow brief is the foundation of the entire pilot. Without one, the vendor will test their workflow, not yours.

Component

What to specify

Trigger

The event that starts the workflow — inbound message, form submission, calendar event, data update

Steps

The sequence of actions from trigger to completion, in order

Systems touched

Every tool, database, or platform the agent reads from or writes to

Output

The specific end state that constitutes a completed workflow

Volume

How many instances of this workflow run per day and per week

Exceptions

Known scenarios where the standard steps do not apply

Owner

The person on your team responsible for reviewing agent outputs during the pilot

A workflow brief should be one to two pages. If it cannot be written in that length, the workflow is not yet scoped tightly enough to pilot.


Step 2: Define the success threshold

The success threshold is a single sentence with three variables. Fill it in before the pilot starts.

The pilot succeeds if accuracy exceeds [X]%, escalation rate stays below [Y]%, and average completion time is under [Z] minutes, measured over [N] live workflow runs. This is important to be ruthlessly clear in your checklist.

Guidance on setting each variable:

Variable

How to set it

Common mistake

Accuracy [X]

Match the accuracy standard you hold your current human process to

Setting it higher than your current process, which creates an unfair bar

Escalation [Y]

Based on team capacity to handle escalated reviews during the pilot period

Setting it at zero, which penalises appropriate caution

Completion time [Z]

Based on the time requirement your customers or operations actually impose

Using average demo speed rather than your operational requirement

Sample size [N]

Minimum 200 instances for most business workflows

Ending the pilot before enough volume has run to produce reliable results


Step 3: Define the failure threshold

Failure thresholds are separate from success thresholds. They are conditions that trigger an immediate pause, not performance metrics to trend over time.

Hard stops. Define each of the following before the pilot begins:

  • The agent produces a customer-facing error that cannot be reversed

  • The agent accesses data or systems outside its documented scope

  • The agent produces a compliance or policy violation

  • Accuracy drops more than 50%

    in a single week with no explanation from the vendor

  • The vendor makes a change to the agent without advance notice

Share the hard-stop list with the vendor in writing before the pilot starts. A vendor who objects to any item on this list is providing useful information about their confidence in the product.


Step 4: Build the test dataset

Before the agent runs on live inputs, assemble a test dataset with three categories:

Category

Volume

Description

Standard inputs

20 to 30

Typical workflow instances — the cases that make up the majority of daily volume

Edge cases

5 to 10

Unusual but real inputs your team encounters regularly that a vendor demo never includes

Known-bad inputs

3 to 5

Inputs the agent should escalate or reject rather than process

Run the agent against the full test dataset in a non-live environment before it touches real operations. Any failure on a known-bad input at this stage is a finding to address before the live pilot begins, not a surprise to manage after.


Step 5: Define the review process

Every pilot needs a structured review cadence. Without one, issues accumulate silently.

Review type

Frequency

Owner

Purpose

Output sample review

Daily, first two weeks

Operations or team lead

Identify systematic errors before they scale

Edge case audit

Weekly

Department head

Confirm unusual inputs are handled correctly

Vendor sync

Weekly

Business lead and vendor CSM

Surface issues before they compound

Performance snapshot

Weekly

Both parties

Track accuracy, escalation rate, and completion time against threshold

End-of-pilot report

Final day of pilot

Both parties

Compare measured results against the pre-defined success threshold

The review cadence does not need to be heavy. A fifteen-minute daily sample check in week one catches most systematic issues before they become significant.


Step 6: Define the exit outcomes

At the end of the pilot, one of three outcomes applies. Define what each one means before the pilot starts.

Expand: Results met or exceeded the success threshold across the defined sample size. The agent moves to full deployment with a defined monitoring process in place.

Extend: Results were mixed. Specific underperforming areas are documented. The vendor has a written remediation plan. A second pilot phase is scoped with a revised success threshold and a shorter timeline.

Exit: Results fell below the success threshold, or a hard stop was triggered. The deployment ends. A post-mortem documents what failed and why before any alternative vendor is evaluated.

Pre-defined exit outcomes convert the end of a pilot from a negotiation into a decision.


Pilot timeline

Phase

Duration

Primary activity

Pre-pilot setup

1 week

Workflow brief, test dataset, success threshold, failure threshold, and exit outcomes — all documented and shared with vendor

Non-live validation

2 to 3 days

Agent runs against test dataset; findings addressed before go-live

Live pilot

3 to 4 weeks

Agent runs on real or parallel workflows with daily and weekly review cadence

Assessment

3 to 5 days

Results measured against pre-defined success threshold

Decision

1 week

Expand, extend, or exit based on written criteria

A pilot shorter than three weeks does not generate enough volume to produce reliable results on most business workflows. A pilot longer than six weeks without a structured review process becomes an informal continuation rather than an active evaluation.


Common failure patterns

The following are the most frequent causes of inconclusive or failed pilots. Each is preventable at the pre-pilot stage.

Failure pattern

Root cause

Prevention

"It mostly worked"

Success threshold was never defined

Complete Step 2 before go-live

Pilot drifts past deadline

Exit outcomes were not specified

Complete Step 6 before go-live

Vendor disputes the results

Metrics were not agreed on in writing

Share the full checklist with the vendor before the pilot starts

Hard stop triggered with no plan

Failure threshold was not defined

Complete Step 3 before go-live

Edge cases discovered during live pilot

Test dataset was not built

Complete Step 4 in non-live environment

Issues found too late to fix

No review cadence in place

Complete Step 5 before go-live

Sample size too small to conclude

Volume threshold was not defined

Specify [N] in the success threshold before the pilot starts


What Signal does with this

When a business submits a workflow to Signal, this checklist is the starting point.

Signal works through the workflow brief with you, defines realistic success thresholds based on evidence from comparable deployments, builds the test dataset including edge cases and known-bad inputs, and runs the agent evaluation before you commit to a live pilot.

The trust profile Signal produces at the end maps directly to the criteria above, accuracy, escalation rate, failure modes, and edge case handling documented against your specific workflow.

If you are preparing for a pilot and want a structured evaluation before you commit, submit your workflow here.