What is a 2-week AI pilot process?

A 2-week AI pilot process is a 10-business-day delivery cycle for one prepared workflow. It starts after scope and baseline approval, builds the smallest useful path, tests it against representative records, and ends with a documented go, fix, or stop decision. Zyphh's Custom Build option uses this cadence for a narrow app, portal, agent, or automation.

It does not compress discovery, data repair, security review, and company-wide rollout into a fortnight. The Australian Government's AI transition guidance separates proof of concept, pilot, and production. A pilot tests real operating fit with a limited group and controlled data; production is a later stage with broader integration and support.

The schedule below turns Zyphh's public scope into a planning model. It is not a record of a named client engagement or a promise that every workflow can launch in ten days.

What must be ready before day 1?

Day 1 should begin with a build-ready decision, not an open question about where AI might help. The workflow needs one trigger, one main outcome, one accountable owner, a baseline source, and an agreed boundary for data and permissions. If those inputs are missing, the team needs discovery before the clock starts.

Preflight itemReady meansPause the pilot when
Workflow boundaryOne start event, finish event, owner, and excluded workSeveral teams expect different outcomes
BaselineA named source can show current volume, time, quality, or errorsThe target relies on recollection alone
Test recordsNormal, incomplete, duplicate, and failure cases are availableOnly a polished demo example exists
AccessTest identities and minimum required scopes are approvedThe build depends on a shared admin account
Decision rightsAn owner can approve scope, risk, launch, and stop choicesNobody can accept or reject residual risk

NIST's AI RMF Core says the Map function should create enough context for an initial go or no-go decision, including intended use, business value, scope, human oversight, risks, and suitable benchmarks. A two-week build is downstream of that decision.

What happens across the ten business days?

The pilot should produce a decision trail, not ten disconnected demos. Each day has a primary output that the workflow owner can inspect. The order can move when an integration or approval demands it, but scope freeze, failure testing, limited use, and a final decision should remain visible.

DayPrimary workInspectable output
1Kickoff and scope freezeAction contract, target metric, exclusions, stop rule
2Trace records, fields, identities, and permissionsData map, access plan, redacted test set
3Build the normal pathOne trigger reaches one verifiable outcome
4Add validation and exception handlingMissing or invalid work enters an owned queue
5Connect the first end-to-end sliceMidpoint review with logs and open risks
6Run acceptance and failure casesResults for normal, edge, timeout, and retry tests
7Start shadow, draft-only, or limited useControlled records and reviewer feedback
8Inspect exceptions and correct narrow faultsIssue ledger, changes, and repeated tests
9Rehearse ownership, alerts, and recoveryRunbook, handoff, rollback, and support owner
10Compare evidence with the baselineGo, fix, or stop decision with next boundary

What is built during days 1 through 5?

The first week turns the approved workflow into one inspectable slice. Day 1 records the action contract: allowed inputs, intended output, prohibited actions, success measure, stop condition, and work deliberately left out. Day 2 confirms the source records, stable identifiers, fields, retention limits, test environment, and service identities.

Days 3 and 4 implement the normal route and its first failure paths together. Validation, deterministic business rules, model-assisted steps, approvals, downstream writes, notifications, and exception ownership should appear in the same flow. OWASP's Excessive Agency guidance recommends minimum tool functionality, minimum permissions, downstream authorization, and human approval for high-impact actions.

Day 5 is a midpoint review, not a sales demo. The owner should trace one record from trigger to outcome, inspect the evidence used, explain every write, and show where a bad record stops. Any request that changes the success metric or adds another workflow goes into a later backlog.

How do days 6 through 8 test real operating behavior?

The second week starts by trying to break the agreed path. Test missing fields, duplicates, stale records, conflicting instructions, unavailable dependencies, expired credentials, timeouts, retries, partial writes, low-confidence output, and a replay of the same event. A clean normal case is necessary, but it says little about recovery.

Day 7 moves only the tested boundary into shadow, draft-only, or limited live use. The Australian Government's AI assurance guidance says a pilot should state scope, duration, objectives, key results, participants, consent, and risk controls, then monitor its acceptance criteria and incidents. Higher-risk work may need a longer test or an earlier stop.

Day 8 reviews exceptions with the people who do the work. Fix only faults inside the agreed scope, then repeat the affected test set. A new product feature, second team, or broad data-cleanup project is not a pilot correction. It is evidence that the next scope needs a separate decision.

What should happen on days 9 and 10?

Day 9 proves that the workflow can be operated when its builder is not standing beside it. The owner should receive the event map, field rules, service identities, approval points, alert destinations, exception definitions, test cases, rollback steps, change process, and known limits. The team then rehearses one alert and one recovery path.

CISA's logging guidance calls for choosing events to record, centralizing logs, alerting on high-risk activity, protecting log access, setting retention, and assigning incident roles. For an AI workflow, the log should connect the source event, decision inputs, version, approval, action, downstream response, retry, and manual repair.

Day 10 compares the observed evidence with the baseline and acceptance rules. Go means the bounded use can continue at its current authority. Fix means a specific fault deserves another test. Stop means the evidence does not justify more complexity or risk. None of the three decisions should silently widen access or audience.

How should the pilot be measured?

Use one primary business measure and a small set of operating safeguards. NIST says performance and risk tests should use documented metrics, representative deployment conditions, uncertainty, repeatable methods, and monitoring. A faster average is weak evidence if errors, manual review, or unowned exceptions rise.

MeasureQuestion it answersDo not hide
Primary outcomeDid the selected workflow improve the result named before build?Baseline gaps or a changed definition
Quality and errorWere outputs correct enough for the stated use?Corrections, overrides, and silent failures
Human workDid review and exception handling fall, move, or grow?Time shifted into another queue
ReliabilityDid retries, duplicate protection, alerts, and recovery work?Manual rescue by the build team
Adoption and costDid intended users use it, and what did each run consume?Unused routes or variable vendor charges

When is two weeks the wrong timeline?

Two weeks is the wrong promise when the workflow is not mapped, the target is disputed, source data needs repair, access or procurement is unresolved, or legal, privacy, security, accessibility, and domain reviews cannot fit the schedule. It is also wrong for several workflows, a full platform migration, organization-wide training, or an unrestricted autonomous system.

A pilot is not production in miniature. Official Australian guidance treats production as a later stage with full integration, support, and governance. The pilot can leave behind useful code and operating documents, but its main product is evidence about the next decision.

If the result earns expansion, add one source, user group, action, or permission boundary at a time and keep the tested controls. Zyphh's AI automation guardrails guide explains the permissions, approvals, logs, stop rules, and recovery paths that should survive beyond the pilot. If it does not earn expansion, preserve the decision record and stop cleanly.

Turn this into your own build plan.

Run the Workflow Opportunity Score or book a strategy call. Bring one repeated workflow, the tools involved, and the number that should move.

Run the score

Sources and further reading

  1. Australian Government: Guidance for AI proof of concept to scale
  2. Australian Government: AI transition stages and dimensions
  3. Australian Government AI assurance guidance: Reliability and safety
  4. NIST AI Risk Management Framework Core
  5. OWASP LLM06:2025 Excessive Agency
  6. CISA: Use Logging on Business Systems

FAQ

Is two weeks enough for an AI pilot?

Two weeks can be enough for one prepared workflow with an approved scope, accessible data, minimum permissions, a named owner, and clear acceptance rules. It is not enough for open-ended discovery, major data repair, broad security approval, several integrations, and an organization-wide production launch.

What is the difference between an AI proof of concept and a pilot?

A proof of concept asks whether an approach is technically feasible in a constrained setting. A pilot tests operating fit with a limited group, representative or controlled data, real ownership, and stated measures. Production follows only after the team accepts the wider integration, support, governance, and risk.

Should a two-week AI pilot use production data?

Use the least sensitive data that can answer the pilot question. Some pilots can use synthetic, historical, or redacted records. Others need a limited set of live records under approved access, retention, consent, monitoring, and rollback rules. The risk and decision context should determine the boundary.

What happens if the AI pilot misses its target?

The team should fix a specific, testable fault or stop. Keep the baseline, test results, issue ledger, cost record, and decision notes. Do not expand the scope to protect the original idea. A stopped pilot can still prevent a larger unsupported investment.