What is a 2-week AI pilot process?
A 2-week AI pilot process is a 10-business-day delivery cycle for one prepared workflow. It starts after scope and baseline approval, builds the smallest useful path, tests it against representative records, and ends with a documented go, fix, or stop decision. Zyphh's Custom Build option uses this cadence for a narrow app, portal, agent, or automation.
It does not compress discovery, data repair, security review, and company-wide rollout into a fortnight. The Australian Government's AI transition guidance separates proof of concept, pilot, and production. A pilot tests real operating fit with a limited group and controlled data; production is a later stage with broader integration and support.
The schedule below turns Zyphh's public scope into a planning model. It is not a record of a named client engagement or a promise that every workflow can launch in ten days.
What must be ready before day 1?
Day 1 should begin with a build-ready decision, not an open question about where AI might help. The workflow needs one trigger, one main outcome, one accountable owner, a baseline source, and an agreed boundary for data and permissions. If those inputs are missing, the team needs discovery before the clock starts.
| Preflight item | Ready means | Pause the pilot when |
|---|---|---|
| Workflow boundary | One start event, finish event, owner, and excluded work | Several teams expect different outcomes |
| Baseline | A named source can show current volume, time, quality, or errors | The target relies on recollection alone |
| Test records | Normal, incomplete, duplicate, and failure cases are available | Only a polished demo example exists |
| Access | Test identities and minimum required scopes are approved | The build depends on a shared admin account |
| Decision rights | An owner can approve scope, risk, launch, and stop choices | Nobody can accept or reject residual risk |
NIST's AI RMF Core says the Map function should create enough context for an initial go or no-go decision, including intended use, business value, scope, human oversight, risks, and suitable benchmarks. A two-week build is downstream of that decision.
What happens across the ten business days?
The pilot should produce a decision trail, not ten disconnected demos. Each day has a primary output that the workflow owner can inspect. The order can move when an integration or approval demands it, but scope freeze, failure testing, limited use, and a final decision should remain visible.
| Day | Primary work | Inspectable output |
|---|---|---|
| 1 | Kickoff and scope freeze | Action contract, target metric, exclusions, stop rule |
| 2 | Trace records, fields, identities, and permissions | Data map, access plan, redacted test set |
| 3 | Build the normal path | One trigger reaches one verifiable outcome |
| 4 | Add validation and exception handling | Missing or invalid work enters an owned queue |
| 5 | Connect the first end-to-end slice | Midpoint review with logs and open risks |
| 6 | Run acceptance and failure cases | Results for normal, edge, timeout, and retry tests |
| 7 | Start shadow, draft-only, or limited use | Controlled records and reviewer feedback |
| 8 | Inspect exceptions and correct narrow faults | Issue ledger, changes, and repeated tests |
| 9 | Rehearse ownership, alerts, and recovery | Runbook, handoff, rollback, and support owner |
| 10 | Compare evidence with the baseline | Go, fix, or stop decision with next boundary |
What is built during days 1 through 5?
The first week turns the approved workflow into one inspectable slice. Day 1 records the action contract: allowed inputs, intended output, prohibited actions, success measure, stop condition, and work deliberately left out. Day 2 confirms the source records, stable identifiers, fields, retention limits, test environment, and service identities.
Days 3 and 4 implement the normal route and its first failure paths together. Validation, deterministic business rules, model-assisted steps, approvals, downstream writes, notifications, and exception ownership should appear in the same flow. OWASP's Excessive Agency guidance recommends minimum tool functionality, minimum permissions, downstream authorization, and human approval for high-impact actions.
Day 5 is a midpoint review, not a sales demo. The owner should trace one record from trigger to outcome, inspect the evidence used, explain every write, and show where a bad record stops. Any request that changes the success metric or adds another workflow goes into a later backlog.
How do days 6 through 8 test real operating behavior?
The second week starts by trying to break the agreed path. Test missing fields, duplicates, stale records, conflicting instructions, unavailable dependencies, expired credentials, timeouts, retries, partial writes, low-confidence output, and a replay of the same event. A clean normal case is necessary, but it says little about recovery.
Day 7 moves only the tested boundary into shadow, draft-only, or limited live use. The Australian Government's AI assurance guidance says a pilot should state scope, duration, objectives, key results, participants, consent, and risk controls, then monitor its acceptance criteria and incidents. Higher-risk work may need a longer test or an earlier stop.
Day 8 reviews exceptions with the people who do the work. Fix only faults inside the agreed scope, then repeat the affected test set. A new product feature, second team, or broad data-cleanup project is not a pilot correction. It is evidence that the next scope needs a separate decision.
What should happen on days 9 and 10?
Day 9 proves that the workflow can be operated when its builder is not standing beside it. The owner should receive the event map, field rules, service identities, approval points, alert destinations, exception definitions, test cases, rollback steps, change process, and known limits. The team then rehearses one alert and one recovery path.
CISA's logging guidance calls for choosing events to record, centralizing logs, alerting on high-risk activity, protecting log access, setting retention, and assigning incident roles. For an AI workflow, the log should connect the source event, decision inputs, version, approval, action, downstream response, retry, and manual repair.
Day 10 compares the observed evidence with the baseline and acceptance rules. Go means the bounded use can continue at its current authority. Fix means a specific fault deserves another test. Stop means the evidence does not justify more complexity or risk. None of the three decisions should silently widen access or audience.
How should the pilot be measured?
Use one primary business measure and a small set of operating safeguards. NIST says performance and risk tests should use documented metrics, representative deployment conditions, uncertainty, repeatable methods, and monitoring. A faster average is weak evidence if errors, manual review, or unowned exceptions rise.
| Measure | Question it answers | Do not hide |
|---|---|---|
| Primary outcome | Did the selected workflow improve the result named before build? | Baseline gaps or a changed definition |
| Quality and error | Were outputs correct enough for the stated use? | Corrections, overrides, and silent failures |
| Human work | Did review and exception handling fall, move, or grow? | Time shifted into another queue |
| Reliability | Did retries, duplicate protection, alerts, and recovery work? | Manual rescue by the build team |
| Adoption and cost | Did intended users use it, and what did each run consume? | Unused routes or variable vendor charges |
When is two weeks the wrong timeline?
Two weeks is the wrong promise when the workflow is not mapped, the target is disputed, source data needs repair, access or procurement is unresolved, or legal, privacy, security, accessibility, and domain reviews cannot fit the schedule. It is also wrong for several workflows, a full platform migration, organization-wide training, or an unrestricted autonomous system.
A pilot is not production in miniature. Official Australian guidance treats production as a later stage with full integration, support, and governance. The pilot can leave behind useful code and operating documents, but its main product is evidence about the next decision.
If the result earns expansion, add one source, user group, action, or permission boundary at a time and keep the tested controls. Zyphh's AI automation guardrails guide explains the permissions, approvals, logs, stop rules, and recovery paths that should survive beyond the pilot. If it does not earn expansion, preserve the decision record and stop cleanly.
Turn this into your own build plan.
Run the Workflow Opportunity Score or book a strategy call. Bring one repeated workflow, the tools involved, and the number that should move.
Run the scoreSources and further reading
FAQ
Is two weeks enough for an AI pilot?
Two weeks can be enough for one prepared workflow with an approved scope, accessible data, minimum permissions, a named owner, and clear acceptance rules. It is not enough for open-ended discovery, major data repair, broad security approval, several integrations, and an organization-wide production launch.
What is the difference between an AI proof of concept and a pilot?
A proof of concept asks whether an approach is technically feasible in a constrained setting. A pilot tests operating fit with a limited group, representative or controlled data, real ownership, and stated measures. Production follows only after the team accepts the wider integration, support, governance, and risk.
Should a two-week AI pilot use production data?
Use the least sensitive data that can answer the pilot question. Some pilots can use synthetic, historical, or redacted records. Others need a limited set of live records under approved access, retention, consent, monitoring, and rollback rules. The risk and decision context should determine the boundary.
What happens if the AI pilot misses its target?
The team should fix a specific, testable fault or stop. Keep the baseline, test results, issue ledger, cost record, and decision notes. Do not expand the scope to protect the original idea. A stopped pilot can still prevent a larger unsupported investment.