Home  /  Insights  /  Week-8 gate

The week-8 go/no-go gate: the exact thresholds and why each one exists

Most AI programmes never define what "working" means before they start. A gate does. These are the thresholds SEAS applies at the end of week 8 and week 24, with the reason behind each number.

By Lalit Kumar · Published 20 September 2026 · 8 minute read

What a phase gate is, and why week 8

A gate is a scheduled decision with pre-agreed criteria: either the pilot clears every threshold and the next phase is funded, or it pauses for a fixed remediation period and is re-tested. The alternative, an open-ended pilot, is how artificial intelligence (AI) initiatives drift for a year without ever producing a number a board can act on. In the SEAS implementation system the first gate falls at the end of week 8 because eight weeks is long enough for a single workflow to run in parallel with the manual process on real volume, and short enough that a pause costs a fortnight rather than a quarter. Remediating a flawed pilot at week 8 costs about one tenth of remediating it after it has been scaled.

The five performance metrics

MetricThresholdWhy this number
Agent accuracyAbove 90%, manually validated on 100 or more transactionsBelow 90% the exception queue grows faster than the team can clear it; the 100-transaction floor stops a good week from masking a bad month
Cycle timeDown 25% or more against the recorded baselineSmaller gains are usually the manual workaround improving, not the agent
Cost per transactionDown 15% or more against baselineThe floor at which savings survive the vendor licence and the exception handling
Exception rateUnder 8%Above 8% the humans "approving every action" are doing the work again
User satisfactionAbove 3.5 out of 5 in a team surveyAdoption fails silently; a score below 3.5 predicts workarounds in Phase 2

The thresholds only mean something against a baseline recorded before the pilot. The Project Manager's Runbook records baseline cycle time, cost per transaction and error rate in the pre-launch week for exactly that reason.

The five operational criteria

Performance is half the gate. A pilot that hits every number but runs on a fragile integration is not ready to scale. The Runbook adds:

  1. Stable operation for four weeks or more at above 95 percent uptime
  2. Unresolved data quality issues under 2 percent
  3. Zero unplanned downtime in the last two weeks
  4. The team trained and confident in agent oversight
  5. Budget variance under 10 percent from plan

Pass threshold: seven of ten. Below seven, the pilot pauses for two weeks, the failed criteria get named owners, and the gate is re-run. The playbook's Phase 1 success metrics are the first five in the table; the Runbook adds the five operational criteria and sets the seven-of-ten rule for the printed checklist a project manager uses.

What "accuracy" means, and the 85 percent question

Accuracy is measured by a human reviewing a random sample of at least 100 agent decisions against the correct answer, not by the vendor's dashboard. Two numbers appear in the SEAS materials and they play different roles: 85 percent is the floor the vendor guarantees contractually by week 8, the point at which the SLA addendum's refund triggers activate; 90 percent is the gate the company applies to its own decision to scale. A vendor can meet its contract and still not clear your gate. The Runbook's troubleshooting note applies when accuracy stalls at 85 percent: 85 percent on 95 percent of volume beats 100 percent on none, provided the remaining 15 percent has an escalation path.

What the numbers look like in a healthy pilot

The Runbook's tracking exhibits show the trajectory a pilot should follow on a $125-per-transaction baseline: accuracy at 87 percent in week 4, 93 percent in week 8, 96 percent in week 12; cycle time from 5.2 days to 3.4 days by week 8; cost per transaction from $125 to $82 by week 4. On 1,200 transactions a week, the last figure is $51,600 of weekly savings, $2,683,200 annualized at full volume, and $670,800 on a pilot covering a quarter of volume. Those are the numbers that go on slide 4 of the board deck.

The second gate: production readiness at week 24

Clearing week 8 earns controlled production, not enterprise rollout. The week-24 gate is stricter and operational: uptime of 99 percent or better for four weeks, exception rate under 2 percent and stable, accuracy consistently above 98 percent, cost savings consistent week to week, written recovery procedures for every documented exception, backup and fallback procedures tested, the support team resolving 95 percent of issues without the vendor, monitoring and alerting configured, escalation procedures tested, and a business continuity plan on file. Pass threshold: eight of ten. The vendor SLA runs at 99.5 percent uptime, separately, with its own contractual remedies.

Who decides

The gate is a steering committee decision, not the vendor's and not the project manager's alone. The Runbook's escalation matrix puts scope changes, timeline shifts over two weeks and exception rates above 5 percent at Level 2 (chief financial officer or chief operating officer with the chief information officer and the vendor project manager), and dropping a target process or an EBITDA shortfall of more than half the target at Level 3, the board. A gate that fails twice is a Level 3 conversation.

Frequently asked questions

What are the go/no-go criteria for an AI pilot?
Five performance metrics (accuracy above 90 percent on at least 100 validated transactions, cycle time down 25 percent, cost per transaction down 15 percent, exceptions under 8 percent, user satisfaction above 3.5 of 5) plus five operational criteria (four weeks above 95 percent uptime, unresolved data issues under 2 percent, no unplanned downtime for two weeks, team trained, budget variance under 10 percent). Pass at seven of ten.
Why is the gate at week 8 and not week 12?
Eight weeks gives one workflow enough real volume to measure against baseline while keeping the cost of a pause to two weeks. The Runbook estimates fixing a pilot at this stage costs one tenth of fixing it after scaling.
What happens if the pilot fails the gate?
It pauses for two weeks. Each failed criterion gets a named owner and a remediation action, then the gate is re-run. Two consecutive failures escalate to the board under the three-level escalation matrix.
Is 85 percent accuracy acceptable?
As a vendor contractual floor by week 8, yes; that is where the SLA refund triggers sit. As the company's own scaling gate, no; the threshold is 90 percent, manually validated. The two numbers do different jobs.
What is the week-24 production readiness gate?
Ten operational criteria including 99 percent uptime for four weeks, exceptions under 2 percent, accuracy above 98 percent, tested backup and escalation procedures, and 95 percent support independence from the vendor. Pass at eight of ten.

The gate checklists, the weekly metrics dashboard, the escalation matrix and the exception protocols are one printable document in SEAS, the Project Manager's Runbook.

See what ships in SEAS   Get the free FLOAT diagnostic book