03 · Gate

Turn comparison evidence into a clear ship or hold decision.

Bring baseline, candidate, critical regressions, prompt and model versions, reviewer notes, cost deltas, and acceptance policy into one release record.

Release Controlsupport-agent / rc-28 Candidate review
CANDIDATE

gpt-5.6 · prompt v19

HOLD 2 critical gates
Quality86.4+2.8
P95 latency2.74s-8.1%
Cost / success$0.084+14.6%
Coverage94%1 missing
Baseline comparison482 cases
Help
RAG
Tools
Tone
Policy
Long
Edge
Release gates6 / 8 pass
Task successPASS
Policy adherenceFAIL
Tool accuracyPASS
Cost ceilingFAIL
Latency budgetPASS
MKMira K.Reviewer

Policy refusals improved, but fallback routing raises successful-task cost above the approved ceiling.

STAGE OUTPUTS

Leave this stage with evidence the next one can use.

Each stage is configured around the customer's application, risk, quality standard, operational ownership, and release policy.

01

Explicit release policy

Define required metrics, critical failure modes, segment thresholds, and minimum coverage.

02

Versioned decision record

Preserve the prompt, model, configuration, dataset, evaluators, and evidence used to decide.

03

Human review where needed

Route material regressions, policy exceptions, and uncertain cases to named reviewers.

04

Controlled promotion

Record approve, hold, exception, rollback, and rollout status without losing the rationale.

CONNECTED RELEASE RECORD

No stage operates in isolation.

The evidence links traces, versions, evaluation runs, release gates, production monitors, and cost attribution.

Inspect architecture →
ENGINEERING NOTES

Practical thinking for production AI quality.

View the research library →
MAKE THE DECISION EXPLICIT

Define what must be true before the candidate ships.

Bring critical regressions, reviewer judgment, acceptance policy, and promotion state into one accountable release record.

Define a release gate