POC Results

Strong judgment misses less. Strong judgment reduces rework.

Two bounded POC signals: missed-risk suppression in the public v12 benchmark, and workflow-loop reduction in a separate synthetic workflow POC.

Public v12 POC0.0249%Missed-risk rate
Defined POC proxy. Not general accuracy, uptime, or a safety guarantee.
Synthetic workflow POC10 → ~2Rework loops to convergence
Synthetic result. Not a production guarantee and not a promise of two-loop convergence.
JudgmentMissed RiskConvergenceReworkHOLDHuman Gate
Two evidence signals

Miss less. Converge with less rework.

These numbers come from different bounded POC settings and are labeled separately to avoid implying a single universal benchmark.

Public v12 benchmark

Strong judgment misses less.

0.0249%

Missed-risk rate proxy for P7 Tail-Zoom AI in the defined chronological five-regime POC. Equivalent miss-avoidance: 99.975%.

Synthetic workflow POC

Strong judgment reduces rework.

10 → ~2

Rework loops versus the baseline workflow. The synthetic sales-stage POC described about one-fifth the baseline loop count, including repair plus re-evaluation / confirmation.

P7 miss-avoidance99.975%Defined irreversible-miss avoidance proxy in the chronological five-regime POC.
Tail-good miss vs Safeguarded-32.54%Relative reduction reported by the public paper.
Emerging-good miss vs Safeguarded-54.60%Largest incremental P7 effect in the public summary.
Overall miss vs Safeguarded-10.66%Aggregate incremental improvement is smaller than the tail/emerging effect.
Review-load delta+0.002Proxy increase per episode relative to Safeguarded AI.
Workflow interpretation≈1/5Synthetic rework-loop ratio: about 10 loops to about 2. Separate from the public v12 benchmark.
Evidence boundary: 0.0249% belongs to the defined public v12 POC proxy. The 10 → ~2 loop result belongs to a separate synthetic workflow POC. Neither is a production guarantee, a universal detection rate, a safety guarantee, or a claim that every workflow converges in two loops.
Four-tier comparison

Public chronological v12 summary.

Tiermiss_ratemiss_avoidancesuppression vs standardtail_good_missemerging_good_missreview_loadlatency_ms_estcomplexity_x
Standard AI0.12168987.831%1.0x0.5466170.4746470.000421.0
Deferred AI0.00044099.956%276.7x0.0598250.0135577.900651.5
Safeguarded AI0.00027999.972%436.2x0.0255080.0082008.0501182.4
P7 Tail-Zoom AI0.00024999.975%488.2x0.0172080.0037238.0521212.8

Source: Table 1 of the public v0.5 POC preprint. “latency_ms_est” and “complexity_x” are the paper’s benchmark fields and should not be read as production guarantees.

Interpretation

Judgment value appears in both misses and rework.

Miss fewer critical cases

The public v12 benchmark supports selective preservation and re-evaluation of low-frequency, boundary-near, and emerging candidates rather than indiscriminate review of everything.

Reduce unresolved residue

The synthetic workflow POC links fewer missed issues with fewer unresolved residues and fewer correction loops. It is a mechanism-level sales signal, not a production guarantee.

Keep the tradeoff visible

Review load, latency, retained state, and Human Gate requirements remain part of the decision. Strong judgment is not free and does not eliminate final accountability.

POC preprint cover
Beyond 99.9% Miss-Avoidance under Regime Shift
P7 Tail-Zoom Mapping for Reducing Miss Surveillance · public preprint draft v0.5 · April 2026
Download PDF