Per-launch host observability inside P8
Synchronizing after each of eight kernels can restore one-launch host observability while retaining the confirmed P8 throughput gain.
Diagnostic work, exact algebra and local capabilities are not treated as end-to-end mining advantage.
What was tested?
Synchronizing after each of eight kernels can restore one-launch host observability while retaining the confirmed P8 throughput gain.
Why the test is meaningful
Copy frequency and work-change observability are separate controls: a synchronization can expose completion without copying results, at a measurable scheduling cost.
R_P1=throughput(P8-S1)/throughput(P1)retention=throughput(P8-S1)/throughput(P8-S0)observation interval=batch time per launchHow it was tested
Keep TPB 512, NPT 4, eight kernels per reset/copy and exact SHA-256d; add only one synchronization after each kernel and test four new headers in 24 runs.
What happened
P8-S1 was 1.002597× versus P1 (CI95 1.001831–1.003364, 4/4) and retained 0.999752 of P8-S0, a ~0.025% cost. Host observation fell to 0.020269 s.
Exactness and statistical controls
All binaries found B32, matched compiled factors, used 44 registers and had zero spills.
What the result means
Per-launch observability is a plausible low-cost latency mechanism, but this discovery campaign requires frozen independent replication.
Limitations
- Discovery rather than independent replication.
- Energy and thermal behavior were not measured.
- Synchronization semantics and cost are platform-specific.
Evidence trail
Freeze P1, P8-S0 and P8-S1 before unseen headers; replicate wall throughput, retention and host-observation interval.
Canonical variants
CANONICAL-EXP-155PREREGISTERED-CAMPAIGNSEALED-AUDITSource: internally audited canonical reports. Local filesystem structure, private headers and operational identifiers are excluded from publication.