EXP-079EXPERIMENTAL
Canonical GPU engineering / SHA-256d

Per-launch host observability inside P8

Synchronizing after each of eight kernels can restore one-launch host observability while retaining the confirmed P8 throughput gain.

Claim statusNO END-TO-END MINING ADVANTAGE DEMONSTRATED

Diagnostic work, exact algebra and local capabilities are not treated as end-to-end mining advantage.

Date2026-08-16
HardwareNVIDIA GeForce RTX 5080 · exact CUDA SHA-256d verifier
Run scopePreregistered paired campaign with sealed validation and audit
ReproducibilityFreeze P1, P8-S0 and P8-S1 before unseen headers; replicate wall throughput, retention and host-observation interval.
01 / Question & hypothesis

What was tested?

Synchronizing after each of eight kernels can restore one-launch host observability while retaining the confirmed P8 throughput gain.

02 / Scientific basis

Why the test is meaningful

Copy frequency and work-change observability are separate controls: a synchronization can expose completion without copying results, at a measurable scheduling cost.

R_P1=throughput(P8-S1)/throughput(P1)retention=throughput(P8-S1)/throughput(P8-S0)observation interval=batch time per launch
Per-launch host observability inside P8Visual reading of the published metrics and gates for EXP-079; it summarizes the registered result, not a mining advantage.EXP-079 / PER-LAUNCH HOST OBSERVABILITY INSIDE P8VERSUS P11.002597×P8 RETENTION0.999752HOST OBSERVATION0.020269 s
FIGURE / RESULT READINGVisual reading of the published metrics and gates for EXP-079; it summarizes the registered result, not a mining advantage.
03 / Method

How it was tested

Keep TPB 512, NPT 4, eight kernels per reset/copy and exact SHA-256d; add only one synchronization after each kernel and test four new headers in 24 runs.

04 / Observed result

What happened

Versus P11.002597×
P8 retention0.999752
Host observation0.020269 s

P8-S1 was 1.002597× versus P1 (CI95 1.001831–1.003364, 4/4) and retained 0.999752 of P8-S0, a ~0.025% cost. Host observation fell to 0.020269 s.

05 / Validation

Exactness and statistical controls

All binaries found B32, matched compiled factors, used 44 registers and had zero spills.

06 / Interpretation

What the result means

Per-launch observability is a plausible low-cost latency mechanism, but this discovery campaign requires frozen independent replication.

Limitations

  • Discovery rather than independent replication.
  • Energy and thermal behavior were not measured.
  • Synchronization semantics and cost are platform-specific.
07 / Reproduction

Evidence trail

Freeze P1, P8-S0 and P8-S1 before unseen headers; replicate wall throughput, retention and host-observation interval.

Canonical variants

CANONICAL-EXP-155PREREGISTERED-CAMPAIGNSEALED-AUDIT

Source: internally audited canonical reports. Local filesystem structure, private headers and operational identifiers are excluded from publication.