Forced 2×/4× inter-nonce unrolling
Forcing two or four nonces into one unrolled body can expose enough instruction-level parallelism to exceed the baseline.
Diagnostic work, exact algebra and local capabilities are not treated as end-to-end mining advantage.
What was tested?
Forcing two or four nonces into one unrolled body can expose enough instruction-level parallelism to exceed the baseline.
Why the test is meaningful
ILP is useful only if it outweighs duplicated instructions and register pressure under identical arithmetic and nonce assignment.
speedup_Uk=Tbaseline/TUkSASS_Uk≈k·SASS_basepromote on unseen headers onlyHow it was tested
Choose U2 on two discovery headers, freeze it, and evaluate four unseen headers in 44 balanced runs of 134,217,728 × 48 hashes.
What happened
U2 lost on 4/4 replication headers: kernel ratio 0.996673 and 95% interval 0.995713–0.997634. U2/U4 raised SASS from 2,507 to 4,968/9,874 and registers from 44 to 64.
Exactness and statistical controls
All variants recovered exact B32 hits with no discrepancies or spills. Selection was completed before replication.
What the result means
The compiler materialized duplication, but extra ILP did not repay footprint and register pressure.
Limitations
- Only forced factors 2 and 4 were tested.
- The conclusion is architecture-specific.
- Larger factors were closed by the negative mechanism.
Evidence trail
Freeze U2 after discovery, use unseen headers and audit exact work, SASS and register counts.
Canonical variants
CANONICAL-EXP-139PREREGISTERED-CAMPAIGNSEALED-AUDITSource: internally audited canonical reports. Local filesystem structure, private headers and operational identifiers are excluded from publication.