Exact lazy evaluation of the B8 predicate
The exact B8 predicate can reject on the first required-bit mismatch without evaluating every demanded bit.
- Variants
- 6
- Hardware
- Classical deterministic verifier
Seventy-five Advantage Decomposition families are consolidated from canonical reports. Every new entry includes its scientific basis, mathematical model, validation controls and explicit limitations. Some investigations require weeks or months of collection, replication and analysis.
6 published entries remain experimental or in progress.
“Local gain” is deliberately narrow: a measured reduction in diagnostic work or an exact avoided-enumeration capability inside a registered subproblem. It does not mean faster full SHA-256, higher hashrate, lower energy per valid Bitcoin block or an economic mining advantage.
The exact B8 predicate can reject on the first required-bit mismatch without evaluating every demanded bit.
The exact lazy mechanism remains valid before R56, while measured savings reveal where dependency cost begins to dominate.
Only the demanded semantic interface needs to be transported between checkpoints to reproduce the exact progressive decision.
Context classes and quotient representations can compress repeated semantic behavior without changing exact decisions.
A portfolio of exact specialized representations can cover more contexts than one global grammar.
Explicit carry state can factor contextual behavior that appears opaque in a coarse word-level representation.
A partial engine can expose only two safe online outcomes: exact rejection or defer-to-baseline.
Pruning evidence can be separated from the discovery engine and checked through a small abstract certificate interface.
Reachable Bitcoin states occupy structured local contexts that can select exact reusable capabilities.
For a fixed job, the early nonce-dependent lane begins from one 32-bit source W3 with exact A4/E4 translations.
Exact context inherited from the previous round can guide which specialized representation is applicable next.
Offline relation discovery can be separated from a small exact online motor with explicit provenance.
States labeled GENERAL by a coarse grammar may still admit exact, structured low-bit representations.
R4.T2 low four bits can be classified into exact affine context formulas over the registered source.
The context rule behind the V14 affine cases can be promoted from sampled observation to an exact fixed-context theorem.
Legal nTime values can dispatch exact R4 low-bit capability classes while version and merkle context remain fixed.
An exact M4 representation remains useful when inherited into the M5 transition even if the coarse next-stage label becomes GENERAL.
The registered R5 low-four-bit GENERAL cases can be represented as an affine function plus at most one x2·x3 term.
AQ23 and inherited borrow metadata can preserve sparse exact structure into later low-bit transitions.
Global low-bit complexity at R5 can be localized into carry boundaries and conditional Majority branches instead of treated as one opaque wall.
The R5 Σ0/carry and Majority low-bit wall can be replaced by finite exact program families, and any candidate simplification that survives into R6 can be detected without approximation.
Carry propagation across selected A5 intervals can be represented by exact Kill/Propagate/Generate transfer states, and an interval with identically zero propagation makes its incoming carry irrelevant.
Frequent P-zero carry jumps can be explained by exact single-bit resets or small certificates that prove a propagation chain cannot span the entire interval.
Hard resets at bits 16–18 arise from compact exact constancy or pair-relation certificates in the T1 and T2 operands, rather than from an irreducible black-box event.
The dominant bit-18 freeze can be reconstructed from the inner carry of Σ0(a4)+Maj4, and the original source-predicate shortcut must be rejected if independent review does not support it.
For bits 15–18, Σ0(a4) freezes exactly when one low rotated source bit freezes, because the other rotated source bits and Majority inputs are already constant in the registered context.
The V26 freeze and carry certificate can be derived from one baseline source trace and exact modular boundaries, without enumerating all 16 W3.low4 children online.
The complete outer T1/T2 bit-18 hard reset can be predicted exactly with 16-lane truth masks and ripple carry from one inherited M4 baseline state.
The exact bit-18 reset certificate can skip the preceding dependency and propagate through bits 19–21 to produce the true carry-22 mask on the same survivor population.
If the connected reset path retains a useful simplification, it should appear in the exact a5[22] mask or in the four low lanes of Σ0(a5) for the actual survivors.
The attenuation observed at Σ0(a5) lane 0 can be attributed exactly to cancellation among rotated input monomials rather than described only by an aggregate complexity score.
If attenuation is diffuse, a complete monomial-by-monomial overlap map should show broadly distributed small differences rather than one dominant reproducible cancellation term.
The replicated suppression pattern can be localized to terms containing source variable x3 and traced to the exact rotated input components that create or cancel them.
Any suppression attributed to x3 may be erased when modular addition regenerates nonlinear terms through carry, so the next connected carry frontier must be profiled exactly.
The exact source of carry-21 regeneration can be separated into Generate at bit 20 and the product of Propagate with incoming carry, revealing which primitive restores x3 complexity.
A local simplification is meaningful only if it survives composition into the complete T2₅ low-four-bit consumer, including Σ0(a5), Majority and inner carry.
The increased Majority complexity seen in V36 may be explained by mixing the c5=0 and c5=1 Boolean branches rather than by one stable within-branch mechanism.
If the Majority contrast is structural, it should remain directionally coherent after conditioning on exact affine truth-table families and c5 branches.
Any reset-derived reduction that matters beyond T2₅ must remain visible after the complete state update into a6 and e6, rather than only inside an operand.
The small e6 attenuation can support the bit18-to-R6 line only if the complete nonlinear monomial union, not a selected subset, replicates across original and fresh cohorts.
A hard reset at bit 17 remains operationally effective only when no later bit-18 reset overwrites it; under that condition it should reconstruct carry 22 exactly.
If the effective bit-17 restart creates a connected simplification, its direction should replicate across carry 22, a5[22], Σ0(a5)[0], T2₅.low4 and a6.low4 against matched no-later-reset controls.
Reset positions 16–18 can be unified by one exact rule: choose the latest active reset and evaluate only the suffix after that position.
A nested decision tree of exact reset certificates can evaluate the R60 target predicate with less diagnostic work than the exact ripple baseline.
A candidate that passes the first target byte can be tested for the second byte incrementally, with total work equal to B8 work plus conditional B16 work and no decision errors.
The B8→B16 nesting rule generalizes byte by byte through B24 and B32 while the comparison remains inside digest word 7.
The exact target ladder can cross from digest word 7 to the first byte of word 6 when the required e62/feedforward identity is included.
Once the cross-word bridge is established, the same byte-wise target family should extend through B48, B56 and B64 within digest word 6.
The target ladder can cross from B64 into digest word 5 by reconstructing the exact e63/feedforward relation and evaluating the next target byte backward from the final digest.
After the B72 prefix, exact lexicographic comparison can classify many states as final ACCEPT or REJECT while safely deferring only prefix ties to the remaining suffix.
For cases deferred by V50, comparison with the next registered target byte 0x35 should again produce exact ACCEPT, REJECT or a much smaller DEFER set.
Applying the next target byte 0x3D to V51 defers should resolve nearly all remaining cases and reveal whether a zero-suffix deferral can still satisfy PoW.
The digest-word-5 target relation can be transported exactly to an R59 predecessor interface using a fixed affine identity and only four reduced T1 primitives.
A complete SHA-256 round is a point-state bijection when the schedule word and round constant are fixed, allowing exact inversion and restart from the recovered predecessor.
Holding a next-round state fixed except for next-a defines a compact one-dimensional predecessor fiber with invariant support, affine d/h laws and exact forward restart.
The local M2 kernel should outperform a fixed, externally maintained CUDA SHA-256d baseline under equivalent work.
A different threads-per-block and nonces-per-thread geometry can improve the frozen external kernel by at least 1%.
Three independent Nsight Compute profiles can expose the bottlenecks of the frozen T512-N32 baseline.
The sm_120 compiler already maps the external SHA-256d arithmetic to Blackwell-native shift, Boolean and add instructions.
A frozen linear score of partial second-SHA state can select nonces that beat uniform completion under equivalent SHA-round work.
Traversing aligned blocks of 32 nonces in Gray order can reuse σ₀ through XOR linearity and improve throughput by at least 1%.
A non-unrolled Gray loop with constant nonce and σ₀ deltas preserves reuse without the code expansion observed in EXP-137.
Forcing two or four nonces into one unrolled body can expose enough instruction-level parallelism to exceed the baseline.
A small pilot can select headers whose disjoint holdout has a higher below-target density than equal-work random allocation.
Pilot B8 density within a high-byte nonce region predicts density in a disjoint low-24-bit holdout from the same region.
Moving invariant σ0(W16) and σ0(W17) from the nonce loop into CPU setup reduces GPU work.
Moving W18…W61 assignments directly before their first consuming rounds reduces live ranges, registers or instructions.
NPT values 1, 2, 4 or 8 at TPB 256/512 can outperform the retained T512-N32 baseline by at least 1%.
The small T512-N4 throughput signal from EXP-144 replicates on unseen headers with frozen binaries.
T512-N4 preserves its speed gain without worsening energy per hash or thermal stability.
Balancing both order sequences within every header confirms T512-N4 energy/thermal noninferiority.
An adjacent integer NPT value 3, 5, 6 or 7 improves the confirmed NPT4 default by at least 0.2%.
Batching multiple launches before result copy/synchronization improves end-to-end throughput by amortizing host overhead.
The P8 host-polling gain replicates on eight unseen headers without exceeding the registered 0.20-second polling interval.
P8 retains its throughput gain while meeting paired energy/thermal noninferiority.