PositionRecorded outputMeasured next-token predictions
Pick a baseline inference run, add as many as four alternatives, and replay what every run wanted to emit from the identical recorded prefix.
Step 1 · establish the reference universe
Every comparison starts from one canonical Triton, BF16-weight, BF16-key/value-cache run.