A benchmark is a measurement system. It can report a precise number while measuring removed work, an unstable sample, or an uncontrolled machine. This talk argues you should trust a Go benchmark only after answering three questions: Is the compiler measuring real work? Is the sample stable enough? Is the difference large relative to the noise?

We work through three layers. Compiler honesty covers dead-code elimination, sinks, constant folding, inlining, timer ordering, and the B.Loop construct, which manages the timer for you and keeps the loop body from being optimized away. Statistical interpretation covers repeated samples, benchstat, coefficient of variation, run-count discipline, and the p-hacking traps that inflate false positives. Environment control covers both local machines and CI: Linux frequency and isolation controls, perflock, benchdiff, and the bare-metal runners that shared CI instances are not.

A war story from dd-trace-go ties the layers together. A benchmark measured code the PR did not touch, the same-machine result reversed the CI sign, and code layout was the likely explanation for the small shift that made the result directionally wrong. The companion repositories ship the captured outputs and the tools that turn those lessons into repeatable checks.

Recording

Links

  • gopherconuk-26 — slides, speaker notes, and demo results
  • benchlab — the honestbench, benchgate, and benchenv CLIs

Events

Related