A benchmark is a measurement system. It can report a precise number while measuring removed work, an unstable sample, or an uncontrolled machine. This talk argues you should trust a Go benchmark only after answering three questions: Is the compiler measuring real work? Is the sample stable enough? Is the difference large relative to the noise?
We work through three layers. Compiler honesty covers dead-code elimination, sinks, constant folding, inlining, timer ordering, and the B.Loop construct, which manages the timer for you and keeps the loop body from being optimized away. Statistical interpretation covers repeated samples, benchstat, coefficient of variation, run-count discipline, and the p-hacking traps that inflate false positives. Environment control covers both local machines and CI: Linux frequency and isolation controls, perflock, benchdiff, and the bare-metal runners that shared CI instances are not.
A war story from dd-trace-go ties the layers together. A benchmark measured code the PR did not touch, the same-machine result reversed the CI sign, and code layout was the likely explanation for the small shift that made the result directionally wrong. The companion repositories ship the captured outputs and the tools that turn those lessons into repeatable checks.
Recording
Links
- gopherconuk-26 — slides, speaker notes, and demo results
- benchlab — the
honestbench,benchgate, andbenchenvCLIs
Events
- GopherCon UK 2026 — Wednesday 12 August 2026, 15:15, The Google Track
Related
- Why Your Go Benchmarks Are Lying — the five-part written series
- The Go Benchmark That Measured Nothing: Compiler Honesty in testing.B — part 1
- A Single Benchmark Number Is a Lie — part 2
- Before CI: Can You Trust a Benchmark on Your Own Laptop? — part 3
- Benchmark CI That Doesn’t Lie — part 4
- Three Questions Before You Trust a Benchmark — part 5
- Companion posts
- Benchmarking Go, quickly — a primer on writing and reading one benchmark, before part 1
- The laptop said 230% — a macrobenchmark where laptop and CI disagreed, between parts 3 and 4
- A PR gate that actually fails — a PR gate catching a slow commit, between parts 4 and 5
- A/B is the wrong model for CI — a benchmark’s history as a time series, between parts 4 and 5
- Measuring Software Performance: Why Your Benchmarks Are Probably Lying — full technical blog post expanding on this talk
