Why Benchmarks Lie More Than Marketing Claims
I spent most of an evening last week explaining to a friend why his benchmark-winning laptop compiles slower than my mid-range one. His synthetic scores were about forty percent higher. His build times were thirty percent longer. He didn’t believe the stopwatch until we ran it three times.
Marketing claims get all the mockery. “Up to 50% faster” is a punchline. “Revolutionary performance” gets an eye-roll. Fair enough, but at least those are visibly suspicious. Benchmarks turn up in the respectable clothing of objectivity, with numbers and graphs and decimal places, and we trust them precisely because they look like science.
My British lilac cat, Pixel, runs her own benchmark suite. She evaluates new furniture by lying on it and refusing to move for four hours. By her metric the cardboard box from a delivery in March comfortably outperforms the £180 cat tree. She isn’t wrong. She’s measuring what matters to her rather than what the manufacturer wanted measured, which is exactly the thing synthetic benchmarks fail to do for you.
The choice happens before the test runs
Every benchmark makes decisions about what to measure, how, and how to present the result. Those decisions are invisible to almost everyone reading the output, and they determine the winner before a single test executes.
A CPU benchmark can emphasise single-thread performance, multi-thread throughput, specific instruction sets, thermal behaviour or power efficiency. Each emphasis produces a different ranking of the same chips. None of them is dishonest. They’re answers to different questions, presented as if they answered yours.
Ask something apparently simple, like which phone is faster, and it falls apart immediately. Faster at launching apps? Rendering video? Loading pages? Switching tasks? A handset tuned for burst performance wins synthetic tests and loses sustained ones. A handset with conservative thermal management scores lower and feels quicker all day because it never throttles.
There’s an incentive problem stacked on top. Plenty of popular benchmarks come from companies selling hardware, software or advertising, and even the independent ones face pressure to produce interesting results. “Everything in this price bracket performs about the same” is the honest finding most of the time and it generates no clicks at all. The structural pull is toward tests that separate products dramatically, whether or not the separation exists in use.
How they mislead without lying
The effective deceptions never involve fabricated numbers. They involve genuinely measured results from carefully chosen conditions.
Cherry-picked conditions are the most common. A laptop hits its score plugged in, fans at maximum, in a climate-controlled room. You’ll run it on battery, in quiet mode, in a café, in whatever the weather is doing. Identical hardware, unrecognisable performance.
Optimising for the test is now an industry. When manufacturers know which benchmarks reviewers run, firmware gets tuned for those specific workloads, and a few phone makers have been caught detecting benchmark apps outright and unlocking behaviour they don’t otherwise permit. Your workload receives none of that attention, because nobody knows what it is.
Then there’s precision that isn’t there. Processor A scores 12,847 and processor B scores 12,651, and the four significant figures imply the gap means something. Run-to-run variance frequently exceeds it. Ambient temperature, background processes and luck all move the number. That’s noise, printed as signal.
And averaging quietly destroys the information you needed. A ninety-five frames per second average sounds excellent right up until you learn that every third second contained a 200-millisecond stall. Your eyes catch the stalls. The mean doesn’t contain them.
The workload mismatch nobody can fix
The deepest problem isn’t manipulation. It’s that synthetic workloads don’t resemble real use, and that isn’t anyone’s fault. A universal benchmark that predicts every user’s experience is not a hard engineering problem, it’s an impossible one.
Real work is filthy. You run a browser with forty-odd tabs, a chat client, something playing music, and whatever you’re actually meant to be doing, all at once, switching unpredictably, idle for ten minutes then demanding everything. No test captures that, because the shape of the chaos is different for every person.
The benchmark environment is surgically clean by comparison. Fresh install, minimal background processes, tuned settings, controlled temperature. Your machine has four years of accumulated software, a drive that’s eighty percent full, an antivirus scan starting at the worst possible moment, and a desk in direct sunlight in July.
Which is why benchmark leaders so often disappoint. The device tuned for clean conditions may have traded away the robustness that carries you through dirty ones.
Thermals: the variable that decides everything
No factor illustrates this better than heat. Modern processors adjust performance continuously based on temperature, delivering rated speed only while cool and slowing down as they warm. That gap between peak and sustained is where benchmark scores and lived experience separate.
Benchmarks run for minutes. Your work runs for hours. A laptop can post a spectacular score during a short test and then throttle hard when you ask for the same load across an afternoon. The test measured peak capability honestly. Sustained capability might be a third lower.
Thin machines suffer worst. The identical chip in a thicker chassis with real cooling benchmarks about the same and outperforms it substantially over an afternoon. The numbers match. The experience doesn’t.
I learned this expensively. Years ago I bought a beautiful thin laptop with impressive scores and discovered within a fortnight that it couldn’t compile for more than ten minutes before throttling below the older, fatter machine it replaced. The benchmark told the truth about peak performance and lied by omission about everything else.
Storage is the worst offender
Storage benchmarks might be the most divorced from reality of any category, and they’re the ones people quote most confidently.
A modern SSD claims 7,000 MB/s sequential reads. That number is real and almost entirely irrelevant, because sequential speed describes reading and writing large continuous files, which is what you do when copying video. Most computer use is small random access: loading a program, reading settings, opening a document. Random performance typically runs at a tiny fraction of the sequential figure, and the impressive number describes maybe a few percent of your actual disk activity.
Even the random figures come from empty drives in controlled conditions. A drive at eighty percent capacity behaves differently from one at twenty. A drive that’s been rewritten for two years behaves differently from a fresh one. A drive at 70°C in a thin laptop behaves differently again.
For most people the gap between a decent modern SSD and the fastest available is imperceptible. The benchmark difference might be threefold. The difference you can notice when an application launches is measured in milliseconds you cannot perceive.
What I do instead
After enough of these mistakes I’ve settled on something slower and more accurate. It isn’t a methodology and I’m not going to pretend it’s rigorous. It’s five habits.
Write down your actual workload first, before looking at any number. Not what you might do, not what sounds impressive. Mine for a laptop: writing in a text editor, a browser with too many tabs, video calls, occasional photo work. That list determines which numbers can possibly matter.
Then go looking for tests that match it. Generic synthetic scores matter far less than something resembling your use. For a writing machine I care about keyboard feel, screen legibility and battery under light load, and 3D rendering throughput tells me nothing at all.
Hunt specifically for sustained-performance data. Some reviewers now run long tests and report what happens after an hour, which is worth more than every peak figure combined.
Weight experience reports above specifications, especially from people whose usage resembles yours. A gaming reviewer’s frustration is not information about your writing laptop.
And use the return policy as a test rig where you can. A week of your real work reveals more than any published number, and the friction of sending something back is cheaper than three years of quiet disappointment.
Why numbers get past your defences
Benchmarks work on you because of some fairly well-documented biases, and knowing the names helps a little.
Anchoring means the first number sticks. Once you’ve seen that a phone scored 800,000 on some suite, that figure shapes your judgment even though you have no idea what the scale means or whether it relates to anything you do.
Precision bias means detail reads as credibility. “Twelve percent faster in JavaScript execution” sounds more trustworthy than “feels snappier when browsing”, even in cases where the vague subjective report is the better predictor of your satisfaction.
And comparison bias makes relative gaps feel urgent regardless of whether either option is already sufficient. Once you know A beats B by fifteen percent you want A, even when B clears every threshold you actually have. Benchmarks trigger this constantly by framing everything as a competition with a winner.
Where they genuinely help
Having spent two thousand words on the problem, fairness requires the other half.
Benchmarks are good at identifying inadequacy. A machine scoring far below its peers at compilation will disappoint someone who compiles all day, and the number tells you that reliably. They’re weak at choosing between adequate options and strong at eliminating inadequate ones.
They’re also useful as a tiebreaker once you’ve narrowed to a few similar devices at a similar price, where the differences are small but real.
And they’re genuinely valuable pointed at your own machine over time. Benchmark your laptop annually and you’ll catch a degrading drive, a developing thermal fault, or accumulated software rot months before it becomes obvious. The same test that misleads across devices illuminates changes within one.
So the fix isn’t ignoring numbers. It’s asking five questions before you use one: what exactly did this measure, how long did it run, under what conditions, who made it and what do they sell, and does it test anything I actually do. That last one eliminates most of what you’ll encounter.
Next time you’re choosing between devices, try running it backwards. Read the experience reports, handle the hardware if you can, ask two people who own one what they hate about it, decide, and only then look at the benchmarks to see whether they agree. You’ll be surprised how often the loser was your better option.
Pixel has just walked across the keyboard to demonstrate her own metric, speed at interrupting productivity, on which she scores extremely well. The test is rigged in her favour and I’ve stopped contesting the result.
One email a month: the upcoming live event + free recording access for subscribers. No spam, unsubscribe anytime.

