The Brainteaser Interview: What Hiring Got Wrong About Puzzles

October 4, 20266 min readBen Miller

Why are manhole covers round? How would you move Mount Fuji? How many piano tuners work in Chicago? If you interviewed at a technology company anytime between the late 1980s and the 2010s, questions like these may have stood between you and a paycheck. Microsoft made the style famous; half an industry copied it; the journalist William Poundstone's 2003 bestseller How Would You Move Mount Fuji? codified it for anxious candidates everywhere. The logic seemed self-evident: we want smart people, puzzles reveal smartness, therefore ambush candidates with puzzles.

We run a puzzle site. We are, professionally and personally, as pro-puzzle as it is possible to be. Which is exactly why we want to walk through what happened when someone finally checked whether any of it worked — because the answer clarifies what puzzles are actually for.

The audit

The someone was Google — by reputation the brainteaser's natural habitat. In a 2013 New York Times interview, Laszlo Bock, then the company's senior vice president of people operations, reported what emerged when Google studied its own hiring data at scale, comparing interview assessments against how hires actually performed on the job. His verdict on brainteasers has become the canonical obituary: "a complete waste of time. They don't predict anything. They serve primarily to make the interviewer feel smart."

That last sentence deserves its fame. It names the mechanism honestly: the gotcha question is theater, and the audience is the person asking.

The finding shouldn't have surprised anyone, because the broader science of selection was already decades old. A landmark 1998 meta-analysis by Frank Schmidt and John Hunter, synthesizing eighty-five years of research on hiring methods, found the strongest predictors of job performance to be work samples — watching someone do a piece of the actual job — alongside general ability measured properly and structured interviews: the same job-relevant questions, asked of every candidate, scored against the same rubric. What fails is the unstructured conversation steered by vibes, and the clever ambush lives at the far vibes end of that spectrum.

Why the gotcha can't measure

It's worth being precise about why brainteasers fail, because the failure is instructive to solvers.

They mostly measure prior exposure. Interview riddles circulate. The candidate who "brilliantly" derives the manhole answer often heard it in a dorm room in 2019; the one who stumbles simply hadn't. You are not sampling reasoning — you are sampling gossip networks. (Notice the deep contrast with a well-made logic puzzle, where the entire point is that everything needed is on the board.)

One trial, no feedback, maximal stress. Real reasoning — the kind we celebrate in every essay here — is iterative: hypothesize, test, contradict, revise. The interview brainteaser allows one attempt, under adrenaline, before a judge with the answer in his pocket. Psychology's findings on stereotype threat and evaluation anxiety all point the same way: high-stakes ambush conditions measure composure under ambush, a skill with almost no daily relevance to most jobs.

No connection to the work. The deepest failure is the simplest: knowing why manholes are round tells you nothing about whether someone can design a schema, comfort a patient, or ship a feature. Prediction comes from sampling the behavior you care about — which is why work samples top every meta-analysis.

The counterexample that proves the rule

Here's the twist regular readers will enjoy: history's most successful puzzle-based hiring program is one we've already written about. In 1942, the Daily Telegraph's crossword speed contest quietly fed codebreakers to Bletchley Park — puzzle-selected candidates who helped shorten a world war.

Was that a brainteaser interview? No — and the difference is the whole lesson. Bletchley's task was pattern-wrangling under time pressure; a timed crossword was, for that job, a work sample, not a gotcha. It was administered identically to all comers (structure!), scored objectively in minutes, and no interviewer got to feel clever. The Telegraph contest succeeded for precisely the reasons the Mount Fuji question fails. Puzzles weren't the problem in the brainteaser era. Irrelevance was — puzzles deployed as IQ theater for jobs that weren't puzzles.

What puzzles are actually for

So where does this leave those of us who spend our mornings on grids? In a better place, honestly, than the brainteaser era ever put us.

The interview fad treated a puzzle as an oracle — a one-shot verdict on the quality of a mind. That framing was always corrupt, and everything we know about solving says so. Performance on any single puzzle is noise: it depends on familiarity with the type (Einstellung cuts both ways), sleep, stress, and whether you happened to meet this trick before. What compounds into something real is the practice — hundreds of low-stakes reps in elimination, hypothesis-testing, contradiction-hunting, and the emotional skill of staying methodical while stuck. A puzzle is not a thermometer. It is a gym.

And gyms don't work as ambushes. Every finding we've toured this year — desirable difficulties, the testing effect, incubation, the value of explaining your stuckness aloud — depends on conditions the interview brainteaser inverts: safety to fail, feedback, iteration, return visits. Same artifact, opposite deployment, opposite results.

There is one more quiet lesson in Bock's data, and it's the one we'd frame on the wall. Google found that what did work looked humble: consistent questions, real tasks, evidence over impressions. The interviewers' beloved flourishes added nothing but noise. That is — solvers will recognize — exactly the ethic a good grid teaches. Showy leaps feel like intelligence; methodical verification is intelligence. The brainteaser interview died because it flattered the wrong thing.

The manhole covers, for the record, are round so they can't fall through their own holes — and so they can be rolled by one worker, and re-seated without alignment. It's a genuinely lovely piece of design thinking. It just never belonged in a job interview. It belonged here, among people solving for the only stakes that improve a mind: none at all.

← Back to Blog