Here is a result that should be printed inside the cover of every textbook, and isn't. In 2006, the psychologists Henry Roediger and Jeffrey Karpicke had students study short prose passages, then either study them again or take a practice test — closing the book and recalling what they could. Five minutes later, the re-studiers looked better. A week later, everything had flipped, and it wasn't close: the students who had been tested retained dramatically more, while the re-studiers' fluent familiarity had evaporated.
The kicker is what the students believed. The re-studiers predicted they'd do best — rereading feels smooth, and smoothness feels like learning. The tested group, who had struggled and stumbled, underestimated themselves. The feeling of learning and the fact of learning had come apart completely.
The UCLA psychologist Robert Bjork has a name for this whole family of effects: desirable difficulties. Conditions that make practice feel harder and look slower — testing yourself, spacing sessions out, mixing problem types — reliably produce more durable, more flexible skill. Conditions that feel efficient produce performance that evaporates. It is among the most replicated findings in the learning sciences, and almost everything about it is counterintuitive — which makes it exactly the kind of thing a puzzle person should want to understand, because a puzzle habit is a learning machine you get to design yourself.
The engine: retrieval, not review
Why would struggling to remember beat looking at the answer? Because memory is not a shelf; it is a muscle-and-path system. Retrieving something — effortfully, without the page in front of you — strengthens the retrieval route itself. Re-exposure merely re-lights the room, and recognizing a lit room is easy in a way that finding it in the dark never is. Bjork's crucial distinction is between performance (how you're doing right now, during practice) and learning (what will survive until you need it). The two are not just different; they are negatively correlated across many conditions. Smooth practice inflates performance while starving learning. Rough practice depresses performance while feeding it.
If that sounds abstract, every solver has lived it. The technique you were shown slips away by Thursday. The technique you fought for — hunted, half-invented, confirmed through your own contradiction — is yours for years. We've written about why hard puzzles feel good; this is the companion fact: they also work better, and the ache is not the price of the learning. The ache is the learning.
Interleaving: the virtuous mess
The second difficulty worth engineering is the one practice schedules avoid by instinct: mixing. Doug Rohrer's studies of mathematics practice make the pattern brutally clear. Students who practiced problems in tidy blocks — ten of this type, ten of that — outperformed mixed-practice students during the session, then collapsed on later tests, where the interleaved group roughly doubled them in some experiments.
The mechanism is almost embarrassingly simple. Blocked practice answers the hardest question for you: what kind of problem is this? Ten cage-arithmetic drills in a row require no diagnosis — you apply the day's hammer to everything. Mixed practice forces the diagnostic act every single time, and the diagnosis is the skill. Real grids never announce their technique; they simply sit there, and the entire game is recognizing what the moment calls for.
Translated to puzzling: a Sudoku-only diet builds beautiful Sudoku grooves and a quiet Einstellung risk — the favorite-hammer problem we've covered before. Rotating puzzle types — a Nonogram today, KenKen tomorrow, Tents when you're brave — feels choppier and less masterful, which is precisely the tell that the deeper skill (read the constraints, pick the tool) is the thing being trained.
The struggle before the hint
Third difficulty: generation. People remember what they produce better than what they receive — an effect documented since the 1970s and robust across materials. The application is a rule about hints: earn your stuckness first. A hint taken after ten genuine minutes of attack lands on prepared ground — you know the terrain, you feel exactly which door it opens, and the pattern encodes deeply. The same hint taken at first resistance is rain on pavement. This is also the honest defense of not fearing hints entirely: research on learning from problem-solving suggests that failed attempts followed by the solution can outteach polished instruction alone. Struggle first; then let the answer teach.
Desirable is a load-bearing word
Now the crucial caveat, because "embrace difficulty" curdles fast into bad advice. Bjork's phrase is desirable difficulties, and the desirability condition is specific: the difficulty must be one you can eventually overcome with effort, and it must exercise the actual skill. A puzzle two tiers above you that reduces you to blind guessing trains guessing. Frustration past the edge of ability isn't noble; it's noise. The target is the rim of your competence — hard enough that errors happen and diagnosis is required, close enough that the struggle resolves. (Our difficulty tiers exist to make that rim findable; the honest advice has never been "play the hardest," but "play the one you can barely beat.")
And some days, comfort is the correct prescription. An easy grid after a brutal week is medicine of a different kind, and a habit that survives is worth more than an optimized one that quits. Difficulty is a training variable, not a moral one.
Designing your own gymnasium
The practical program, assembled from fifty years of evidence, is short:
Prefer recalling to reviewing. Before looking up a technique, try to reconstruct it. Wrong reconstructions, corrected, outteach smooth rereads.
Mix your grids. Rotate puzzle types across the week. The diagnostic stumble at each switch is the curriculum, not the tax.
Struggle, then hint, then continue. Ten honest minutes buys you the right to look — and makes the look worth something.
Ignore how practice feels. Fluency lies in both directions: today's smoothness predicts forgetting, today's fumbling predicts growth. Judge your habit by next month's solving, not this evening's grace.
The gym metaphor earns its keep one last time: nobody expects to grow by lifting weights that feel light. The strange, liberating news from the learning labs is that your ten daily minutes of honest difficulty were never inefficient. They were the only part that counted.