Sudoku Paradise Benchmark

Hardest Sudoku Compared: AI Escargot vs. Everest vs. Golden Nugget

Three enduring Sudoku Monsters were submitted to the same logic-first solver library. All three were solved without blind Trial & Error, but each resisted in a remarkably different way.

Three Monsters, three kinds of difficulty

AI Escargot, created by Arto Inkala in 2006; the Golden Nugget by tarek in 2007; and Arto Inkala’s 2012 puzzle known as Everest or AI Maze all appear in discussions of the world’s hardest Sudoku. AI Escargot and the 2012 Inkala puzzle are sometimes confused because both were promoted as the world’s hardest. They are separate grids with very different logical personalities.

The current verification runs used the evolving Sudoku Paradise Human Solver and its 54 human-style methods. Every puzzle received an Insane rating. Every puzzle was logic-solved. None required blind Trial & Error.

Puzzle Givens Logical actions Human Solver time* Defining resistance
AI Escargot 23 112 48.18 sec. ALS-XZ, contradiction proofs, SET, and finned fish
Golden Nugget 21 95 53.00 sec. Immediate high-level chains and dense advanced resistance
Everest / AI Maze 21 149 20.95 sec. Large SET deductions followed by broad subset logic

*Times are machine-dependent and are not universal performance claims. Logical-action counts describe the recorded path produced by this solver deck and dispatch order; they are not interchangeable with human minutes.

What the benchmark does—and does not—prove

The Sudoku Paradise deduction solvers do not use brute force or derive deductions from a stored solution. Candidate eliminations are produced by identifiable logical strategies, and the final log records a logical narrative rather than an opaque computational answer.

Important: a Sudoku’s measured difficulty depends partly upon the grid and partly upon the logical vocabulary available to the solver. The puzzle does not change when a new strategy is learned; our understanding of its shortest logical path does.

AI Escargot

AI Escargot is comparatively compact, but several critical deductions must be proven before it yields. In the current run, ALS-XZ, Conjugate Contradiction Scout, SET Equivalence, a Finned Swordfish, and several Locked Candidate interactions all contributed to the break-in.

100007090030020008009600500005300900010080002600004000300000010040000007007000300
DifficultyInsane
Log entries112
Trial & ErrorNo
Logic solvedYes
Selected recorded strategies
  • Naked Single: 49
  • Hidden Single: 9
  • ALS-XZ: 9
  • Conjugate Contradiction Scout: 6
  • Locked Candidates Claiming Col: 6
  • Naked Triples: 6
  • Set Equivalence Theory: 6
  • Finned Swordfish: 5
  • Locked Candidates Pointing Col: 4
  • Unique Rectangle Type 4: 3
  • X-Wing: 2
  • W-Wing: 1

The Golden Nugget

The Golden Nugget produced fewer recorded actions than the other two current runs, but that number understates its human hostility. It demands high-level relationships almost immediately and remains densely resistant.

000000012000003004001040500030600000005020007800009000002050040060800000900000700
DifficultyInsane
Log entries95
Trial & ErrorNo
Logic solvedYes
Selected recorded strategies
  • Naked Single: 45
  • Hidden Single: 15
  • Naked Pairs: 8
  • Conjugate Contradiction Scout: 7
  • Hidden Pairs: 5
  • Locked Candidates Pointing Row: 4
  • ALS-XZ: 2
  • Empty Rectangle: 2
  • Locked Candidates Claiming Row: 2
  • Locked Candidates Pointing Col: 2
  • 2D-Kite: 1
  • W-Wing: 1
  • XY-Wing: 1

Everest, also called AI Maze

Everest produced the largest logical record—149 actions—but completed in the shortest machine time. That apparent contradiction is the most instructive result in the comparison.

800000000003600000070090200050007000000045700000100030001000068008500010090000400
DifficultyInsane
Log entries149
Trial & ErrorNo
Logic solvedYes
Selected recorded strategies
  • Naked Single: 44
  • Set Equivalence Theory: 22
  • Naked Pairs: 18
  • Hidden Pairs: 17
  • Locked Candidates Claiming Row: 6
  • Locked Candidates Pointing Row: 4
  • Naked Triples: 4
  • ALS-XZ: 3
  • Conjugate Contradiction Scout: 3
  • Locked Candidates Claiming Col: 3
  • Unique Rectangle Type 2B-Oddball: 3
  • Finned Jellyfish: 1

SET changed the logical landscape

The latest important addition to the solver deck was Set Equivalence Theory. Removing its former cap of 12 eliminations allowed Human Solver to recognize the full value of the pattern on Everest. SET removed 22 candidates and displaced 21 steps previously attributed to Forcing Nets.

That result did more than improve the running time. It demonstrated that the “Three Easy Steps” approach to Everest—SET, focused contradiction, and solving to the end through cascaded Singles—provides a genuine strategic advantage. Everest did not become a different puzzle. The available logical understanding changed.

So which puzzle is hardest?

“Hardest” can mean the most logical actions, the greatest number of advanced strategies, the most contradiction proofs, the fewest accessible openings, or the puzzle that forces an experienced human to request the most assistance. Those measurements do not always choose the same winner.

My conclusion: under the present Sudoku Paradise methods and benchmarks, I place the Golden Nugget first among these three. Everest demonstrates the most sustained recorded work, but SET transformed its break-in and made it the fastest machine benchmark. AI Escargot remains formidable and substantially slower than Everest under the current deck. The Golden Nugget remains the puzzle that caused me the most human difficulty.

The larger lesson is that no “hardest Sudoku” title is permanent. Difficulty is partly a property of the puzzle and partly a measure of the strategies we have not yet learned to see.

Try the evidence yourself

  1. Copy one of the 81-character puzzle strings above.
  2. Open Sudoku Paradise and choose the Custom Puzzle option.
  3. Paste the grid, analyze it, and work as far as your own logic carries you.
  4. Request a Hint only when you need a reproducible continuation.

This comparison was written from the Sudoku Paradise verification audits and Andy Beaudry’s manual field tests. The solver library was actively evolving during the study, so the latest detailed audits supersede the shorter action counts in the original forum opening. Solver timing varies by machine; the repeatable evidence is the recorded logical path.