Skip to content

Search Roadmap & Evaluation

The engine includes five single-turn heuristic bots and the rollout-based MonteCarloSearch. The Scala search API also provides configurable two- or three-ply expectimax with Star1/Star2 pruning, Zobrist hashing and transposition caching. These are implemented; they are no longer future stages of the roadmap.

See Expectimax Search for the algorithm and Project Status & Roadmap for the broader delivery overview. The default npm bot roster does not include expectimax.

CapabilityStatusReference
Full-turn heuristic botsImplementedPrimitive bots
Monte-Carlo pre-roll equity and rollout botImplementedMonte-Carlo equity
Clock-aware search budgetsImplementedTime management
Expectimax, Star1/Star2 and transposition tablesImplemented, configurable through the Scala APIExpectimax
ONNX evaluation and model serving contractImplemented on JVM; models supplied by the hostONNX integration, model contract
Parallel chance-node evaluation with OxProposed, not implementedIssue #61

Implementation and evaluation answer different questions. A configurable deeper search must still demonstrate useful move quality within a target application’s time and memory budget. Do not infer playing strength, deployment status or hardware requirements from a completed implementation alone.

Future changes should preserve full-turn legality, reproducible testing and bounded runtime. Compare both move quality and the cost in time, memory and implementation complexity.

GreedySearch (Level 3) remains a simple reference bot with random tie-breaking. Every experiment should explicitly name its baseline, even when comparing against a stronger configured search. The protocol below applies to changes in existing algorithms as well as new ones.

Every new algorithm or optimization must be validated at three levels.

These tests prove that the algorithm respects core game semantics and obvious tactical priorities.

  • immediate king capture is always preferred over any non-terminal material sequence
  • the maximum micro-moves rule is never violated
  • castling still consumes the correct dice and is scored correctly
  • promotion branches are ranked consistently

This level protects correctness and prevents regressions that would otherwise be hidden inside large simulations.

The implemented scenario suite is a small, versioned catalog of fixed positions designed to expose evaluator and shallow-search differences without the noise of a full match. The bundled catalog lives at arena/src/main/resources/search-evaluation/core-v1.json and covers:

  • tactical decisions, including immediate king captures and material choices
  • defensive decisions around king exposure and blocking attacks
  • endgame decisions, including promotions and sparse captures
  • forced passes when none of the remaining dice can move

Each scenario has a stable ID, category, rationale, DFEN (including the remaining dice pool), and either a set of allowed turn paths or an expected pass. The catalog also names an explicit seed set. Every candidate and baseline decision is checked for legality and expectation matching, so a report distinguishes improvements, regressions, shared matches, and shared misses. The focused fixture and reproducibility tests run as part of mise run check.

Run the default comparison (aggressive against greedy) with:

Terminal window
mise run arena:evaluate

The task accepts positional bot, baseline, and fixtures arguments:

Terminal window
mise run arena:evaluate aggressive greedy arena/src/main/resources/search-evaluation/core-v1.json

The runner prints a human-readable row for every scenario and seed. For an additive-stable, machine-readable JSON report, invoke the runner directly with --json:

Terminal window
sbt 'arena/runMain dicechess.engine.bench.SearchEvaluationRunner --bot aggressive --baseline greedy --json target/search-evaluation.json'

The JSON report records the fixture and seed-set identities, both bot identities, per-run decisions and scores, legality and expectation results, comparison classifications, and aggregate totals. Preserve fixture IDs and seed-set IDs when comparing archived reports; add a new version when their meaning changes.

This is the acceptance layer for comparing playing strength. Each challenger algorithm plays a large head-to-head match against the baseline:

  • baseline as White, challenger as Black
  • challenger as White, baseline as Black
  • fixed seed lists for reproducibility
  • repeated across multiple starting positions, tactical middlegames, and simplified endgames

The benchmark harness should follow a stable protocol so that results remain comparable over time.

  • deterministic PRNG seeding
  • explicit algorithm identity and Git commit hash in output reports
  • symmetric color assignment (equal games as White and Black)
  • turn/time budget settings recorded in the report
  • fixed resignation, repetition, and 50-move limit rules

At minimum, collect:

  • win, loss, and draw rates
  • average game length (turns)
  • average decision time per turn
  • average number of candidate paths or nodes evaluated
  • runtime distribution (percentiles), not just mean runtime

A challenger should replace the baseline only when all of the following are true:

  • it passes unit correctness and scenario-suite checks
  • it shows a statistically meaningful improvement in head-to-head results (Win Rate > 50%)
  • it does not introduce an unacceptable runtime regression for the target environment (e.g. browser/JS vs server/JVM)

Each experiment should produce a concise report:

================================================================================
🎲♟️ Dice Chess Bot Arena - JVM Match Runner
Baseline Bot: [Name] (ID)
Games per Color: [N] (Total [2N] games per match)
================================================================================
Opponent Bot | Total | Wins (W/B) | Losses (W/B) | Draws (W/B) | Win Rate | Time
----------------------------------------------------------------------------------------
[Bot A] | 1600 | ... | ... | ... | ...% | ...s

This makes search changes reviewable inside pull requests instead of relying on anecdotal observations.

Search evaluation reports and fixtures are implemented. Use the open issues for current work rather than recreating the completed expectimax, pruning or hashing stages. The project status page links the release history and live milestones.