What it takes

Every bot, advisor or analysis tool for Slay the Spire is built from the same layers. This page walks through them in order, with the projects that have tried each one and the numbers they report.

1. Read the game

A program has to see what a player sees and act the way a player acts. Neither game ships an API, so this layer is always a mod.

Slay the Spire

ModTheSpire loads mods and BaseMod provides the modding API. On top of them, Communication Mod is the standard bridge: it starts your program, sends it the game state as JSON whenever a decision is due, and reads commands like play, choose and end back. spirecomm wraps that protocol in Python objects. Nearly every STS1 bot in this index runs on these four pieces.

Playing the real game is slow, so there are mods for that too: SuperFastMode skips animations, and LudicrousSpeed with STSStateSaver let a search try a line in the real game and rewind it. MCPTheSpire exposes the same state to language-model agents over MCP.

Slay the Spire 2

The sequel runs on Godot with its game logic in C#, and it ships a native mod loader plus an official Workshop uploader. There is no single standard bridge yet. STS2MCP (REST plus MCP) is the most used, STS2-Agent offers HTTP and MCP with an in-game model runner, and several bots carry their own forks. Because the game is in Early Access, each bridge is tested against specific game versions and can break on patch day.

2. Represent the state

The bridge hands over a snapshot: hand, draw and discard pile contents, energy, HP, block, powers, enemy intents, relics, potions, the map and the current screen. A rules-based bot can read that directly. Anything that learns has to turn it into numbers. LearnTheSpire encoded a deck as a 1×234 card vector; sts2-rl-agent uses a 131-dimensional combat observation.

What the snapshot leaves out matters as much as what it includes. A player doesn't know the draw order, future card rewards or event outcomes, but a simulator seeded with the run's RNG does. A bot that uses that knowledge is playing a different, easier game. Projects should say which one they play, and the better ones do.

3. Decide

A full run is a chain of different decisions, and projects often use a different technique for each.

DecisionWhat's been tried
Combat Search dominates: Bottled AI simulates hand orderings and scores outcomes, sts_lightspeed and CombatSolver search inside simulators. Learning combat end to end has mostly failed, and sts-rl-agent documents its own failed attempts.
Card rewards, shops, upgrades Priority lists (Bottled AI), learned policies (sts-rl-agent, sts-ironclad-agent) and models of human picks (LearnTheSpire).
Map path Dynamic programming over the map (STS2FableBot), learned policies, and studies of how winners path (map entropy paper, pathing thesis).
Events, campfires, Neow Hand-written conditions in most bots; learned policies in the hybrid agents above.
Everything, by language model Prompting a model with the state and a list of legal actions (AgenticSTS, STS2-Agent, slay_the_spire_agent). Given typed options to choose from, small models reach floor 14 at A0, and fine-tuning on a search bot's choices takes them to 19 (typed-decision).

The strongest results so far pair a learned or hand-tuned out-of-combat policy with classical search in combat. See the approaches page for every project grouped by technique.

4. Test faster than real time

Search needs to try lines, and learning needs millions of games. Both need a simulator, and every simulator trades speed against fidelity.

  • sts_lightspeed (STS1, C++) is the reference: built to be RNG-accurate, 1M random playouts in 5 seconds on 16 threads per its README, but card coverage centers on Ironclad and colorless.
  • decapitate-the-spire (Python) and conquer-the-spire (C++) are partial reimplementations aimed at reinforcement learning.
  • sts2-cli sidesteps fidelity entirely for STS2 by running the real game engine headless, which needs a copy of the game and pins you to its version.
  • sts2-sim ports STS2 line by line with matching RNG and cloneable states for search; sts2-rl-agent trades fidelity for speed at about 1,200 combats a second in Python.

Whatever the simulator, check it against the real game. sts-ironclad-agent ships a parity checker, and better-run-logs records the RNG state of every action, which makes a good test set.

5. Measure honestly

  • State the character, ascension level and game version. A win rate means nothing without them.
  • Use a fixed set of seeds the bot never trained on, and report how many runs. Fifty runs leaves a 20% win rate anywhere between about 9% and 31%.
  • Say whether the bot can see anything a player can't, such as RNG state or upcoming draws.
  • Compare with people. In the 2020 slice of Mega Crit's run dump, 9% of 18 million runs were wins (Fox Row's analysis).

Where results stand

Self-reported by each project, without independent verification. Settings differ, so compare with care.

ProjectSettingReported resultApproach
sts-ironclad-agentSTS1, Ironclad, A20, real game50.1% Heart wins over 501 runsLearned policy + combat search
Bottled AISTS1, all four, ascension not stated20–52% wins, ~50 runs eachRules + combat search
sts-rl-agentSTS1, Ironclad, A0, simulator14% wins, floor 42.5, 50 seedsLearned policy + MCTS
STS2FableBotSTS2, A0 and A120.7% (343 runs) and 19.4% (175 runs)Search + scoring
AgenticSTSSTS2, A06 of 10 runsLanguage model + memory
typed-decisionSTS1, Ironclad, A0, 200 seedsFloor 19.5 after fine-tuningSmall language model

The hard problems

  • Hidden information. Draw order, rewards and events are random. Search with the real RNG is easier than the game; search without it needs sampling.
  • Long horizon. A run is 50-plus floors with one win or loss at the end. Card choices in Act 1 decide fights in Act 3.
  • Combat branching. Hands with many cards and targets grow the search tree quickly, which is why speed matters.
  • Simulator fidelity. Every reimplementation drifts somewhere, and a policy trained on the drift learns the wrong game.
  • Moving target. Slay the Spire 2 is in Early Access. Content and balance change between patches, and bridges and simulators pin versions.

Where to start

If you want to…Start with
Get a bot playing STS1 this weekendCommunication Mod + spirecomm, then read Bottled AI
Do the same for STS2STS2MCP or sts2-cli
Try searchsts_lightspeed (STS1) or sts2-sim (STS2)
Try reinforcement learningsts-rl-agent and sts2-rl-agent
Point a language model at the gameMCPTheSpire (STS1) or STS2MCP (STS2), then AgenticSTS
Analyze how people playDatasets
Build a seed or run toolSeedSearch, sts_map_oracle, Spire Codex