AscensionBench
AscensionBench measures how well language models play real Slay the Spire under one frozen interface, on private seeds, with replication.
Slay the Spire 1 · Ironclad, Ascension 0 · text interface · model only · edition 1 · fingerprint dd510451…
Results
| configuration | mean floor (95%) | win rate | cost per run | seconds per decision |
|---|---|---|---|---|
| DeepSeek V4.1 Flash (high) | 39.0 [35.6, 42.5] | 18% | $1.01 | 5.1 |
| GLM 5.3 Flash (high) | 29.9 [25.5, 34.2] | 5% | $0.17 | 5.3 |
| GPT-5.6 Luna (high) | 29.8 [26.4, 33.2] | 2% | $0.53 | 1.9 |
| MiMo V2.6 Flash (high) | 29.1 [24.6, 33.7] | 5% | $0.20 | 9.7 |
| GPT-6 Luna (high) | 20.6 [15.6, 25.7] | 5% | $0.29 | 2.2 |
Twenty private seeds, two attempts each, every attempt scored. Mean floor with its 95% interval; cost per run and seconds per decision (measured median) are per attempt.
Floor against cost
Each point is one configuration; the bar is its 95% interval. Cost is per run on a log scale.
In plain words
What the model gets
The text a human would see on screen at that moment: the fight, the cards, the map, the shop. The list of actions the game allows. One note it writes to itself and reads back next turn. Nothing else: no hints, no history, no replay.
Why the numbers hold
The seeds are held out and private. Every model plays each seed twice, served by its own vendor, through one frozen interface whose fingerprint is on this page. Every attempt counts, including the ones that stall. Intervals say what twenty seeds can and cannot separate.