Projects / Research

Slay the Spire

typed-decision-slay-the-spire

Benchmark where small language models choose among typed options on the same 200 A0 Ironclad seeds. Nimble-9B averaged floor 14.1 and Laya English 10.4; fine-tuning on a search bot's decisions raised them to 19.5 and 17.1.

View on GitHub

Each decision is shown in several option orders to check for position bias. The fine-tuning set is 39,884 decisions from the search bot. The README reports the gains and the hardware cost of each model.