Play. Verify. Evolve.
"A poem is never finished, only abandoned." — Paul Valéry
Neither is a generated game. RSIGame keeps developing what it creates — autonomously exploring, improving, and verifying the game until progress saturates, then preserving the best version or opening a new stage of evolution through high-level guidance.
rsigame_demo_78s_1080p.mp4
Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions.
We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development.
Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training. Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.
Left is the initial generation, right is what RSIGame delivers. Same game, same scripted inputs, replayed on both builds.
The inner loop, once per round:
| Explore | an agent plays the executable game and records what actually happens — logs, metrics, screenshots |
| Diagnose | that evidence becomes one concrete objective: the place with the most improvement headroom, not a guess |
| Edit | a repair agent changes code or content against that single objective |
| Verify | an independent agent replays the game: did the change land, and did anything that worked before break |
The outer loop, across rounds. A Global Quality Monitor scores progress, holds the globally best checkpoint, and detects saturation — three consecutive checkpoints where nothing beat the held build. Saturation is not the end: sparse high-level guidance, from a person or a stronger model, opens the next stage.
Mean Overall of the build a budget of k rounds delivers. A free-form playtest-and-revise baseline finishes a strong base where it started; the Global Quality Monitor is what keeps late rounds from undoing earlier gains.
git clone <this repo> && cd rsigame
pip install -e .GameCraft-Bench supplies the tasks, rubrics and replay harness. It is a separate checkout; install it into the same environment and point the config at it:
pip install -e ../gamecraft-benchYou also need Godot 4 for the Godot line, and Node 20+ with an OpenGame checkout for the Phaser line.
cp .env.example .env # one model API key is enough
$EDITOR configs/default.toml # [paths]: the bench checkout and the Godot binary
rsigame config # what a run would actually use, resolved
rsigame check # paths, keys and models, before anything runs
rsigame develop <task> -c paper_godot_gptrsigame check is worth running first: a missing bench path or key used to
surface at round 12, not at round 0.
Three layers, each overriding the one above:
configs/default.toml— every knob and the default the code itself uses-c configs/experiment/<name>.toml— what one experiment changes- environment variables — a temporary override mid-run
Every variable this project owns is RSIGAME_<SECTION>_<KEY>, matching
[section] key in the toml exactly. Secrets are read from .env only, never
from a toml.
The monitor decides which build is the best seen so far, and the saturation stop decides
when development has stopped paying: K consecutive checkpoints with an unchanged
champion. [monitor] live picks how it runs:
| what happens | what a run costs | |
|---|---|---|
live = false (default) |
every round runs; scripts/vm_replay.py replays the checkpoints afterwards and python -m rsigame.monitor.value_stop reads where the run would have stopped |
the full budget, every arm the same — this is what the paper reports |
live = true |
the monitor judges each checkpoint as the run goes and ends the run when the saturation stop fires | fewer rounds, and arms no longer share a budget |
A live monitor never kills a round mid-flight, never rolls the tree back, and drops itself for the rest of the run if the estimator fails — an outage must not look like saturation.
Every model call the run makes goes through a loopback proxy, so the accounting covers the whole loop — Controller, Explorer, Verifier and Monitor, not only the repair subprocess. Two files land beside the rounds:
| file | what is in it |
|---|---|
all_token_calls.jsonl |
one line per call: endpoint, model, prompt/completion/cached/reasoning tokens, cost |
all_token_totals.json |
the same summed, written when the run ends — including a run that stopped early |
The proxy forwards to one upstream and signs with one key, so it only takes over an endpoint that already points at that upstream; anything left unmetered (a monitor pointed at its own provider, say) is named in the run log.
Both engines are scored by the same GameCraft-Bench rubric, Overall = BUILD × (0.15·Mechanics + 0.35·Depth + 0.15·Visuals + 0.35·Art), by replaying scripted demos and judging the recordings.
python -m rsigame.eval.score_game --project <tree> --game <task> --output <dir> # Godot
scripts/score_checkpoints.py <run root> # Godot, every Nth round
scripts/score_checkpoints_web.py <run root> # Phaser, every Nth roundThe engines differ in about 900 lines out of 48k: the build gate (Godot's
headless import against the web build check on dist/), how a demo is replayed,
how a project tree is read, and the engine-specific half of the repair prompt.
Everything else is shared.
src/rsigame/
loop.py one development round: explore → diagnose → edit → verify
project.py a game project: copy it, build it, import it, read it
cli.py config.py paths.py preflight.py
planning/ what a round is for, and what it will work on
checklist/ what the task asked for, and what is true of it now
view/ what the build contains, read from the tree
controller/ how the work is chosen and how the repair is asked for
evidence/ what was observed, frozen so a later round can cite it
polish/ the quality pass: the largest remaining bottleneck
verify/ replay a repair and decide whether it earned its commit
monitor/ the champion across rounds, and the saturation stop
eval/ scoring, the proxy estimator, pairwise judging
agent/ the exploration arm: sessions, probes, the repair agent
review/ the human review tool used for outer guidance
data_pipeline/ recorded sessions -> a supervised fine-tuning corpus
training/ the corpus -> a LoRA adapter, merged for serving
configs/ every knob, with its default and a comment
scripts/ scoring and replay entry points
patches/ the changes this work applies to GameCraft-Bench
Base games, run trees, recordings and scores are not in this repo. The scoring artefacts every number in the paper is read from are on Hugging Face at RSIGame/RSIGame-TableArtifacts: for every row of Tables 1 and 2, the base game, the build that was scored, and the judge's records.
Coming soon. The base games, run trees, supervised corpus and model weights are going through our company's internal review, and will be released under RSIGame as soon as it is approved.
Until then, data_pipeline/ documents how the corpus is built from recordings
and training/ how one arm is trained, so both can be rebuilt rather than taken
on trust.
tests/test_no_leaks.pyruns on every commit: no credentials, no absolute machine paths, no internal endpoints, no personal emails, no legacy variable names.tests/test_imports_resolve.pyresolves every internal import at any nesting depth, including imports inside functions — the kind a passing test suite otherwise hides until a real run reaches that line.- Two prompt strings in
eval/estimator/are Chinese and stay that way: they are sent to the judge, and the paper's numbers were produced with that wording.
Apache-2.0. This work builds on OpenGame
(Apache-2.0) for the Phaser line and on
GameCraft-Bench for
the tasks and the scoring harness; see NOTICE. No source from either is
redistributed here.
@misc{wu2026rsigame,
title = {RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement},
author = {Wu, Wenyi and Fu, Minghao and You, Jieyu and Zhou, Kun and Liu, Siqi and
Salvi, Aayush and Lin, Yiheng and Zhang, Ce and Lan, Xiaohan and Zhu, Jiahui and
Zhong, Yujie and She, Qi and Huang, Biwei},
year = {2026},
eprint = {2609.39045},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.39045}
}







