GPT-5.6 Terra’s final Rewind Runner build had zero page errors and passed its rewind verification 3/3 times. Then a first-time player spent several minutes unable to clear the first hazard.
That is the useful, slightly humbling distinction in this small game-building case study: an automated checker can confirm that time moves backward. It cannot tell whether a person can see where to jump next.
Disclosure: I work with OrcaRouter and used it to run this evaluation. One OpenAI-compatible key gave me access to every model in this test, without changing how any model answered.
AI-generated illustration featuring official model logos; logos and model names are used descriptively and remain the property of their respective owners.
Explore GPT-5.6 Terra on OrcaRouter.
What the verifier saw
The controlled portion covered GPT-5.6 Terra and Kimi K3 only. Each received one run per round across four ordered prompts, using vendor-default parameters. This was n=1 per run or attempt: a case study, not a benchmark or a general ranking.
The rewind test was deliberately mechanical. A headless browser sent fixed key presses, fingerprinted changing pixels, and checked whether matching pre-rewind frames moved backward. A build passed when three replays cleared a monotonic score of 0.90. In Terra’s final game, the check found no page errors and all three rewind replays passed.
That result matters—but narrowly. It shows that the rewind mechanic responded and retraced game mechanics under the test. It does not test whether a player understands the level, spots a safe route, or finds the timing forgiving enough to execute.
What a player saw
The human review supplied precisely that missing perspective, though it was much weaker evidence in another sense. It involved one non-blinded player who knew the model identities, played for a few minutes, and used no formal rubric.
That player could not get past Terra’s first hazard and described it as a difficult, fast two-part jump. They also explicitly acknowledged it could have been a skill issue. So this is a play report, not proof that Terra generated an unbeatable level. The original brief requested a beatable route; the failed attempt is therefore a useful follow-up question about readability and tuning, not a verdict on whether a route exists.
Automation and play are measuring different things
All automated metrics treated the three final builds in the broader review as equivalent. The same player, however, subjectively preferred the builds in the order Opus 5, Kimi K3, then Terra.
That preference should not become a leaderboard. Claude Opus 5 was tested separately through Anthropic’s native messages endpoint, with the entire brief provided at once and a time-boxed feedback loop. Its setup was not comparable with the controlled Terra/Kimi runs, so Opus 5 should not be ranked against either of them.
The more modest conclusion is practical: use a behavioral verifier to catch broken mechanics, then let people play before calling a generated game finished. The former can establish that rewind works under fixed inputs; the latter can reveal that the first challenge feels like a wall.
This is a first-party Rewind Runner test record, not a universal ranking of models or a measure of whether every player will enjoy or complete these games.
Practical takeaway
For builders evaluating AI-made interactive work:
- Treat a passing automated replay as evidence that a specific mechanic works under a specific test.
- Add human play sessions to check route discovery, pacing, and difficulty spikes.
- Keep player impressions separate from proof: one stuck player can identify a problem worth investigating, not settle whether a level is possible.
Limitations
This was a small case study: one run or attempt per model and prompt in the controlled test, not a benchmark. The human review was one non-blinded player without a rubric. The rewind checker tested state reversal rather than enjoyment or solvability. And Opus 5’s separately run setup is non-comparable to the controlled Terra and Kimi K3 work.
Sources
First-party Rewind Runner evaluation records: controlled model-run records, rewind-verification records, and the first-player play report. No external sources were used.
This evaluation was run through OrcaRouter. The author works with OrcaRouter; model access does not imply affiliation with, endorsement by, or sponsorship from model providers. Model names and logos are used descriptively. All trademarks belong to their respective owners.