An experiment · October 2026

Two models. Two prompts. Four slot machines.

How detailed markdowns impact outcomes in a slot machine game.

I wanted to explore how much detailed markdown files impact the outcome of games so... I gave Claude Opus 5.5 and Qwen3.8-max-0902 the same job twice:
1) simple prompt: "build a fun slot machine in three.js."
2) second prompt: prompt + a design doc about what makes slot-style games compelling and told them to treat it as authoritative.

You can play all four of them below, compare them side by side, see how they were built, and read the doc yourself.

There's a lot of "one shot" experiments as well as "heavy iteration" experiments. "One shots" are visually impressive because large error towards an objective goal appears to automate workflow - but one shots are often not very fun and therefore require massive rework which results in increased problems with scalability and customization. Heavily iterated projects with a human-in-the-loop allow for more control, but as the agent's context window fills up and is compressed we see a lot of strange decisions, hallucinations, poor technical decisions and a host of other problems.

I wanted to explore detailed game design PRIOR to handing an agent effort. Depending on the outcome it may change how we can approach game development harnesses and shift iteration to the documentation stage or, perhaps, separate game design iteration from agentic development.

Prompt A "build a fun slot machine in three.js"
Prompt B Prompt A + design reference
What jumped out

The document dramatically impacted game design

Caveat up front: this is one run per model per prompt... so n=1. In no way is this some kind of "benchmark." Unsurprisingly, the addition of specific documentation dramatically impacted what was built.

I am also curious to try meta-iteration (I made that up), essentially after iteratively building something for a few days/weeks it tends to degrade in terms of code quality and design intent. It gets much harder to add "simple" things because of all kinds of problems in code structure, agent documentation (memory), etc. Meta-iteration would be taking what was finished, turning it BACK into a markdown (or many markdowns), saving the assets used and remaking the game with a fresh session.

01

Experience and Production Quality were higher than I expected.

What was more surprising (to me anyway) was that the details and level of "fun" achieved by the first iteration was quite good. I've been thinking about/planning on taking an open source model and adding design training data to it in order to improve outcomes, but I suspect better planning documents could do as much or more. It's also A LOT easier and cheaper.

02

Focusing on Documentation vs Better tools/harnesses is likely higher value (in most cases).

There's TONS of "use our tools to make the game of your dreams" services/apps. I haven't used them all and I'm sure there's high variation, but I think most developers could get MUCH further improving documentation and then running a frontier model than they do today.

03

Qwen and Claude made very different games with the same information.

Both directed builds landed on Lemons that feed off Cherries, Bells that grow, wilds, reel locks and drafted upgrades. Claude kept going: a symbol bag you add to and trim, Bombs, Stars that copy their neighbor, Keys that open a vault, bosses and six rounds. Qwen's is a tighter 25-spin sprint. More isn't automatically better.

04

Documentation details can matter A LOT: e.g. Near-misses

A big part of slot machine design is the "near miss." Clever designers have gone to great lengths to make it look like you "almost won" even though nothing of the kind happened. In contrast to this diabolical behavior, the doc is pretty clear that near-misses have to come from the real game state. Claude's directed build slows the last reel based only on reels that have ALREADY stopped. Qwen's adds the extra spin only when it already knows you won big... so the pause gives the result away. I won't go so far as to say that Claude is "better" but it certainly executed against what I "wanted" more accurately.

05

Code size/complexity increased dramatically, but efficiently.

Claude went from 876 to 1,941 lines (44 → 115 KB). Qwen went from 674 to 858 (25 → 45 KB). The doc also says to keep the prototype small: 5–8 symbols, 5–10 upgrades. This is far less bloat than I get when I "iterate rapidly" with an agent one step at a time.

06

The one-liners have weird economies

Doing the math on the paytables: Qwen's Lucky Reels pays back about 125%, so the house LOSES and your credits drift up forever (unless you drop below 5, then you're just stuck). Claude's Lucky Sevens pays back about 82% but has a free, unlimited Refill button. Real casino slot machines are around 90%-98%, so out of the box Claude/Qwen with a simple "one shot" didn't really build a useful slot machine :). Either way, there's no real reason to keep spinning.

The arcade

Play them

Click a game to start it (heads up: they all make sound). Hit Side by side to put any two next to each other. You'll need to click inside a game before its keyboard shortcuts work.

Side by side runs two 3D scenes at once, so on a phone or an older laptop it might chug.

Head to head

How the four builds differ

Here's what's actually in each file. The return-to-player numbers come from the paytables and reel weights. The directed builds have no betting at all, so RTP doesn't apply.

Code review

Technical comparison: Claude one-line vs directed

I had a separate Claude instance read both Claude builds in full and compare them purely as code. (Qwen's builds weren't part of this pass.) Its short version: claude_basic is a well-built 3D toy. claude_basic_directed is a properly engineered game: about 2.5× the code, with a much cleaner structure, and it's mostly CHEAPER to render, not more expensive.

Size and structure

Claudeone-line (claude_basic)Claudedirected (claude_basic_directed)
File876 lines, 44 KB1,941 lines, 115 KB
CSS / HTML~60 lines~320 lines (HUD rails, side panels, pop-up windows, mobile layouts)
Game rulesMixed into the render code~370 lines of pure rules (SlotLogic, no rendering code)
Presentation~770 lines in one module~1,235 lines
Game stateOne global state objectRun state object, seeded RNG, event list

Basic is one script with everything mixed in. spin() picks the results, starts the animations and runs the tease check. evaluate() scores the lines and fires the celebrations. The reels are fixed 12-symbol strips, and an outcome is just three random numbers. It's simple and easy to follow, and each feature costs about as little code as it can. Its paytable works out to 82.4% return to player with a 37.2% hit rate (counting every possible outcome).

Directed has a real architecture:

  • The rules are kept apart from the rendering. resolveSpin() returns an ordered list of events (mimic → bomb → cell scores → echo → line → prism), and the presentation plays them back with async/await. That's how it can show cause and effect one step at a time, and it means the rules can be tested without a browser.
  • Seeded RNG (mulberry32), so a run can be replayed exactly. Basic uses Math.random().
  • Reels draw from a symbol bag instead of fixed strips. To make that work, each reel keeps a virtual endless strip in a Map keyed by position. It reuses 12 panels, swaps textures from a cache, and splices the result column in just before the stop. This is the hardest piece of engineering in either file, and the bag mechanic needs it.
  • There are also bosses, relics, drafting with rerolls, near-miss detection, an end-of-run analysis that tells you what went wrong, and one-time tips.

Runtime cost

Directed has far more going on, but basic makes the more expensive rendering choices:
Claudeone-lineClaudedirected
ShadowsPCF soft shadows with a 2048² map, so the scene is rendered an extra timeNone
Marquee bulbs~110 separate meshes, each with its own material (~110 draw calls)1 InstancedMesh
Coins / particlesOne mesh per coin, up to ~380 shadow-casting meshes on a jackpot1 Points draw call using a fixed 2,400-slot buffer
Reel faces36 tiles sharing 7 materials60 panels (5 reels × 12) with 60 materials, needed for the per-panel curvature dimming

Directed does waste some work:

  • updateParticles() (line 1221) loops over all 2,400 slots and re-uploads all four attribute buffers every frame, even with no particles on screen. Color and size only change when particles are emitted.
  • updateHoldBar() (line 1284) projects 7 points to the screen and writes styles on 5 DOM buttons every frame, because the camera drifts constantly.

Basic has one nice efficiency touch: the LED display only redraws its texture when the credits, bet or win value change (line 385).

Small issues in directed

  • teaseFor() (line 1456) counts a Star as breaking a line. Stars copy their left neighbor before lines are scored, so the game sometimes skips a tease it had earned. It errs toward under-teasing, so the near-miss stays honest.
  • A comment says the rules are "also loaded by the balance simulator," but there's no simulator in the folder, so the round targets look hand-tuned.

Verdict

Basic gets maximum polish for each line of code, but it's a dead end: rules, rendering and state are all tangled together. Directed spends its extra ~1,100 lines on things that matter: rules separated from rendering, reproducible runs, and an event list that drives the presentation. It also renders more cheaply, thanks to instancing, particles in a single draw call and no shadows. If I had to extend one of them, it would be directed without hesitation.
Section 21 of the reference

The design-test scorecard

The doc ends with ten questions you're supposed to answer "before declaring the game finished." So I answered them for all four builds. These are judgment calls and you may disagree; hover or tap a rating to see why.

Under the hood

Build details

The rules, symbols, upgrades and feedback for each build... and the bugs.

Inputs

The two prompts

Each model got each prompt once. For prompt B the doc was sitting in the model's working folder.

Prompt A · one line
build a fun slot machine in three.js
Prompt B · with design reference
build a fun slot machine in three.js Use compelling_slot_machine_design_reference.md as an authoritative design reference. Apply its principles when making gameplay, pacing, feedback, progression, and player-agency decisions. Do not merely summarize the document; use it to guide the implementation.
The document

compelling_slot_machine_design_reference.md

The full doc both models got for prompt B, unedited. Sections tagged "cited above" are the ones I refer to in the comparison and scorecard.

Loading the document…