Skip to content
codebiy
← Back to the writing desk

Balancing a game economy with six bots and Swift assertions

The first balance report for our market-stall game said no bot could pay off the starting debt. The cause was the bots: they bought too little and paid the asking price. Stock ended up dearer.

This article gives away how some of Stallkeeper's endings are reached.

Stallkeeper is our turn-based game for iPhone and iPad about running a stall in a Turkish neighbourhood market: haggling at the wholesale hall at dawn, laying out the stall, selling through the day, closing the ledger at night. The campaign lasts 24 days, the player starts with a debt of ₺60,000 that falls due in four instalments, and there are five endings. Whether an honest, average player should be able to pay that debt is a design decision, and nobody plays a 24-day campaign forty times after every change. Six bots do. Their first report said that no play style could clear the debt. The cause was the bots more than the game: they bought too little stock and paid the price the supplier asked.

The game is written in Swift 6 with SwiftUI, for iOS 18 and later. The tuning described here was done on 28 September 2026. Version 1.0 went to App Review the next day and is on the App Store.

Stallkeeper's ledger for day 1: ₺670 net to cash, an ₺8,000 instalment due in 5 days and ₺60,000 of debt left. The ledger that closes day 1. Every amount is in-game money.

A rules engine that runs without the interface

The rules live in a core engine that knows nothing about the interface. A run is a value, a move is a command, and Engine.apply(command, to: &state) returns the events that followed. The in-app purchases are cosmetic and never reach that layer, so the bots play the game every buyer gets. Money is an integer, typealias Kurus = Int, counted in kuruş, a hundredth of a lira: ₺1,800 is written 180_000, and a balance never has a fraction.

Because of that separation the whole campaign compiles and runs from the command line. One script hands the core, the content and the bots to swiftc and runs the result; no Xcode project and no simulator are involved. Our thriller Upstairs is built the same way, which is how we checked a puzzle lock without opening the app.

A bot is one small protocol. It looks at the state of the run and picks the next command, or returns nil when it has nothing to do.

protocol Bot {
    var name: String { get }
    func next(_ s: RunState) -> Command?
}

A driver starts a run from a seed, the number that fixes every random draw, and asks the bot for commands until the days are played. The same seed and the same bot give the same campaign.

The six bots and their play styles

Each bot stands for a way of playing, and none of them is a difficulty level.

  • PassiveBot sets up an honest stall and then only waits. It serves nobody.
  • HonestBot is the honest, average player. It serves the queue and haggles with evidence, but it buys fixed quantities and never buys an upgrade.
  • PlannerBot is the good player. It checks the bottom of large lots and uses the rot it finds as evidence, restocks by what sold, buys the sign upgrade and pays the instalment the game suggests.
  • SelinBot is PlannerBot following one character's storyline to its ending.
  • CheaterBot cheats: a heavy scale, the best produce on top, made-up origin labels.
  • KenanBot earns like PlannerBot and accepts every shady offer on the way to the dark ending.

The first report and what we changed

The first report, on 20 seeds, said two things. No bot cleared the debt in any seed. And the good bot saved the market in 18 of 20 runs: on the petition that decides whether the market stays where it is, it collected 240 to 325 signatures against a target of 150.

First we closed three holes in the haggle, because a number tuned around an exploit has to be tuned again once the exploit is gone. Walking away and coming back no longer resets a negotiation. Showing a supplier the same evidence twice no longer moves their trust twice. And the price offered to a returning player now rounds against the player in integer math.

Then four things changed in the game:

  • The wholesale average went from 0.55 to 0.60 of the fair price.
  • On the petition, a buying customer now signs with a probability of reputation/1,100, down from reputation/200.
  • The visit rate of two regular customers went from 50% to 30%.
  • The suggested instalment now leaves ₺5,000 of working capital in the till. The earlier one paid the whole till into the debt.

The bots changed too. Until then PlannerBot had played like HonestBot with a cash reserve: it bought the same fixed quantities every day, at the price the supplier asked. That day it learned to play like someone who knows the game: haggling with evidence, restocking by sales, the sign, the suggested instalment.

So the two reports are not a controlled comparison. The rules, the best bot and the number of seeds all changed between them. The second report, on 40 seeds on 28 September 2026:

Bot Debt cleared Other result
PlannerBot 28 of 40 market saved in 21
SelinBot 29 of 40 her ending in 27
HonestBot 16 of 40 market saved in 12
PassiveBot 0 of 40
CheaterBot 0 of 40 fined in all 40
KenanBot 1 of 40 dark ending in 17

What made the debt payable

None of the four rule changes did. Dearer stock, a slower petition and rarer regulars make the game harder, and the new instalment suggestion does not help the good bot either: with the old one it clears the debt in 30 of 40 seeds, with the new one in 28. To find out what did, we ran the report again for this article in October 2026, on the rules of the released game, with PlannerBot's skills switched off one at a time and, in the last row, with the wholesale price put back to 0.55:

PlannerBot, 40 seeds Debt cleared
As released 28
Without the sign upgrade 23
Paying the asking price, no evidence 1
Stocking up only to the first report's fixed quantities 0
As released, with stock at the old 0.55 38

A bot that stops at the first report's quantities does not clear the debt in a single seed, even when it haggles well. What holds it back is the size of the order: HonestBot also buys fixed quantities, one and a half times as large, and clears the debt in 16 of 40. And a bot that restocks by sales clears it in one seed of forty when it pays the asking price. The first report's zero was a statement about the bots.

Illustration: six wind-up toy robots, one of them red, stand on one pan of a balance scale, and a stack of plain coins on the other pan keeps the beam level. Six bots, one per play style, show whether the debt can be paid.

The last row is why stock became dearer. Once PlannerBot played properly, it cleared the debt in 38 of 40 seeds at the old wholesale price, which was too easy. The plan had asked for the honest, average player to clear the debt in at least 75% of campaigns. With the numbers in front of us we set a different target that day: a good player clears it in 50–80%, and an honest, average one sometimes goes under. At 0.60 the good bot lands on 28 of 40.

Targets written as assertions

The report prints a table. What keeps the game from drifting is the block after the table, where each design goal is a range, and a range that does not hold fails the check. In the sample, share(n) is n divided by the number of runs, reached is the set of endings that any bot arrived at, and expect stops the check with its message when the condition is false. The messages are Turkish in the source and translated here.

expect((0.50...0.85).contains(planner.share(planner.cleared)),
       "PlannerBot clears the debt 50–85% (\(planner.cleared)/\(planner.runs))")
expect((0.30...0.75).contains(planner.share(planner.saved)),
       "PlannerBot saves the market 30–75% (\(planner.saved)/\(planner.runs))")
// …
expect(honest.saved <= planner.saved,
       "the petition is easier for the one who invests in reputation (PlannerBot)")
// …
for id: EndingID in EndingID.allCases {
    expect(reached.contains(id.rawValue), "ending \(id) is reached by the bots")
}

The target for the debt is 50–80%. The assertion allows 85%, a margin of two runs in forty. The block goes on:

  • The good player's median daily net is ₺1,800–3,200 in the first chapter and ₺3,000–5,000 in the last.
  • The passive player clears the debt in at most 35% of seeds. The honest, average one clears it in at least two seeds, in at most 70%, and never more often than the good player.
  • The cheater is fined in at least 80% of seeds, finishes at least 90% of them with a conscience score (the game's measure of honest trading) below 30, and never clears the debt.
  • A rare ending shows up in 5–20% of the good player's runs, and no bot stalls before day 24.

The released game still runs on these numbers: in October 2026 the report prints the same table on the current code.

A fine for the honest bot that came from a rule

HonestBot lied about nothing, and it was still fined about three times per campaign for a misleading origin label: ₺1,000 and a seizure each time. No constant was wrong. The quick-setup helper put yesterday's leftover lot and today's lot of the same product into one slot under one label, and when their origins differed, the inspector's rule counted that as a false label. The fix was a rule about slots: one slot holds one origin, and the helper lays out by product and origin. The inspector's rule did not change. On the released rules HonestBot is not fined for an origin label in any of its 40 campaigns.

An ending no bot could reach

A rare ending had the opposite problem. It asked for ₺25,000 in the till once the debt was cleared. In the 28 seeds where the good bot cleared the debt, its final cash ran from ₺994 to ₺21,111, with a median near ₺10,400.

Our first decision was to leave the threshold as a goal for a master player and to test the ending with a constructed state. We reversed that the same afternoon, because an ending that no simulated player reaches is a design error. The threshold went to ₺13,000, roughly the top 15% of the good bot's results. PlannerBot then reached that ending in 6 of 40 seeds, and the 5–20% band holds it there.

Where the checks themselves misled us

Two layers that measured the same thing drifted apart. A daily loss floor moved from −₺2,000 to −₺4,000 in the check script, while the XCTest that measured the same limit stayed at −₺2,000. The test broke one stage later, in the middle of interface work. The rule: when a threshold changes in the check script, the test that measures the same thing changes in the same commit.

The simulation was deterministic; the save file was not. The “same seed, same save” check failed because Swift encodes a Set to JSON in a different order in every process. Sets that enter the save are now stored as sorted arrays and the encoder uses .sortedKeys. The check script runs the same bot campaign in two separate processes and compares the hashes of the two saves.

A failing check was silent. The script is compiled with -O, so a failing precondition printed no message, only “Trace/BPT trap”. Failures now go to stderr through fputs, and the process leaves with exit(1).

What the bots cannot show

The bots prove numbers: that a careful strategy can clear the debt, that a cheating one is caught, that every ending can be reached. They do not show that haggling is fun, or that a person finds the strategy we gave PlannerBot.

The order we would follow again

  1. Keep the rules runnable without the interface, and keep money an integer.
  2. Write one bot per play style, including the one who does nothing and the one who cheats.
  3. Close the exploits before tuning any number.
  4. Make the best bot play well before trusting what the report says about the game.
  5. State each goal as a range and make the range an assertion.
  6. Assert that every ending is reached. An unreachable ending is a bug, not a goal.
  7. When a bot loses, read the rule and the helper before the constant.
  8. Make failures print their reason, and compare the same run across two processes.
KEEP READINGPuzzle-game bug in Swift: the lock that refused the right code ↗Adaptive music with AVAudioPlayer: three loops on one clock ↗