Skip to content
codebiy
← Back to the writing desk

What a prompt cannot guarantee: three LLM rules we enforce in code

Field lengths that must pass a save schema, an #ad label that must be on every sponsored caption, and a public demo that must not invent claims. None could be left to a prompt. This is where each one is enforced now.

Spoolt reads a website and writes short-video scripts from it in seven formats. The scripts come from OpenAI's GPT-5.4 family (gpt-5.4-nano) through the Responses API, with a Zod schema as the output format (openai SDK 6, Zod 4). The system prompt behind them is a short block of rules, and every rule in it is a request: the model follows it most of the time.

For three things, most of the time was not acceptable. A saved brand profile has to pass our own schema. A caption that features a sponsored product has to carry its label. The public demo must not make a claim that the page does not make. Code now guarantees all three, and the prompt keeps its sentence only to make the enforced case rare.

From a website to a script

We load the page with a fetch that refuses internal addresses and hand at most 12,000 characters of its text to the model. It returns a brand profile as structured output, which the user reviews and can edit during onboarding.

Each script is then one call: a fixed system block, a short brief for the format, a target length (30 seconds by default, at about 2.5 words a second) and the brand and product data. A batch of drafts is several of these calls, run with Promise.allSettled. One failed call does not discard the drafts that came back and were already paid for; the batch fails only when none did. A second call gives each draft a score, which the app shows as a badge. If that call fails, the draft has no badge: the user paid for the drafts, not the badge.

Every answer comes back through a schema, and we parse it once more with the same schema on our side. A schema describes shape. It does not know what is true or what has to be disclosed, and it did not settle length for us either.

Field length: clamped in code, not in the output schema

The first failure looked like a form bug. In August 2026 the review step of onboarding tried to save a brand profile that the analysis call had just produced, and the save was refused with 400 Too big: expected string to have <=100 characters. The field was brandValues, which the user had never typed.

The analysis call's output schema had no limits, while the schema the review step saves with capped every field. Any talkative answer from the model produced a row that could not be edited afterwards.

The obvious repair is to copy the caps into the output schema. OpenAI's guide does not promise that they would hold there. Structured Outputs support a subset of JSON Schema, and the guide's supported string keywords are pattern and format: no length keyword, as of October 2026. Our script call's schema has carried a maxLength on its caption field since September 2026. We have not tested what the API does with it, so nothing of ours depends on it. The fit is made at our boundary instead:

export function clampText(value: string, max: number): string {
  if (value.length <= max) return value;
  const cut = value.slice(0, max);
  const lastSpace = cut.lastIndexOf(" ");
  return (lastSpace > max * 0.8 ? cut.slice(0, lastSpace) : cut).trimEnd();
}

The limits now have one definition per field, used by both the save schema and the clamp. The cut lands on a word boundary when one is close to the limit, so a trimmed value still reads as language. The prompt states the budget too (“up to 20 short phrases, each under 160 characters”), which makes the clamp the backstop. One cap moved from 100 to 160 characters, because real values carry a parenthetical and 100 cut them mid-phrase.

Illustration: a blank folded note beside a shape-sorting board, where a round and a square block sit in their cut-outs and a red star lies on top. A prompt can ask for a rule, but only code makes it hold.

The 400 cannot come back on this path: a check builds an over-long profile, clamps it and asserts that the save schema accepts the result. The clamp does not log, so we cannot say how often it cuts in production.

Whatever a provider does with a length keyword, the row has to pass our own save schema, and only our own code can promise that. A bedtime-story app of ours draws the same line after its own model calls: what code checks after structured output.

The #ad label: appended in code

A script can feature a product, and a product can be sponsored. The prompt covers that in one sentence: when a product has a sponsor name, disclose the sponsorship naturally in the caption. In July 2026 we saw the model skip that sentence often enough that a legal disclosure could not ride on prompt compliance. We did not count the misses.

So the tag is appended in code:

export function withSponsorDisclosure(
  caption: string,
  sponsored: boolean,
): string {
  if (!sponsored) return caption;
  if (/#ad\b/i.test(caption)) return caption;
  return caption ? `${caption} #ad` : "#ad";
}

It runs when drafts are saved, for every draft of a batch that features a sponsored product, and it does nothing when the model already wrote #ad. The sentence in the prompt stays, because a caption that names the sponsor in its own words reads better than a bare tag. The tag no longer depends on it.

This is the label in the caption's text. TikTok's own commercial content disclosure is a separate setting that the user chooses for each post, as described in what TikTok's Direct Post audit required of us. Which label a given market expects is not something this article settles.

The public demo: no model, so no invented claim

The landing page lets a visitor paste a URL and see three sample openings before signing up. When it shipped in August 2026 it was a model call with everything a public endpoint needs around it: a CAPTCHA, per-IP and global daily budgets that refuse the request when they cannot be checked, bounded input and output, a 20-second timeout and the same fetch the app uses.

It failed in a way none of those cover. A free, unauthenticated demo pays a provider per visitor, so the page was only as available as the provider account behind it. When that account could not serve requests, OpenAI answered 429, and our route mapped every 429 to its “capacity” state. Visitors read “The studio is at capacity” next to a countdown: “Try again in 53s”. No retry could clear it. The limit was on the account, not on the rate of requests.

In September 2026 we took the model out of the demo. The three openings are composed from the scraped page: the brand name from the title, the summary and captions from the site's own sentences, the opening lines from fixed templates per language. The response schema, the budgets and the CAPTCHA did not change. The demo now costs nothing per visitor. The loading text used to promise up to 40 seconds and now promises a few, and the “capacity” message can only come from our own budget.

One side effect mattered to us more than the fix. Text assembled from the page's own sentences cannot make a claim the page does not make, and a prompt could only ask for that. Composed text has failure modes of its own: on the first day real pages gave us a label list scraped as one line (“Passport Passport Pasaporte Pasaporte”), a navigation menu glued to the front of a sentence, and pages made only of unpunctuated fragments. Each got its own filter the same day.

A line under the result tells the visitor that the openings came from the site's own sentences, and that a model writes the script once they are in their account.

What the prompt still decides

Everything else is still a request, and most of it is about evidence. Use only supplied product facts, and if evidence is missing, omit the claim. The speaker is “a presenter reading brand copy, not a verified customer”: offering to show something is fine, claiming to have bought, tried or used the product is not.

We also changed what we ask for. A batch of several drafts gives each draft a different opening direction, so the batch is a real variant test and not several samples of one prompt. The first list of six directions, from July 2026, included “open with a surprising fact or specific number” and “open with a relatable first-person (POV) moment”. A page does not always contain a number, and the presenter has no first-person moment to relate, so both asked for the thing the evidence rules forbid. Since September 2026 every direction changes the presentation and never the available evidence. One of them reads: “open with one concrete product detail present in the supplied evidence”.

A website is someone else's text, and it can contain a sentence addressed to the model. In the script call the brand and product data travel as a JSON user message and are never spliced into the instructions. The system block tells the model to treat those fields as material to write from, and to leave alone anything in them that reads like a command. That is a request like the others. What bounds a hostile page is the call itself: neither the analysis call nor the script call is given tools, nothing secret is in either input, and what comes back is text in fixed fields that a person reviews.

Where a rule goes: prompt or code

  • Style, voice and structure go in the prompt, and we expect them to hold most of the time.
  • If a save, a bill or a disclosure depends on a rule, code enforces it after the model. The prompt keeps the sentence only to make the enforced case rare.
  • Someone else's text travels as data, and the prompt says so. The limit on the damage is in the call: no tools, no secrets, a draft that a person reviews.
  • A free, public feature gets one question first: does it need a model at all?
  • When a call fails, the user keeps what already succeeded.

Drawing this line in an existing product is the work we describe under AI integration.

KEEP READINGChecking AI agents' front-end work: green build, wrong page ↗Twenty OpenAI calls behind one button, and a deadline for each ↗