Spoolt, our web app that turns a website into short videos for TikTok, renders each script into a 1080×1920 video with Remotion on a server we run. The first version rendered inside the Next.js web process: submit() ran the whole render and returned the finished result inline. Five different failures followed, three of them before Spoolt had a production image. The fix for the last one, on 24 August 2026, changed the design: one serial worker with a PostgreSQL table as its queue.
All of this ran on Remotion 4 and Next.js 16, with PostgreSQL 17 as the database. The production image, which dates from 9 August 2026, is Node.js 22 on Debian bookworm.
Why the render started inside the web process
Remotion became the default renderer because it is self-hosted. OpenAI writes the script (what that prompt is and is not trusted with is a separate article), Remotion renders the pixels, and no render API charges per video. A hosted render API stayed in the code as an alternative that an environment variable selects.
Remotion's page on using @remotion/renderer in Next.js calls the combination “a bit tricky” and, on a server you host yourself, “not officially supported”. Our next.config.ts lists @remotion/bundler, @remotion/renderer and remotion under serverExternalPackages, so Next.js loads them with Node's own require and does not bundle them.
The rest of the pipeline had been written for the hosted kind of provider: submit a job, receive a job id, then wait for a webhook or poll. An in-process render does not have that shape. It is one long call to renderMedia(). So we let submit() return the terminal result, and gave the provider a poll() that always answers “failed”: if anyone has to ask, the process that was rendering must have died.
The five failures of the in-process render
A render is paid for in credits, which Spoolt reserves when the render is requested and gives back if it fails. Two pieces of code besides the renderer look at unfinished jobs: a scheduled sweep that runs every minute, and a reconcile that runs when a user opens their library.
- The poller failed renders that were still running.
poll()cannot see a render in flight, so it reported failure, and the reconcile called it. Opening the library while a render ran was enough to fail a render that was about to complete, with the errorRemotion render did not complete. The fix was a grace period: the sweep and the reconcile leave a self-hosted render alone until it is fifteen minutes old. - A retry paid for the same images twice. The story format generates its scene images before it renders. When the render failed after that step, the failure result came back without the image URLs, so a retry generated and paid for every image again. The failure path now returns the scenes too, and image generation keeps partial results: a retry after a failed render pays only for the scenes that failed.
- Our own SSRF guard blocked us. A finished file is re-hosted to object storage by fetching it from its media URL. Our fetch guard refuses internal addresses, as it should. On a local deployment with object storage configured, every render therefore failed at the last step, while fetching the mp4 it had just written. Own-origin media URLs are now read from disk, behind the same path check the media proxy uses.
- Crashed renders became zombies that kept the credit. A self-hosted render gets its provider job id only when
submit()returns, so a render that died with its process never received one. The sweep and the reconcile both skipped rows without an id. The draft stayed ingeneratingfor good and the credit was never given back. The sweep now fails such a job once it is past the grace period, through the normal refund path. The rule looks only at state and age, so rows that were already stuck needed no separate repair: the first sweep after the deploy applied to them too. - A deploy killed whatever was rendering. The render lived in the web process, and a deploy replaces the web process. On 24 August 2026 one such render sat in
generatingfor thirty minutes, until the grace period (raised to thirty minutes that day) ran out, and then failed.
The first three surfaced on the night of 2 July 2026, within an hour of adding the story format, whose renders run for minutes. Spoolt's production image dates from 9 August 2026, so no user met them. We dealt with the other two on 24 August 2026, when we went through the whole render path.
One render worker and a PostgreSQL table as the queue
The web route now does two things. It flips the draft to generating with a compare-and-swap, so two simultaneous requests cannot both queue a paid render, and it inserts a queued row into the jobs table.
A second container runs the same image with a different command. It is the only process that renders, and it claims work like this:
/** Atomically claim the oldest queued job; null when the queue is empty. */
async function claimNext(): Promise<string | null> {
const [claimed] = await db
.update(generationJob)
.set({
status: "processing",
startedAt: new Date(),
attempts: sql`${generationJob.attempts} + 1`,
})
.where(
and(
eq(generationJob.status, "queued"),
eq(
generationJob.id,
sql`(select id from generation_job
where status = 'queued'
order by created_at limit 1
for update skip locked)`,
),
),
)
.returning({ id: generationJob.id });
return claimed?.id ?? null;
}
The claim is a single UPDATE, so it is atomic without help. FOR UPDATE SKIP LOCKED is there for a second claimer: PostgreSQL's documentation of the locking clause says that it skips rows it cannot lock at once, and that this can be used “to avoid lock contention with multiple consumers accessing a queue-like table”. With one worker there is no second consumer, so today the clause is insurance. The queue is a table in the database we already had, and the worker polls it every two seconds when it is idle.
One worker claims the oldest queued job and renders one at a time.
When the worker starts, it looks for jobs that are processing without a provider job id. There is exactly one worker and this is it, so nothing can still be rendering them. A job gets two attempts: the first interruption puts it back in the queue, the second fails it and refunds the credit, so a render that kills the process cannot put the worker into a crash loop. The reclaim assumes a single worker replica. A second one needs a lease column first.
That is also how the fifth failure ended. The worker is built from the same Dockerfile as the web container, so a deploy that rebuilds the image restarts it too, and a render in flight is still cut off. It is no longer lost: the new worker queues the job again and renders it from the start. One job was stuck in production on the day the worker went out, and we left it to these rules.
Both containers mount the same media volume, so a file the worker writes is served by the web process and survives a redeploy of either.
Which process may fail a job, and when
Moving the render out of the web process made another question sharper. Three parties can now end a job: the worker, the scheduled sweep, and the reconcile on read. If their rules overlap, one of them fails a job that another is still working on.
The difficulty is the one behind the fourth failure. From outside, a render that is still running and a render whose worker died look the same: processing, no provider job id. Call that an unfinished self-hosted render. Only its age tells the two apart.
| Who | Job | May fail it |
|---|---|---|
| Worker | the job it is rendering | when the render itself fails |
| Worker, on start | unfinished self-hosted render | on its second interruption; the first re-queues it |
| Sweep | unfinished self-hosted render | only once it is 30 minutes old |
| Sweep and reconcile | processing with a provider job id |
when a poll of the provider says it failed; a self-hosted render is not polled until it is 30 minutes old |
| Sweep | processing, any provider |
once it is 24 hours old |
| Sweep | queued |
once it is 30 minutes old, and only if no job at all was claimed in that time |
The thirty minutes began as fifteen. We raised it on 24 August 2026: a story render on a CPU-bound server can run past fifteen, and failing it would refund the credit and then discard the file that finishes minutes later. Thirty is a margin we chose. We have not measured our render times.
The last row is the one we got wrong first. For nine days the rule was simpler: a queued job that nobody has claimed for half an hour is dead. With one serial worker it usually is not, because a healthy backlog waits longer than that. Since 2 September 2026 the sweep asks whether the worker claimed anything at all in the window. If it did, it is busy. If it did not, it is down, and the queued jobs are failed so that the credits come back.
Each of the sweep's checks is a compare-and-swap pinned to the status it observed, so a job that the worker claims in between is not failed underneath it. Every path ends in the same function, which flips the draft and refunds the credit.
What Remotion needs from the process and the image
One render at a time. While renders still ran in the web process, two requests could start two of them at once, so we chained every render behind the previous one in a single promise. The rule was prevention: we have no out-of-memory incident on record and no memory measurements. One render is already parallel inside, because renderMedia() by default starts render processes on half of the machine's CPU threads. Now that a single worker takes one job at a time, the chain has nothing left to hold back. If throughput ever matters, the way out is a limit of two or more per host.
No standalone output. The image does not use Next.js output: "standalone". A standalone build copies only the files that tracing finds, and the Next.js documentation says that tracing can miss required files. @remotion/renderer takes its compositor, ffmpeg and ffprobe from one package per platform, such as @remotion/compositor-linux-x64-gnu, and finds them at run time by joining a file name onto that package's directory. We have no failed standalone build to show for this, and so no error message: the image has shipped the complete node_modules since its first version. The worker needs it in any case, because it runs as a plain script outside Next.js.
Chromium's libraries and two fonts. The runtime image, node:22-bookworm-slim, installs the shared libraries Chrome Headless Shell needs on Linux. Fonts are not on that list. We add fonts-liberation and fonts-noto-color-emoji so that captions do not come out as empty boxes.
What we would keep and what we would change
- A table is enough of a queue for one worker. The hard part was deciding who may fail a job, and those rules need writing down before the second process exists.
- Every failure path ends in a refund. The user never sees the queue. They see whether the credit came back.
- Each shortcut has its limit written in a comment next to it: one worker replica, one render at a time, the complete
node_modules.
What we would change is the order. The in-process render saved one container and cost five fixes. For a job that spends minutes inside a headless browser, we would start with the worker.