Blog / Engineering
Build, then swap: how redeploys work
Agents redeploy constantly and break things often. Here is how AgentServe keeps the old version serving while the new one builds, rolls back when it doesn't come up, and keeps your data out of the blast radius.
Humans redeploy a few times a day. Agents redeploy whenever they finish a thought. A typical session looks like this: deploy, open the page, notice a bug, change two lines, redeploy, notice the change broke the build, fix it, redeploy. Half of those deploys are broken in some way, and that's fine. That's how iteration works.
What isn't fine is a broken deploy taking down the version that worked, especially when a human is clicking around in it. So redeploys in AgentServe follow one rule: build the new generation while the old one keeps serving, and only swap once the new one is built. If anything goes wrong after that, put the old one back.
This post walks through how that works in runtime.py, including the parts that are less clever than
they sound.
Generations on disk#
Each stack has a directory, and each deploy gets a numbered generation inside it:
stacks/<id>/
gen-3/src/ # the extracted upload for generation 3
gen-4/src/ # generation 4, being built
data/<service>/ # $DATA_DIR, one per service, outside every generation
logs/<service>.build.log
logs/<service>.run.log
A redeploy first reserves its generation number in the database, conditionally:
if not db.run("UPDATE stacks SET deploying = ?, deploy_error = NULL WHERE id = ? AND deploying IS NULL "
"AND status IN ('live', 'failed')", gen, stack["id"]):
raise Rejected(409, "a build is already running")
If two redeploys arrive at once, one gets the row and the other gets a 409. Agents are enthusiastic, and sometimes
they fire the same command twice. A conditional UPDATE is the cheapest lock we know of.
Build without touching anything#
The upload is extracted into gen-N/src, and every service is built there, in dependency order. Python
services get a fresh uv venv. Node and static services get npm ci (or
npm install without a lockfile), with dev dependencies, because that's where Vite lives. Static services
are built once and then served directly by the proxy, so they never run a process at all.
None of this touches the running generation. The old processes keep serving from gen-(N-1) on their
ports. That matters because the build is the slow part: installing dependencies and running a bundler takes far
longer than starting a process. Doing it in the background means the stack stays up for most of the deploy.
Environment references like ${api.url} are resolved before the build, not at runtime. A Vite frontend
bakes VITE_API_URL into its bundle at build time, so it needs the real URL then. Services keep their
hostnames across generations, so the URL a frontend baked in during generation 3 is still correct in generation 4.
The swap#
Once every service has built, the deploy swaps. Here's the core of it, lightly trimmed:
try:
built = _build(before, generation, m)
_swap(stack_id, generation, m, built)
swapped = True
_bring_up(stack_id)
except (BuildError, manifest.ManifestError, OSError) as e:
if rollback_to is None:
_fail(stack_id, str(e))
return
_roll_back(stack_id, generation, before, services_before, swapped, str(e))
return
_swap stops the old processes, records the new generation and manifest, and drops services that the
new manifest removed. _bring_up then starts each service in dependency order on its port.
We want to be honest about one thing: this is a stop-then-start swap, not a blue-green one. Each service keeps its port across generations, so the old process has to stop before the new one can bind. Between those two moments, a request to that service fails. That window is a process start plus a health check, usually a second or two, not the minutes of a build. A real zero-downtime cutover needs two sets of ports and a proxy that flips between them, and it's on our list. For agent-built apps being poked at by one or two people, this was the right first trade.
Health checks decide#
A process that started isn't a process that works. After starting each service, AgentServe polls it until it answers or a timeout runs out (60 seconds by default):
while time.time() < deadline:
if proc.poll() is not None:
tail = rlog.read_text(errors="ignore")[-1500:]
raise BuildError(f"process exited with code {proc.returncode} before becoming healthy:\n{tail}")
try:
r = httpx.get(url, timeout=2)
if (health and r.status_code < 300) or (not health and r.status_code < 500):
return
except httpx.HTTPError:
pass
time.sleep(0.3)
If the manifest declares health: /healthz, that path has to return a 2xx. If it doesn't declare one,
any response below 500 on / counts. A 404 still proves something is listening. A process that exits early
fails immediately, and the last part of its log ends up in the error. That's deliberate: the error message is what the
agent reads next, and "exited with code 1" plus a traceback is far more useful to it than "unhealthy".
Rolling back#
If the build fails, nothing was swapped. The deploy deletes the new generation's directory, clears the reservation, and the old generation never noticed.
If the build succeeded but the new generation didn't come up healthy, it's more work. Before the deploy started, we
took a snapshot of the stack and its service rows. Rollback stops whatever half-started, restores those rows, prunes
the failed generation, and runs _bring_up again on the old one. The stack stays live, with
deploy_error set, and emits an event:
{"type": "stack.deploy_failed", "data": {"generation": 5, "serving_generation": 4, "error": "…"}}
The CLI exits with code 1 in that case, so an agent running it in a loop notices. The usual next step is
agentserve logs --build or --service api, a fix, and another deploy.
There's one honest failure mode left: if the old generation also refuses to come back and fails its own health check
on the way up, the stack is marked failed with both errors. We'd rather say that than
pretend it's serving.
On success, older generation directories are pruned before the stack is marked live. Only one generation is ever kept around, so there's no "roll back to three deploys ago". Redeploying the old code does the same job.
$DATA_DIR survives all of it#
None of the above would be acceptable if a redeploy wiped the database. So persistent data lives outside the
generations, in data/<service>, and each process gets it as $DATA_DIR. The example
app's backend does exactly what we'd hope agents do:
db = sqlite3.connect(os.path.join(os.environ.get("DATA_DIR", "."), "todos.db"), check_same_thread=False)
Pruning only ever deletes gen-* directories, so that file survives redeploys, rollbacks, expiry and
claiming. Anything a service writes into its own source directory is gone at the next deploy, and the skill file tells
agents so in plain words.
Supervision, backoff and crash loops#
Once a generation is live, a background loop checks on it every couple of seconds. Each service runs under
/bin/sh -c "exec …" in its own process group, so signals reach the real process and stopping it takes the
whole group down: SIGTERM, five seconds of patience, then SIGKILL.
When a process has exited and nobody asked it to, it's handed to a restart job:
recent = [t for t in _crashes.get((stack_id, sname), []) if now - t < CRASH_WINDOW] + [now]
events.emit(stack_id, "service.crashed", service=sname, exit_code=exit_code, crashes=len(recent))
if len(recent) > MAX_RESTARTS:
_fail(stack_id, f"{sname} crashed {len(recent)} times in {CRASH_WINDOW // 60} min ...")
return
time.sleep(min(2 ** (len(recent) - 1), 30))
The backoff goes 1s, 2s, 4s and so on, capped at 30 seconds. Once a service has crashed more than five times in ten
minutes, we stop restarting it and mark the stack failed with a message telling the agent to check the logs
and redeploy. A service that crashes on every request will never fix itself, and restarting it forever only hides the
problem from the one reader who could fix it.
A successful deploy resets the crash count, and so does a manual restart from the console. Run logs are capped at 5 MB with one rotated copy, so a chatty crash loop can't fill the disk while it's at it.
What happens when the server restarts#
Builds in flight live in a thread pool, so a server restart loses them. On startup, AgentServe marks those deploys as failed ("the server restarted during the build; redeploy") and brings every live stack back from its current generation's built artifacts, without rebuilding. Expired stacks that get claimed come back the same way.
What's next#
The obvious gaps are the ones above: a true zero-downtime cutover, and running all of this inside real isolation
instead of plain processes. If you want the user-facing side of this, the
manifest reference covers health, depends_on and references,
and Limits lists the timeouts.