Blog / Engineering

Build, then swap: how redeploys work

Agents redeploy constantly and break things often. Here is how AgentServe keeps the old version serving while the new one builds, rolls back when it doesn't come up, and keeps your data out of the blast radius.

Humans redeploy a few times a day. Agents redeploy whenever they finish a thought. A typical session looks like this: deploy, open the page, notice a bug, change two lines, redeploy, notice the change broke the build, fix it, redeploy. Half of those deploys are broken in some way, and that's fine. That's how iteration works.

What isn't fine is a broken deploy taking down the version that worked, especially when a human is clicking around in it. So redeploys in AgentServe follow one rule: build the new generation while the old one keeps serving, and only swap once the new one is built. If anything goes wrong after that, put the old one back.

This post walks through how that works in runtime.py, including the parts that are less clever than they sound.

Generations on disk#

Each stack has a directory, and each deploy gets a numbered generation inside it:

stacks/<id>/
  gen-3/src/        # the extracted upload for generation 3
  gen-4/src/        # generation 4, being built
  data/<service>/   # $DATA_DIR, one per service, outside every generation
  logs/<service>.build.log
  logs/<service>.run.log

A redeploy first reserves its generation number in the database, conditionally:

if not db.run("UPDATE stacks SET deploying = ?, deploy_error = NULL WHERE id = ? AND deploying IS NULL "
              "AND status IN ('live', 'failed')", gen, stack["id"]):
    raise Rejected(409, "a build is already running")

If two redeploys arrive at once, one gets the row and the other gets a 409. Agents are enthusiastic, and sometimes they fire the same command twice. A conditional UPDATE is the cheapest lock we know of.

Build without touching anything#

The upload is extracted into gen-N/src, and every service is built there, in dependency order. Python services get a fresh uv venv. Node and static services get npm ci (or npm install without a lockfile), with dev dependencies, because that's where Vite lives. Static services are built once and then served directly by the proxy, so they never run a process at all.

None of this touches the running generation. The old processes keep serving from gen-(N-1) on their ports. That matters because the build is the slow part: installing dependencies and running a bundler takes far longer than starting a process. Doing it in the background means the stack stays up for most of the deploy.

Environment references like ${api.url} are resolved before the build, not at runtime. A Vite frontend bakes VITE_API_URL into its bundle at build time, so it needs the real URL then. Services keep their hostnames across generations, so the URL a frontend baked in during generation 3 is still correct in generation 4.

The swap#

Once every service has built, the deploy swaps. Here's the core of it, lightly trimmed:

try:
    built = _build(before, generation, m)
    _swap(stack_id, generation, m, built)
    swapped = True
    _bring_up(stack_id)
except (BuildError, manifest.ManifestError, OSError) as e:
    if rollback_to is None:
        _fail(stack_id, str(e))
        return
    _roll_back(stack_id, generation, before, services_before, swapped, str(e))
    return

_swap stops the old processes, records the new generation and manifest, and drops services that the new manifest removed. _bring_up then starts each service in dependency order on its port.

We want to be honest about one thing: this is a stop-then-start swap, not a blue-green one. Each service keeps its port across generations, so the old process has to stop before the new one can bind. Between those two moments, a request to that service fails. That window is a process start plus a health check, usually a second or two, not the minutes of a build. A real zero-downtime cutover needs two sets of ports and a proxy that flips between them, and it's on our list. For agent-built apps being poked at by one or two people, this was the right first trade.

Health checks decide#

A process that started isn't a process that works. After starting each service, AgentServe polls it until it answers or a timeout runs out (60 seconds by default):

while time.time() < deadline:
    if proc.poll() is not None:
        tail = rlog.read_text(errors="ignore")[-1500:]
        raise BuildError(f"process exited with code {proc.returncode} before becoming healthy:\n{tail}")
    try:
        r = httpx.get(url, timeout=2)
        if (health and r.status_code < 300) or (not health and r.status_code < 500):
            return
    except httpx.HTTPError:
        pass
    time.sleep(0.3)

If the manifest declares health: /healthz, that path has to return a 2xx. If it doesn't declare one, any response below 500 on / counts. A 404 still proves something is listening. A process that exits early fails immediately, and the last part of its log ends up in the error. That's deliberate: the error message is what the agent reads next, and "exited with code 1" plus a traceback is far more useful to it than "unhealthy".

Rolling back#

If the build fails, nothing was swapped. The deploy deletes the new generation's directory, clears the reservation, and the old generation never noticed.

If the build succeeded but the new generation didn't come up healthy, it's more work. Before the deploy started, we took a snapshot of the stack and its service rows. Rollback stops whatever half-started, restores those rows, prunes the failed generation, and runs _bring_up again on the old one. The stack stays live, with deploy_error set, and emits an event:

{"type": "stack.deploy_failed", "data": {"generation": 5, "serving_generation": 4, "error": "…"}}

The CLI exits with code 1 in that case, so an agent running it in a loop notices. The usual next step is agentserve logs --build or --service api, a fix, and another deploy.

There's one honest failure mode left: if the old generation also refuses to come back and fails its own health check on the way up, the stack is marked failed with both errors. We'd rather say that than pretend it's serving.

On success, older generation directories are pruned before the stack is marked live. Only one generation is ever kept around, so there's no "roll back to three deploys ago". Redeploying the old code does the same job.

$DATA_DIR survives all of it#

None of the above would be acceptable if a redeploy wiped the database. So persistent data lives outside the generations, in data/<service>, and each process gets it as $DATA_DIR. The example app's backend does exactly what we'd hope agents do:

db = sqlite3.connect(os.path.join(os.environ.get("DATA_DIR", "."), "todos.db"), check_same_thread=False)

Pruning only ever deletes gen-* directories, so that file survives redeploys, rollbacks, expiry and claiming. Anything a service writes into its own source directory is gone at the next deploy, and the skill file tells agents so in plain words.

Supervision, backoff and crash loops#

Once a generation is live, a background loop checks on it every couple of seconds. Each service runs under /bin/sh -c "exec …" in its own process group, so signals reach the real process and stopping it takes the whole group down: SIGTERM, five seconds of patience, then SIGKILL.

When a process has exited and nobody asked it to, it's handed to a restart job:

recent = [t for t in _crashes.get((stack_id, sname), []) if now - t < CRASH_WINDOW] + [now]
events.emit(stack_id, "service.crashed", service=sname, exit_code=exit_code, crashes=len(recent))
if len(recent) > MAX_RESTARTS:
    _fail(stack_id, f"{sname} crashed {len(recent)} times in {CRASH_WINDOW // 60} min ...")
    return
time.sleep(min(2 ** (len(recent) - 1), 30))

The backoff goes 1s, 2s, 4s and so on, capped at 30 seconds. Once a service has crashed more than five times in ten minutes, we stop restarting it and mark the stack failed with a message telling the agent to check the logs and redeploy. A service that crashes on every request will never fix itself, and restarting it forever only hides the problem from the one reader who could fix it.

A successful deploy resets the crash count, and so does a manual restart from the console. Run logs are capped at 5 MB with one rotated copy, so a chatty crash loop can't fill the disk while it's at it.

What happens when the server restarts#

Builds in flight live in a thread pool, so a server restart loses them. On startup, AgentServe marks those deploys as failed ("the server restarted during the build; redeploy") and brings every live stack back from its current generation's built artifacts, without rebuilding. Expired stacks that get claimed come back the same way.

What's next#

The obvious gaps are the ones above: a true zero-downtime cutover, and running all of this inside real isolation instead of plain processes. If you want the user-facing side of this, the manifest reference covers health, depends_on and references, and Limits lists the timeouts.

Shameless plug.

Point your agent at AgentServe. The next thing it builds gets a live URL, and you get the claim link.