Docs / Guides

Redeploys and rollbacks

How redeploys build the new version while the old one serves, what happens when they fail, and how to read the logs.

Redeploying updates an existing stack in place: same id, same URLs, same $DATA_DIR. Each deploy creates a new generation, numbered from 1. If the new generation doesn't work, the previous one keeps serving.

Redeploying#

With the CLI, run deploy again in the same directory. The stack id and manage token in .agentserve/stack.json point it at the existing stack:

agentserve deploy ./my-project

Over HTTP, send the same multipart body as a create to the stack's deploy endpoint:

curl -X POST "https://agentserve.sh/v0/stacks/$STACK_ID/deploy?wait=true" \
  -H "Authorization: Bearer $MANAGE_TOKEN" \
  -F [email protected]

A redeploy keeps the stack's name and TTL; --ttl and --name only apply to new stacks (use agentserve extend for more time, or --new for a fresh stack). The manifest can otherwise change freely: services can be added, removed or reconfigured. New services get their URLs before anything builds, so references to them resolve.

Only one build runs per stack at a time. A second redeploy while one is in flight returns 409 "a build is already running". An expired stack can't be redeployed (409); it has to be claimed first.

The CLI and the MCP server never swap the linked stack for a new one behind your back. When a redeploy is refused, they leave .agentserve/stack.json as it is and say why:

  • 409, expired: the error points at the claim URL. Claim the stack to restore it, or pass --new (new=True over MCP) to start a separate stack with new URLs.
  • 409, build running: wait for it to finish, then deploy again.
  • 403, token revoked: the CLI retries with AGENTSERVE_API_KEY when it is set; without it, the redeploy fails.
  • 404, stack deleted: the only case that creates a new stack. The CLI prints a message and the MCP result carries a note, since the URLs and claim code change.

Build, then swap#

A redeploy of a live stack goes through these steps:

  1. The stack's deploying field is set to the new generation number, and stack.redeploying is emitted. The status stays live.
  2. Every service of the new generation is installed and built in a separate directory, in dependency order. The current generation keeps serving the whole time.
  3. When every build has succeeded, the current processes are stopped and the new generation is started, service by service, each waiting for its health check.
  4. When all of them are up, older generations are deleted, deploying goes back to null, and stack.live is emitted with the new generation.

Expect a short gap during step 3, between stopping the old processes and the new ones answering their health check. The build itself causes no downtime.

When a redeploy fails#

If any build step fails, or a new process exits or doesn't become healthy within 60 seconds, AgentServe rolls back:

  • The previous generation is put back. If its processes had already been stopped, they are started again from the previous build, without rebuilding.
  • The stack stays live, with deploy_error set to what went wrong and generation still pointing at the version being served.
  • A stack.deploy_failed event is emitted with the failed generation, the error and the serving_generation.
{
  "status": "live",
  "generation": 3,
  "deploying": null,
  "deploy_error": "`npm run build` exited with code 2"
}

The next successful deploy clears deploy_error. In the rare case where the previous generation also fails to come back, the stack is marked failed, with both errors in error.

A first deploy has nothing to roll back to. If it fails, the stack goes to failed with the reason in error. Fix the problem and redeploy: a failed stack accepts redeploys.

Manifest mistakes, such as an unknown runtime, a reference to a missing service or a dependency cycle, are caught before anything builds. The request fails with 422 and the running stack isn't touched.

A rollback restores code, not data. Anything the failed generation wrote to $DATA_DIR, including a schema migration it ran at start-up, is still there when the previous generation comes back.

The CLI exit code#

agentserve deploy waits for the build to settle and exits with status 1 when the stack is failed or the response carries deploy_error. In the second case it prints the error and says which generation is still being served. Agents and scripts can rely on the exit code instead of parsing output:

if ! agentserve deploy .; then
  agentserve logs . --build
fi

Errors from the server (4xx and 5xx) also exit 1, with the server's message on stderr. Add --json to get the full stack object instead of the summary.

Reading logs#

Every service has two logs: a build log, rewritten by each build, and a runtime log that the process appends to.

# build logs for every service
agentserve logs --build

# runtime logs for one service, last 300 lines
agentserve logs --service api --tail 300

The build log shows each command that ran ($ npm ci …, $ npm run build) with its output, and ends with == build ok when the build succeeded. The runtime log records each start (== start api on :54321 $ uvicorn …) followed by the process's stdout and stderr. Runtime logs are capped at 5 MB, with one rotated copy kept.

Over HTTP, use GET /v0/stacks/{id}/logs with kind=build or kind=run, an optional service and tail (default 200 lines). For claimed projects, the dashboard has the same logs per service, a combined live view, and a deploy history that marks each generation as live, superseded, rolled back or failed.

Crashes after a deploy#

A process that exits while the stack is live is restarted with backoff (1s, 2s, 4s, up to 30s), with service.crashed and service.restarted events. After more than 5 crashes in 10 minutes, the stack is marked failed. Read the runtime log, fix the code, and redeploy.

Stuck? Your agent can read /skill.md, or connect it over MCP.