DevOps · Linux

Zero-downtime deploys with systemd and a health check

· 2 min read

You do not need an orchestrator to deploy without dropping requests. On a single host, two instances of the app, a health endpoint and a short script are enough.

Run the app as a template unit

A systemd template unit lets one file describe any number of instances. The text after the @ becomes %i, and we use it as the port.

# /etc/systemd/system/app@.service
[Unit]
Description=App instance on port %i
After=network.target

[Service]
User=app
ExecStart=/opt/app/current/bin/app --port %i
Restart=on-failure
KillSignal=SIGTERM
TimeoutStopSec=30

[Install]
WantedBy=multi-user.target

Now systemctl start app@8081 and systemctl start app@8082 run two copies side by side.

Drain on SIGTERM

When the app receives SIGTERM it should stop accepting new connections, finish the requests in flight and exit. TimeoutStopSec is how long systemd waits before it sends SIGKILL, so set it higher than your slowest request.

A health endpoint that tells the truth

Return 200 from /healthz only when the instance can actually serve traffic: configuration loaded, database reachable, migrations applied. Keep it cheap. A health check that calls five other services will fail for reasons that have nothing to do with this deploy.

The switch-over script

Start the new instance on the idle port, wait until it is healthy, point traffic at it, then stop the old one.

#!/usr/bin/env bash
set -euo pipefail

OLD=8081
NEW=8082

systemctl start "app@${NEW}"

for _ in $(seq 1 30); do
  curl -fsS "http://127.0.0.1:${NEW}/healthz" >/dev/null && break
  sleep 1
done

# Fail the deploy if the new instance never became healthy.
curl -fsS "http://127.0.0.1:${NEW}/healthz" >/dev/null

# Switch the web server's upstream to port ${NEW} and reload it here.

systemctl stop "app@${OLD}"

Because systemctl stop sends SIGTERM and waits, the old instance finishes its in-flight requests before it goes away. Swap the two port numbers on the next deploy.

Migrations and rollbacks

For a short time the old and new versions run together, so the schema must work for both. Ship additive changes first, deploy the code, and remove old columns in a later release. Rollback is the same script in reverse: start the previous release on the idle port, check it, switch back. Keep the last release directory in place for exactly this.

← All notes