Health checks
Raytha exposes two unauthenticated endpoints for orchestrators and monitors: /healthz answers whether the process is up, and /healthz/ready answers whether it can reach PostgreSQL and its file storage.
The two endpoints
| Endpoint | Use it for | What it checks |
|---|---|---|
/healthz | Liveness: "restart me if this fails" | Nothing. If the process can answer, it returns 200. A broken database never makes it fail, so an orchestrator will not restart a healthy app over a database outage. |
/healthz/ready | Readiness: "send me traffic" | PostgreSQL and the file storage provider. Returns 200 when both are healthy and 503 otherwise. |
The checks inside /healthz/ready:
postgresopens a connection with the configured connection string and runs a trivial query, with a 5 second timeout.storagedepends on the provider. ForLocalit writes a small.healthz-*file into the upload directory and deletes it, so it proves the directory exists and is writable. ForAzureBlobandS3it checks that the provider is configured and can generate a signed URL. It makes no network call to the bucket, so it does not prove your credentials or CORS rule are right.
Responses
Both endpoints return JSON with Cache-Control: no-store. This is a healthy readiness response:
{
"version": "2.0.0",
"environment": "Production",
"status": "Healthy",
"totalDuration": "00:00:00.0123456",
"checks": [
{
"name": "postgres",
"status": "Healthy",
"description": null,
"error": null,
"duration": "00:00:00.0089001",
"data": null
},
{
"name": "storage",
"status": "Healthy",
"description": "Local storage directory is writable.",
"error": null,
"duration": "00:00:00.0004110",
"data": { "provider": "local", "directory": "/app/user-uploads" }
}
]
}
With PostgreSQL stopped, /healthz/ready answers 503 Service Unavailable, status is Unhealthy, and the postgres entry carries the driver's message:
{
"name": "postgres",
"status": "Unhealthy",
"description": "Failed to connect to 127.0.0.1:15432",
"error": "Failed to connect to 127.0.0.1:15432",
"duration": "00:00:00.0243716",
"data": null
}
In the same situation /healthz still returns 200. When PostgreSQL comes back, /healthz/ready returns 200 again with no restart. A Local storage problem shows as "description": "Storage probe failed." with the exception message in error, or "Local storage directory does not exist.".
curl -s -o /dev/null -w "%{http_code}\n" http://localhost:5001/healthz
curl -s http://localhost:5001/healthz/ready | jq '{status, checks: [.checks[] | {name, status}]}'
Docker
The image already ships a health check against the liveness endpoint: curl -fsS http://127.0.0.1:8080/healthz every 30 seconds, with a 40 second start period and 3 retries. See the state with:
docker inspect --format '{{.State.Health.Status}}' raytha-app
To make the container report "unhealthy" when the database or storage is down, override the check with the readiness endpoint. In Compose:
services:
app:
healthcheck:
test: ["CMD", "curl", "-fsS", "http://127.0.0.1:8080/healthz/ready"]
interval: 30s
timeout: 10s
start_period: 40s
retries: 3
Plain Docker and Compose only record the status; they do not restart an unhealthy container. Tools that act on it, such as an autoheal sidecar or Docker Swarm, will restart the app over a database outage, which does not fix the database. Prefer the liveness endpoint for anything that restarts containers.
Kubernetes
apiVersion: apps/v1
kind: Deployment
metadata:
name: raytha
spec:
replicas: 1
selector:
matchLabels:
app: raytha
template:
metadata:
labels:
app: raytha
spec:
containers:
- name: raytha
image: raythahq/raytha:2.0.0
ports:
- containerPort: 8080
envFrom:
- secretRef:
name: raytha-env
startupProbe:
httpGet:
path: /healthz
port: 8080
periodSeconds: 5
failureThreshold: 36
livenessProbe:
httpGet:
path: /healthz
port: 8080
periodSeconds: 10
timeoutSeconds: 5
readinessProbe:
httpGet:
path: /healthz/ready
port: 8080
periodSeconds: 10
timeoutSeconds: 10
failureThreshold: 3
- The readiness
timeoutSecondsis above the 5 second PostgreSQL timeout, so a slow database shows up as503rather than a probe timeout. - The startup probe gives the app up to three minutes to come up, which covers applying migrations on first start.
- The
Localstorage provider writes to the container's disk, which replicas cannot share. Stay at one replica, or use S3 or Azure Blob. See File storage. - The image comes from Docker Hub; pin the version tag rather than
latestso a rollout is deliberate. See Deploy with Docker.
Load balancers and uptime monitors
Point a load balancer's target health check at /healthz/ready. For an external uptime monitor, /healthz/ready also catches database outages, but alert on several consecutive failures: a deployment that applies migrations is briefly unreachable.
If you need an application-level check that exercises authentication, GET /raytha/api/v1/ping with an X-API-KEY header returns {"success": true, "version": "...", "organizationName": "..."}. It needs a valid API key, so it is not a replacement for the probes above.
Gotchas
- The port is closed during start-up. If PostgreSQL is unreachable when Raytha boots, it never opens its port (it needs the database for its data-protection keys, and for migrations when
APPLY_PENDING_MIGRATIONSis on). You get a refused connection, not a503. Once the database is reachable the app finishes starting on its own, with no restart. Use a start-up probe or a start period. - The readiness endpoint is public and detailed. It shows the version, the environment name, the storage directory and database error text to anyone who can reach it. If that matters, do not route it through your public proxy. With Caddy:
raytha.example.com { @health path /healthz/ready respond @health 404 reverse_proxy app:8080 } - Probe the container port directly. Probes that go through the public proxy depend on HTTPS settings and proxy trust. See Running behind a proxy.
- The tracing configuration skips
/healthzrequests, so probes do not flood your OpenTelemetry traces.
Next steps
- Running behind a proxy.
- Backups and restores.
- Configuration, including the observability variables.