Platform
Learn
Developer docs User guide Quickstart Blog
Company
Services About Contact Links Get started

Health checks

Updated

Raytha exposes two unauthenticated endpoints for orchestrators and monitors: /healthz answers whether the process is up, and /healthz/ready answers whether it can reach PostgreSQL and its file storage.

The two endpoints

EndpointUse it forWhat it checks
/healthzLiveness: "restart me if this fails"Nothing. If the process can answer, it returns 200. A broken database never makes it fail, so an orchestrator will not restart a healthy app over a database outage.
/healthz/readyReadiness: "send me traffic"PostgreSQL and the file storage provider. Returns 200 when both are healthy and 503 otherwise.

The checks inside /healthz/ready:

  • postgres opens a connection with the configured connection string and runs a trivial query, with a 5 second timeout.
  • storage depends on the provider. For Local it writes a small .healthz-* file into the upload directory and deletes it, so it proves the directory exists and is writable. For AzureBlob and S3 it checks that the provider is configured and can generate a signed URL. It makes no network call to the bucket, so it does not prove your credentials or CORS rule are right.

Responses

Both endpoints return JSON with Cache-Control: no-store. This is a healthy readiness response:

{
  "version": "2.0.0",
  "environment": "Production",
  "status": "Healthy",
  "totalDuration": "00:00:00.0123456",
  "checks": [
    {
      "name": "postgres",
      "status": "Healthy",
      "description": null,
      "error": null,
      "duration": "00:00:00.0089001",
      "data": null
    },
    {
      "name": "storage",
      "status": "Healthy",
      "description": "Local storage directory is writable.",
      "error": null,
      "duration": "00:00:00.0004110",
      "data": { "provider": "local", "directory": "/app/user-uploads" }
    }
  ]
}

With PostgreSQL stopped, /healthz/ready answers 503 Service Unavailable, status is Unhealthy, and the postgres entry carries the driver's message:

{
  "name": "postgres",
  "status": "Unhealthy",
  "description": "Failed to connect to 127.0.0.1:15432",
  "error": "Failed to connect to 127.0.0.1:15432",
  "duration": "00:00:00.0243716",
  "data": null
}

In the same situation /healthz still returns 200. When PostgreSQL comes back, /healthz/ready returns 200 again with no restart. A Local storage problem shows as "description": "Storage probe failed." with the exception message in error, or "Local storage directory does not exist.".

curl -s -o /dev/null -w "%{http_code}\n" http://localhost:5001/healthz
curl -s http://localhost:5001/healthz/ready | jq '{status, checks: [.checks[] | {name, status}]}'

Docker

The image already ships a health check against the liveness endpoint: curl -fsS http://127.0.0.1:8080/healthz every 30 seconds, with a 40 second start period and 3 retries. See the state with:

docker inspect --format '{{.State.Health.Status}}' raytha-app

To make the container report "unhealthy" when the database or storage is down, override the check with the readiness endpoint. In Compose:

services:
  app:
    healthcheck:
      test: ["CMD", "curl", "-fsS", "http://127.0.0.1:8080/healthz/ready"]
      interval: 30s
      timeout: 10s
      start_period: 40s
      retries: 3

Plain Docker and Compose only record the status; they do not restart an unhealthy container. Tools that act on it, such as an autoheal sidecar or Docker Swarm, will restart the app over a database outage, which does not fix the database. Prefer the liveness endpoint for anything that restarts containers.

Kubernetes

apiVersion: apps/v1
kind: Deployment
metadata:
  name: raytha
spec:
  replicas: 1
  selector:
    matchLabels:
      app: raytha
  template:
    metadata:
      labels:
        app: raytha
    spec:
      containers:
        - name: raytha
          image: raythahq/raytha:2.0.0
          ports:
            - containerPort: 8080
          envFrom:
            - secretRef:
                name: raytha-env
          startupProbe:
            httpGet:
              path: /healthz
              port: 8080
            periodSeconds: 5
            failureThreshold: 36
          livenessProbe:
            httpGet:
              path: /healthz
              port: 8080
            periodSeconds: 10
            timeoutSeconds: 5
          readinessProbe:
            httpGet:
              path: /healthz/ready
              port: 8080
            periodSeconds: 10
            timeoutSeconds: 10
            failureThreshold: 3
  • The readiness timeoutSeconds is above the 5 second PostgreSQL timeout, so a slow database shows up as 503 rather than a probe timeout.
  • The startup probe gives the app up to three minutes to come up, which covers applying migrations on first start.
  • The Local storage provider writes to the container's disk, which replicas cannot share. Stay at one replica, or use S3 or Azure Blob. See File storage.
  • The image comes from Docker Hub; pin the version tag rather than latest so a rollout is deliberate. See Deploy with Docker.

Load balancers and uptime monitors

Point a load balancer's target health check at /healthz/ready. For an external uptime monitor, /healthz/ready also catches database outages, but alert on several consecutive failures: a deployment that applies migrations is briefly unreachable.

If you need an application-level check that exercises authentication, GET /raytha/api/v1/ping with an X-API-KEY header returns {"success": true, "version": "...", "organizationName": "..."}. It needs a valid API key, so it is not a replacement for the probes above.

Gotchas

  • The port is closed during start-up. If PostgreSQL is unreachable when Raytha boots, it never opens its port (it needs the database for its data-protection keys, and for migrations when APPLY_PENDING_MIGRATIONS is on). You get a refused connection, not a 503. Once the database is reachable the app finishes starting on its own, with no restart. Use a start-up probe or a start period.
  • The readiness endpoint is public and detailed. It shows the version, the environment name, the storage directory and database error text to anyone who can reach it. If that matters, do not route it through your public proxy. With Caddy:
    raytha.example.com {
    	@health path /healthz/ready
    	respond @health 404
    	reverse_proxy app:8080
    }
  • Probe the container port directly. Probes that go through the public proxy depend on HTTPS settings and proxy trust. See Running behind a proxy.
  • The tracing configuration skips /healthz requests, so probes do not flood your OpenTelemetry traces.

Next steps