Skip to main content

Cluster

Use this page when more than one Fire Arrow Server instance serves traffic against the same PostgreSQL database. Cluster mode ships in Fire Arrow Server 2.5.0. PostgreSQL is the only coordination: no Redis, message broker, or clustered scheduler.

Property names and defaults are in Configuration. This page is the procedure: turn the cluster on, check it, roll an instance, and respond when it is not healthy.

What a cluster does​

Every instance opens three database sessions on top of the connection pool: one holds the leader lock, one receives cluster events, and one confirms that the instance has caught up with the others.

  • Background jobs run on the leader only. HAPI scheduled jobs, including subscription processing, and the CarePlan scheduled jobs run on the current leader. If the leader stops, another instance takes over after about 10 seconds. After a network partition, detecting the lost connection takes up to about 25 seconds more (TCP keepalive: idle + interval × count, defaults 10 + 5 × 3).
  • Subscription and search changes apply on every instance when the writing transaction commits. That covers Subscription, SubscriptionTopic, and SearchParameter.
  • Permission changes apply on every instance as part of the commit. Deleting a PractitionerRole, for example, drops the cached authorization entries everywhere without waiting for the cache time-to-live.
  • A client can read a write it just made on another instance by sending Cache-Control: no-cache or no-store. See Client behavior.
  • While an instance is not receiving cluster events, it authorizes without its caches. fire_arrow_cluster_bus_ready is 0 in that state.

A single PostgreSQL instance that passes the startup check also runs in cluster mode. It is the leader, and it uses the three extra sessions. H2 never joins a cluster.

On PostgreSQL, when fire-arrow.cluster.enabled is not false and hapi.fhir.subscription.immediately_queued is unset, the server sets that subscription flag to true. A synchronous CarePlan/$subscribe-due-events or $renew-due-events handled by an instance that is not the leader can then finish without waiting for the leader to drain the subscription outbox. Set hapi.fhir.subscription.immediately_queued yourself to keep another value.

Before you add a second instance​

Do this on every instance before the replica count goes above one.

  • Database is PostgreSQL, and every instance uses the same database and schema.
  • The connection is direct, or through a pooler in session mode. Transaction mode and statement mode fail the startup check (PgBouncer pool_mode=transaction is the usual case). LISTEN and session-level locks do not survive those modes.
  • max_connections has room for the Hikari pool plus three sessions per instance.
  • fire-arrow.mutex.provider is jdbc (the default). local only locks inside one process.
  • fire-arrow.graphql.cursor-hmac-secret is set to the same value on every instance, if GraphQL is enabled. Otherwise a cursor from one instance is rejected by the others.
  • Full-text search uses Elasticsearch. A Lucene index lives on the instance that built it and diverges.
  • Binaries are in Azure Blob, the database, or a filesystem path mounted on every instance. hapi.fhir.binary_storage_mode=FILESYSTEM on a local disk stores the file on one instance only.
  • Clients use REST hook subscriptions. WebSocket subscriptions are delivered only to clients connected to the instance that processed the change.
  • The load balancer sticks a client to one instance for composite search paging ($everything and similar). Authorization does not need sticky sessions.
  • Webhook consumers tolerate a duplicate delivery. A normal shutdown preserves in-flight delivery; an abrupt stop can drop one and a retry can deliver one twice.
  • The orchestrator stop timeout is longer than the preStop delay + HTTP shutdown grace (spring.lifecycle.timeout-per-shutdown-phase, default 30 seconds) + fire-arrow.cluster.quiesce-timeout (default 30 seconds). With those defaults, Kubernetes terminationGracePeriodSeconds: 75 is enough.

Turn cluster mode on​

fire-arrow.cluster.enabled defaults to auto. On PostgreSQL, auto turns cluster mode on when the startup check passes. If the check fails, auto logs an error and runs that instance on its own. More than one replica in that state is unsafe: each instance schedules jobs and serves authorization as if it were alone.

Set enabled to true on every instance of a deployment with more than one replica. A failed check then stops the process instead of serving traffic uncoordinated.

fire-arrow:
cluster:
enabled: true
FIRE_ARROW_CLUSTER_ENABLED=true
  1. Apply the preconditions and set enabled: true.

  2. Start one instance. In its log, confirm:

    Cluster mode active:

    A later line names that instance as leader (Cluster node … is leader).

  3. Start the remaining instances. Each one logs Cluster mode active. Exactly one instance logs that it is leader.

  4. Scrape each instance (see Deployment):

    curl -s http://localhost:8080/actuator/prometheus | grep '^fire_arrow_cluster_'

    fire_arrow_cluster_leader is 1 on one instance and 0 on the others. fire_arrow_cluster_bus_ready is 1 on every instance.

  5. Read the startup log for Multi-node readiness. With enabled: true, each finding is a warning. Fix the finding before sending production traffic. The checks are the preconditions above, plus a degraded-mode warning if the startup check failed.

The shipped configuration sets server.shutdown: graceful. Leave it. immediate skips the subscription drain in Roll an instance.

Roll an instance​

Shut an instance down through the orchestrator, not with SIGKILL.

  1. The instance stops accepting new HTTP requests and finishes the ones already in flight (up to spring.lifecycle.timeout-per-shutdown-phase, default 30 seconds).
  2. It pauses background jobs and waits for in-flight subscription processing, up to fire-arrow.cluster.quiesce-timeout (default 30 seconds).
  3. It releases the leader lock. Another instance becomes leader after about one lease (two heartbeats, default 10 seconds).

A shutdown that finishes inside the quiesce timeout keeps the subscription delivery behavior you already have. If the wait hits the timeout, the log contains Subscription channel quiesce timed out and a delivery that was in flight can be lost. The quiesce counter is usually gone by the time metrics are scraped; use that log line.

After the new leader starts, CarePlan reconcile backlog processing runs there. The other instances do not repeat the startup scan. The leader rebuilds the backlog when it takes over, and scans again every fire-arrow.cluster.backlog-rescan-interval (default 15 minutes).

During the handover window, expect:

  • one leader, then a gap of about 10 seconds, then a different leader (fire_arrow_cluster_leader);
  • subscription and CarePlan jobs paused for that gap;
  • possible duplicate webhook deliveries. Consumers must be idempotent.

Adding a replica is the same as Turn cluster mode on from step 3: start it with enabled: true and the same database, and confirm fire_arrow_cluster_bus_ready is 1 and it is not a second leader.

Client behavior​

Send Cache-Control: no-cache or no-store on a request that must reflect a write just made through a different instance. The instance waits until it has applied every authorization change committed before the request arrived, then answers using its caches. The wait is usually a few milliseconds. If it exceeds fire-arrow.cluster.read-fence-timeout (default 1 second), that request is authorized without the caches.

Without the header, other instances usually apply the change within milliseconds, but a request that arrives in that window can still see the previous answer.

On a single instance, Cache-Control does not change authorization caching. The no-cache the server adds internally for search does not trigger this wait.

Session affinity is still required for composite search paging. It is not required for authorization.

Signals​

Scrape /actuator/prometheus on every instance.

SignalHealthyWhat a bad value means
fire_arrow_cluster_leader1 on exactly one instance0 everywhere: no leader, background jobs are paused. 1 on two instances: they are not in the same cluster (different database or schema).
fire_arrow_cluster_bus_ready1 on every instance0: this instance is not receiving cluster events and is authorizing without its caches.
fire_arrow_outbox_depth, fire_arrow_outbox_oldest_age_secondsDepth stays boundedThe leader is not keeping up with subscription processing.
fire_arrow_outbox_drain_age_secondsStableTime from commit until the leader picks up an outbox entry.
fire_arrow_cluster_watchdog_failuresFlatThe instance cannot confirm its own event path and will reconnect. bus_ready drops while it reconnects.
fire_arrow_authz_cache_discarded_storesOccasional increments are normalA permission change raced a request that had already read the old answer. The old answer was discarded. A steady high rate means permission changes are overlapping lookups continuously.

Alert on:

  • sum(fire_arrow_cluster_leader) != 1 for more than a minute
  • fire_arrow_cluster_bus_ready == 0 on any instance for more than a minute
  • fire_arrow_outbox_depth rising for 10 minutes
  • log line Running DEGRADED or Subscription channel quiesce timed out

Troubleshooting​

Startup fails: cluster self-test failed​

enabled is true and the process exits with Cluster self-test failed. The detail either says pg_backend_pid() changed within one session (a transaction-mode pooler) or NOTIFY loopback not received.

Connect the instance directly to PostgreSQL, or through a session-mode pooler. Confirm the database user can use the session, then start again. Do not flip enabled back to auto to get the pod up: auto will serve traffic as a single node.

Startup fails: enabled=true requires a PostgreSQL datasource​

enabled: true was set on an H2 URL. Point spring.datasource.url at PostgreSQL, or set enabled: false for a single-node H2 experiment.

Log says Running DEGRADED​

enabled is auto (or unset) and the startup check failed. The instance behaves as a single node. The log is Cluster self-test failed … Running DEGRADED.

If a second degraded instance of the same database becomes visible, the newer one refuses readiness (Another degraded Fire Arrow node … refuses readiness traffic) and keeps background jobs paused. Readiness stays down until the older degraded instance is gone. That guard is not reliable behind a transaction-mode pooler, because the other session may be hidden.

Set FIRE_ARROW_CLUSTER_ENABLED=true, fix the connection so the startup check passes, and roll the instances. A healthy start logs Cluster mode active and does not log DEGRADED.

Readiness flips down after a second replica starts​

You are in the degraded case above. The newer instance is refusing traffic on purpose. Fix the startup check; do not raise the replica count further until every instance logs Cluster mode active.

Two leaders, or jobs running on every instance​

fire_arrow_cluster_leader is 1 on more than one instance, or CarePlan and subscription jobs are clearly running everywhere.

Each leader is coordinating with a different database identity. Compare spring.datasource.url and the schema. Instances must share both. Also confirm neither instance logged running single-node or DEGRADED.

A write on one instance is not visible on another​

  1. Confirm fire_arrow_cluster_bus_ready is 1 on the instance that served the read. If it is 0, that instance is authorizing without its caches only while the event path is down; wait for it to return to 1, or send Cache-Control: no-cache.
  2. If the read was a follow-up to a write on a different instance, send Cache-Control: no-cache or no-store on that read.
  3. If bus_ready stays 0, look for Cluster event bus watchdog failure and Cluster event bus reconnect requested. Two watchdog failures reconnect the event session. Check that the three extra sessions are not exhausted (max_connections, pooler limits) and that a firewall is not cutting idle connections faster than keepalive (keepalive.idle-seconds, default 10).

GraphQL cursor from one instance is rejected by another​

fire-arrow.graphql.cursor-hmac-secret differs between instances, or it is unset and each process generated its own. Set the same secret everywhere and restart. Clients restart pagination; old cursors do not survive the change.

WebSocket client misses an event​

WebSocket delivery stays on the instance that matched the change. A client connected to another instance does not get it. Use a REST hook subscription for multi-instance deployments.

Subscription delivery missing after a rollout​

Look for Subscription channel quiesce timed out on the instance that stopped. The stop took longer than quiesce-timeout, so an in-flight delivery can be gone. Lengthen terminationGracePeriodSeconds (preStop + HTTP grace + quiesce timeout) or raise fire-arrow.cluster.quiesce-timeout. Webhook handlers should treat a repeated event as the same event.

Outbox depth keeps growing​

fire_arrow_outbox_depth and fire_arrow_outbox_oldest_age_seconds climb, and only the leader drains them. Subscription processing is not spread across instances. Give the leader more database and CPU, or reduce delivery cost (slow webhook endpoints hold the leader). A growing backlog is not fixed by adding Fire Arrow instances.