Skip to main content

Fire Arrow Server 2.5.0

· 9 min read

Fire Arrow Server 2.5.0 has been released.

  • (security) A permission change can no longer stay cached when a concurrent request stores an older answer after the change has committed
  • (feature) Several instances that share one PostgreSQL database now coordinate through that database
  • (bugfix) Replacing a CarePlan via PlanDefinition/$apply now removes Tasks for activities the new plan dropped, including past-due Tasks
  • (bugfix) Orphan-blob cleanup recognizes FHIR Binary blobs, counts them separately, and never deletes them

Permission Changes Could Stay Cached Past Commit​

Authorization results are cached so a search does not re-check the same access on every row. A permission change — granting access, or revoking it — evicts the cached answer so the next request sees the new rule.

That eviction could lose a race with a request already in flight. The request read the old answer, the change committed and cleared the cache, and the request then stored the old answer again. Until that entry expired, the server kept using it. A revocation could still allow a Patient search, a clinical read or write, or a token operation that should have been refused. A new grant could still return 403 Forbidden. The overlap was one request against one commit, on a single instance as well as in a cluster, for every role and validator.

Starting with 2.5.0, an answer read before the change and stored after it is discarded. The next lookup reads the current permissions. A discard is counted on fire_arrow_authz_cache_discarded_stores. General-practitioner patient sets use this same cache: a client Cache-Control: no-cache no longer makes that lookup skip the cache. In a cluster, that header waits until this instance has applied authorization changes committed before the request, then uses the cache (see below). On a single instance the header does not change authorization caching.

No configuration change is required.

PostgreSQL Multi-Node Coordination​

Several Fire Arrow Server instances that share one PostgreSQL database and schema can now run as one cluster. PostgreSQL is the only coordination. One instance is the leader, and the database notifies the others when subscriptions, search parameters, or permissions change. No Redis, message broker, or clustered scheduler is required.

In a cluster:

  • Background jobs run on the leader only. HAPI scheduled jobs, including subscription processing, and the CarePlan scheduled jobs run on the current leader. If the leader stops or crashes, another instance takes over after about 10 seconds. After a network partition, detecting the lost connection can take up to about 25 seconds more.
  • Subscription and search changes apply everywhere. A Subscription, SubscriptionTopic, or SearchParameter written on one instance is active on the others when that transaction commits.
  • Access changes apply everywhere. A write that changes authorization — for example deleting a PractitionerRole — evicts the affected cache entries on every instance as part of the commit, without waiting for the cache time-to-live.
  • A client can read its own writes. Send Cache-Control: no-cache or no-store on a request that must reflect a write just made through another instance. This instance waits until it has applied every authorization change committed before the request arrived, then answers using its caches. The wait is usually a few milliseconds. If it exceeds fire-arrow.cluster.read-fence-timeout (default 1 second), that request is authorized without the caches. Without the header, other instances usually apply the change within milliseconds of the commit, but a request that arrives in that window may not see it yet.
  • An instance that is not receiving cluster events authorizes without its caches until it is receiving them again. fire_arrow_cluster_bus_ready is 0 in that state.

fire-arrow.cluster.enabled defaults to auto. On PostgreSQL, cluster mode turns on when a startup check passes. On H2 it stays off. If the check fails, auto logs an error and runs that instance on its own, which is unsafe when more than one replica is serving traffic.

On every deployment with more than one replica, set fire-arrow.cluster.enabled to true (FIRE_ARROW_CLUSTER_ENABLED=true). A failed startup check then stops the instance instead of leaving the replicas uncoordinated.

The check fails behind a connection pooler in transaction or statement mode (for example PgBouncer with pool_mode=transaction), because coordination needs session-level PostgreSQL features. Connect directly to PostgreSQL, or through a session-mode pooler. Each instance opens three database sessions on top of the connection pool (leader lock, cluster events, and the read-your-writes check). Allow those three connections per instance in max_connections.

SettingWhy
fire-arrow.cluster.enabled: trueA failed cluster check stops the instance.
fire-arrow.mutex.provider: jdbcLocks are shared through the database. This is the default.
fire-arrow.graphql.cursor-hmac-secret (the same value on every instance)A GraphQL cursor issued by one instance is accepted by the others.
Elasticsearch for full-text searchA local Lucene index exists only on the instance that built it.
Shared storage for filesystem binaries, or Azure Blob / database storageOtherwise a binary stored on one instance is missing on the others.

Each instance logs a warning at startup for any of these that is not set safely. WebSocket subscriptions reach only clients connected to the instance that processed the change; use REST hook subscriptions when more than one instance is serving traffic. Session affinity is still required for composite search paging (for example $everything), whose pages stay in memory on the instance that ran the search. It is not required for authorization.

On shutdown, an instance finishes in-flight HTTP requests, pauses background jobs, and waits for in-flight subscription processing to finish before it gives up leadership. That wait is bounded by fire-arrow.cluster.quiesce-timeout (FIRE_ARROW_CLUSTER_QUIESCE_TIMEOUT, default 30 seconds). Set the orchestrator stop timeout higher than the preStop delay plus the HTTP shutdown grace (spring.lifecycle.timeout-per-shutdown-phase, default 30 seconds) plus the quiesce timeout — for example Kubernetes terminationGracePeriodSeconds of 75 seconds with the defaults.

Normal operation and a shutdown that finishes within the quiesce timeout keep existing subscription delivery behavior. An abrupt stop, or a quiesce that hits its timeout, can drop a delivery that was in flight. Duplicate deliveries can occur, so webhook consumers must be idempotent.

CarePlan reconcile backlogs are processed on the leader. The other instances do not repeat the startup reconcile scan.

A single PostgreSQL instance that passes the startup check also runs in cluster mode: it is the leader, and it uses the three extra database sessions. Client Cache-Control does not change authorization caching on one instance. H2 is unchanged. On PostgreSQL, hapi.fhir.subscription.immediately_queued defaults to true when it is unset and cluster mode is not disabled, so a synchronous CarePlan subscribe or renew served by a non-leader instance does not wait for the leader to drain subscriptions. Set that property yourself to keep another value.

Further detail, including the cluster metrics, is in the Cluster runbook.

Bug Fixes​

Replaced CarePlans kept Tasks for removed activities​

A CarePlan materialized in one shape, replaced with the output of PlanDefinition/$apply, and then renewed with $renew-due-events was supposed to match the new shape: Tasks for removed activities gone, Tasks for new activities created. The renewed plan instead kept both. Tasks for the removed activities stayed beside the new ones.

This showed up when the replacement had no CarePlan.period. Reconcile then moved the scheduling window start to the current time, and any Task due before that start was kept — including Tasks whose activity the replacement had dropped. Later reconciles did not delete them, because the window start does not move backward. Tasks for activities that are still on the plan stay under that same protection.

Starting with 2.5.0, reconcile deletes a non-terminal Task whose activity is not on the CarePlan as it is stored now, even when the Task is already past due or outside the scheduling horizon. Tasks for activities that are still defined are unchanged: a past-due Task, or one beyond the horizon, is still kept. Tasks already in a terminal status are not deleted. A Task that does not record which activity it came from, or whose activity is on a RequestGroup the server cannot read, is still kept or deleted only by its due time.

A replace that assigns new ids to contained RequestGroup resources makes the previous Tasks look like they belong to activities that no longer exist. The next reconcile removes those non-terminal Tasks, including past-due ones.

After upgrading, the next reconcile removes the leftover non-terminal Tasks. That reconcile is $renew-due-events, $subscribe-due-events, $materialize, or a CarePlan update. No configuration change is required.

Orphan-blob cleanup counted FHIR Binary blobs as foreign​

With fire-arrow.binary-storage.enabled=true, FHIR Binary content is stored separately from attachment blobs. The orphan-blob garbage collector did not recognize that layout. Every such blob was counted as unparseable (fire_arrow.binary.gc.skipped.unparseable), so on a deployment that offloads Binary resources the counter sat at the whole Binary population on every cycle. It could no longer mean "a blob this server does not own," and the dry-run review before turning deletion on was not usable. A Binary blob id that contained a . could also be read as an attachment name and deleted when no resource listed it.

Starting with 2.5.0, those blobs are counted on fire_arrow.binary.gc.skipped.hapi_managed and are never deleted by this collector. The server still removes a Binary blob when that Binary resource is expunged. skipped.unparseable again means a blob the collector does not recognize. Attachment cleanup and deletion are unchanged.

If you alert on fire_arrow.binary.gc.skipped.unparseable, re-baseline it. After upgrade it should fall to about zero when storage holds only Fire Arrow attachments and FHIR Binary content; the Binary population moves to fire_arrow.binary.gc.skipped.hapi_managed. Deployments that leave the garbage collector off are unaffected. No configuration change is required.