Skip to main content
Version: 0.6.0

Worker

The worker is a second container from the same image (ri-worker in Docker Compose; pnpm start:worker in a checkout) with its own entrypoint and no port. The web process only enqueues background jobs; the worker claims and settles them.

Its jobs are the credential verifications behind the library and the issuance of credential batch items. When a credential is registered, issued or re-verified, the web process records the request and queues a job. The worker fetches the credential, checks its digest and signature, its status and validity window, and its schema conformance, extracts the details the library shows, and writes the result to the record's verification envelope. A scheduled reconciliation sweep settles jobs a crashed worker abandoned, so a record cannot stay pending indefinitely. A batch submitted to POST /api/v1/credentials/batches is issued here too, one item at a time, with the same issuance code as the single credential route. Without a running worker, registrations and re-verifications stay pending, and batch items are never issued.

Worker Boot​

The worker's wrapper passes --process-role=worker as the first argument to the shared entrypoint. The shared entrypoint runs the worker preflight before database URL construction, migrations and the worker process. The worker issues batch items with the same code as POST /api/v1/credentials, so its preflight requires RI_APP_URL and checks the issuance settings those items read, with the same rules as the web process. The settings table below lists them. The worker does not check settings only the web process reads, such as MAX_REQUEST_BODY_BYTES and IDEMPOTENCY_STALE_CLAIM_MINUTES. An environment that omits those still starts a worker. A boot failure is one line on stderr, Worker boot failed: <message> [<code>], with the cause chain beneath it, and exit code 1. An invalid LOG_REDACT_PATHS fails the same way with the logger's own message naming the path.

Reading the queue​

The application uses five pg-boss queues: library.verify-generation carries library verification work, library.reconcile-pending-runs carries the verification reconciliation sweep, credentials.issue-batch carries sequential batch issuance, credentials.reconcile-batches recovers batches whose issuance job vanished, and credentials.expire-batches removes retained batch item data. BATCH_JOB_CONCURRENCY controls parallel batches in one worker process and defaults to one, and it never makes items within one batch concurrent. A pre-dispatch fault is requeued for its item with a doubling delay starting at BATCH_JOB_RETRY_BACKOFF_SECONDS (default 30 seconds) and capped at BATCH_JOB_RETRY_BACKOFF_MAX_SECONDS (default 600 seconds), with BATCH_JOB_RETRY_LIMIT (default 4) setting the total attempts for that item, so never-attempted items can run before deferred retries; when the batch has a cancellation request, a faulted item that still has attempts left is cancelled instead of requeued, while an item that exhausts its attempts still becomes FAILED. A post-dispatch fault still throws for queue recovery. The batch issuance runbook covers what to do when a batch needs an operator. A growing backlog means work is arriving faster than the worker can finish it, or jobs are retrying or waiting for unavailable dependencies.

Worker settings​

The worker reads these settings from its environment. Unset and blank values use the documented defaults where shown.

VariableDefaultBoundsReaderValidation behaviour
WORKER_JOB_TIMEOUT_SECONDS300 secondsInteger from 30 seconds to 86400 secondsreadWorkerJobTimeoutSecondsUnset or blank uses the default. A non-integer, fractional, or out-of-range value fails worker configuration before the queue starts.
LIBRARY_RECONCILE_PENDING_RUNS_CRON*/10 * * * *Any expression accepted by cron-parser with UTC and non-strict parsingreadReconcilePendingRunsCronUnset or blank uses the default. A value the parser rejects fails worker configuration and names the variable.
LIBRARY_RECONCILE_PENDING_RUNS_BATCH_SIZE500Positive integer no greater than 10000readReconcilePendingRunsBatchSizeUnset or blank uses the default. A non-positive, fractional, or over-limit value fails worker configuration before the queue starts.
BATCH_JOB_CONCURRENCY1Positive integerreadBatchJobConcurrencyUnset or blank uses the default. Any other value fails worker configuration before the queue starts. It sets how many batches run at once, never how items inside a batch run.
BATCH_EXPIRY_SWEEP_MINUTES60 minutesPositive integer under 60, or a multiple of 60readBatchExpirySweepCronUnset or blank uses the default. Any other value fails the boot preflight in the worker role, because the value becomes the sweep's cron expression.
BATCH_RETENTION_DAYS30 daysPositive integerreadBatchRetentionDaysUnset or blank uses the default. Any other value fails the boot preflight in the worker role. It is applied when a batch settles, so a change does not move an existing deadline.
MAX_BATCH_ITEMS500Positive integerreadMaxBatchItemsUnset or blank uses the default. Any other value fails the boot preflight in both roles. Only the web process enforces it, on submission.
WORKER_HEARTBEAT_PATH/tmp/worker-heartbeatA writable file path shared by the worker and its healthcheckheartbeat.ts and docker-worker-healthcheck.shNo application-level parsing is applied. The worker removes and atomically writes the path; a path that cannot be used leaves the healthcheck without a valid heartbeat.
WORKER_HEARTBEAT_MAX_AGE_SECONDS30 secondsNo application-defined ceiling; the healthcheck expects a numeric age in secondsdocker-worker-healthcheck.shThe shell healthcheck rejects a non-numeric comparison and reports unhealthy when the heartbeat is missing, future-dated, or older than the configured limit.
OTEL_WORKER_SERVICE_NAMEreference-implementation-worker in ComposeA non-blank service name; whitespace is trimmedCompose maps it to the worker process's OTEL_SERVICE_NAME, read by resolveServiceNameBlank or unset uses the worker default. The process emits the resolved value as service.name.
DATA_ENCRYPTION_KEYNone64-character hexadecimal application encryption keyresolveDataEncryptionKey and validateConfiguredEncryptionKeyRequired for the worker. Boot fails when it is absent, has the invalid key format, is the development placeholder outside local development, or cannot validate against existing encrypted data.
SCHEMA_CACHE_TTL_MS3600000 msNon-negative finite numberreadSchemaCacheTtlMs in schema-loader.tsUnset or blank uses one hour. An invalid or negative value logs a warning and falls back to one hour.
CONTEXT_CACHE_TTL_MS3600000 msNon-negative finite numberreadContextCacheTtlMs in context-cache.tsUnset or blank uses one hour. An invalid or negative value logs a warning and falls back to one hour.
CACHE_MAX_ENTRIES1000 per cachePositive integerreadCacheMaxEntriesUnset or blank uses the default. An invalid value throws when a document cache is constructed; the web process also checks it at Node boot.
RI_HTTP_USER_AGENTBuilt-in uncefact-untp-utils valueNon-blank Latin-1 text without control characters when setThe guarded fetch's document resolverUnset or blank uses the built-in value. The web process validates a configured override at boot; the worker's guarded fetch uses the same value and an invalid header value cannot be sent.
RI_APP_URLNoneAn http(s) URL without a username or passwordresolveAppUrlRequired. Use the same value as the web process. Boot fails when it is unset, is not a valid http(s) URL, or carries a username or password. A batch item published without humanVerificationUrl takes its default verification link from it.
FETCH_ALLOW_PRIVATE_URLSfalseExact lowercase true permits private, loopback and reserved addressesreadFetchAllowPrivateUrlsUse the same value as the web process. The old name VERIFY_ALLOW_PRIVATE_URLS is still read, with a rename warning. Set only one of the two names: setting both fails boot, even when the values are equal. Any value other than true keeps private-address protection on and is not an error.
CREDENTIAL_STATUS_DEFAULT_PURPOSESrevocationComma-separated unique values from revocation and suspension, or none alonereadDefaultStatusPurposesUse the same value as the web process. An invalid value fails boot. A value naming more than one purpose fails boot unless CREDENTIAL_STATUS_MULTIPLE_PURPOSES_ENABLED is true. none logs a warning that credentials issued without explicit statusPurposes carry no status entry.
CREDENTIAL_STATUS_MULTIPLE_PURPOSES_ENABLEDfalseExact true or falsereadStatusMultiplePurposesEnabledUse the same value as the web process. Unset or empty means false. Any other value fails boot. Unless it is true, a batch item that asks for more than one status purpose is refused.
CREDENTIAL_STATUS_LOCK_ACQUIRE_MS2000 msPositive integerreadStatusLockAcquireMsUse the same value as the web process. It bounds the wait for the status-list lock when the worker signs a batch item. A value that begins with a positive integer is used as that integer without a warning, so 1e3 and 1.5 mean 1 ms and 1000ms means 1000 ms. Any other non-blank value, such as 0, -5 or abc, logs a warning and uses the 2000 ms default.

The worker needs outbound HTTPS access to the UNTP and VC Data Model artefact hosts: untp.unece.org, vocabulary.uncefact.org, test.uncefact.org, www.w3.org, and w3c.github.io. It honours the same SCHEMA_CACHE_TTL_MS, CONTEXT_CACHE_TTL_MS, CACHE_MAX_ENTRIES, and RI_HTTP_USER_AGENT settings as the web process. When an enabled bundled-artefact fallback has a matching copy and a host is unreachable, that bundled copy stands in for the remote artefact.

The Compose worker blocks pass most of these settings through only when they are set. RI_APP_URL and the private-URL setting are given the same expressions as the web service, so the worker runs with the web's values. The application defaults and validation remain the source of truth.

If you override RI_APP_URL, the private-URL setting or a status issuance setting for ri, override it for ri-worker with the same value. Compose merges environment per service, so an override file that changes ri alone leaves the worker on the value in docker-compose.yml. Both services set RI_APP_URL to the literal http://localhost:3003 there, and a value in .env does not change it. A worker left on that value gives a batch item published without humanVerificationUrl a verification link on localhost.

LIBRARY_STORED_COPY_READ_TIMEOUT_MS is no longer read. If it remains set, the application ignores it. Use WORKER_JOB_TIMEOUT_SECONDS for the shared job expiry and working budget.

The verification job uses its remaining budget for a second provider call with the validity window skipped when the first call stops at that window, so proof and status can still be established.

Issuance settings in your own deployment​

A batch item and a request to POST /api/v1/credentials go through the same issuance code, and that code reads RI_APP_URL, the private-URL setting and the three status issuance settings from the environment of the process running it. The Compose files in this repository give the worker the same value for each as the web service. A deployment that defines its own containers has to do the same: give the worker each of these settings with the value the web process actually runs with. That includes a setting the web process leaves unset. For example, if the web process has no CREDENTIAL_STATUS_DEFAULT_PURPOSES, it issues with the built-in revocation default, so leave the setting unset on the worker too. A worker given suspension alone would issue batch items with a different status purpose from the single route. Set the private-URL setting under one name only, preferably FETCH_ALLOW_PRIVATE_URLS. The status operation settings (CREDENTIAL_STATUS_OPERATION_BUDGET_MS, CREDENTIAL_STATUS_RECONCILE_GRACE_MS and CREDENTIAL_STATUS_MUTATION_ENABLED) stay on the web process only, because the worker does not change issuer status.

Stopping​

Three numbers govern a stop, each with its own job. On SIGTERM or SIGINT the worker stops taking jobs and gives a running one 30 seconds to finish (the drain); a job still running then is failed by the queue and retried later. The whole shutdown (drain, queue release, database disconnect, telemetry flush with 5 seconds to reach the collector) must finish inside a 45-second process deadline, after which the worker exits non-zero. The container's grace period, 60 seconds, sits above that deadline so the runtime never kills the process before it has exited on its own terms: stop_grace_period: 60s in Compose (already set on ri-worker), terminationGracePeriodSeconds: 60 in Kubernetes, docker stop -t 60 by hand. So a job that needs 40 seconds is interrupted by the drain even though it fits the container's grace period. The worker exits non-zero if the queue release or the database disconnect fails or the deadline passes; a second signal exits at once; a telemetry flush that fails, which is the normal case when no collector is running, is logged and does not change the exit code. After an abrupt kill the interrupted job is not handed over at once: it stays claimed until its WORKER_JOB_TIMEOUT_SECONDS attempt expiry, five minutes by default, and the queue's maintenance sweep notices, then waits out the retry backoff (30 seconds as the base, growing and randomised on later attempts), so the next attempt comes minutes later on the same job id, provided the job has attempts left (four retries). A job that exhausts them settles its generation as a retryable failure on that final attempt, and a job that never reports at all is settled by the reconciliation sweep above. Either way the caller re-verifies the record to run the check again.

Health​

The worker serves no HTTP, so it proves itself instead. Every 10 seconds it runs a probe through the job queue's own connection pool and checks that a consumer has fetched within the last 30 seconds or is inside a job it started within WORKER_JOB_TIMEOUT_SECONDS + 60 seconds (360 seconds by default; the queue keeps a consumer's job count after a settlement that threw, so an older count is not work); when that holds it publishes /tmp/worker-heartbeat (atomically, so a failed write never leaves an empty but fresh file). The container health check (docker-worker-healthcheck.sh, wired in the compose files) reads that file's age and reports unhealthy when it is missing, older than 30 seconds, or stamped in the future. That catches a wedged event loop, a queue pool that is down or exhausted, and a consumer that has stopped fetching, while an idle worker stays healthy. The worker does not exit on a failed probe: both of its database pools recover a lost connection on their own, and exiting would restart every worker on a shared outage without repairing it. Unhealthy is the signal. An orchestrator's liveness probe restarts on it; under Compose it is what docker compose ps shows, and restart: always still replaces a worker that exits. From the last successful beat, Compose reports unhealthy after roughly 50 to 60 seconds (the age limit plus three failed checks); count that, plus the 45-second shutdown deadline, into any orchestrator's timings. During a stop the last proof stays in place rather than being removed, so the health check keeps passing for as long as that proof is inside its age limit and then needs three consecutive failures before the container is unhealthy. With the numbers above that is usually enough to cover the 30-second drain, though not by a wide margin, and it depends on how old the proof was when the signal arrived and where the check's schedule falls; an orchestrator that restarts on liveness should carry the 60-second termination grace so a drain in progress is left to finish. Nothing yet raises an alert on a growing backlog; the queue tables hold that and the health check does not read them.

Running it elsewhere​

The container must replace the image's entrypoint with /app/docker-worker-entrypoint.sh (it sets the flags that skip schema convergence and passes --process-role=worker as the first argument before the shared entrypoint reads them), replace the image's HTTP health check with /app/docker-worker-healthcheck.sh, and carry a restart policy. It also needs the worker's environment, including the issuance settings described under Issuance settings in your own deployment. Bare Docker: docker run --entrypoint /app/docker-worker-entrypoint.sh --health-cmd /app/docker-worker-healthcheck.sh --health-interval 10s --health-timeout 5s --health-retries 3 --health-start-period 40s --restart always --stop-timeout 60 <image>. Kubernetes: command: ["/app/docker-worker-entrypoint.sh"], a livenessProbe that execs /app/docker-worker-healthcheck.sh (period 10 s, failure threshold 3, initial delay 40 s), terminationGracePeriodSeconds: 60, no Service. The heartbeat file lives under /tmp, which the image's nextjs user can write; a read-only root filesystem needs a writable mount there, and WORKER_HEARTBEAT_PATH moves the file for both the worker and the check. WORKER_HEARTBEAT_MAX_AGE_SECONDS raises the check's 30-second age limit for a deployment whose probe schedule needs more room. The worker's own 10-second beat is not configurable, so a limit set below three beats will report a healthy worker unhealthy.