455 lines
24 KiB
Markdown
455 lines
24 KiB
Markdown
# Remote Scan Worker Operations
|
|
|
|
This runbook covers the opt-in remote worker boundary. Production defaults remain
|
|
disabled: `supervisor.worker_api.enabled` and
|
|
`supervisor.worker_api.admin.enabled` are both `false` in
|
|
`app/config.linux.yaml`. Enabling either one, applying schema changes, or starting
|
|
the production edge requires a separate reviewed rollout.
|
|
|
|
Remote workers are trusted clients. Their only operator-authored runtime settings
|
|
are the HTTPS server origin, one opaque device token, and a positive slot count
|
|
`N`. Target/source settings, immutable plans, scanner policy, limits, and only the
|
|
credentials needed for an assignment come from the server. Protocol-2 package
|
|
schema 3 manifests advertise exact `(source, platform, planning_kind)`
|
|
capabilities. The distributed core package contains GitLab `exact_git_v1`,
|
|
DockerHub `docker_direct_v1`, and HuggingFace `huggingface_space_v1`; GitHub is a
|
|
legacy optional capability and is not in the core profile. Detailed keycheck
|
|
remains server-only after bundle ingestion; clients must not receive keycheck
|
|
configuration or run keycheckers.
|
|
|
|
## Isolated verification
|
|
|
|
Run each verifier independently from the root of the isolated
|
|
`D:\truf-workers` checkout. Do not run plain `docker compose`, production
|
|
Compose files, import overrides, unrestricted pytest, or broad Docker cleanup.
|
|
Do not run these verifiers concurrently. They create random, ownership-labelled
|
|
resources and never use production data, credentials, volumes, or image tags.
|
|
|
|
### Main Docker E2E
|
|
|
|
From Linux or WSL, with the already-built local images
|
|
`truf-worker-test:runtime` and `truf-worker-test:test` and an already ignored
|
|
`docker/test-results/latest.json`:
|
|
|
|
```sh
|
|
python3 -I -S -B docker/verify.py
|
|
```
|
|
|
|
The verifier invokes only `compose.e2e.yaml`, does no build or pull, and normally
|
|
removes only its ownership-verified resources. It retains artifacts after a
|
|
failure. Use `--keep` only when a reviewed investigation needs stopped artifacts;
|
|
record the printed project name and never substitute prune, broad `down`, or
|
|
`down --volumes` commands.
|
|
|
|
### Packaged Windows/Linux client E2E
|
|
|
|
From Windows, with Windows Python, `wsl.exe`, passwordless `sudo -n docker` in the
|
|
selected WSL distribution, the built portable directory, and the already-built
|
|
Linux worker and test images:
|
|
|
|
```powershell
|
|
python -I -S -B docker/verify_packaged_workers.py `
|
|
--windows-artifact dist/truf-worker-windows-x86_64 `
|
|
--linux-image truf-remote-worker:linux-x86_64 `
|
|
--test-image truf-worker-test:test `
|
|
--wsl-distro Ubuntu-24.04
|
|
```
|
|
|
|
This single gate runs real packaged Windows and Linux scanners at `N=2`, tests a
|
|
server outage and restart recovery, and compares normalized cross-platform
|
|
evidence. It does not use Compose. Successful resources are removed unless
|
|
`--keep` is supplied; failures retain the labelled resources and the reported
|
|
`build/pwe-*` evidence directory.
|
|
|
|
### Edge E2E
|
|
|
|
From Windows with the same WSL Docker access and already-built
|
|
`truf-worker-test:test` and `truf-edge-e2e:test` images:
|
|
|
|
```powershell
|
|
python -I -S -B docker/verify_edge_e2e.py --wsl-distro Ubuntu-24.04
|
|
```
|
|
|
|
This gate uses only random labelled resources and a private synthetic backend. It
|
|
checks authenticated routes, two-failure admin bans, restart persistence,
|
|
automatic expiry, SSH-equivalent unban behavior, forwarded-header handling, and
|
|
worker availability from the banned admin IP. Success removes its resources;
|
|
failure reports the retained owned inventory and evidence path.
|
|
|
|
## Client bootstrap
|
|
|
|
Use a separately issued token for every device. The server stores only its
|
|
SHA-256 digest. Do not place a real token in documentation, source control,
|
|
shell transcripts, support output, or process diagnostics. The client validates
|
|
the server certificate and accepts only an HTTPS origin without credentials,
|
|
path, query, or fragment. `N` must be between 1 and 128; it bounds local occupied
|
|
slots but never overrides the user's server cap across devices.
|
|
|
|
Generated state, pending bundles, and work directories are recovery data, not
|
|
additional source/provider configuration. Preserve them across restarts until
|
|
the server authoritatively resolves the corresponding slots.
|
|
|
|
### Artifact source and distribution
|
|
|
|
Worker tools come from a reviewed release checkout; they are not installed
|
|
piecemeal on each client. The Windows builder and Linux worker package/image
|
|
targets assemble and verify the complete worker authority, Python runtime where
|
|
applicable, Git helper, TruffleHog binary, detector policy, CA roots, and
|
|
hash-locked dependencies. They deliberately exclude PostgreSQL tools,
|
|
server/runtime authority, provider implementations, detailed keycheck code,
|
|
server credentials, and database credentials.
|
|
|
|
There is no public worker download, image registry, installer, or automatic
|
|
updater. Build each release once in a controlled release environment, retain its
|
|
generated manifest/release metadata, and distribute the exact ZIP or image digest
|
|
through a trusted artifact channel. Do not rebuild independently on every worker.
|
|
Install the corresponding trusted package manifest beneath
|
|
`/etc/truf/worker-packages` on the server before allowing that artifact to claim
|
|
protocol-2 work.
|
|
|
|
Onboard a new worker in this order:
|
|
|
|
1. Create or select its server-side user and start with active-assignment cap `1`.
|
|
2. Create a distinct device identity and issue its token. The plaintext token is
|
|
shown once; never reuse it for another device.
|
|
3. Deliver the exact reviewed Windows ZIP, native Linux package, or Linux image,
|
|
and compare its package identity with the registered server manifest.
|
|
4. Create a private three-field YAML document with the HTTPS origin, token, and
|
|
conservative parallelism such as `1`, then run `install --config`. Delete the
|
|
input YAML after installation. This writes the existing private local
|
|
configuration; lifecycle commands never need the token on their command line.
|
|
5. Run `doctor`, start the supervisor, inspect `status` and `attach`, and confirm
|
|
server `last_contact_at`.
|
|
6. Reconcile one real assignment through accepted receipt, ingestion, settlement,
|
|
and projection before increasing either the user cap or local parallelism.
|
|
7. Preserve the private state tree across restart or outage. Drain work before
|
|
token rotation, revocation, artifact replacement, or state removal.
|
|
|
|
For a routine new device, artifact delivery and token issuance are the only
|
|
installation work. Building the artifact and registering its trusted manifest are
|
|
release-management operations and should not be delegated to the device operator.
|
|
|
|
The verified copy-paste command sheets are:
|
|
|
|
- `docs/remote-worker-cheatsheet-windows-ru.md`
|
|
- `docs/remote-worker-cheatsheet-linux-ru.md`
|
|
- `docs/remote-worker-cheatsheet-docker-ru.md`
|
|
|
|
### Windows portable client
|
|
|
|
Build the pinned amd64 package in the isolated checkout when producing a release:
|
|
|
|
```powershell
|
|
New-Item -ItemType Directory -Path dist -Force | Out-Null
|
|
python -B app/worker_package_builder.py windows `
|
|
--project-root . `
|
|
--output dist/truf-worker-windows-x86_64 `
|
|
--archive dist/truf-worker-windows-x86_64.zip `
|
|
--cache build/worker-cache
|
|
```
|
|
|
|
Verify the ZIP and adjacent release JSON through the release process. On the
|
|
client, extract to a private local directory and run `prepare-worker.ps1` once to
|
|
replace inherited ACLs. Create `worker-install.yaml` in that private directory and
|
|
put only `server`, `token`, and `parallelism` in it:
|
|
|
|
```powershell
|
|
.\prepare-worker.ps1
|
|
.\truf-worker.cmd install --config .\worker-install.yaml
|
|
Remove-Item -LiteralPath .\worker-install.yaml
|
|
.\truf-worker.cmd doctor
|
|
.\truf-worker.cmd start --startup-timeout 30
|
|
.\truf-worker.cmd status
|
|
```
|
|
|
|
By default, state and data are below the current user's `LOCALAPPDATA`. The
|
|
package verifies its manifest and application files before launch and bundles
|
|
Python, Git, TruffleHog, detector policy, and locked dependencies. There is no
|
|
automatic updater. Use the same Windows account for installation and operation;
|
|
the supervisor instance and private state belong to that account.
|
|
|
|
### Native Linux portable client
|
|
|
|
Install the reviewed package beneath a root-owned path such as
|
|
`/opt/truf-worker`, then run `sudo ./prepare-worker.sh` from that directory as the
|
|
worker OS user. The preparation keeps native Git and TruffleHog immutable and
|
|
root-owned while making application code exact-private to the worker user. It
|
|
does not install a systemd unit. Use the private YAML installation and lifecycle
|
|
commands in `docs/remote-worker-cheatsheet-linux-ru.md` from that package root.
|
|
|
|
### Linux client image
|
|
|
|
Build the worker-only image for the target architecture through the reviewed
|
|
`worker` target. For x86-64:
|
|
|
|
```sh
|
|
docker build --target worker -t truf-remote-worker:linux-x86_64 .
|
|
```
|
|
|
|
Install one device into a private persistent volume. Put the server, token, and
|
|
parallelism in a mode-0600 `worker-install.yaml`, pass it on standard input to the
|
|
short-lived install container, and delete it after success:
|
|
|
|
```sh
|
|
docker volume create truf-worker-device-a-data
|
|
chmod 600 worker-install.yaml
|
|
docker run --rm -i \
|
|
--mount type=volume,source=truf-worker-device-a-data,target=/data \
|
|
--env XDG_DATA_HOME=/data/client --env XDG_STATE_HOME=/data/state-base \
|
|
truf-remote-worker:linux-x86_64 install \
|
|
--config - < worker-install.yaml
|
|
rm -- worker-install.yaml
|
|
|
|
docker run --rm \
|
|
--mount type=volume,source=truf-worker-device-a-data,target=/data \
|
|
--env XDG_DATA_HOME=/data/client --env XDG_STATE_HOME=/data/state-base \
|
|
truf-remote-worker:linux-x86_64 doctor --json
|
|
```
|
|
|
|
Run the installed supervisor in the foreground under Tini while Docker supplies
|
|
detachment and restart policy:
|
|
|
|
```sh
|
|
docker run --detach --name truf-worker-device-a --restart unless-stopped \
|
|
--read-only --cap-drop ALL --security-opt no-new-privileges --pids-limit 256 \
|
|
--tmpfs /tmp:rw,nosuid,nodev,noexec,size=128m,mode=1777 \
|
|
--mount type=volume,source=truf-worker-device-a-data,target=/data \
|
|
--env XDG_DATA_HOME=/data/client --env XDG_STATE_HOME=/data/state-base \
|
|
truf-remote-worker:linux-x86_64 run
|
|
```
|
|
|
|
The image entrypoint supplies `python -u -I -S -B` and the integrity-checking
|
|
bootstrap. It contains no PostgreSQL client, server runtime, provider
|
|
implementations, or detailed keycheck code.
|
|
|
|
Ordinary private OS storage is supported on both platforms. Application-layer
|
|
encryption of the local workspace is not required. Operators may still use host
|
|
full-disk encryption according to their own endpoint policy; this is not a TRUF
|
|
protocol requirement.
|
|
|
|
## Daily worker operation
|
|
|
|
The commands below use `truf-worker.cmd` on Windows. On Linux without Docker, use
|
|
the generated `truf-worker` launcher. For a running Docker worker, define a local
|
|
helper that executes the same package bootstrap inside its container:
|
|
|
|
```sh
|
|
workerctl() {
|
|
docker exec truf-worker-device-a /usr/local/bin/python3 -u -I -S -B \
|
|
/opt/truf-worker/app/remote_worker_bootstrap.py -- "$@"
|
|
}
|
|
```
|
|
|
|
Use these commands for normal operation:
|
|
|
|
| Command | Purpose |
|
|
| --- | --- |
|
|
| `truf-worker status` | One current human-readable supervisor and slot snapshot. |
|
|
| `truf-worker status --json` | Versioned snapshot for automation. |
|
|
| `truf-worker attach --follow-seconds 300` | Follow the verified running instance for a bounded interval; detaching does not stop it. |
|
|
| `truf-worker watch --follow-seconds 300` | Human live-status alias over the same verified attach stream. |
|
|
| `truf-worker attach --ndjson --follow-seconds 300` | Machine-readable bounded event stream. |
|
|
| `truf-worker logs --tail 200` | Read bounded rotating supervisor logs. |
|
|
| `truf-worker logs --follow --follow-seconds 300` | Follow logs for a bounded interval. |
|
|
| `truf-worker history --limit 50` | Show terminal assignment outcomes and durations. |
|
|
| `truf-worker history --reservation <id> --json` | Retrieve one assignment's terminal local record. |
|
|
| `truf-worker doctor --json` | Validate package identity, directories, configuration, retention, instance state, and runtime prerequisites. |
|
|
|
|
Run `status` first when investigating. It reports configured/occupied slots,
|
|
current phase, phase age, scan deadline, assignment time remaining, child state,
|
|
last progress age, backoff/idle reason, pending retention data, and progress-outbox
|
|
cursor. It does not invent percentage completion.
|
|
|
|
### Phases and deadlines
|
|
|
|
| Phase | Operator interpretation |
|
|
| --- | --- |
|
|
| `idle`, `claiming` | Slot is available or asking the server for work. |
|
|
| `assigned` | Immutable assignment identity and deadlines are persisted locally. |
|
|
| `waiting_permit` | Assignment owns a slot but is waiting for the shared scanner permit. This time counts against the scan-stage deadline. |
|
|
| `preparing`, `resolving`, `downloading`, `cloning` | Source-specific preparation before scanning. |
|
|
| `scanning` | Scanner process tree is active. |
|
|
| `filtering`, `cleaning`, `bundling` | Findings are converted, work is cleaned, and deterministic result bytes are staged. These phases remain inside the hard scan-stage deadline. |
|
|
| `uploading`, `awaiting_receipt` | Staged result is being transferred or waiting for authoritative server acknowledgement. |
|
|
| `backoff` | A bounded retry delay is active; inspect the reason and next-claim time. |
|
|
| `draining`, `stopped` | No new local work is starting; existing work is resolving or shutdown completed. |
|
|
|
|
The scan-stage deadline starts before `waiting_permit` and covers preparation,
|
|
provider access, scanning, filtering, cleanup, and bundle staging. Crossing it
|
|
terminates the contained runner tree and produces a normal phase-specific timeout
|
|
result while the assignment upload window remains available. The server's
|
|
assignment deadline is fixed at issue time and is never renewed by progress,
|
|
polling, restart, or upload retries. `status` therefore presents both deadlines
|
|
separately.
|
|
|
|
### Diagnostics and retained evidence
|
|
|
|
`history` identifies the terminal receipt, prebundle, timeout, stale, or recovery
|
|
outcome. `logs` shows supervisor operation; diagnostic records carry the stable
|
|
phase/category/code and optional body/log material. The private state tree stores:
|
|
|
|
```text
|
|
events/worker-events.jsonl
|
|
history/worker-history.jsonl
|
|
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.json
|
|
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.body
|
|
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.log
|
|
logs/worker.log
|
|
```
|
|
|
|
Diagnostic JSON records original/stored sizes, hash, encoding, and truncation
|
|
state. A missing or truncated body must not be described as complete. The server
|
|
admin detail view separately shows the ordered progress/receipt/ingestion/
|
|
settlement/projection timeline and canonical diagnostic snapshot. Assignment
|
|
transport outcome, scan outcome, and diagnostics are independent fields.
|
|
|
|
### Graceful stop and drain
|
|
|
|
Request local stop before maintenance. The supervisor closes new claims for this
|
|
worker, leaves authentication and the Worker API available while existing work and
|
|
pending uploads resolve, and requires a clean drain receipt:
|
|
|
|
```powershell
|
|
.\truf-worker.cmd stop --timeout 120 --json
|
|
```
|
|
|
|
For Docker, run `workerctl stop --timeout 120 --json`; the foreground supervisor
|
|
then exits and Tini returns the shutdown code to Docker. A successful shutdown
|
|
receipt reports `drained: true` and `exit_code: 0`. Do not interpret `docker stop`,
|
|
process termination, or loss of contact as assignment cancellation.
|
|
|
|
### Recovery
|
|
|
|
After an OS restart, network outage, server outage, or unclean process exit:
|
|
|
|
1. Preserve the entire private state/volume; do not remove work, bundle, event,
|
|
history, or control files.
|
|
2. Run `doctor --json` and `status --json` from the exact installed package.
|
|
3. Start the same package with `start` on Windows/Linux, or restart the same Docker
|
|
container/volume. The supervisor replays its journal and recovers assigned,
|
|
staged, upload-retry, and awaiting-receipt slots.
|
|
4. Use `attach --ndjson --follow-seconds 300` and server assignment detail to
|
|
distinguish active recovery from backoff.
|
|
5. Keep the same device identity and state until recovered slots and
|
|
accepted-but-not-ingested bundles are reconciled. Escalate only if the instance
|
|
is unverifiable, a fixed assignment deadline has passed without server recovery,
|
|
or repeated startup validation fails.
|
|
|
|
Re-uploading identical accepted bytes returns the original receipt. Conflicting
|
|
bytes or stale ownership are rejected; never delete local bytes merely to silence
|
|
that signal.
|
|
|
|
### Update and rollback
|
|
|
|
1. Request local graceful stop, reconcile unresolved/precommit work, and obtain a
|
|
successful clean drain receipt.
|
|
2. Retain the complete state tree and previous exact artifact/digest.
|
|
3. Verify the new ZIP/image and its registered server manifest. On Windows run
|
|
`prepare-worker.ps1`, then `doctor`; for Docker recreate only the container and
|
|
mount the same volume.
|
|
4. Start at parallelism `1`, confirm package identity/contact and one complete
|
|
accepted-ingested-projected assignment, then restore the intended cap.
|
|
5. If validation fails, stop and return to the previous exact artifact with the
|
|
same state. Additive server progress/diagnostic records need no rollback.
|
|
|
|
Do not replace binaries beneath a running supervisor or switch packages while an
|
|
assignment runner is active.
|
|
|
|
### Device removal
|
|
|
|
1. Stop every device that uses the identity locally; leave authentication valid
|
|
while all unresolved assignments and precommit bundles reach zero.
|
|
2. Stop gracefully and retain the shutdown receipt and terminal history.
|
|
3. Revoke the device and disable its user only if that user is not shared by an
|
|
active device.
|
|
4. Confirm the old token no longer authenticates and no authoritative recovery
|
|
remains.
|
|
5. Remove the container/package. Remove its private volume/state only after the
|
|
server reconciliation evidence is retained and no rollback requires it.
|
|
|
|
## Server tuning
|
|
|
|
Make tuning changes in the reviewed private runtime configuration, not on the
|
|
client. Keep the raw worker service on loopback/private addressing and expose it
|
|
only through certificate-validating Caddy HTTPS.
|
|
|
|
| Control | Meaning |
|
|
| --- | --- |
|
|
| Client `--parallelism N` | Maximum locally occupied slots, one claim per free slot. |
|
|
| Typed admin assignment cap | Atomic positive active-assignment cap for one user across all devices; lowering it does not cancel existing assignments. |
|
|
| `assignment_ttl_seconds` | Fixed server-clock lifetime covering download, scan, and upload; default `86400`. API contact and restart do not renew it. |
|
|
| `bundle_body_timeout_seconds` | Upload body deadline; default `1800`. The assignment lifetime must exceed the largest configured source scan timeout plus this value plus 60 seconds. |
|
|
| `json_body_timeout_seconds` / `body_idle_timeout_seconds` | Request and idle transport bounds; defaults `60` and `30`. They do not renew ownership. |
|
|
| `reaper_interval_seconds` / `reaper_batch_size` | Expired-assignment recovery cadence and bounded batch; defaults `60` and `1000`. |
|
|
| `limit_concurrency` | Worker API request concurrency bound, not a replacement for user quotas; default `64`. |
|
|
|
|
Shortening the fixed lifetime can reject a valid long scan or upload and allow a
|
|
second physical execution after recovery. Lengthening it holds user quota,
|
|
target ownership, and dependent leases longer after a lost client. Tune it from
|
|
observed end-to-end duration plus upload headroom, not from HTTP polling cadence.
|
|
|
|
## Typed administration
|
|
|
|
Use only the authenticated random-prefix admin page configured by the production
|
|
edge. It provides CSRF/Origin-checked typed operations to create/enable/disable
|
|
users, set assignment caps, issue/rotate/revoke/unrevoke device tokens, and
|
|
requeue selected deferred queue IDs. It deliberately provides no shell or generic
|
|
supervisor command.
|
|
|
|
An issued or rotated token is shown once. Rotation replaces the stored digest, so
|
|
the old token stops authenticating; update that device without copying the token
|
|
to other devices. Rotation does not clear an existing revoked state. Revocation
|
|
blocks further API authentication but does not invent cancellation for assigned
|
|
work; plan for outstanding work to be completed before revocation or recovered at
|
|
its fixed expiry. Use local graceful stop to close claims before planned device
|
|
maintenance or removal while preserving authentication for pending work.
|
|
|
|
For an admin IP ban, use SSH and fail2ban first so fail2ban and Caddy agree:
|
|
|
|
```sh
|
|
sudo fail2ban-client set truf-admin-auth unbanip 203.0.113.10
|
|
sudo /usr/local/sbin/truf-caddy-admin-denylist status
|
|
sudo /usr/local/sbin/truf-caddy-admin-denylist expire
|
|
```
|
|
|
|
If fail2ban is unavailable, use the explicit admin-only updater:
|
|
|
|
```sh
|
|
sudo /usr/local/sbin/truf-caddy-admin-denylist unban 203.0.113.10
|
|
```
|
|
|
|
These commands change only the admin-route matcher. Never replace them with a
|
|
global port 443 firewall unban/ban. The full damaged-snippet recovery procedure
|
|
is in `deploy/edge/README.md`.
|
|
|
|
## Accepted custody and expiry
|
|
|
|
`accepted` means the server validated the canonical v2 `.trb`, durably published
|
|
it, and persisted ready/recovery state before returning a receipt. It does not
|
|
mean ingestion, projection, candidate handling, or detailed keycheck has
|
|
finished. Track `accepted` and `ingested` separately in the typed admin view.
|
|
|
|
The client keeps assignment identity and pending bytes until authoritative
|
|
acknowledgement. Retrying identical accepted bytes returns the original receipt,
|
|
including after ingestion, spool cleanup, restart, or the former deadline;
|
|
conflicting bytes are rejected. An interrupted, invalid, expired-before-first-
|
|
acceptance, or stale upload is not a successful scan.
|
|
|
|
An unfinished assignment expires at the original server-set deadline, normally
|
|
24 hours after issue. The periodic recovery pass reconciles quota, credits,
|
|
target ownership, and dependent plan/blob leases through existing retry policy.
|
|
Already accepted ready bundles are not requeued as unfinished. A crashed client
|
|
can therefore delay work for about a day, while a partitioned client can continue
|
|
physical scanning after the server has expired and reissued the target. Ownership
|
|
fencing guarantees one authoritative acceptance, not exactly-once physical work.
|
|
|
|
## Staged rollout and rollback
|
|
|
|
1. Obtain separate review for production schema/deployment changes. Drain protocol-1 work before replacing packages. Keep worker API and admin disabled while configuring exact protocol-2 capability profiles, task-specific auth entries, trusted package manifests, edge origin/marker, and per-user caps.
|
|
2. Pass the main Docker, packaged Windows/Linux, and edge gates with their pinned artifacts. Do not infer readiness from unit mocks or one platform.
|
|
3. Enable a small canary: one user, one device, cap `1`, and client `N=1`. Confirm assignment, acceptance, ingestion, projection, detailed server keycheck, expiry, and safe logs before increasing either cap or device count.
|
|
4. Expand caps and clients in stages while comparing unfinished/completed/failed/expired counts, last authenticated contact, accepted-versus-ingested state, durations, capacity, and stale/duplicate events. Contact age alone is not a liveness failure while a client holds long-running work and has not yet returned to claim polling.
|
|
5. To drain, request local graceful stop on every affected device. Leave authentication and the Worker API available so pending uploads and terminal reports can resolve. Wait for unfinished assignments to complete or pass through fixed-expiry recovery, and separately reconcile accepted bundles awaiting ingestion.
|
|
6. After the drain is authoritative, disable remote admission and return scheduling to local-only execution. Then revoke unused device tokens if required. Retain reservation metadata, accepted receipts, bundles/recovery state, and schema until reviewed reconciliation is complete.
|
|
7. Do not drop worker metadata, clear spool state, rotate/revoke tokens before a drain, switch server binaries underneath unfinished work, or treat a service stop as cancellation. The local execution path remains available and must not depend on a remote worker.
|