Files
truf-server/docs/remote-worker-operations.md
2026-09-30 20:30:56 +03:00

455 lines
24 KiB
Markdown

# Remote Scan Worker Operations
This runbook covers the opt-in remote worker boundary. Production defaults remain
disabled: `supervisor.worker_api.enabled` and
`supervisor.worker_api.admin.enabled` are both `false` in
`app/config.linux.yaml`. Enabling either one, applying schema changes, or starting
the production edge requires a separate reviewed rollout.
Remote workers are trusted clients. Their only operator-authored runtime settings
are the HTTPS server origin, one opaque device token, and a positive slot count
`N`. Target/source settings, immutable plans, scanner policy, limits, and only the
credentials needed for an assignment come from the server. Protocol-2 package
schema 3 manifests advertise exact `(source, platform, planning_kind)`
capabilities. The distributed core package contains GitLab `exact_git_v1`,
DockerHub `docker_direct_v1`, and HuggingFace `huggingface_space_v1`; GitHub is a
legacy optional capability and is not in the core profile. Detailed keycheck
remains server-only after bundle ingestion; clients must not receive keycheck
configuration or run keycheckers.
## Isolated verification
Run each verifier independently from the root of the isolated
`D:\truf-workers` checkout. Do not run plain `docker compose`, production
Compose files, import overrides, unrestricted pytest, or broad Docker cleanup.
Do not run these verifiers concurrently. They create random, ownership-labelled
resources and never use production data, credentials, volumes, or image tags.
### Main Docker E2E
From Linux or WSL, with the already-built local images
`truf-worker-test:runtime` and `truf-worker-test:test` and an already ignored
`docker/test-results/latest.json`:
```sh
python3 -I -S -B docker/verify.py
```
The verifier invokes only `compose.e2e.yaml`, does no build or pull, and normally
removes only its ownership-verified resources. It retains artifacts after a
failure. Use `--keep` only when a reviewed investigation needs stopped artifacts;
record the printed project name and never substitute prune, broad `down`, or
`down --volumes` commands.
### Packaged Windows/Linux client E2E
From Windows, with Windows Python, `wsl.exe`, passwordless `sudo -n docker` in the
selected WSL distribution, the built portable directory, and the already-built
Linux worker and test images:
```powershell
python -I -S -B docker/verify_packaged_workers.py `
--windows-artifact dist/truf-worker-windows-x86_64 `
--linux-image truf-remote-worker:linux-x86_64 `
--test-image truf-worker-test:test `
--wsl-distro Ubuntu-24.04
```
This single gate runs real packaged Windows and Linux scanners at `N=2`, tests a
server outage and restart recovery, and compares normalized cross-platform
evidence. It does not use Compose. Successful resources are removed unless
`--keep` is supplied; failures retain the labelled resources and the reported
`build/pwe-*` evidence directory.
### Edge E2E
From Windows with the same WSL Docker access and already-built
`truf-worker-test:test` and `truf-edge-e2e:test` images:
```powershell
python -I -S -B docker/verify_edge_e2e.py --wsl-distro Ubuntu-24.04
```
This gate uses only random labelled resources and a private synthetic backend. It
checks authenticated routes, two-failure admin bans, restart persistence,
automatic expiry, SSH-equivalent unban behavior, forwarded-header handling, and
worker availability from the banned admin IP. Success removes its resources;
failure reports the retained owned inventory and evidence path.
## Client bootstrap
Use a separately issued token for every device. The server stores only its
SHA-256 digest. Do not place a real token in documentation, source control,
shell transcripts, support output, or process diagnostics. The client validates
the server certificate and accepts only an HTTPS origin without credentials,
path, query, or fragment. `N` must be between 1 and 128; it bounds local occupied
slots but never overrides the user's server cap across devices.
Generated state, pending bundles, and work directories are recovery data, not
additional source/provider configuration. Preserve them across restarts until
the server authoritatively resolves the corresponding slots.
### Artifact source and distribution
Worker tools come from a reviewed release checkout; they are not installed
piecemeal on each client. The Windows builder and Linux worker package/image
targets assemble and verify the complete worker authority, Python runtime where
applicable, Git helper, TruffleHog binary, detector policy, CA roots, and
hash-locked dependencies. They deliberately exclude PostgreSQL tools,
server/runtime authority, provider implementations, detailed keycheck code,
server credentials, and database credentials.
There is no public worker download, image registry, installer, or automatic
updater. Build each release once in a controlled release environment, retain its
generated manifest/release metadata, and distribute the exact ZIP or image digest
through a trusted artifact channel. Do not rebuild independently on every worker.
Install the corresponding trusted package manifest beneath
`/etc/truf/worker-packages` on the server before allowing that artifact to claim
protocol-2 work.
Onboard a new worker in this order:
1. Create or select its server-side user and start with active-assignment cap `1`.
2. Create a distinct device identity and issue its token. The plaintext token is
shown once; never reuse it for another device.
3. Deliver the exact reviewed Windows ZIP, native Linux package, or Linux image,
and compare its package identity with the registered server manifest.
4. Create a private three-field YAML document with the HTTPS origin, token, and
conservative parallelism such as `1`, then run `install --config`. Delete the
input YAML after installation. This writes the existing private local
configuration; lifecycle commands never need the token on their command line.
5. Run `doctor`, start the supervisor, inspect `status` and `attach`, and confirm
server `last_contact_at`.
6. Reconcile one real assignment through accepted receipt, ingestion, settlement,
and projection before increasing either the user cap or local parallelism.
7. Preserve the private state tree across restart or outage. Drain work before
token rotation, revocation, artifact replacement, or state removal.
For a routine new device, artifact delivery and token issuance are the only
installation work. Building the artifact and registering its trusted manifest are
release-management operations and should not be delegated to the device operator.
The verified copy-paste command sheets are:
- `docs/remote-worker-cheatsheet-windows-ru.md`
- `docs/remote-worker-cheatsheet-linux-ru.md`
- `docs/remote-worker-cheatsheet-docker-ru.md`
### Windows portable client
Build the pinned amd64 package in the isolated checkout when producing a release:
```powershell
New-Item -ItemType Directory -Path dist -Force | Out-Null
python -B app/worker_package_builder.py windows `
--project-root . `
--output dist/truf-worker-windows-x86_64 `
--archive dist/truf-worker-windows-x86_64.zip `
--cache build/worker-cache
```
Verify the ZIP and adjacent release JSON through the release process. On the
client, extract to a private local directory and run `prepare-worker.ps1` once to
replace inherited ACLs. Create `worker-install.yaml` in that private directory and
put only `server`, `token`, and `parallelism` in it:
```powershell
.\prepare-worker.ps1
.\truf-worker.cmd install --config .\worker-install.yaml
Remove-Item -LiteralPath .\worker-install.yaml
.\truf-worker.cmd doctor
.\truf-worker.cmd start --startup-timeout 30
.\truf-worker.cmd status
```
By default, state and data are below the current user's `LOCALAPPDATA`. The
package verifies its manifest and application files before launch and bundles
Python, Git, TruffleHog, detector policy, and locked dependencies. There is no
automatic updater. Use the same Windows account for installation and operation;
the supervisor instance and private state belong to that account.
### Native Linux portable client
Install the reviewed package beneath a root-owned path such as
`/opt/truf-worker`, then run `sudo ./prepare-worker.sh` from that directory as the
worker OS user. The preparation keeps native Git and TruffleHog immutable and
root-owned while making application code exact-private to the worker user. It
does not install a systemd unit. Use the private YAML installation and lifecycle
commands in `docs/remote-worker-cheatsheet-linux-ru.md` from that package root.
### Linux client image
Build the worker-only image for the target architecture through the reviewed
`worker` target. For x86-64:
```sh
docker build --target worker -t truf-remote-worker:linux-x86_64 .
```
Install one device into a private persistent volume. Put the server, token, and
parallelism in a mode-0600 `worker-install.yaml`, pass it on standard input to the
short-lived install container, and delete it after success:
```sh
docker volume create truf-worker-device-a-data
chmod 600 worker-install.yaml
docker run --rm -i \
--mount type=volume,source=truf-worker-device-a-data,target=/data \
--env XDG_DATA_HOME=/data/client --env XDG_STATE_HOME=/data/state-base \
truf-remote-worker:linux-x86_64 install \
--config - < worker-install.yaml
rm -- worker-install.yaml
docker run --rm \
--mount type=volume,source=truf-worker-device-a-data,target=/data \
--env XDG_DATA_HOME=/data/client --env XDG_STATE_HOME=/data/state-base \
truf-remote-worker:linux-x86_64 doctor --json
```
Run the installed supervisor in the foreground under Tini while Docker supplies
detachment and restart policy:
```sh
docker run --detach --name truf-worker-device-a --restart unless-stopped \
--read-only --cap-drop ALL --security-opt no-new-privileges --pids-limit 256 \
--tmpfs /tmp:rw,nosuid,nodev,noexec,size=128m,mode=1777 \
--mount type=volume,source=truf-worker-device-a-data,target=/data \
--env XDG_DATA_HOME=/data/client --env XDG_STATE_HOME=/data/state-base \
truf-remote-worker:linux-x86_64 run
```
The image entrypoint supplies `python -u -I -S -B` and the integrity-checking
bootstrap. It contains no PostgreSQL client, server runtime, provider
implementations, or detailed keycheck code.
Ordinary private OS storage is supported on both platforms. Application-layer
encryption of the local workspace is not required. Operators may still use host
full-disk encryption according to their own endpoint policy; this is not a TRUF
protocol requirement.
## Daily worker operation
The commands below use `truf-worker.cmd` on Windows. On Linux without Docker, use
the generated `truf-worker` launcher. For a running Docker worker, define a local
helper that executes the same package bootstrap inside its container:
```sh
workerctl() {
docker exec truf-worker-device-a /usr/local/bin/python3 -u -I -S -B \
/opt/truf-worker/app/remote_worker_bootstrap.py -- "$@"
}
```
Use these commands for normal operation:
| Command | Purpose |
| --- | --- |
| `truf-worker status` | One current human-readable supervisor and slot snapshot. |
| `truf-worker status --json` | Versioned snapshot for automation. |
| `truf-worker attach --follow-seconds 300` | Follow the verified running instance for a bounded interval; detaching does not stop it. |
| `truf-worker watch --follow-seconds 300` | Human live-status alias over the same verified attach stream. |
| `truf-worker attach --ndjson --follow-seconds 300` | Machine-readable bounded event stream. |
| `truf-worker logs --tail 200` | Read bounded rotating supervisor logs. |
| `truf-worker logs --follow --follow-seconds 300` | Follow logs for a bounded interval. |
| `truf-worker history --limit 50` | Show terminal assignment outcomes and durations. |
| `truf-worker history --reservation <id> --json` | Retrieve one assignment's terminal local record. |
| `truf-worker doctor --json` | Validate package identity, directories, configuration, retention, instance state, and runtime prerequisites. |
Run `status` first when investigating. It reports configured/occupied slots,
current phase, phase age, scan deadline, assignment time remaining, child state,
last progress age, backoff/idle reason, pending retention data, and progress-outbox
cursor. It does not invent percentage completion.
### Phases and deadlines
| Phase | Operator interpretation |
| --- | --- |
| `idle`, `claiming` | Slot is available or asking the server for work. |
| `assigned` | Immutable assignment identity and deadlines are persisted locally. |
| `waiting_permit` | Assignment owns a slot but is waiting for the shared scanner permit. This time counts against the scan-stage deadline. |
| `preparing`, `resolving`, `downloading`, `cloning` | Source-specific preparation before scanning. |
| `scanning` | Scanner process tree is active. |
| `filtering`, `cleaning`, `bundling` | Findings are converted, work is cleaned, and deterministic result bytes are staged. These phases remain inside the hard scan-stage deadline. |
| `uploading`, `awaiting_receipt` | Staged result is being transferred or waiting for authoritative server acknowledgement. |
| `backoff` | A bounded retry delay is active; inspect the reason and next-claim time. |
| `draining`, `stopped` | No new local work is starting; existing work is resolving or shutdown completed. |
The scan-stage deadline starts before `waiting_permit` and covers preparation,
provider access, scanning, filtering, cleanup, and bundle staging. Crossing it
terminates the contained runner tree and produces a normal phase-specific timeout
result while the assignment upload window remains available. The server's
assignment deadline is fixed at issue time and is never renewed by progress,
polling, restart, or upload retries. `status` therefore presents both deadlines
separately.
### Diagnostics and retained evidence
`history` identifies the terminal receipt, prebundle, timeout, stale, or recovery
outcome. `logs` shows supervisor operation; diagnostic records carry the stable
phase/category/code and optional body/log material. The private state tree stores:
```text
events/worker-events.jsonl
history/worker-history.jsonl
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.json
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.body
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.log
logs/worker.log
```
Diagnostic JSON records original/stored sizes, hash, encoding, and truncation
state. A missing or truncated body must not be described as complete. The server
admin detail view separately shows the ordered progress/receipt/ingestion/
settlement/projection timeline and canonical diagnostic snapshot. Assignment
transport outcome, scan outcome, and diagnostics are independent fields.
### Graceful stop and drain
Request local stop before maintenance. The supervisor closes new claims for this
worker, leaves authentication and the Worker API available while existing work and
pending uploads resolve, and requires a clean drain receipt:
```powershell
.\truf-worker.cmd stop --timeout 120 --json
```
For Docker, run `workerctl stop --timeout 120 --json`; the foreground supervisor
then exits and Tini returns the shutdown code to Docker. A successful shutdown
receipt reports `drained: true` and `exit_code: 0`. Do not interpret `docker stop`,
process termination, or loss of contact as assignment cancellation.
### Recovery
After an OS restart, network outage, server outage, or unclean process exit:
1. Preserve the entire private state/volume; do not remove work, bundle, event,
history, or control files.
2. Run `doctor --json` and `status --json` from the exact installed package.
3. Start the same package with `start` on Windows/Linux, or restart the same Docker
container/volume. The supervisor replays its journal and recovers assigned,
staged, upload-retry, and awaiting-receipt slots.
4. Use `attach --ndjson --follow-seconds 300` and server assignment detail to
distinguish active recovery from backoff.
5. Keep the same device identity and state until recovered slots and
accepted-but-not-ingested bundles are reconciled. Escalate only if the instance
is unverifiable, a fixed assignment deadline has passed without server recovery,
or repeated startup validation fails.
Re-uploading identical accepted bytes returns the original receipt. Conflicting
bytes or stale ownership are rejected; never delete local bytes merely to silence
that signal.
### Update and rollback
1. Request local graceful stop, reconcile unresolved/precommit work, and obtain a
successful clean drain receipt.
2. Retain the complete state tree and previous exact artifact/digest.
3. Verify the new ZIP/image and its registered server manifest. On Windows run
`prepare-worker.ps1`, then `doctor`; for Docker recreate only the container and
mount the same volume.
4. Start at parallelism `1`, confirm package identity/contact and one complete
accepted-ingested-projected assignment, then restore the intended cap.
5. If validation fails, stop and return to the previous exact artifact with the
same state. Additive server progress/diagnostic records need no rollback.
Do not replace binaries beneath a running supervisor or switch packages while an
assignment runner is active.
### Device removal
1. Stop every device that uses the identity locally; leave authentication valid
while all unresolved assignments and precommit bundles reach zero.
2. Stop gracefully and retain the shutdown receipt and terminal history.
3. Revoke the device and disable its user only if that user is not shared by an
active device.
4. Confirm the old token no longer authenticates and no authoritative recovery
remains.
5. Remove the container/package. Remove its private volume/state only after the
server reconciliation evidence is retained and no rollback requires it.
## Server tuning
Make tuning changes in the reviewed private runtime configuration, not on the
client. Keep the raw worker service on loopback/private addressing and expose it
only through certificate-validating Caddy HTTPS.
| Control | Meaning |
| --- | --- |
| Client `--parallelism N` | Maximum locally occupied slots, one claim per free slot. |
| Typed admin assignment cap | Atomic positive active-assignment cap for one user across all devices; lowering it does not cancel existing assignments. |
| `assignment_ttl_seconds` | Fixed server-clock lifetime covering download, scan, and upload; default `86400`. API contact and restart do not renew it. |
| `bundle_body_timeout_seconds` | Upload body deadline; default `1800`. The assignment lifetime must exceed the largest configured source scan timeout plus this value plus 60 seconds. |
| `json_body_timeout_seconds` / `body_idle_timeout_seconds` | Request and idle transport bounds; defaults `60` and `30`. They do not renew ownership. |
| `reaper_interval_seconds` / `reaper_batch_size` | Expired-assignment recovery cadence and bounded batch; defaults `60` and `1000`. |
| `limit_concurrency` | Worker API request concurrency bound, not a replacement for user quotas; default `64`. |
Shortening the fixed lifetime can reject a valid long scan or upload and allow a
second physical execution after recovery. Lengthening it holds user quota,
target ownership, and dependent leases longer after a lost client. Tune it from
observed end-to-end duration plus upload headroom, not from HTTP polling cadence.
## Typed administration
Use only the authenticated random-prefix admin page configured by the production
edge. It provides CSRF/Origin-checked typed operations to create/enable/disable
users, set assignment caps, issue/rotate/revoke/unrevoke device tokens, and
requeue selected deferred queue IDs. It deliberately provides no shell or generic
supervisor command.
An issued or rotated token is shown once. Rotation replaces the stored digest, so
the old token stops authenticating; update that device without copying the token
to other devices. Rotation does not clear an existing revoked state. Revocation
blocks further API authentication but does not invent cancellation for assigned
work; plan for outstanding work to be completed before revocation or recovered at
its fixed expiry. Use local graceful stop to close claims before planned device
maintenance or removal while preserving authentication for pending work.
For an admin IP ban, use SSH and fail2ban first so fail2ban and Caddy agree:
```sh
sudo fail2ban-client set truf-admin-auth unbanip 203.0.113.10
sudo /usr/local/sbin/truf-caddy-admin-denylist status
sudo /usr/local/sbin/truf-caddy-admin-denylist expire
```
If fail2ban is unavailable, use the explicit admin-only updater:
```sh
sudo /usr/local/sbin/truf-caddy-admin-denylist unban 203.0.113.10
```
These commands change only the admin-route matcher. Never replace them with a
global port 443 firewall unban/ban. The full damaged-snippet recovery procedure
is in `deploy/edge/README.md`.
## Accepted custody and expiry
`accepted` means the server validated the canonical v2 `.trb`, durably published
it, and persisted ready/recovery state before returning a receipt. It does not
mean ingestion, projection, candidate handling, or detailed keycheck has
finished. Track `accepted` and `ingested` separately in the typed admin view.
The client keeps assignment identity and pending bytes until authoritative
acknowledgement. Retrying identical accepted bytes returns the original receipt,
including after ingestion, spool cleanup, restart, or the former deadline;
conflicting bytes are rejected. An interrupted, invalid, expired-before-first-
acceptance, or stale upload is not a successful scan.
An unfinished assignment expires at the original server-set deadline, normally
24 hours after issue. The periodic recovery pass reconciles quota, credits,
target ownership, and dependent plan/blob leases through existing retry policy.
Already accepted ready bundles are not requeued as unfinished. A crashed client
can therefore delay work for about a day, while a partitioned client can continue
physical scanning after the server has expired and reissued the target. Ownership
fencing guarantees one authoritative acceptance, not exactly-once physical work.
## Staged rollout and rollback
1. Obtain separate review for production schema/deployment changes. Drain protocol-1 work before replacing packages. Keep worker API and admin disabled while configuring exact protocol-2 capability profiles, task-specific auth entries, trusted package manifests, edge origin/marker, and per-user caps.
2. Pass the main Docker, packaged Windows/Linux, and edge gates with their pinned artifacts. Do not infer readiness from unit mocks or one platform.
3. Enable a small canary: one user, one device, cap `1`, and client `N=1`. Confirm assignment, acceptance, ingestion, projection, detailed server keycheck, expiry, and safe logs before increasing either cap or device count.
4. Expand caps and clients in stages while comparing unfinished/completed/failed/expired counts, last authenticated contact, accepted-versus-ingested state, durations, capacity, and stale/duplicate events. Contact age alone is not a liveness failure while a client holds long-running work and has not yet returned to claim polling.
5. To drain, request local graceful stop on every affected device. Leave authentication and the Worker API available so pending uploads and terminal reports can resolve. Wait for unfinished assignments to complete or pass through fixed-expiry recovery, and separately reconcile accepted bundles awaiting ingestion.
6. After the drain is authoritative, disable remote admission and return scheduling to local-only execution. Then revoke unused device tokens if required. Retain reservation metadata, accepted receipts, bundles/recovery state, and schema until reviewed reconciliation is complete.
7. Do not drop worker metadata, clear spool state, rotate/revoke tokens before a drain, switch server binaries underneath unfinished work, or treat a service stop as cancellation. The local execution path remains available and must not depend on a remote worker.