Files
truf-server/docs/remote-worker-operations.md
T
2026-09-30 20:30:56 +03:00

24 KiB

Remote Scan Worker Operations

This runbook covers the opt-in remote worker boundary. Production defaults remain disabled: supervisor.worker_api.enabled and supervisor.worker_api.admin.enabled are both false in app/config.linux.yaml. Enabling either one, applying schema changes, or starting the production edge requires a separate reviewed rollout.

Remote workers are trusted clients. Their only operator-authored runtime settings are the HTTPS server origin, one opaque device token, and a positive slot count N. Target/source settings, immutable plans, scanner policy, limits, and only the credentials needed for an assignment come from the server. Protocol-2 package schema 3 manifests advertise exact (source, platform, planning_kind) capabilities. The distributed core package contains GitLab exact_git_v1, DockerHub docker_direct_v1, and HuggingFace huggingface_space_v1; GitHub is a legacy optional capability and is not in the core profile. Detailed keycheck remains server-only after bundle ingestion; clients must not receive keycheck configuration or run keycheckers.

Isolated verification

Run each verifier independently from the root of the isolated D:\truf-workers checkout. Do not run plain docker compose, production Compose files, import overrides, unrestricted pytest, or broad Docker cleanup. Do not run these verifiers concurrently. They create random, ownership-labelled resources and never use production data, credentials, volumes, or image tags.

Main Docker E2E

From Linux or WSL, with the already-built local images truf-worker-test:runtime and truf-worker-test:test and an already ignored docker/test-results/latest.json:

python3 -I -S -B docker/verify.py

The verifier invokes only compose.e2e.yaml, does no build or pull, and normally removes only its ownership-verified resources. It retains artifacts after a failure. Use --keep only when a reviewed investigation needs stopped artifacts; record the printed project name and never substitute prune, broad down, or down --volumes commands.

Packaged Windows/Linux client E2E

From Windows, with Windows Python, wsl.exe, passwordless sudo -n docker in the selected WSL distribution, the built portable directory, and the already-built Linux worker and test images:

python -I -S -B docker/verify_packaged_workers.py `
  --windows-artifact dist/truf-worker-windows-x86_64 `
  --linux-image truf-remote-worker:linux-x86_64 `
  --test-image truf-worker-test:test `
  --wsl-distro Ubuntu-24.04

This single gate runs real packaged Windows and Linux scanners at N=2, tests a server outage and restart recovery, and compares normalized cross-platform evidence. It does not use Compose. Successful resources are removed unless --keep is supplied; failures retain the labelled resources and the reported build/pwe-* evidence directory.

Edge E2E

From Windows with the same WSL Docker access and already-built truf-worker-test:test and truf-edge-e2e:test images:

python -I -S -B docker/verify_edge_e2e.py --wsl-distro Ubuntu-24.04

This gate uses only random labelled resources and a private synthetic backend. It checks authenticated routes, two-failure admin bans, restart persistence, automatic expiry, SSH-equivalent unban behavior, forwarded-header handling, and worker availability from the banned admin IP. Success removes its resources; failure reports the retained owned inventory and evidence path.

Client bootstrap

Use a separately issued token for every device. The server stores only its SHA-256 digest. Do not place a real token in documentation, source control, shell transcripts, support output, or process diagnostics. The client validates the server certificate and accepts only an HTTPS origin without credentials, path, query, or fragment. N must be between 1 and 128; it bounds local occupied slots but never overrides the user's server cap across devices.

Generated state, pending bundles, and work directories are recovery data, not additional source/provider configuration. Preserve them across restarts until the server authoritatively resolves the corresponding slots.

Artifact source and distribution

Worker tools come from a reviewed release checkout; they are not installed piecemeal on each client. The Windows builder and Linux worker package/image targets assemble and verify the complete worker authority, Python runtime where applicable, Git helper, TruffleHog binary, detector policy, CA roots, and hash-locked dependencies. They deliberately exclude PostgreSQL tools, server/runtime authority, provider implementations, detailed keycheck code, server credentials, and database credentials.

There is no public worker download, image registry, installer, or automatic updater. Build each release once in a controlled release environment, retain its generated manifest/release metadata, and distribute the exact ZIP or image digest through a trusted artifact channel. Do not rebuild independently on every worker. Install the corresponding trusted package manifest beneath /etc/truf/worker-packages on the server before allowing that artifact to claim protocol-2 work.

Onboard a new worker in this order:

  1. Create or select its server-side user and start with active-assignment cap 1.
  2. Create a distinct device identity and issue its token. The plaintext token is shown once; never reuse it for another device.
  3. Deliver the exact reviewed Windows ZIP, native Linux package, or Linux image, and compare its package identity with the registered server manifest.
  4. Create a private three-field YAML document with the HTTPS origin, token, and conservative parallelism such as 1, then run install --config. Delete the input YAML after installation. This writes the existing private local configuration; lifecycle commands never need the token on their command line.
  5. Run doctor, start the supervisor, inspect status and attach, and confirm server last_contact_at.
  6. Reconcile one real assignment through accepted receipt, ingestion, settlement, and projection before increasing either the user cap or local parallelism.
  7. Preserve the private state tree across restart or outage. Drain work before token rotation, revocation, artifact replacement, or state removal.

For a routine new device, artifact delivery and token issuance are the only installation work. Building the artifact and registering its trusted manifest are release-management operations and should not be delegated to the device operator.

The verified copy-paste command sheets are:

  • docs/remote-worker-cheatsheet-windows-ru.md
  • docs/remote-worker-cheatsheet-linux-ru.md
  • docs/remote-worker-cheatsheet-docker-ru.md

Windows portable client

Build the pinned amd64 package in the isolated checkout when producing a release:

New-Item -ItemType Directory -Path dist -Force | Out-Null
python -B app/worker_package_builder.py windows `
  --project-root . `
  --output dist/truf-worker-windows-x86_64 `
  --archive dist/truf-worker-windows-x86_64.zip `
  --cache build/worker-cache

Verify the ZIP and adjacent release JSON through the release process. On the client, extract to a private local directory and run prepare-worker.ps1 once to replace inherited ACLs. Create worker-install.yaml in that private directory and put only server, token, and parallelism in it:

.\prepare-worker.ps1
.\truf-worker.cmd install --config .\worker-install.yaml
Remove-Item -LiteralPath .\worker-install.yaml
.\truf-worker.cmd doctor
.\truf-worker.cmd start --startup-timeout 30
.\truf-worker.cmd status

By default, state and data are below the current user's LOCALAPPDATA. The package verifies its manifest and application files before launch and bundles Python, Git, TruffleHog, detector policy, and locked dependencies. There is no automatic updater. Use the same Windows account for installation and operation; the supervisor instance and private state belong to that account.

Native Linux portable client

Install the reviewed package beneath a root-owned path such as /opt/truf-worker, then run sudo ./prepare-worker.sh from that directory as the worker OS user. The preparation keeps native Git and TruffleHog immutable and root-owned while making application code exact-private to the worker user. It does not install a systemd unit. Use the private YAML installation and lifecycle commands in docs/remote-worker-cheatsheet-linux-ru.md from that package root.

Linux client image

Build the worker-only image for the target architecture through the reviewed worker target. For x86-64:

docker build --target worker -t truf-remote-worker:linux-x86_64 .

Install one device into a private persistent volume. Put the server, token, and parallelism in a mode-0600 worker-install.yaml, pass it on standard input to the short-lived install container, and delete it after success:

docker volume create truf-worker-device-a-data
chmod 600 worker-install.yaml
docker run --rm -i \
  --mount type=volume,source=truf-worker-device-a-data,target=/data \
  --env XDG_DATA_HOME=/data/client --env XDG_STATE_HOME=/data/state-base \
  truf-remote-worker:linux-x86_64 install \
  --config - < worker-install.yaml
rm -- worker-install.yaml

docker run --rm \
  --mount type=volume,source=truf-worker-device-a-data,target=/data \
  --env XDG_DATA_HOME=/data/client --env XDG_STATE_HOME=/data/state-base \
  truf-remote-worker:linux-x86_64 doctor --json

Run the installed supervisor in the foreground under Tini while Docker supplies detachment and restart policy:

docker run --detach --name truf-worker-device-a --restart unless-stopped \
  --read-only --cap-drop ALL --security-opt no-new-privileges --pids-limit 256 \
  --tmpfs /tmp:rw,nosuid,nodev,noexec,size=128m,mode=1777 \
  --mount type=volume,source=truf-worker-device-a-data,target=/data \
  --env XDG_DATA_HOME=/data/client --env XDG_STATE_HOME=/data/state-base \
  truf-remote-worker:linux-x86_64 run

The image entrypoint supplies python -u -I -S -B and the integrity-checking bootstrap. It contains no PostgreSQL client, server runtime, provider implementations, or detailed keycheck code.

Ordinary private OS storage is supported on both platforms. Application-layer encryption of the local workspace is not required. Operators may still use host full-disk encryption according to their own endpoint policy; this is not a TRUF protocol requirement.

Daily worker operation

The commands below use truf-worker.cmd on Windows. On Linux without Docker, use the generated truf-worker launcher. For a running Docker worker, define a local helper that executes the same package bootstrap inside its container:

workerctl() {
  docker exec truf-worker-device-a /usr/local/bin/python3 -u -I -S -B \
    /opt/truf-worker/app/remote_worker_bootstrap.py -- "$@"
}

Use these commands for normal operation:

Command Purpose
truf-worker status One current human-readable supervisor and slot snapshot.
truf-worker status --json Versioned snapshot for automation.
truf-worker attach --follow-seconds 300 Follow the verified running instance for a bounded interval; detaching does not stop it.
truf-worker watch --follow-seconds 300 Human live-status alias over the same verified attach stream.
truf-worker attach --ndjson --follow-seconds 300 Machine-readable bounded event stream.
truf-worker logs --tail 200 Read bounded rotating supervisor logs.
truf-worker logs --follow --follow-seconds 300 Follow logs for a bounded interval.
truf-worker history --limit 50 Show terminal assignment outcomes and durations.
truf-worker history --reservation <id> --json Retrieve one assignment's terminal local record.
truf-worker doctor --json Validate package identity, directories, configuration, retention, instance state, and runtime prerequisites.

Run status first when investigating. It reports configured/occupied slots, current phase, phase age, scan deadline, assignment time remaining, child state, last progress age, backoff/idle reason, pending retention data, and progress-outbox cursor. It does not invent percentage completion.

Phases and deadlines

Phase Operator interpretation
idle, claiming Slot is available or asking the server for work.
assigned Immutable assignment identity and deadlines are persisted locally.
waiting_permit Assignment owns a slot but is waiting for the shared scanner permit. This time counts against the scan-stage deadline.
preparing, resolving, downloading, cloning Source-specific preparation before scanning.
scanning Scanner process tree is active.
filtering, cleaning, bundling Findings are converted, work is cleaned, and deterministic result bytes are staged. These phases remain inside the hard scan-stage deadline.
uploading, awaiting_receipt Staged result is being transferred or waiting for authoritative server acknowledgement.
backoff A bounded retry delay is active; inspect the reason and next-claim time.
draining, stopped No new local work is starting; existing work is resolving or shutdown completed.

The scan-stage deadline starts before waiting_permit and covers preparation, provider access, scanning, filtering, cleanup, and bundle staging. Crossing it terminates the contained runner tree and produces a normal phase-specific timeout result while the assignment upload window remains available. The server's assignment deadline is fixed at issue time and is never renewed by progress, polling, restart, or upload retries. status therefore presents both deadlines separately.

Diagnostics and retained evidence

history identifies the terminal receipt, prebundle, timeout, stale, or recovery outcome. logs shows supervisor operation; diagnostic records carry the stable phase/category/code and optional body/log material. The private state tree stores:

events/worker-events.jsonl
history/worker-history.jsonl
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.json
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.body
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.log
logs/worker.log

Diagnostic JSON records original/stored sizes, hash, encoding, and truncation state. A missing or truncated body must not be described as complete. The server admin detail view separately shows the ordered progress/receipt/ingestion/ settlement/projection timeline and canonical diagnostic snapshot. Assignment transport outcome, scan outcome, and diagnostics are independent fields.

Graceful stop and drain

Request local stop before maintenance. The supervisor closes new claims for this worker, leaves authentication and the Worker API available while existing work and pending uploads resolve, and requires a clean drain receipt:

.\truf-worker.cmd stop --timeout 120 --json

For Docker, run workerctl stop --timeout 120 --json; the foreground supervisor then exits and Tini returns the shutdown code to Docker. A successful shutdown receipt reports drained: true and exit_code: 0. Do not interpret docker stop, process termination, or loss of contact as assignment cancellation.

Recovery

After an OS restart, network outage, server outage, or unclean process exit:

  1. Preserve the entire private state/volume; do not remove work, bundle, event, history, or control files.
  2. Run doctor --json and status --json from the exact installed package.
  3. Start the same package with start on Windows/Linux, or restart the same Docker container/volume. The supervisor replays its journal and recovers assigned, staged, upload-retry, and awaiting-receipt slots.
  4. Use attach --ndjson --follow-seconds 300 and server assignment detail to distinguish active recovery from backoff.
  5. Keep the same device identity and state until recovered slots and accepted-but-not-ingested bundles are reconciled. Escalate only if the instance is unverifiable, a fixed assignment deadline has passed without server recovery, or repeated startup validation fails.

Re-uploading identical accepted bytes returns the original receipt. Conflicting bytes or stale ownership are rejected; never delete local bytes merely to silence that signal.

Update and rollback

  1. Request local graceful stop, reconcile unresolved/precommit work, and obtain a successful clean drain receipt.
  2. Retain the complete state tree and previous exact artifact/digest.
  3. Verify the new ZIP/image and its registered server manifest. On Windows run prepare-worker.ps1, then doctor; for Docker recreate only the container and mount the same volume.
  4. Start at parallelism 1, confirm package identity/contact and one complete accepted-ingested-projected assignment, then restore the intended cap.
  5. If validation fails, stop and return to the previous exact artifact with the same state. Additive server progress/diagnostic records need no rollback.

Do not replace binaries beneath a running supervisor or switch packages while an assignment runner is active.

Device removal

  1. Stop every device that uses the identity locally; leave authentication valid while all unresolved assignments and precommit bundles reach zero.
  2. Stop gracefully and retain the shutdown receipt and terminal history.
  3. Revoke the device and disable its user only if that user is not shared by an active device.
  4. Confirm the old token no longer authenticates and no authoritative recovery remains.
  5. Remove the container/package. Remove its private volume/state only after the server reconciliation evidence is retained and no rollback requires it.

Server tuning

Make tuning changes in the reviewed private runtime configuration, not on the client. Keep the raw worker service on loopback/private addressing and expose it only through certificate-validating Caddy HTTPS.

Control Meaning
Client --parallelism N Maximum locally occupied slots, one claim per free slot.
Typed admin assignment cap Atomic positive active-assignment cap for one user across all devices; lowering it does not cancel existing assignments.
assignment_ttl_seconds Fixed server-clock lifetime covering download, scan, and upload; default 86400. API contact and restart do not renew it.
bundle_body_timeout_seconds Upload body deadline; default 1800. The assignment lifetime must exceed the largest configured source scan timeout plus this value plus 60 seconds.
json_body_timeout_seconds / body_idle_timeout_seconds Request and idle transport bounds; defaults 60 and 30. They do not renew ownership.
reaper_interval_seconds / reaper_batch_size Expired-assignment recovery cadence and bounded batch; defaults 60 and 1000.
limit_concurrency Worker API request concurrency bound, not a replacement for user quotas; default 64.

Shortening the fixed lifetime can reject a valid long scan or upload and allow a second physical execution after recovery. Lengthening it holds user quota, target ownership, and dependent leases longer after a lost client. Tune it from observed end-to-end duration plus upload headroom, not from HTTP polling cadence.

Typed administration

Use only the authenticated random-prefix admin page configured by the production edge. It provides CSRF/Origin-checked typed operations to create/enable/disable users, set assignment caps, issue/rotate/revoke/unrevoke device tokens, and requeue selected deferred queue IDs. It deliberately provides no shell or generic supervisor command.

An issued or rotated token is shown once. Rotation replaces the stored digest, so the old token stops authenticating; update that device without copying the token to other devices. Rotation does not clear an existing revoked state. Revocation blocks further API authentication but does not invent cancellation for assigned work; plan for outstanding work to be completed before revocation or recovered at its fixed expiry. Use local graceful stop to close claims before planned device maintenance or removal while preserving authentication for pending work.

For an admin IP ban, use SSH and fail2ban first so fail2ban and Caddy agree:

sudo fail2ban-client set truf-admin-auth unbanip 203.0.113.10
sudo /usr/local/sbin/truf-caddy-admin-denylist status
sudo /usr/local/sbin/truf-caddy-admin-denylist expire

If fail2ban is unavailable, use the explicit admin-only updater:

sudo /usr/local/sbin/truf-caddy-admin-denylist unban 203.0.113.10

These commands change only the admin-route matcher. Never replace them with a global port 443 firewall unban/ban. The full damaged-snippet recovery procedure is in deploy/edge/README.md.

Accepted custody and expiry

accepted means the server validated the canonical v2 .trb, durably published it, and persisted ready/recovery state before returning a receipt. It does not mean ingestion, projection, candidate handling, or detailed keycheck has finished. Track accepted and ingested separately in the typed admin view.

The client keeps assignment identity and pending bytes until authoritative acknowledgement. Retrying identical accepted bytes returns the original receipt, including after ingestion, spool cleanup, restart, or the former deadline; conflicting bytes are rejected. An interrupted, invalid, expired-before-first- acceptance, or stale upload is not a successful scan.

An unfinished assignment expires at the original server-set deadline, normally 24 hours after issue. The periodic recovery pass reconciles quota, credits, target ownership, and dependent plan/blob leases through existing retry policy. Already accepted ready bundles are not requeued as unfinished. A crashed client can therefore delay work for about a day, while a partitioned client can continue physical scanning after the server has expired and reissued the target. Ownership fencing guarantees one authoritative acceptance, not exactly-once physical work.

Staged rollout and rollback

  1. Obtain separate review for production schema/deployment changes. Drain protocol-1 work before replacing packages. Keep worker API and admin disabled while configuring exact protocol-2 capability profiles, task-specific auth entries, trusted package manifests, edge origin/marker, and per-user caps.
  2. Pass the main Docker, packaged Windows/Linux, and edge gates with their pinned artifacts. Do not infer readiness from unit mocks or one platform.
  3. Enable a small canary: one user, one device, cap 1, and client N=1. Confirm assignment, acceptance, ingestion, projection, detailed server keycheck, expiry, and safe logs before increasing either cap or device count.
  4. Expand caps and clients in stages while comparing unfinished/completed/failed/expired counts, last authenticated contact, accepted-versus-ingested state, durations, capacity, and stale/duplicate events. Contact age alone is not a liveness failure while a client holds long-running work and has not yet returned to claim polling.
  5. To drain, request local graceful stop on every affected device. Leave authentication and the Worker API available so pending uploads and terminal reports can resolve. Wait for unfinished assignments to complete or pass through fixed-expiry recovery, and separately reconcile accepted bundles awaiting ingestion.
  6. After the drain is authoritative, disable remote admission and return scheduling to local-only execution. Then revoke unused device tokens if required. Retain reservation metadata, accepted receipts, bundles/recovery state, and schema until reviewed reconciliation is complete.
  7. Do not drop worker metadata, clear spool state, rotate/revoke tokens before a drain, switch server binaries underneath unfinished work, or treat a service stop as cancellation. The local execution path remains available and must not depend on a remote worker.