Files
2026-09-30 20:30:56 +03:00

19 KiB

Truf Runtime Operations

All scanner, keycheck, dashboard, and PostgreSQL lifecycle mutation is owned by supervisor.py. Direct console_runner.py, mutating keycheck_runner.py, and legacy app.py controls are retired.

PostgreSQL is the sole authority for scan results, queue completion, keycheck results, and keycheck current state. Scanner sources publish private version-2 bundles to S:\scanner-result-bundles; the singleton result ingester commits them transactionally. JSONL and status files are asynchronous, rebuildable compatibility projections and may lag without rolling back a committed scan.

The scanner has max_active_scans=3 guaranteed fair permits plus at most one memory-gated non-Docker bonus permit. A permit covers staging, TruffleHog, normalization, bundle fsync, and the atomic ready rename only. PostgreSQL ingestion, JSONL projection, and keychecks do not hold scan permits.

Run commands from D:\truf\app unless a full path is shown.

Start And Stop

Canonical production start:

..\start_runtime.ps1

Canonical full coordinated shutdown:

..\stop_runtime.ps1

Equivalent authenticated launch after cluster identity has been verified:

python -I -S -B runtime_bootstrap.py supervisor -- --runtime-bootstrap-entrypoint D:\truf\app\supervisor.py --config config.yaml --background --no-dashboard --with-postgres

Direct read/control operations remain supported by supervisor.py:

python supervisor.py --config config.yaml --background-status
..\attach_runtime.ps1
python supervisor.py --config config.yaml --cmd "status"
python supervisor.py --config config.yaml --stop-background --with-postgres

In an attached prompt, q only detaches; shutdown requests full coordinated shutdown. Prefer ..\stop_runtime.ps1 for canonical full shutdown.

The dashboard is currently disabled. To use the read-only dashboard, enable it in config.yaml and perform a coordinated runtime restart. The legacy scanner UI is intentionally retired.

Source Commands

Use the foreground supervisor prompt or authenticated --cmd requests:

python supervisor.py --config config.yaml --cmd "status"
python supervisor.py --config config.yaml --cmd "start github"
python supervisor.py --config config.yaml --cmd "once gitlab"
python supervisor.py --config config.yaml --cmd "restart dockerhub"
python supervisor.py --config config.yaml --cmd "pause npm"
python supervisor.py --config config.yaml --cmd "resume npm"
python supervisor.py --config config.yaml --cmd "stop package_git"
python supervisor.py --config config.yaml --cmd "stop all"
python supervisor.py --config config.yaml --cmd "start all"
python supervisor.py --config config.yaml --cmd "logs pypi 80"
python supervisor.py --config config.yaml --cmd "command github"

stop all stops managed children while the supervisor and PostgreSQL remain running. start all includes pypi; do not use it when pypi must remain stopped.

Configure discovery mode, queries, custom target files, timeouts, workers, and source-specific arguments in config.yaml before starting or restarting a source. Do not pass tokens or mutable scan options through a direct runner command.

Keychecks

The supervisor manages keychecks as the keychecks pseudo-source:

python supervisor.py --config config.yaml --cmd "start keychecks"
python supervisor.py --config config.yaml --cmd "recheck all network"
python supervisor.py --config config.yaml --cmd "recheck gemini all --max-keys 100"
python supervisor.py --config config.yaml --cmd "recheck replicate valid --max-keys 25"
python supervisor.py --config config.yaml --cmd "recheck all --summary-only"
python supervisor.py --config config.yaml --cmd "logs keychecks 80"

Provider probe arguments belong under keychecks.service_args in config.yaml. Normal providers claim fenced PostgreSQL keycheck_candidates and commit keycheck_results plus keycheck_current_state directly. Files under D:\truf\runtime\keychecks are compatibility projections, not current-state authority. At most four provider children run concurrently, each with a bounded candidate slice so later services cannot starve.

Configuration And Secrets

Primary files:

D:\truf\app\config.yaml
D:\truf\app\secrets.yaml
D:\truf\.env.postgres
D:\truf\runtime\proxy.txt

config.yaml contains paths, source settings, and auth-pool names. Actual source tokens belong in secrets.yaml; PostgreSQL credentials belong in .env.postgres. Managed children receive one canonical loopback PostgreSQL DSN after all configurable environment overrides.

Important runtime paths:

D:\truf\runtime\results
S:\scanner-result-bundles
D:\truf\runtime\result_spool  (legacy import compatibility only)
D:\truf\runtime\queues
D:\truf\runtime\state
D:\truf\runtime\logs
D:\truf\runtime\control
D:\truf\runtime\keychecks
D:\truf\runtime\postman_cache
D:\truf\runtime\postgres\data
D:\truf\tmp

Runtime startup performs read-only ACL/owner/reparse preflight and never repairs paths.

Remote Assignment Capacity

global.result_bundle_max_event_bytes and supervisor.worker_api.max_bundle_bytes are hard per-bundle limits and remain 64 MiB. They are not admission reservations. Each unresolved remote assignment instead charges the persisted global.remote_assignment_reserve_bytes baseline of 2 MiB on both the bundle and projection byte axes; local scans retain their existing worst-case reservation behavior.

Remote admission enforces global.remote_assignment_max_active: 50 atomically in addition to each user's typed active_assignment_cap. The intended 50-assignment user must therefore have its cap set to 50 through the authenticated worker administration path. Lower either cap to reduce concurrency; do not raise the global cap above the validated maximum.

A valid remote bundle larger than 2 MiB atomically expands its persisted bundle charge to actual bytes before the server returns an acceptance receipt. Temporary aggregate bundle exhaustion returns retryable capacity backpressure, and the worker must retry the identical durable upload. Projection serialization similarly expands a leased job to exact aggregate bytes before any append. If projection capacity is unavailable, the untouched job returns to pending; capacity backpressure alone never quarantines it.

The production keycheck limits of 131,072 items and 128 MiB cover fifty baseline candidate reservations. Bundle and projection aggregate capacities and projection headroom remain independent safety bounds. Runtime-document validation rejects a hard bundle limit above 64 MiB, a remote baseline below 2 MiB or above the hard limit, a global cap above 50, and any aggregate axis that cannot hold all configured baselines.

Authority Model

One cross-session lock is derived only from the canonical bundled PostgreSQL data directory. It is held by every lifecycle-owning foreground/background supervisor and by PostgreSQL bootstrap/verification, migration, reconciliation, and offline hardening. Changing control directory, instance file, or port cannot split authority; different data directories have independent locks.

The configurable control-directory lock remains a secondary per-instance safety layer. Duplicate launch failure never sends coordinated shutdown to a different owner.

Before spawn, the launcher captures exact config, supervisor, and code-manifest authority. The manifest also covers the result bundle, ingester, projector, keycheck candidate, and isolated janitor modules.

The child completes canonical DSN validation and all controller/backend/source/keycheck/dashboard construction before publishing ACTIVATING. PostgreSQL, dashboard, and source ticks are forbidden until authenticated parent activation and exact post-activation recheck publish ACTIVE.

If activation becomes uncertain, rollback uses authenticated shutdown and the full configured deadline. It never force-terminates an exact published candidate that may be active. An unconfirmed exact candidate is left running and reported rather than risking an orphaned PostgreSQL tree.

Authenticated shutdown atomically publishes STOPPING and closes all start gates under the control lock before acknowledgement. Mutating commands and snapshot-triggered polls cannot start children after this transition.

Code/config drift closes start gates and triggers safe shutdown. Authenticated shutdown remains available through private instance credentials, endpoint binding, and retained process identity even when on-disk code changed. Tokens, raw DSNs, and secret-derived hashes are not written to logs.

Child Authentication

Managed scanner, result-ingester, JSONL-projector, keycheck, janitor, and dashboard children must prove all of the following before mutation. The janitor receives no database capability and persists only an exact-private bounded local cursor:

  1. Private per-instance metadata matches inherited instance credentials.
  2. The retained supervisor process and config command line match metadata.
  3. The supervisor handshake reports ACTIVE.
  4. Config, supervisor, and complete code-manifest hashes match.
  5. SCANNER_DB_URL, DATABASE_URL, and the inherited managed DSN name the same canonical loopback PostgreSQL authority.

An environment marker by itself grants no authority. Empty or foreign DSNs fail before application files, ScannerDB, dependency checks, provider checks, or scanning.

PostgreSQL Maintenance

One-time offline cluster identity binding, with all runtime processes stopped:

python postgres_runtime.py bootstrap --config config.yaml

Read-only identity verification:

python postgres_runtime.py verify --config config.yaml

Identity-verified maintenance start/stop, without scanner or keycheck children:

python postgres_runtime.py maintenance-start --config config.yaml
python postgres_runtime.py maintenance-stop --config config.yaml

Docker Depth Experiment Operator

Keep sources.dockerhub.docker_depth_experiment.enabled: false while reviewing and applying the cohort and hold. From D:\truf\app, use the existing private runtime\state directory and always stop maintenance PostgreSQL in a finally step if an operator command fails:

..\stop_runtime.ps1
python postgres_runtime.py maintenance-start --config config.yaml
python docker_depth_operator.py --config config.yaml --status
python docker_depth_operator.py --config config.yaml --generate-cohort-manifest D:\truf\runtime\state\docker-depth-cohort-review.json
$cohortSha256 = Read-Host 'Reviewed cohort SHA-256'
python docker_depth_operator.py --config config.yaml --apply-cohort-manifest D:\truf\runtime\state\docker-depth-cohort-review.json --approve-sha256 $cohortSha256 --apply --sources-stopped
python docker_depth_operator.py --config config.yaml --generate-hold-manifest D:\truf\runtime\state\docker-depth-hold-review.json
$holdSha256 = Read-Host 'Reviewed hold SHA-256'
python docker_depth_operator.py --config config.yaml --apply-hold-manifest D:\truf\runtime\state\docker-depth-hold-review.json --approve-sha256 $holdSha256 --apply --sources-stopped
python postgres_runtime.py maintenance-stop --config config.yaml

Review the private files out of band and approve exactly the SHA-256 printed by their generation commands. The operator has no DSN option, never prints queries or targets, and never starts or stops PostgreSQL. After the hold apply and maintenance stop, changing only enabled from false to true is the separate activation decision; its semantic config_sha256 must remain unchanged. Run --status between another maintenance start/stop pair to verify that hash before a separately approved ..\start_runtime.ps1.

An attempt-limit hold caused by the retired zero-graph resolver defect has a separate one-time reviewed recovery. Keep enabled config, stop sources, start maintenance PostgreSQL, review the private manifest, and approve only its exact printed SHA-256:

python docker_depth_operator.py --config config.yaml --generate-resolver-refund-manifest D:\truf\runtime\state\docker-depth-resolver-refund-review.json
$refundSha256 = Read-Host 'Reviewed resolver refund SHA-256'
python docker_depth_operator.py --config config.yaml --apply-resolver-refund-manifest D:\truf\runtime\state\docker-depth-resolver-refund-review.json --approve-sha256 $refundSha256 --apply --sources-stopped

A genuine remote resolver_attempt_limit hold uses a separate reviewed disposition. The manifest deterministically selects the next fresh repository or records terminal remote-unavailable scarcity when none remains:

python docker_depth_operator.py --config config.yaml --generate-resolver-disposition-manifest D:\truf\runtime\state\docker-depth-resolver-disposition-review.json
$dispositionSha256 = Read-Host 'Reviewed resolver disposition SHA-256'
python docker_depth_operator.py --config config.yaml --apply-resolver-disposition-manifest D:\truf\runtime\state\docker-depth-resolver-disposition-review.json --approve-sha256 $dispositionSha256 --apply --sources-stopped

Release is reviewed only after the experiment reaches completed under enabled config:

..\stop_runtime.ps1
python postgres_runtime.py maintenance-start --config config.yaml
python docker_depth_operator.py --config config.yaml --generate-reactivation-manifest D:\truf\runtime\state\docker-depth-reactivation-review.json
$reactivationSha256 = Read-Host 'Reviewed reactivation SHA-256'
python docker_depth_operator.py --config config.yaml --apply-reactivation-manifest D:\truf\runtime\state\docker-depth-reactivation-review.json --approve-sha256 $reactivationSha256 --apply --sources-stopped
python postgres_runtime.py maintenance-stop --config config.yaml

Runtime-safety schema migration, with supervisor, dashboard, scanner, and keycheck sessions stopped:

python migrate_runtime_safety.py --config config.yaml --apply --sources-stopped

Final cutover refuses a nonempty legacy result spool, any scan_publication_outbox row, any legacy raw_result_json row, or a prepared legacy JSONL-ledger append. Runtime workers refuse to start until the migration records the singleton PostgreSQL cutover marker.

Import bounded batches from the retired durable spool until remaining=0:

python migrate_runtime_safety.py --config config.yaml --import-legacy-spool --max-rows 1000 --apply --sources-stopped

Convert bounded legacy raw rows to normalized-v2 data, then rerun the normal migration to restore the cutover marker:

python migrate_runtime_safety.py --config config.yaml --backfill-normalized-results --max-rows 1000 --max-bytes 201326592 --max-seconds 30 --apply --sources-stopped

If the cutover gate reports legacy outbox rows, project only that bounded backlog offline before retrying migration:

python migrate_runtime_safety.py --config config.yaml --drain-legacy-outbox --max-rows 1000 --apply --sources-stopped

If a legacy event is too expensive for the per-finding ledger drain, first apply the additive schema (the command remains nonzero while the gate is closed), then transfer exact outbox references into the singleton projector queue without rewriting historical payloads:

python migrate_runtime_safety.py --config config.yaml --import-legacy-outbox-to-projection --max-rows 1000 --apply --sources-stopped

Build a full bounded, resumable PostgreSQL-derived projection in a dedicated empty directory. Repeat until completed=true; live projection files are never overwritten:

python migrate_runtime_safety.py --config config.yaml --rebuild-jsonl-output "D:\truf\runtime\rebuild" --max-rows 1000 --max-bytes 201326592 --max-seconds 30 --apply --sources-stopped

Quarantine review accepts only a private truf-pipeline-quarantine-review-v1 manifest with exact id, reason_code, payload_sha256, and action (discard, deterministic retry, or Docker layer rescan) entries. rescan retires the stale bundle and returns its immutable target to the normal fresh-claim path:

python migrate_runtime_safety.py --config config.yaml --review-pipeline-quarantine "D:\review\quarantine.json" --max-rows 1000 --apply --sources-stopped

Bounded todo reconciliation:

python migrate_runtime_safety.py --config config.yaml --apply --sources-stopped --todo "D:\truf\runtime\queues\todo_github.txt" --source github --platform github --max-rows 1000 --max-bytes 4194304 --max-seconds 5

Offline layout hardening creates required directories parent-first, recursively hardens runtime-owned trees, and hardens existing config, secrets, proxy, detector config, and PostgreSQL environment files:

python migrate_runtime_safety.py --config config.yaml --harden-runtime --apply --sources-stopped

Maintenance and runtime exclude each other through the same cluster authority lock. PostgreSQL sessions use search_path=public; pg_catalog keeps implicit precedence, authority built-ins are explicitly qualified, and PUBLIC CREATE on public fails closed.

D:\truf\docker-compose.postgres.yml is a disabled, noncanonical manual-recovery fixture. It is behind the noncanonical-manual-recovery profile, has no restart policy, requires an explicit unused TRUF_DOCKER_POSTGRES_PORT, and binds data under D:\truf; never use it to start or replace the supervisor-owned cluster. Any legacy Docker named volume is intentionally left untouched.

Dashboard Secrecy

Default dashboard queries and frames do not contain raw credentials. Raw scanner and validation values are available only behind explicit default-false per-session reveal controls with a local warning. Treat the database itself as sensitive because persisted findings still contain raw values.

The dashboard binds to loopback only. Do not expose it through 0.0.0.0, a reverse proxy, screenshots, or shared logs.

Cleanup

The authenticated janitor child is the only stale-tree recovery worker. It does not import scanner, and it enforces exact PID/creation-time/executable identities plus enumeration, candidate, entry, byte, time, and depth budgets. Its exact-private local cursor rotates layouts after every inspected entry so a huge first layout cannot starve later trees. Unknown identity retains the tree. Sources may only attempt bounded cleanup of a directory they just used; there is no startup sweep, low-space sweep, periodic supervisor scanner import, or atexit cleanup.

Definitively aborted unreferenced admission intents and deleted artifacts are retired in bounded keyset batches after 30 days. Retirement updates an aggregate SHA-256 chain and count before deleting exact rows; pending, committed/referenced, active, and open-quarantine authority is never eligible.

Do not manually delete results, result spool events, queue state, PostgreSQL data, or control metadata. Use authenticated coordinated shutdown before offline maintenance.