Files
2026-09-30 20:30:56 +03:00

208 lines
8.9 KiB
Markdown

# Remote Worker Operator Experience Findings
Recorded: 2026-09-23.
This file preserves the investigation that led to `add-worker-operator-experience` so implementation does not have to rediscover current behavior and terminology.
## Production Validation Facts
- Native Windows protocol-2 worker reached exact concurrency 3 and never exceeded it.
- Its bounded production cohort accepted 180 assignments: DockerHub 63, GitLab 65, HuggingFace 52.
- All 180 accepted results ingested, queue-settled, projected, and reconciled without global lineage violations or quarantine.
- A later dual-worker run proved Windows max 1, WSL max 1, combined max 2.
- Two WSL DockerHub assignments in separate runs stayed unresolved until fixed two-hour lease expiry and produced no accepted bundle.
- The server expired/refunded those reservations correctly; the missing information is where the worker spent the time before expiry.
Durable validation report:
`docs/worker-parallelism-validation-2026-09-23.md`
## Timeout and Lease Map
### Assignment deadline
Production explicitly used:
```yaml
supervisor:
worker_api:
assignment_ttl_seconds: 7200
```
- Production value: 7,200 seconds.
- Code/template fallback: 86,400 seconds.
- Managed validation bounds: 60 through 604,800 seconds.
- Validation requires the effective TTL to cover the maximum configured source scan timeout, bundle upload deadline, and 60 seconds of handoff margin.
- The server commits one immutable expiry at assignment issuance.
- Ordinary worker contacts and progress do not renew it.
- Relevant code: `app/runtime_document.py`, `app/worker_api.py`, `app/worker_assignment.py`, `app/scanner_db.py`.
### Docker target timeout
Canonical managed full-runtime and production value:
```yaml
sources:
dockerhub:
timeout: 600
```
- Managed legacy Windows/full runtime: 600 seconds.
- Current production profile: 600 seconds.
- Direct CLI/function fallback without managed config: 1,800 seconds.
- Legacy Windows Job containment terminates the owned TruffleHog process tree at the subprocess deadline.
- Historical evidence: `tests/test_scanner_queue_high_fixes.py`, `tests/test_validated_high_scanner_fixes.py`, and `openspec/changes/stabilize-dockerhub-trufflehog-lifecycle/`.
### Other relevant timers
- Worker API client socket timeout: 120 seconds.
- Result upload absolute body deadline: 1,800 seconds.
- Result upload idle timeout: 30 seconds.
- Remote assignment expiry reaper interval: 60 seconds.
- Empty-claim `Retry-After`: normally 5 seconds.
- Result-ingester/projector leases: 300 seconds after upload, unrelated to pre-upload scanning.
### Why ten minutes became two hours
The 600-second budget hard-preempts the TruffleHog subprocess path, but it is not one hard preemptive boundary around every surrounding Python operation. Permit acquisition, process startup/termination recovery, cleanup, filtering, serialization, bundle staging/fsync, and handoff can outlive the subprocess deadline. The remote client checks assignment expiry before synchronous execution and then has no phase heartbeat or cancellation loop.
The evidence does not prove which phase stalled. Increasing the assignment TTL would only make the unknown stall retain ownership longer. The implementation needs phase instrumentation and a complete supervised execution-unit deadline.
## Current Worker Experience
Current supported CLI flags in `app/remote_worker_client.py`:
```text
--server
--token
--parallelism
```
There are no worker `start`, `stop`, `status`, `attach`, `logs`, `history`, `doctor`, or JSON-output commands.
Current behavior:
- one daemon thread per slot;
- slot recovery state in `slot-N.json`;
- ready bundles retained until an authoritative receipt;
- scanner call is synchronous from the slot controller;
- state remains broadly `assigned` until bundle readiness;
- client loop failures print generic lines without a structured timeline;
- successful claim/scan/upload/receipt is mostly silent;
- scanner stdout/stderr is captured in temporary files and returned only after process completion;
- exact Git execution can emit no useful start line;
- concurrent messages can interleave;
- server `last_contact_at` does not prove or disprove active scanner progress.
Default paths:
- Windows: `%LOCALAPPDATA%/TRUF/RemoteWorker`.
- Linux: `$XDG_STATE_HOME/truf/remote-worker` plus `$XDG_DATA_HOME/truf/remote-worker`.
- Docker production convention: persistent `/data` volume with separate state/data roots.
## Legacy Attach Pattern
The old full-runtime supervisor has an `--attach` implementation in `app/supervisor.py` and a wrapper `attach_runtime.ps1`. The wrapper is currently disabled and the supervisor is not in the remote-worker package.
Useful semantics to reuse:
- exact background-instance identity;
- startup and loopback control handshake;
- initial status table;
- interactive `attach>` prompt;
- alternate-screen `watch` table;
- bounded log tail;
- `q`/EOF/Ctrl-C detach without worker shutdown;
- explicit coordinated shutdown command.
The worker needs a smaller implementation over its own event/status model, not a copy of the complete server supervisor.
## Current Error Model
The admin `Error category` column is rendered in `app/admin_api.py` from:
```sql
result_reservations.last_error_code AS error_category
```
It therefore describes assignment/transport errors such as remote prebundle failure or assignment expiry. It is not `errors.category`, scanner `error_class`, or `source_failure_category`.
Consequences:
- an accepted bundle is shown as assignment `completed` even when scan status is `error`;
- accepted scan errors usually leave assignment `Error category` blank;
- process versus storage prebundle failure survives in resolution JSON but is not shown;
- permanent provider skips can exist in metadata/warnings without an `errors` row;
- provider response bodies are inconsistently reduced or discarded.
Useful data already persisted but not presented together:
- `result_reservations`: resolution kind/JSON, receipt, issue/expiry/resolve timestamps, last error code/detail;
- `target_scans`: status, error count, skipped reason, first error summary, start/end/duration;
- `errors`: category, summary, raw selected error line;
- `scan_result_compat.metadata_json`: error class, retryability, source failure category, warnings, degraded/skipped flags, process return/timeout/output metadata;
- `keycheck_results`: provider status group/message/metadata.
## Diagnostic Direction
Use separate dimensions rather than one overloaded category:
```text
assignment outcome: accepted | prebundle_failed | expired | unfinished
scan outcome: clean | found | degraded | error | skipped | unavailable
phase: scanning | cleaning | bundling | uploading | ...
kind: provider_http | scanner_process | exception | storage | protocol | ...
category: authorization | rate_limit | timeout | network | scanner | ...
code: stable concrete identifier
retryable: true | false
```
Candidate transmitted limits from the investigation:
- provider body material: 16 KiB;
- process log head/tail: 32 KiB combined;
- one diagnostic envelope: 64 KiB;
- at most 32 diagnostics and 256 KiB total per assignment;
- prebundle envelope profile sized to fit the existing Worker API JSON limit.
The local worker archive can retain larger/full artifacts under configurable age and byte rotation. Every transport transformation must state original size, stored size, hash, encoding, and truncation state.
## Progress and Duration Direction
Minimum phase transitions needed to explain the two-hour event:
```text
assignment_received
scan_permit_acquired
runner_started
source_prepare_started/completed
scanner_started/exited
filtering_started/completed
cleanup_started/completed
bundle_started/ready
upload_started/acknowledged
```
Progress is evidence only and does not renew the immutable lease.
Duration percentiles:
- p50: median duration;
- p95: 95 percent of observations finish at or below this duration;
- p99: 99 percent finish at or below it.
Compute them separately by source, phase, outcome, and time window, with sample counts. They describe observed behavior and inform policy; they do not silently set policy.
## Product Decisions
- Build one final operator architecture rather than a temporary admin patch.
- Implement it in large vertical chunks that each remain part of the final system.
- Use a worker-specific supervisor and local event stream.
- Split scanner execution into a supervised per-assignment runner process so the complete scan stage can be hard-preempted.
- Keep immutable server assignment ownership and make progress non-renewing.
- Add global fallback plus per-source assignment TTL policy.
- Use one diagnostic envelope across local files, terminal reports, bundles, PostgreSQL, API, and UI.
- Preserve diagnostic fidelity and expose all explicit transformations.
- Separate assignment outcome, scan outcome, and diagnostics in the admin UI.
- Include from-zero operator documentation and fault injection in the same change.