## Context DockerHub TruffleHog commands currently run under the project's supervisor, scan-slot lease, and Windows Job containment while TruffleHog also starts its own overseer process. Runtime evidence shows frequent exit-code-1 runs that emit `running source` but not `finished scanning`; the current fallback classifies these runs as non-retryable `command_exit` failures. Controlled runs of previously affected images completed with exit code 0 under the same Windows Job when TruffleHog's embedded overseer was bypassed with `--local-dev`. The flag changes process lifecycle only; it does not enable verification, alter detectors, or relax containment. ## Goals / Non-Goals **Goals:** - Give the existing external supervisor sole ownership of Docker scan lifecycle. - Require positive completion evidence for a successful Docker TruffleHog process. - Retry incomplete Docker runs through the existing bounded target retry policy. - Preserve findings emitted before an incomplete process exit. - Roll out and replay historical failures in bounded, observable stages. **Non-Goals:** - Upgrade or replace the installed TruffleHog binary. - Change non-Docker TruffleHog commands. - Change detector selection, provider routing, or keycheck behavior. - Treat every Docker diagnostic as retryable. - Replay the full historical failure set before the canary is healthy. ## Decisions ### Bypass the embedded overseer only for DockerHub `scan_docker_image` will add `--local-dev` while retaining `--no-update`. The project already provides restart policy, process-tree containment, timeout enforcement, and immutable executable authority, so the embedded updater/overseer is redundant. A Docker-only rollout limits behavioral scope and makes the canary attributable. Alternative considered: upgrade TruffleHog first. Rejected for this change because an official current binary also retained the overseer behavior in controlled tests, while an upgrade changes detectors and Docker internals at the same time. ### Make normal completion explicit Diagnostic parsing will record whether the exact JSON message `finished scanning` was observed. For Docker, exit code 0 without this marker is an incomplete run rather than success. A nonzero unexplained exit remains a failure even if the marker exists, but it is retryable because the wrapper lifecycle did not terminate cleanly. Alternative considered: trust exit code alone. Rejected because the observed coverage gap is specifically caused by ambiguous process exits and partial output. ### Reuse the existing bounded retry state machine An unexplained Docker exit without completion evidence will use a distinct `command_incomplete` class with `retryable=true`. An unexplained exit after a completion marker will use `wrapper_exit` with `retryable=true`. Existing `target_retry_max_attempts=3` and exponential delay remain authoritative; no unbounded or immediate retry loop is introduced. Alternative considered: retry every Docker exit code 1 without changing command lifecycle. Rejected because controlled retries were inconsistent and repeatedly downloaded/scanned the same image without removing the triggering lifecycle race. ### Preserve partial findings and fail closed Findings parsed before an incomplete exit remain in the durable result. The target is not marked clean or done until a complete run succeeds. This preserves useful evidence without claiming full image coverage. ### Bound internal Docker parallelism DockerHub will pass a source-configured TruffleHog concurrency of 4 and use a 600-second target timeout. The existing two Docker workers and 6 GiB per-process Windows Job limit remain unchanged. This replaces up to 32 aggregate internal workers across two image processes with at most 8, reducing decompression and chunking pressure while preserving source-level parallelism. Alternative considered: increase the memory cap. Rejected because a production OOM image completed within the existing 6 GiB Job in 239 seconds at concurrency 4; increasing the cap would raise host-wide risk without addressing amplification. ### Keep detector context timeouts target-scoped and nonfatal The exact TruffleHog diagnostic `a detector ignored the context timeout` will be retained as a nonfatal `detector_timeout` warning. A completed RC=0 scan containing only these diagnostics is degraded, not failed. Other timeout diagnostics keep their existing retryable error behavior. ### Separate canary, replay, and binary upgrade The runtime will first deploy the command and classification changes. The canary will compare unexplained Docker exit rate, completion-marker rate, source restarts, and queue health against the existing baseline. Historical exact-signature failures will be requeued in bounded batches only after the canary is healthy. A pinned official TruffleHog upgrade remains a follow-up change. ## Risks / Trade-offs - [The hidden `--local-dev` flag changes in a future binary] -> Keep executable hash pinning and add command-level regression coverage before any binary upgrade. - [Completion logging changes upstream] -> Treat a missing marker as retryable and bounded, not as success or an infinite retry. - [More retries increase Docker traffic] -> Keep the existing three-attempt cap and delay policy; replay historical failures in small batches. - [Lower concurrency increases scan duration] -> Raise the Docker target timeout to 600 seconds and retain two source workers. - [A detector context timeout omits some detector coverage] -> Persist it as degraded status rather than silently calling the image clean. - [Partial findings are duplicated across attempts] -> Rely on existing finding and credential identities for deduplication while preserving each scan event. - [Docker-only behavior diverges from other sources] -> Use the canary to validate the lifecycle decision before considering a common TruffleHog command policy. ## Migration Plan 1. Add focused command, diagnostic, and queue-policy tests. 2. Deploy the Docker-only lifecycle change and restart the managed runtime so immutable code authority is refreshed. 3. Observe at least 100 completed Docker attempts or two hours, whichever is longer. 4. Require no stale scan leases, no new unexplained terminal RC=1 failures, and normal pipeline drain before replay. 5. Requeue a small exact-signature batch, verify completion and deduplication, then increase batches conservatively. 6. Roll back by removing the Docker-only flag/classification change and restarting the supervisor; replayed rows remain ordinary auditable scan events. ## Open Questions - What batch size gives acceptable registry traffic during historical replay? Determine from the canary's average duration and Docker rate-limit headroom. - Should `--local-dev` later become the common policy for every externally supervised TruffleHog source? Decide in a separate change using per-source evidence.