Initial server source import
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-08-19
|
||||
@@ -0,0 +1,85 @@
|
||||
## Context
|
||||
|
||||
DockerHub TruffleHog commands currently run under the project's supervisor, scan-slot lease, and Windows Job containment while TruffleHog also starts its own overseer process. Runtime evidence shows frequent exit-code-1 runs that emit `running source` but not `finished scanning`; the current fallback classifies these runs as non-retryable `command_exit` failures.
|
||||
|
||||
Controlled runs of previously affected images completed with exit code 0 under the same Windows Job when TruffleHog's embedded overseer was bypassed with `--local-dev`. The flag changes process lifecycle only; it does not enable verification, alter detectors, or relax containment.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Give the existing external supervisor sole ownership of Docker scan lifecycle.
|
||||
- Require positive completion evidence for a successful Docker TruffleHog process.
|
||||
- Retry incomplete Docker runs through the existing bounded target retry policy.
|
||||
- Preserve findings emitted before an incomplete process exit.
|
||||
- Roll out and replay historical failures in bounded, observable stages.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Upgrade or replace the installed TruffleHog binary.
|
||||
- Change non-Docker TruffleHog commands.
|
||||
- Change detector selection, provider routing, or keycheck behavior.
|
||||
- Treat every Docker diagnostic as retryable.
|
||||
- Replay the full historical failure set before the canary is healthy.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Bypass the embedded overseer only for DockerHub
|
||||
|
||||
`scan_docker_image` will add `--local-dev` while retaining `--no-update`. The project already provides restart policy, process-tree containment, timeout enforcement, and immutable executable authority, so the embedded updater/overseer is redundant. A Docker-only rollout limits behavioral scope and makes the canary attributable.
|
||||
|
||||
Alternative considered: upgrade TruffleHog first. Rejected for this change because an official current binary also retained the overseer behavior in controlled tests, while an upgrade changes detectors and Docker internals at the same time.
|
||||
|
||||
### Make normal completion explicit
|
||||
|
||||
Diagnostic parsing will record whether the exact JSON message `finished scanning` was observed. For Docker, exit code 0 without this marker is an incomplete run rather than success. A nonzero unexplained exit remains a failure even if the marker exists, but it is retryable because the wrapper lifecycle did not terminate cleanly.
|
||||
|
||||
Alternative considered: trust exit code alone. Rejected because the observed coverage gap is specifically caused by ambiguous process exits and partial output.
|
||||
|
||||
### Reuse the existing bounded retry state machine
|
||||
|
||||
An unexplained Docker exit without completion evidence will use a distinct `command_incomplete` class with `retryable=true`. An unexplained exit after a completion marker will use `wrapper_exit` with `retryable=true`. Existing `target_retry_max_attempts=3` and exponential delay remain authoritative; no unbounded or immediate retry loop is introduced.
|
||||
|
||||
Alternative considered: retry every Docker exit code 1 without changing command lifecycle. Rejected because controlled retries were inconsistent and repeatedly downloaded/scanned the same image without removing the triggering lifecycle race.
|
||||
|
||||
### Preserve partial findings and fail closed
|
||||
|
||||
Findings parsed before an incomplete exit remain in the durable result. The target is not marked clean or done until a complete run succeeds. This preserves useful evidence without claiming full image coverage.
|
||||
|
||||
### Bound internal Docker parallelism
|
||||
|
||||
DockerHub will pass a source-configured TruffleHog concurrency of 4 and use a 600-second target timeout. The existing two Docker workers and 6 GiB per-process Windows Job limit remain unchanged. This replaces up to 32 aggregate internal workers across two image processes with at most 8, reducing decompression and chunking pressure while preserving source-level parallelism.
|
||||
|
||||
Alternative considered: increase the memory cap. Rejected because a production OOM image completed within the existing 6 GiB Job in 239 seconds at concurrency 4; increasing the cap would raise host-wide risk without addressing amplification.
|
||||
|
||||
### Keep detector context timeouts target-scoped and nonfatal
|
||||
|
||||
The exact TruffleHog diagnostic `a detector ignored the context timeout` will be retained as a nonfatal `detector_timeout` warning. A completed RC=0 scan containing only these diagnostics is degraded, not failed. Other timeout diagnostics keep their existing retryable error behavior.
|
||||
|
||||
### Separate canary, replay, and binary upgrade
|
||||
|
||||
The runtime will first deploy the command and classification changes. The canary will compare unexplained Docker exit rate, completion-marker rate, source restarts, and queue health against the existing baseline. Historical exact-signature failures will be requeued in bounded batches only after the canary is healthy. A pinned official TruffleHog upgrade remains a follow-up change.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [The hidden `--local-dev` flag changes in a future binary] -> Keep executable hash pinning and add command-level regression coverage before any binary upgrade.
|
||||
- [Completion logging changes upstream] -> Treat a missing marker as retryable and bounded, not as success or an infinite retry.
|
||||
- [More retries increase Docker traffic] -> Keep the existing three-attempt cap and delay policy; replay historical failures in small batches.
|
||||
- [Lower concurrency increases scan duration] -> Raise the Docker target timeout to 600 seconds and retain two source workers.
|
||||
- [A detector context timeout omits some detector coverage] -> Persist it as degraded status rather than silently calling the image clean.
|
||||
- [Partial findings are duplicated across attempts] -> Rely on existing finding and credential identities for deduplication while preserving each scan event.
|
||||
- [Docker-only behavior diverges from other sources] -> Use the canary to validate the lifecycle decision before considering a common TruffleHog command policy.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add focused command, diagnostic, and queue-policy tests.
|
||||
2. Deploy the Docker-only lifecycle change and restart the managed runtime so immutable code authority is refreshed.
|
||||
3. Observe at least 100 completed Docker attempts or two hours, whichever is longer.
|
||||
4. Require no stale scan leases, no new unexplained terminal RC=1 failures, and normal pipeline drain before replay.
|
||||
5. Requeue a small exact-signature batch, verify completion and deduplication, then increase batches conservatively.
|
||||
6. Roll back by removing the Docker-only flag/classification change and restarting the supervisor; replayed rows remain ordinary auditable scan events.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- What batch size gives acceptable registry traffic during historical replay? Determine from the canary's average duration and Docker rate-limit headroom.
|
||||
- Should `--local-dev` later become the common policy for every externally supervised TruffleHog source? Decide in a separate change using per-source evidence.
|
||||
@@ -0,0 +1,30 @@
|
||||
## Why
|
||||
|
||||
DockerHub scans frequently terminate with exit code 1 after emitting `running source` but before `finished scanning`. These incomplete runs are currently treated as terminal failures after one attempt, which leaves a material coverage gap even though the scanner runtime and target are often healthy.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Run DockerHub TruffleHog scans without TruffleHog's redundant embedded overseer while retaining the existing external supervisor and Windows Job containment.
|
||||
- Record whether TruffleHog emitted its normal completion marker.
|
||||
- Classify an unexplained Docker exit without the completion marker as an incomplete transient run instead of a permanent target failure.
|
||||
- Reuse the existing bounded target retry policy for incomplete runs.
|
||||
- Bound Docker's internal TruffleHog concurrency and allow enough time for a contained full-image scan.
|
||||
- Treat TruffleHog's exact detector context-timeout diagnostic as degraded detector coverage rather than a failed image scan.
|
||||
- Add a controlled replay path for historical failures matching this exact signature after the canary is healthy.
|
||||
- Keep the TruffleHog binary upgrade out of this change so lifecycle behavior can be measured independently.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `docker-scan-lifecycle`: Defines completion, containment, retry, and replay behavior for DockerHub TruffleHog scans.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects Docker command construction and TruffleHog diagnostic classification in `app/scanner.py`.
|
||||
- Affects Docker target completion disposition in the existing PostgreSQL queue flow.
|
||||
- Adds focused scanner policy tests and runtime canary checks.
|
||||
- Does not change provider keycheck behavior, non-Docker scan commands, or the installed TruffleHog binary.
|
||||
+74
@@ -0,0 +1,74 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: External Docker scan lifecycle ownership
|
||||
The system SHALL bypass TruffleHog's embedded overseer for DockerHub image scans while retaining the existing external supervisor, scan-slot lease, timeout, output bounds, and Windows Job containment.
|
||||
|
||||
#### Scenario: Docker command construction
|
||||
- **WHEN** the system constructs a TruffleHog command for a DockerHub image
|
||||
- **THEN** the command includes both `--local-dev` and `--no-update`
|
||||
|
||||
#### Scenario: Non-Docker command construction
|
||||
- **WHEN** the system constructs a TruffleHog command for a non-Docker source
|
||||
- **THEN** this Docker-only capability does not add `--local-dev`
|
||||
|
||||
### Requirement: Explicit Docker scan completion
|
||||
The system SHALL record whether TruffleHog emitted the exact normal completion message `finished scanning`, and SHALL require both that marker and exit code 0 before treating process execution as complete.
|
||||
|
||||
#### Scenario: Normal completion
|
||||
- **WHEN** a Docker TruffleHog process exits with code 0 after emitting `finished scanning`
|
||||
- **THEN** diagnostic metadata records completed execution and no lifecycle error is added
|
||||
|
||||
#### Scenario: Missing completion marker
|
||||
- **WHEN** a Docker TruffleHog process exits without emitting `finished scanning`
|
||||
- **THEN** the result is classified as an incomplete retryable run and is not treated as clean or done
|
||||
|
||||
#### Scenario: Nonzero exit after completion marker
|
||||
- **WHEN** a Docker TruffleHog process emits `finished scanning` but exits nonzero without a more specific diagnostic
|
||||
- **THEN** the result is classified as a retryable wrapper exit rather than successful execution
|
||||
|
||||
### Requirement: Bounded retry for incomplete Docker runs
|
||||
The system SHALL route incomplete Docker lifecycle failures through the existing bounded target retry policy and SHALL preserve the terminal attempt limit.
|
||||
|
||||
#### Scenario: Retry remains available
|
||||
- **WHEN** an incomplete Docker run occurs before the configured maximum target attempt
|
||||
- **THEN** the queue defers the target using the configured retry delay
|
||||
|
||||
#### Scenario: Attempt limit is reached
|
||||
- **WHEN** an incomplete Docker run occurs at the configured maximum target attempt
|
||||
- **THEN** the queue records a terminal failed target and does not create an unbounded retry loop
|
||||
|
||||
### Requirement: Partial finding preservation
|
||||
The system SHALL retain findings emitted before an incomplete Docker process exit without representing the target as fully scanned.
|
||||
|
||||
#### Scenario: Findings precede incomplete exit
|
||||
- **WHEN** TruffleHog emits one or more findings and then exits before complete execution is confirmed
|
||||
- **THEN** those findings remain durable while the target receives retryable incomplete disposition
|
||||
|
||||
### Requirement: Bounded Docker internal parallelism
|
||||
The system SHALL pass source-configured internal concurrency to Docker TruffleHog commands while retaining the existing source worker and Windows Job limits.
|
||||
|
||||
#### Scenario: Docker canary resource settings
|
||||
- **WHEN** the configured DockerHub source starts an image scan
|
||||
- **THEN** TruffleHog runs with internal concurrency 4 and a target timeout of 600 seconds
|
||||
|
||||
### Requirement: Nonfatal detector context timeout
|
||||
The system SHALL retain the exact diagnostic `a detector ignored the context timeout` as degraded detector coverage rather than a fatal image-scan error.
|
||||
|
||||
#### Scenario: Completed scan with detector timeout
|
||||
- **WHEN** a Docker scan emits the detector context-timeout diagnostic, emits `finished scanning`, and exits with code 0
|
||||
- **THEN** the target result contains a `detector_timeout` warning and no lifecycle error
|
||||
|
||||
#### Scenario: Other timeout diagnostic
|
||||
- **WHEN** a Docker scan emits a different timeout diagnostic
|
||||
- **THEN** the existing retryable timeout error policy remains in effect
|
||||
|
||||
### Requirement: Controlled historical replay
|
||||
The system SHALL replay historical Docker failures matching the exact incomplete-exit signature only in bounded batches after lifecycle canary criteria pass.
|
||||
|
||||
#### Scenario: Canary has not passed
|
||||
- **WHEN** lifecycle health has not met the defined canary criteria
|
||||
- **THEN** historical terminal failures are not mass-requeued
|
||||
|
||||
#### Scenario: Canary has passed
|
||||
- **WHEN** lifecycle health meets the defined canary criteria and a bounded replay batch is selected
|
||||
- **THEN** only exact-signature Docker failures in that batch are returned to the pending queue
|
||||
@@ -0,0 +1,22 @@
|
||||
## 1. Docker Lifecycle Policy
|
||||
|
||||
- [x] 1.1 Add `--local-dev` only to Docker TruffleHog command construction while retaining existing containment and `--no-update`.
|
||||
- [x] 1.2 Record the `finished scanning` marker and classify missing completion, unexplained RC=1, and wrapper exit outcomes with explicit retryable error classes.
|
||||
- [x] 1.3 Confirm incomplete outcomes use the existing three-attempt queue policy and preserve partial findings.
|
||||
- [x] 1.4 Add source-configured Docker internal concurrency 4 and target timeout 600 seconds.
|
||||
- [x] 1.5 Classify the exact detector context-timeout message as nonfatal degraded coverage.
|
||||
|
||||
## 2. Regression Coverage
|
||||
|
||||
- [x] 2.1 Add command-construction tests proving Docker receives `--local-dev` and non-Docker commands do not.
|
||||
- [x] 2.2 Add diagnostic tests for complete RC=0, missing-marker RC=0, incomplete RC=1, wrapper-exit RC=1, and unchanged non-Docker behavior.
|
||||
- [x] 2.3 Add queue-policy coverage for deferred attempts, terminal attempt exhaustion, and partial-finding preservation.
|
||||
- [x] 2.4 Run focused scanner tests and the relevant broader regression suite with bytecode writes disabled.
|
||||
- [x] 2.5 Add resource-plumbing and detector-timeout regression coverage, then rerun affected tests.
|
||||
|
||||
## 3. Validation And Rollout
|
||||
|
||||
- [x] 3.1 Run strict OpenSpec validation and verify the implementation against the change artifacts.
|
||||
- [x] 3.2 Restart the managed runtime and confirm PostgreSQL, pipeline workers, scan slots, and Docker source health.
|
||||
- [x] 3.3 Observe the Docker lifecycle canary for at least 100 attempts or two hours before historical replay.
|
||||
- [x] 3.4 Requeue one bounded batch of exact-signature historical failures and verify completion, deduplication, and queue drain.
|
||||
Reference in New Issue
Block a user