Initial server source import

This commit is contained in:
sashatrask
2026-09-30 20:30:56 +03:00
commit 170dd941b9
498 changed files with 261563 additions and 0 deletions
@@ -0,0 +1,254 @@
## Context
DockerHub currently queues immutable `repository@sha256:<platform-manifest>` targets and gives each target to TruffleHog's Docker source as one indivisible operation. The process must fetch and inspect every layer before the external 600-second deadline. Large model images contain tens of compressed GiB, commonly dominated by one binary/model-weight layer, so two such targets can occupy both Docker workers while still producing only partial findings.
Measured over 48 hours, 196 deadline terminations consumed about 32.7 Docker worker-hours. Every timeout emitted findings, but all timeout findings routed to only 14 provider keys and none was usable. The queue currently resets target attempts after a timeout and the timeout disposition bypasses the maximum-attempt check, so the same immutable target can restart from byte zero indefinitely.
Registry manifest resolution already obtains the exact platform child manifest and ordered layer digests. The missing information is each descriptor's compressed size, media type, image configuration descriptor, durable per-layer coverage, and a scanner capable of processing one bounded content blob independently.
The runtime must retain its authenticated Docker pool, immutable digest identities, slot-first PostgreSQL admission, result-bundle fencing, Windows Job containment, output bounds, private temporary storage, and fail-closed behavior. Bearer tokens and provider keys must remain memory-only or in their existing protected stores.
## Goals / Non-Goals
**Goals:**
- End unbounded timeout retries while preserving findings emitted before a deadline.
- Scan image configuration and useful application layers without downloading giant model/data layers.
- Resume image coverage after failure at layer granularity rather than restarting completed work.
- Scan each immutable content digest once globally and reuse its durable coverage across images.
- Keep download, disk, archive expansion, command execution, and database state bounded and fenced.
- Report exact selected, covered, skipped, failed, and shared-pending scope for every image.
- Prove the strategy against full-image control results before broad production enablement.
**Non-Goals:**
- Reconstruct a runnable merged container root filesystem.
- Claim complete image coverage when configured size or format bounds skip content.
- Download giant layers in a separate long-running production lane in the initial change.
- Implement arbitrary Registry hosts, private registries, cross-host credential forwarding, or resumable CDN downloads.
- Change global guaranteed scan slots, Docker worker count, keycheck classification, or non-Docker scanners.
- Delete the existing full-image implementation or schema on rollback.
## Decisions
### 1. Keep the image target, bind a fenced layer plan after claim
The existing immutable image target remains the queue and reporting identity. After slot-first admission claims an image, a dedicated PostgreSQL connection resolves its exact manifest and binds a canonical `docker_layer_plan` to that result reservation under the queue lease/event fences, analogous to exact Git plan binding.
The plan contains a version, repository, platform manifest digest, configuration descriptor, ordered layer descriptors, configured byte limits, selection reason for each descriptor, and a canonical SHA-256. The reservation stores the canonical JSON and hash. Replay is idempotent; changed replay or manifest mismatch is a conflict.
This avoids multiplying normal target-queue rows and keeps one authoritative parent disposition while still permitting durable child coverage.
Alternatives rejected:
- A separate full-image heavy lane still redownloads all content and cannot resume.
- One target-queue row per layer complicates parent completion, finding attribution, and alternate repository fetch sources.
- Encoding mutable coverage metadata into queue identity would break immutable image deduplication.
### 2. Add durable content and image-coverage tables
Add PostgreSQL-compatible tables through the additive runtime-safety migration:
- `docker_content_blobs`: one row per valid SHA-256 content digest, descriptor kind (`config` or `layer`), declared compressed bytes, media type, state, bounded attempts, active reservation/lease fence, completion metadata, and last bounded error.
- `docker_image_blob_coverage`: one row per platform manifest and content digest with ordered position, selected/skipped reason, plan hash, and covered timestamp.
Successful content state is global because the digest authenticates the bytes. The current image repository is retained only as the bounded fetch source in the reservation plan. If one reservation owns a blob, another image records it as shared-pending and defers without duplicating the download. Expired/refunded reservation leases are reclaimable.
Blob completion changes only during ingestion of a matching reservation plan and matching per-blob execution record. Refund/recovery releases matching blob leases. A stale token cannot mark coverage or insert findings.
### 3. Select content by configurable byte budget, top layers first
The image configuration is always selected within a small independent hard bound because `Env`, labels and history are high-value and cheap.
Layer selection walks the manifest from highest layer to base layer. Already-covered digests require no byte budget. New layers are selected while both the per-layer compressed-byte cap and remaining per-image compressed-byte budget permit them. Non-selected descriptors are recorded with explicit reasons such as `layer_too_large`, `image_budget_exhausted`, or `unsupported_media_type`.
Initial numeric defaults are chosen only after the controlled spike. Configuration always has hard upper bounds; invalid values fail closed to full-image mode while the production gate is disabled.
This is a deliberate partial-coverage policy. It targets application/config layers and globally amortizes common base layers without pretending that skipped model weights were scanned.
Alternatives rejected:
- A whole-image size cutoff can discard a small valuable application layer sitting above giant weights.
- Selecting oldest/base layers first spends budget on widely shared dependencies before application content.
- Inferring binary content without downloading a compressed tar stream is not reliable.
### 4. Stream, verify, scan, and delete one blob at a time
For each blob leased by the reservation:
1. Obtain an account-scoped Registry bearer through the existing Docker account manager.
2. Request the exact Registry blob with a bounded custom redirect policy. Redirects must remain HTTPS, must not contain userinfo, must reject local/private destinations, and must never receive the Registry Authorization header on another host.
3. Stream into a private bounded work file while computing SHA-256. Reject excess bytes, short bodies, digest mismatch, unsupported media types, disk-floor violations, and response deadline exhaustion.
4. Scan configuration JSON directly. Scan supported layer archives with TruffleHog `filesystem` under the existing OwnedProcess Job, timeout, output cap, configured detector policy, archive size/depth/time bounds, and normal process priority.
5. Delete the blob work file before releasing the scan slot. No layer bearer, provider key, or raw result is written outside existing protected result artifacts.
Each layer runs as its own bounded TruffleHog command. This adds small startup overhead but gives exact completion, global deduplication, and restart from the first unfinished layer. Findings are enriched with immutable image, blob digest, kind, and layer position before normal bundle staging.
The controlled spike must prove that the installed TruffleHog build correctly scans supported real layer archive media types. Unsupported formats remain explicit uncovered scope rather than being silently accepted.
### 5. Parent disposition derives from explicit child execution
The result bundle carries the exact bound plan plus one execution record per claimed blob. Ingestion verifies plan hash and lease ownership before updating blob state.
- Successful blob command and digest verification marks that blob covered globally.
- Incomplete blob execution preserves its findings but does not mark it covered.
- Retryable blob failures release it to bounded retry; terminal failures remain explicit uncovered scope.
- A blob active under another valid reservation causes a short parent deferral without charging a content attempt.
- An image whose selected blobs are covered and whose remaining blobs are intentionally skipped completes with `coverage_complete=false` and detailed reasons.
- An image with retryable selected work remains deferred; exhausted selected work becomes terminal failed/degraded according to the bound policy.
The existing full-image path remains available as a rollback/control path.
### 6. Fix timeout accounting before layer rollout
Both production-v2 and legacy completion paths stop resetting attempts for target-scoped timeouts. Timeout disposition checks the configured maximum before returning deferred. Source-wide infrastructure failures may retain their existing attempt-refund semantics.
A stopped-runtime repair reconciles only unfenced Docker queue rows whose latest durable result is a command timeout. It derives prior immutable-target attempt count from durable scans, sets the queue attempt count up to the configured maximum, and terminally closes already exhausted rows. It does not requeue or mutate active leases, reservations, findings, or successful targets.
### 7. Roll out through deterministic modes and evidence gates
Configuration exposes `full`, `canary`, and `layer` modes. Canary eligibility requires a durable previous full-image command timeout, and membership within that eligible set is a stable hash of immutable manifest digest. New images and normally completed controls therefore remain on the full path during canary rollout. Defaults remain `full` until migration and spike criteria pass.
The offline spike compares 20-50 timeout-heavy images and completed controls without printing findings or keys. It records bytes transferred, wall/slot time, peak resource use, distinct detector identities, and routed key identity recall. Production canary additionally tracks coverage reasons, blob reuse, timeout rate, keycheck candidate yield, and strict usable yield.
Broad enablement requires:
- no credential/token persistence regression;
- no stale-fence or duplicate-blob completion;
- exact digest verification for every covered blob;
- material byte and slot-hour reduction on heavy images;
- all routed key identities from completed control images retained, unless an explicitly reviewed coverage bound explains the difference;
- no quarantine growth, projection regression, or source failure increase.
## Risks / Trade-offs
- [Secrets can exist in skipped giant layers] -> Record explicit incomplete scope, keep configurable budgets, compare controls, and retain full mode for targeted replay.
- [Scanning individual layers can report files deleted by later whiteouts] -> Preserve layer provenance and treat this as historical image-content evidence rather than merged-root truth.
- [Archive support differs by media type] -> Verify installed binary in the spike and mark unsupported formats uncovered.
- [Global deduplication can be poisoned by stale completion] -> Require streamed digest verification plus reservation/lease/plan fences in the same ingestion transaction.
- [Registry blob redirects introduce SSRF or credential-forwarding risk] -> Use a bounded validated redirect implementation and strip authorization across hosts.
- [Per-layer process startup adds overhead for tiny layers] -> Batch measurement first; skip already-covered layers and permit a bounded future batching optimization only if needed.
- [A crash can strand blob leases] -> Tie leases to result reservations, release on refund, and permit exact expiry reclamation.
- [Database state grows with layer relationships] -> Enforce descriptor count bounds, compact metadata, indexed identities, and retention metrics.
- [Layer mode can reduce broad secret coverage] -> Report coverage honestly and retain deterministic full-mode controls and rollback.
## Migration Plan
1. Ship and test bounded timeout accounting independently; repair exhausted historical timeout rows with sources stopped.
2. Add nullable reservation plan columns, content/coverage tables, indexes, runtime validation, and additive migration marker. Keep mode `full`.
3. Run a local controlled spike against retained immutable targets and select conservative byte/archive defaults from evidence.
4. Deploy code with `full` mode, migrate offline, restart, and verify no behavior change.
5. Enable deterministic low-percentage canary only for DockerHub images whose latest durable full-image result timed out. Monitor at least one full repository-refresh interval and sufficient heavy-image samples.
6. Increase the timeout-fallback canary only after acceptance gates pass. Keep broad `layer` mode disabled until completed-control routed recall becomes adequate under revised bounds.
7. Roll back by returning mode to `full`. Durable layer tables and nullable columns remain for audit and future resume; no destructive migration is required.
## Capability Spike Evidence
The installed `C:\Tools\trufflehog.exe` development build was exercised through 39 bounded
`filesystem` invocations over synthetic direct tar, gzip-tar, zstd-tar, Docker outer-tar, and OCI
outer-tar fixtures. All invocations exited successfully and all 24 expected synthetic findings were
preserved. Extensionless gzip and zstd blobs were content-sniffed successfully.
The minimum archive depth was two for a direct tar, three for direct compressed layers, and four
for an OCI/Docker archive containing a compressed layer. Shallower bounds produced an explicit
non-fatal `max archive depth reached` diagnostic. `--archive-max-size` was proven to be a per-member
bound rather than a cumulative compressed-image bound, so application-side per-blob and aggregate
byte limits remain mandatory.
Initial conservative implementation bounds are therefore archive depth four, archive member size
256 MiB, archive timeout 30 seconds, one-GiB aggregate selected compressed bytes per image, 256 MiB
per selected layer, and filesystem concurrency two. These are implementation starting points, not
broad-rollout acceptance evidence. The required aggregate timeout-heavy/completed-control image
comparison remains an explicit gate before canary expansion.
The controlled aggregate comparison then ran against ten repeatedly timed-out immutable images and
ten longest completed controls. No target names, findings, keys, or credentials were printed or
persisted. Exact manifest resolution succeeded for all 20 images. The timeout-heavy manifests
contained 341.68 GB of compressed descriptors; the bounded layer policy selected 944.02 MB (0.28%)
and processed 84 blobs in 313.03 worker-seconds versus 6,015.31 historical full-image seconds
(5.2%). All selected timeout-heavy blobs completed. The bounded path produced three routed
credential identities, none overlapping the two identities in the historical incomplete full
results, so it added useful scope while avoiding another byte-zero full-image retry.
The completed controls contained 98.05 GB; the policy selected 1.61 GB (1.64%) and processed 77
blobs in 357.83 worker-seconds versus 5,956.38 historical seconds (6.0%). Five blobs reported
bounded incomplete chunk processing. Only four of 19 historical routed identities were retained
(21.1% recall), and distinct detector-identity recall was 19 of 326 (5.8%). System sampling across
the run observed average CPU 20.27%, peak CPU 44.48%, minimum available physical memory 17.38 GB,
and peak committed memory 30.00 GB; no resource-limit failure occurred.
This evidence accepts the existing 1 MiB config, 256 MiB per-layer, 1 GiB aggregate, eight-layer,
archive-depth-four, 256 MiB archive-member, 30-second archive, 600-second blob, and filesystem
concurrency-two bounds only for a deterministic timeout-fallback canary. It rejects broad random
canary or broad layer mode because completed-control routed recall failed the acceptance gate. Full
mode remains authoritative for new and normally completed images.
The post-refinement regression gate passed 329 related tests, including bounded transfer and content
validation, parent/blob timeout accounting, PostgreSQL migration atomicity, policy-scoped global
deduplication, stale fences, reclaim/quarantine, multi-checkpoint resume, and the durable full-timeout
canary-eligibility transition. Strict OpenSpec validation passed, and no application bytecode was
present after the run.
Rejected approaches from the capability spike are `--force-skip-archives` (it suppresses expected
archive findings), relying on media-type labels without content verification, and relying on
TruffleHog's `--archive-max-size` as an outer download or aggregate expansion bound. The exact
development binary must remain fingerprinted by a capability contract because its reported version
does not identify a stable release.
## Initial Production Canary Evidence
The additive migration and guarded historical timeout repair were applied offline before the
timeout-only canary. Runtime first restarted in full mode with no layer rows, and the bounded canary
was then enabled at 2,500 basis points only for images whose latest durable full-image result was a
command timeout. New and normally completed images remained on the full scanner.
The first selected production image completed nine unique content checkpoints plus one final
no-work completion checkpoint. The nine bounded blobs transferred 327,512 bytes and completed in
33.176 seconds of aggregate parent duration, including 9.358 seconds of transfer and 23.496 seconds
of contained filesystem scanning. All nine blobs reached policy-scoped global coverage, the parent
finished `done`, no selected continuation remained, and quarantine stayed unchanged at 192 items /
337,349,428 bytes. The prior full-image path for this eligibility class reached the 600-second
deadline.
The initial run exposed and then verified a selection-continuity invariant: per-image byte/layer
bounds must apply to the first immutable selection set, not be recomputed after each covered blob.
The binder now reuses the earliest exact `(queue, manifest, coverage policy, position)` selection map
on every later reservation. A PostgreSQL regression proves that layers skipped by the original
count budget remain skipped after selected layers become globally covered. Production replay then
completed without leasing content outside the original config-plus-eight-layer selection.
This single successful image proves the end-to-end checkpoint, resume, bounded selection and final
completion paths, but is not enough evidence to increase the 25% timeout-only canary. Expansion
still requires a longer observation window and more naturally eligible timeout samples.
The following overnight window added a second timeout-only image before a host reboot. Across both
images, twelve unique blobs reached policy-scoped coverage with no retryable or terminal blob
failure. The layer path used 53.192 seconds and transferred 79,614,382 bytes, compared with 1,201.330
seconds consumed by the immediately preceding full-image timeout attempts. One parent completed;
the second retained six exact selected checkpoints for durable resume. No layer findings or routed
candidate identities were produced in this small sample, and quarantine remained unchanged.
Because this installation is an experimental rather than production service, the operator approved
expanding the stable timeout-only cohort from 2,500 to 10,000 basis points. This does not enable broad
layer mode: every new or normally completing image still uses the full scanner, and only an image
with a durable prior full-image timeout may enter the bounded layer fallback. Broad layer mode
remains rejected by the completed-control recall result. Task 7.5 remains open until the expanded
cohort produces additional completed parents and a stable runtime observation window.
The expanded cohort exposed one additional metadata-normalization defect: identical content digests
and sizes can be referenced through equivalent Docker and OCI media-type labels. Treating the label
text itself as immutable metadata caused the Docker source to stop fail-closed before handoff. The
binder now compares the validated semantic content class (`config-json`, `layer-tar`, `layer-gzip`,
or `layer-zstd`) while still rejecting kind, byte-size, and compression-class conflicts. A real
PostgreSQL concurrency regression and the related layer/runtime suites passed 126 tests. After the
restart, 116 layer reservations were acknowledged in the first ten minutes, global covered blobs
grew from 9 to 113, no blob entered a failed state, the source remained running, and quarantine was
unchanged. The immediate digest queue remained intentionally thin (six due and eighteen delayed),
while 12,679 repository anchors remained available to refill it after claimable digest work drains.
## Open Questions
- Which revised per-layer and per-image bounds can improve completed-control routed recall beyond the measured 21.1% without losing the measured slot-hour advantage?
- Which OCI/Docker layer compression media types does the installed TruffleHog filesystem source handle reliably?
- Is one TruffleHog process per selected layer sufficiently efficient, or is a later bounded multi-layer archive batch warranted?
- What short deferral is appropriate when all remaining selected blobs are actively leased by other reservations?