353 lines
22 KiB
Markdown
353 lines
22 KiB
Markdown
## Context
|
|
|
|
Docker scanning currently has two materially different paths. The full-image TruffleHog source is
|
|
authoritative for new and normally completing images, but it treats an immutable image as one
|
|
indivisible operation. The bounded layer path can resume and globally reuse content digests, but its
|
|
fixed highest-eight-layer selector was admitted only for prior full-image timeouts. In the completed
|
|
control cohort it retained 21.1% of routed identities and 5.8% of detector identities, so enabling
|
|
that selector broadly would trade too much useful coverage for speed.
|
|
|
|
The layer path already provides the expensive safety primitives this change needs: exact manifest
|
|
resolution, authenticated bounded Registry transfer, digest verification, private artifacts,
|
|
contained TruffleHog filesystem execution, canonical reservation-bound plans, policy-scoped global
|
|
blob leases, fenced result ingestion, and explicit per-image coverage. The missing pieces are a
|
|
content-aware selector, a reusable execution-policy identity that is independent of selector
|
|
budgets, and trustworthy evidence for deciding whether the adaptive path is safe to broaden.
|
|
|
|
Docker/OCI configuration contains an ordered `history` list that can help identify `COPY`, `ADD`,
|
|
application setup, package installation, generic `RUN`, and likely bulk-data layers. That history is
|
|
untrusted and may itself contain secret material. It is therefore only a bounded selection hint; it
|
|
must never become an authority for content identity or successful coverage and its raw commands
|
|
must never be persisted or logged.
|
|
|
|
The production database already contains durable version-one layer plans and covered blob rows.
|
|
Changing their interpretation in place would invalidate audit and lease fences. The migration must
|
|
be additive, retain version-one validation, and allow only explicitly proven compatible successful
|
|
coverage to enter the new execution namespace.
|
|
|
|
## Goals / Non-Goals
|
|
|
|
**Goals:**
|
|
|
|
- Scan all supported unique content for images that fit conservative bounds.
|
|
- Prioritize likely application, configuration, and source-bearing layers when a large image cannot
|
|
fit those bounds.
|
|
- Reuse successful immutable blob coverage across images and selector revisions when execution
|
|
semantics are unchanged.
|
|
- Freeze the first deterministic selection for an image and selector policy across every checkpoint,
|
|
retry, and execution-policy transition.
|
|
- Reduce per-image checkpoint overhead with bounded multi-blob leases without weakening per-blob
|
|
execution and ingestion fences.
|
|
- Preserve exact selected, reused, skipped, failed, and partial coverage semantics.
|
|
- Produce private, aggregate, non-authoritative shadow evidence against 50-100 completed full-image
|
|
controls before any broad adaptive rollout.
|
|
- Retain full-image scanning and the existing timeout-only layer canary as immediate rollback paths.
|
|
|
|
**Non-Goals:**
|
|
|
|
- Infer arbitrary file contents without downloading a compressed layer.
|
|
- Claim complete coverage for an image with unsupported, oversized, failed, or budget-excluded
|
|
descriptors.
|
|
- Persist raw Docker config history, shadow findings, provider keys, target names, or Registry bearer
|
|
tokens in rollout evidence.
|
|
- Change detector classification, keycheck routing, Docker account ownership, repository resolver
|
|
scheduling, guaranteed scan-slot capacity, worker count, or non-Docker scanners.
|
|
- Reconstruct a merged container filesystem or remove historical whiteout content from individual
|
|
layer evidence.
|
|
- Destructively rewrite or delete version-one plans and coverage rows.
|
|
|
|
## Decisions
|
|
|
|
### 1. Introduce a version-two immutable content plan
|
|
|
|
Adaptive execution uses a canonical version-two plan. It retains the version-one image, repository,
|
|
manifest, platform, media, limits, scan-policy, descriptor, reservation, and plan-hash fences and
|
|
adds:
|
|
|
|
- a versioned selector algorithm and `selection_policy_sha256`;
|
|
- an `execution_policy_sha256` for reusable successful blob evidence;
|
|
- one bounded classifier value per descriptor;
|
|
- exact selection and omission reasons; and
|
|
- bounded checkpoint lease limits that do not alter the frozen selection set.
|
|
|
|
Version-one plan and execution validators remain exact because their JSON is already durable.
|
|
Version-two validation dispatches by the exact integer version and rejects unknown fields, unknown
|
|
classes, malformed hashes, descriptor reordering, and oversized canonical JSON. A reservation can
|
|
bind one canonical plan only; idempotent replay must be byte-identical.
|
|
|
|
After exact manifest resolution and before plan binding, the scanner fetches the configuration blob
|
|
through the existing bounded, authenticated, digest-verifying Registry path. It parses the JSON in
|
|
bounded private memory/storage and maps non-`empty_layer` history entries from base to top onto the
|
|
ordered manifest layers. The optional `rootfs.diff_ids` count and the non-empty history count must
|
|
agree with the layer count. A missing field, malformed value, excess entry count, or alignment
|
|
mismatch classifies every layer as `unknown`; it does not change descriptor identity or fail a
|
|
valid immutable manifest.
|
|
|
|
The classifier uses a small versioned allow-list and emits only these bounded classes:
|
|
`config`, `copy_add`, `app_config_run`, `package_run`, `other_run`, `bulk_data`, and `unknown`.
|
|
Raw `created_by` values are discarded before plan construction and are excluded from logs, result
|
|
metadata, errors, and database rows.
|
|
|
|
Alternatives rejected:
|
|
|
|
- Persisting normalized command text would retain unnecessary secret-bearing input.
|
|
- Treating history as authoritative would let malformed or adversarial metadata hide content.
|
|
- Mutating version-one plans would break exact replay and auditability.
|
|
|
|
### 2. Separate scan, execution, and selection identities
|
|
|
|
Three hashes have distinct responsibilities:
|
|
|
|
- `scan_policy_sha256` remains the installed scanner/detector/config fingerprint.
|
|
- `execution_policy_sha256` hashes the scan policy plus versioned content validation and all archive
|
|
semantics that can change which bytes a successful command examines. It excludes image/layer
|
|
selection budgets, classifier weights, retry counts, lease duration, checkpoint size, and delay.
|
|
- `selection_policy_sha256` hashes the selector version, class ordering, deterministic tie-breaks,
|
|
supported descriptor classes, and all limits that determine the initial selected set. It excludes
|
|
mutable global coverage and execution scheduling.
|
|
|
|
For version-two rows, the existing `docker_content_blobs.coverage_policy_sha256` key stores the
|
|
execution-policy hash. Selector changes therefore do not force an identical successfully scanned
|
|
digest through TruffleHog again, while archive or detector semantic changes still create a separate
|
|
coverage namespace.
|
|
|
|
`docker_image_blob_coverage` receives a non-null `selection_policy_sha256`. Existing rows are
|
|
backfilled from the exact bound plan on their linked reservation and indexed by
|
|
`(queue_id, manifest_digest, selection_policy_sha256, position, reservation_id)`. The earliest full
|
|
position map under that key is the immutable selection baseline. Coverage-policy changes may require
|
|
new execution but cannot expand or contract that baseline silently.
|
|
|
|
Successfully covered version-one evidence may be copied lazily into the version-two execution
|
|
namespace only in the same transaction that validates all of the following:
|
|
|
|
- the old row is durably `covered`, not pending, leased, submitted, failed, or ambiguous;
|
|
- a linked covered image row and reservation contain an exact valid version-one plan;
|
|
- the descriptor digest, kind, declared bytes, and semantic media class match;
|
|
- the old coverage key is exactly the legacy hash derived from that plan; and
|
|
- the scan fingerprint and every scan-affecting archive semantic equal the requested version-two
|
|
execution policy.
|
|
|
|
The alias keeps the original successful reservation, plan, byte count, and completion provenance.
|
|
Legacy rows are never rekeyed or deleted. If any compatibility proof is absent, the new namespace
|
|
starts uncovered and normal fenced execution is required.
|
|
|
|
Alternatives rejected:
|
|
|
|
- Keeping selector limits in the coverage hash defeats global reuse whenever budgets are tuned.
|
|
- Reusing every covered digest across scanner versions can suppress required rescans.
|
|
- Bulk-rekeying legacy rows destroys provenance and races active leases.
|
|
|
|
### 3. Select all-fit images and rank large-image payload deterministically
|
|
|
|
Configuration remains independently eligible under its hard configuration-byte bound. Supported
|
|
unique layer digests are evaluated under the hard per-layer bound. Covered digests in the matching
|
|
execution namespace and duplicate positions in the same image are selected at zero new transfer
|
|
bytes and zero new execution count.
|
|
|
|
If every supported unique descriptor fits the configured aggregate bytes and unique-layer count,
|
|
the selector selects all of them regardless of history class. This is the complete bounded path for
|
|
small images and avoids reducing their coverage merely because history hints are absent.
|
|
|
|
When the complete set does not fit, new unique layer candidates are sorted by:
|
|
|
|
1. class priority: `copy_add`, `app_config_run`, `package_run`, `unknown`, `other_run`, `bulk_data`;
|
|
2. highest manifest position first;
|
|
3. smallest compressed descriptor first; and
|
|
4. lexical digest as the final stable tie-break.
|
|
|
|
The selector greedily admits candidates while both aggregate compressed-byte and unique-layer-count
|
|
limits permit them. Hard per-descriptor limits are never exceeded. Each descriptor records one exact
|
|
reason, including selected class, `already_covered`, `duplicate_digest`, `unsupported_media_type`,
|
|
`config_too_large`, `layer_too_large`, `image_budget_exhausted`, or `layer_limit_exhausted`.
|
|
Changing any class order, classifier rule, supported-media rule, or selection bound changes the
|
|
selector hash.
|
|
|
|
The selection algorithm receives a transactionally consistent coverage snapshot, but mutable
|
|
coverage is not part of its identity. The first complete descriptor-position map is written before
|
|
any new lease and reused exactly on later checkpoints. Consequently a layer skipped by the original
|
|
budget never becomes newly selected merely because an earlier selected layer became globally
|
|
covered.
|
|
|
|
Alternatives rejected:
|
|
|
|
- Fixed highest-first selection has already failed the completed-control recall gate.
|
|
- A whole-image byte cutoff loses small application layers above giant data layers.
|
|
- Selecting only recognized commands lets missing or unusual history hide useful payload.
|
|
- Selecting globally covered content only when it still fits the current budget wastes verified
|
|
immutable evidence.
|
|
|
|
### 4. Lease bounded multi-blob checkpoints
|
|
|
|
The current executor and ingestion format already support more than one leased descriptor, but the
|
|
binder leases one new digest and then defers the parent for 60 seconds. Adaptive execution leases a
|
|
deterministic bounded batch from the frozen selected set under configurable maximum blob count and
|
|
compressed bytes. A first eligible blob larger than the checkpoint-byte target but within its hard
|
|
per-layer bound may be leased alone so it cannot starve indefinitely.
|
|
|
|
Every digest still has its own advisory lock, lease token, attempt count, execution record, digest
|
|
verification, and final state. The executor processes the batch sequentially inside the same owned
|
|
slot and bundle. Ingestion may cover successful earlier blobs while returning a later retryable blob
|
|
to pending. A crash before durable handoff covers none of the un-ingested batch and normal exact
|
|
lease expiry/recovery applies.
|
|
|
|
Checkpoint count, byte target, retry count, lease duration, and continuation delay are scheduling
|
|
controls. They do not enter execution or selector hashes because they cannot turn an incomplete blob
|
|
into successful coverage or change the frozen selected set.
|
|
|
|
### 5. Keep image coverage explicit and policy-specific
|
|
|
|
An image is complete only when its configuration and every manifest layer position are successfully
|
|
covered under the requested execution policy. Reused and duplicate digests count as covered only
|
|
after exact policy-compatible evidence exists. Any unsupported, oversized, budget-excluded, failed,
|
|
or otherwise unselected descriptor makes the image bounded partial coverage.
|
|
|
|
Selected retryable work keeps the parent deferred. Shared active work does not consume another blob
|
|
attempt. Exhausted selected work produces terminal incomplete disposition. Findings from completed
|
|
selected blobs retain image, digest, kind, class, and position provenance and use the existing
|
|
authoritative ingestion, projection, and keycheck paths.
|
|
|
|
Config history classification affects only selection order. It never changes detector output,
|
|
finding authority, digest identity, or completion criteria.
|
|
|
|
### 6. Add adaptive modes without changing legacy rollout semantics
|
|
|
|
Existing `full`, timeout-only `canary`, and legacy `layer` meanings remain available for durable
|
|
version-one work and rollback. Two explicit version-two modes are added:
|
|
|
|
- `adaptive-canary` assigns a configured basis-point cohort across all immutable Docker manifests by
|
|
a stable versioned hash; cohort members use adaptive plans and non-members use full-image scanning.
|
|
- `adaptive` uses adaptive plans for every eligible immutable Docker claim.
|
|
|
|
Neither mode depends on a previous full-image timeout. Retry and checkpoint attempts for the same
|
|
manifest and selector retain the same assignment. Invalid mode, policy, migration, or gate state
|
|
fails closed to full-image execution before any adaptive plan is bound. Returning configuration to
|
|
`full` changes only new claims and leaves adaptive plans and audit rows intact.
|
|
|
|
The existing timeout-only canary remains independent and may continue while adaptive shadow evidence
|
|
is gathered. The broad legacy `layer` mode remains operationally disabled because its selector did
|
|
not pass recall gates.
|
|
|
|
### 7. Gate rollout with non-authoritative aggregate shadow evidence
|
|
|
|
An operator-invoked bounded shadow evaluator selects 50-100 exact immutable images whose authoritative
|
|
full-image scans completed successfully under one scan fingerprint. It executes the candidate
|
|
adaptive policy using the same downloader, validators, process containment, deadlines, and scanner
|
|
fingerprint, but under a shadow authority that cannot call normal result ingestion or mutate target
|
|
status, result reservations, global blob coverage, findings, keycheck candidates, projections, or
|
|
source counters.
|
|
|
|
For each paired control, full routed identities are read from the protected database as
|
|
`(service, provider_key_hash)` and adaptive routed identities are derived in private memory through
|
|
the same candidate normalization. Detector identities use `detector_secret_hash`. Identity sets,
|
|
raw findings, commands, provider material, image names, and bearer tokens are discarded after
|
|
intersection counts are computed.
|
|
|
|
The durable report contains only policy hashes, cohort and completion counts, aggregate full,
|
|
adaptive, and intersection counts, aggregate slot milliseconds, bounded failure counts, threshold
|
|
results, and timestamps. It records no per-image row or identity. Slot timing uses the same outer
|
|
monotonic boundary from admitted work through durable shadow sink completion for both paths; scan
|
|
subprocess duration remains a diagnostic, not the gate denominator.
|
|
|
|
A report passes only when:
|
|
|
|
- 50-100 controls completed both paths without integrity, containment, or fence failure;
|
|
- aggregate routed-identity recall, `intersection / full`, is at least 85%;
|
|
- aggregate adaptive/full slot-time ratio is at most 40%;
|
|
- every adaptive omission is represented in coverage counts; and
|
|
- no credential persistence, quarantine, projection, source-failure, or resource-bound regression
|
|
is observed.
|
|
|
|
Reports are bound to exact selector, execution, and scan policy hashes. Stale or incomplete reports
|
|
cannot authorize another policy. Adaptive canary remains fail-closed until a matching report passes.
|
|
Broad `adaptive` enablement additionally requires a stable low-percentage production canary over at
|
|
least one repository-refresh interval. Operators change rollout configuration explicitly; shadow
|
|
evidence never changes execution mode by itself.
|
|
|
|
Alternatives rejected:
|
|
|
|
- Routing shadow candidates through keycheck would make the experiment authoritative and consume
|
|
external capacity.
|
|
- Persisting per-image shadow identities creates unnecessary sensitive correlation data.
|
|
- Comparing only detector counts does not measure the routed identities the scanner is intended to
|
|
produce.
|
|
- Automatically enabling adaptive mode from a report removes the operational rollback checkpoint.
|
|
|
|
### 8. Preserve incomplete warning semantics without redundant retries
|
|
|
|
The first completed 50-control production shadow report failed closed. Routed recall was 12 of 18
|
|
identities (66.7%), adaptive/full slot time was 46.6%, and the report recorded 43 aggregate
|
|
failures. Its selection evidence showed 230 descriptors omitted by the eight-layer limit and 14
|
|
oversized descriptors, so selector recall remains the primary rollout blocker.
|
|
|
|
The same evidence exposed a separate execution defect. TruffleHog diagnostics such as
|
|
`chunk_processing` and `detector_timeout` are explicitly deterministic, non-retryable warnings.
|
|
Their findings must be retained, but the affected blob cannot establish complete coverage. The
|
|
diagnostic adapter previously discarded the non-retryable bit, causing the layer executor to
|
|
download and scan the same incomplete blob up to three times before reaching the same terminal
|
|
state. The adapter now preserves aggregate warning retryability and the layer executor terminates
|
|
that blob after the first deterministic warning. It does not mark the blob covered or remove the
|
|
report failure.
|
|
|
|
The historical report schema retained only a total failure count, so its 43 failures cannot be
|
|
decomposed exactly after the fact. Future shadow runs keep a fixed allow-list of aggregate-only
|
|
failure categories in protected memory and print their totals in the final operator summary without
|
|
changing report authority or persisting target-level evidence.
|
|
|
|
Shadow execution also suppresses target labels in finding-filter logs. Production scans retain their
|
|
existing target logging, while both private full and private layer paths emit only aggregate filter
|
|
counts. This closes a privacy gap found in the first report log without weakening normal operational
|
|
diagnostics.
|
|
|
|
## Risks / Trade-offs
|
|
|
|
- [History is malformed, misleading, or secret-bearing] -> Bound and validate it, persist only an
|
|
enum, fall back to `unknown`, and keep identity/coverage independent of classification.
|
|
- [Application secrets exist in a low-priority or giant layer] -> Scan all-fit images, keep unknown
|
|
ahead of generic/bulk classes, record partial scope, retain full controls, and enforce the 85%
|
|
routed-recall gate.
|
|
- [Unsafe policy reuse suppresses a required rescan] -> Separate hashes and permit legacy aliasing
|
|
only from exact successful compatible evidence in one fenced transaction.
|
|
- [Selector changes expand coverage during retry] -> Freeze the earliest full position map under the
|
|
selector hash before leases are issued.
|
|
- [Multi-blob checkpoints increase work lost on crash] -> Bound count/bytes and retain independent
|
|
per-blob leases and ingestion records; no pre-handoff result becomes covered.
|
|
- [Config prefetch adds Registry traffic] -> Reuse the already bounded authenticated downloader and
|
|
avoid a second fetch when the config is leased in the same plan.
|
|
- [Individual layer scans expose whiteouted historical files] -> Preserve position provenance and
|
|
describe evidence as image-content coverage, not merged-root state.
|
|
- [Shadow evaluation consumes slots] -> Keep it operator-invoked, bounded, deterministic, and
|
|
subject to the existing slot/resource controls.
|
|
- [Aggregate reports hide individual anomalies] -> Fail the whole report on incomplete paired work
|
|
and retain bounded failure counts without persisting target identity.
|
|
- [Deterministic scanner warnings consume repeated transfer and slot time] -> Preserve their
|
|
non-retryable policy, keep findings and incomplete coverage, and terminate the blob on its first
|
|
bounded attempt.
|
|
|
|
## Migration Plan
|
|
|
|
1. Ship version-one and version-two validators, new modes, and migration code while production stays
|
|
on its existing timeout-only canary.
|
|
2. Stop authoritative runtime and verify no active result, queue, or blob leases remain.
|
|
3. Add the selector-policy coverage column, aggregate shadow-report state, required timing/fence
|
|
columns, indexes, and a new migration marker.
|
|
4. Backfill every historical image-coverage row from its exact valid reservation plan. Abort and
|
|
roll back the migration transaction on an orphan, malformed plan, or invalid hash; then make the
|
|
column non-null and run exact schema validation.
|
|
5. Restart with unchanged mode and verify legacy claims, projection, keycheck, resolver, quarantine,
|
|
and source health before creating version-two work.
|
|
6. Run the private shadow evaluator for 50-100 completed controls. Keep adaptive modes fail-closed if
|
|
the matching report misses recall, timing, safety, or completion gates.
|
|
7. Enable a low deterministic `adaptive-canary`, monitor at least one repository-refresh interval,
|
|
and compare source failures, coverage reasons, routed yield, slot time, and quarantine.
|
|
8. Increase canary basis points and finally enable `adaptive` only after every gate remains satisfied.
|
|
9. Roll back immediately by setting mode to `full` or the existing timeout-only `canary`. Keep all
|
|
version-two plans, reports, and coverage rows for audit and exact future resume.
|
|
|
|
## Open Questions
|
|
|
|
- Which initial bounded checkpoint count and byte target provide the best reduction in continuation
|
|
delay without increasing crash rework materially?
|
|
- Which classifier allow-list revisions improve routed recall in the first 50-100 controls? Every
|
|
revision will receive a new selector-policy hash rather than changing an existing policy.
|
|
- What adaptive-canary basis-point sequence should operators use after the shadow gate passes?
|