Files
truf-server/docs/end-to-end-scanner-validation.md
T
2026-09-30 20:30:56 +03:00

14 KiB

End-to-End Scanner Validation

This runbook validates the real discovery, remote-worker, result-ingestion, and compatibility-projection path with a bounded cohort of 30-40 targets. It is an evidence procedure, not a claim that every repository feature and environment has been proven correct.

The completed 2026-09-22 production evidence is recorded in end-to-end-scanner-validation-2026-09-22.md.

Validation Questions

The run must answer all of the following:

  1. Does bounded discovery create the expected immutable queue identities?
  2. Do remote workers receive each authoritative assignment with the correct source, target identity, execution snapshot, and lease fencing?
  3. Does each worker run the intended TruffleHog scan and upload a canonical protocol-2 result bundle?
  4. Does the server accept a result exactly once and make receipt replay idempotent?
  5. Does the ingester transactionally connect the reservation, queue row, target scan, findings, errors, and bundle record?
  6. Does the JSONL projector reproduce the Windows-compatible scan_results.jsonl and found_secrets.jsonl structures without becoming a second source of truth?
  7. Are naturally found or controlled-canary findings stored with safe identity, location, detector, verification, redaction, and provenance fields?
  8. After the test, are queues settled, projections caught up, no pipeline item quarantined, and the normal production configuration restored?

Authorities and Expected Data Flow

The authoritative sequence is:

discovery cycle
  -> target_queue
  -> result_reservation / immutable remote assignment
  -> worker TruffleHog execution
  -> canonical .trb upload
  -> durable accepted receipt
  -> result ingester transaction
  -> target_scans + findings + errors + queue completion
  -> projection_jobs
  -> scan_results.jsonl + found_secrets.jsonl

PostgreSQL is authoritative. Result bundles are durable pipeline artifacts. JSONL files are rebuildable compatibility projections and may lag briefly. accepted proves durable server receipt; ingested proves the database transaction completed. These states must not be treated as synonyms.

Safety Rules

  • Use only the approved test server and approved worker devices.
  • Never print device tokens, source credentials, raw secrets, active runtime YAML, worker argv, or complete unredacted findings into a terminal/log.
  • Query finding structure using IDs, hashes, redacted values, lengths, booleans, detector names, and location metadata. Review any raw secret only through the already protected admin workflow if explicitly required.
  • Apply and restore configuration through Preview -> Save candidate -> Apply. Do not edit the active runtime document in place.
  • Pause discovery and dispatch and drain before each runtime-document apply.
  • Record the original config SHA-256 and require byte-identical restoration at the end.
  • Use a unique test-run label and database high-water marks. Never infer the cohort from wall-clock time alone.
  • Do not delete queue, bundle, finding, projection, or receipt evidence to make a failed test look clean.

Bounded Test Configuration

Use a temporary candidate derived from the active document. Preserve all secrets and unrelated settings. The exact candidate must be reviewed before it is applied.

DockerHub

  • mode: search
  • pages: 1
  • per_page: 1
  • docker_images_per_repository: 1
  • Use a reviewed finite query list for the test window.

One query is consumed per source cycle. An already-known or unsuitable search result can produce no new queue row, so the number of cycles is not the cohort size.

GitLab

  • mode: search
  • pages: 1
  • per_page: 1
  • Use a reviewed finite query list for the test window.
  • Keep current age, commit-boundary, exact-ref, visibility, and history-depth safety controls unless the test explicitly records a different expectation.

HuggingFace

HuggingFace recent discovery does not support a real per_page: 1 keyword test. Its API runner fetches newest-modified Spaces and the current API page size is fixed at 100; the configured query is only a rotation placeholder.

For a bounded cohort, use mode: custom with a reviewed private target_file containing a small list of Space IDs. Do not claim that changing per_page to 1 bounded this source when it did not.

Target 36 authoritative terminal scans:

  • 16 DockerHub immutable digest targets;
  • 16 GitLab exact-ref/commit-planned targets;
  • 4 HuggingFace custom Space targets.

The exact split may vary between 30 and 40 when discovery deduplicates known targets or a target becomes permanently inaccessible. Continue only until the recorded cohort reaches the agreed bound. Do not inflate discovery simply to hit an exact aesthetic number.

Positive-Finding Requirement

A random public cohort may correctly produce zero findings. Zero findings cannot validate the finding-storage and found_secrets.jsonl path.

Include at least one separately identified, non-live controlled fixture that is expected to trigger an already approved detector. The fixture must contain no usable credential. Record its expected detector and identity before scanning. Do not weaken verification, introduce a new detector, or publish a real secret merely to force a positive result.

If no approved positive fixture is available, report the finding path as unverified by this run even if all zero-finding scans succeed.

Phase 1: Baseline

With runtime healthy, record a secret-safe baseline:

  • active config SHA-256 and semantic config SHA-256;
  • runtime-control revision and open/paused/drain state;
  • enabled source set and active worker package manifests;
  • remote worker/device count, recent contact, and package capability match;
  • high-water IDs for target_queue, result_reservations, target_scans, findings, errors, result_bundles, and projection_jobs;
  • queue counts by source and status;
  • active reservation count and oldest age;
  • pipeline worker readiness and capacity counters;
  • pending/leased/quarantined bundle, projection, and keycheck counts;
  • current projection stream/cursor identity;
  • byte size and final complete-line identity of active JSONL files.

The baseline collector must print aggregates and hashes only. It must not emit targets, assignment payloads, tokens, raw findings, or runtime documents.

Phase 2: Apply the Test Candidate

  1. Pause discovery.
  2. Pause dispatch.
  3. Start drain and wait for blocker count zero and drained.
  4. Preview the bounded candidate and review the semantic diff.
  5. Save and apply the candidate through the host-agent lifecycle.
  6. Require reconciled succeeded, no failed hold, strict runtime health, edge health, admin health, and Worker API health.
  7. Cancel drain, then resume discovery and dispatch in that order.

Do not continue if the lifecycle operation rolls back or enters failed hold.

Phase 3: Build and Freeze the Cohort

Record the baseline target_queue.id high-water mark. Let the bounded sources cycle until 30-40 new eligible queue rows have been created after that mark.

Then:

  1. Pause discovery so the cohort cannot grow.
  2. Leave dispatch open until the selected queue rows settle.
  3. Record cohort queue IDs and only their safe identities: source, normalized target hash, query hash, immutable planning kind, and creation order.
  4. Separate deduplicated, permanently inaccessible, deferred, retried, and actually assigned items. Do not count an API result as a scan.

The authoritative cohort is a fixed set of queue IDs, not "whatever completed during the same hour."

Phase 4: Observe Remote Execution

For every cohort queue ID, verify:

  • no more than one current authoritative reservation;
  • assignment package/platform capability matches the registered worker;
  • lease token and execution snapshot are bound but never printed;
  • Docker targets are immutable repo@sha256 identities;
  • GitLab targets have the intended exact planning/ref identity;
  • HuggingFace targets use the direct Space execution kind;
  • terminal report classification is success, permanent target failure, or retryable provider failure as designed;
  • retries preserve queue identity and increment attempts without creating a second authoritative acceptance;
  • accepted receipt replay returns the same durable result.

Physical work can repeat after a lease expiry or network partition. Correctness means fencing permits one authoritative acceptance and one queue completion, not that duplicate physical execution is impossible.

Phase 5: Validate Bundles and PostgreSQL Structure

For each accepted result, validate without dumping body content:

  • bundle exists at the registered private relative path;
  • bundle byte count and SHA-256 match database metadata;
  • bundle schema/version, event ID/hash, reservation ID, queue ID, source, normalized target identity, execution snapshot identity, and scan policy are internally consistent;
  • result is ingested exactly once;
  • target_queue.target_scan_id references the corresponding target_scans.id;
  • queue completion is applied once with a terminal disposition;
  • target_scans.queue_id and claim lease identity refer back to the cohort row;
  • target_scans.findings_count and error_count equal actual child-row counts;
  • every finding/error references the same target scan, source, cycle, and run;
  • no legacy raw_result_json, publication outbox row, or orphan relation is introduced;
  • ingested bundle credit and pipeline-capacity counters are released exactly according to the durable state machine;
  • no cohort item enters pipeline_quarantine.

Aggregate checks must cover the entire cohort. Additionally inspect a small redacted structural sample from every source and every terminal disposition.

Phase 6: Validate Findings

For every finding in the cohort, inspect structure only:

  • stable finding_uid and finding fingerprint;
  • detector name/type and verification flag;
  • source, target hash, file path, line/commit/source timestamp where applicable;
  • redacted secret and secret/detector hashes;
  • provider and credential-kind enrichment;
  • required-context and raw-payload-omitted flags;
  • bounded source metadata and enrichment JSON decode successfully;
  • no unexpected raw-secret exposure in logs, queue rows, assignment metadata, admin list views, or compatibility scan summaries.

For the controlled positive fixture, require the expected finding to exist in PostgreSQL and to project once to found_secrets.jsonl.

Phase 7: Validate Windows-Compatible Files

Use the active configured global.results_dir. The relevant compatibility outputs are:

  • scan_results.jsonl for one sanitized scan event per projected scan;
  • found_secrets.jsonl for projected finding events;
  • their publication ledgers, active stream metadata, and rotated segments;
  • per-service keycheck result files only if keychecks run for the finding.

For the cohort, verify:

  • every required projection_job reaches completed;
  • projector cursor and append ledger advance monotonically;
  • each cohort scan event appears exactly once by scan_event_id;
  • each cohort finding appears exactly once by finding_uid;
  • JSON lines parse and match the current compatibility schema;
  • scan summaries match PostgreSQL counts and terminal status;
  • finding projections are redacted as designed and preserve safe provenance;
  • active and rotated segments together contain the events; checking only the active file is insufficient when rotation occurs;
  • no torn-tail quarantine, duplicate append, skipped cursor, or unpublished completed job exists.

These files should have the same logical structure as the Windows deployment, but path separators and host/container root paths are platform-specific.

Phase 8: Queue and Pipeline Closure

After all cohort rows settle, require:

  • 30-40 cohort queue rows accounted for by terminal, deferred, or explicitly classified retry state;
  • no expired active reservation remains unreaped;
  • no queue row has multiple authoritative accepted results;
  • no accepted result remains un-ingested beyond the bounded pipeline window;
  • no completed scan remains unprojected beyond the bounded projector window;
  • no stale pipeline lease or capacity leak;
  • no unexpected quarantine;
  • runtime, Worker API, edge, host-agent, host Caddy, and unrelated host service health remain good.

The final report must show counts for discovered, deduplicated, assigned, retried, accepted, ingested, projected, succeeded, skipped/permanent, retryable/deferred, findings, errors, and quarantines.

Phase 9: Restore Production Configuration

  1. Pause discovery and dispatch.
  2. Drain to zero blockers.
  3. Apply the exact original runtime document through the normal lifecycle.
  4. Require byte-identical original config SHA-256, reconciled lifecycle success, strict health, and no failed hold.
  5. Cancel drain, resume discovery, then resume dispatch according to the original control state.
  6. Confirm worker contact and normal post-test assignment flow.

Do not restore by manually editing YAML or replacing files behind the host-agent.

Pass Criteria

The run passes only when:

  • at least 30 and at most 40 fixed-cohort rows are fully accounted for;
  • all accepted cohort results ingest exactly once;
  • all required cohort projections complete exactly once;
  • queue/reservation/scan/finding/error/bundle relationships are consistent;
  • the controlled positive finding reaches PostgreSQL and found_secrets.jsonl, or the report explicitly marks positive-finding validation incomplete because no approved fixture existed;
  • no unexplained retry, orphan, duplicate acceptance, capacity leak, quarantine, failed hold, or projection gap remains;
  • the original production config is restored exactly and services are healthy.

Any failure must retain its operation IDs, queue IDs, reservation IDs, hashes, safe categories, and aggregate evidence for diagnosis. A partial pass must not be reported as "100% scanner correctness."

Evidence Report

Append or link a dated report containing:

  • environment and worker package identities;
  • original/test/restored config hashes;
  • cohort definition and aggregate source split;
  • lifecycle operation IDs for test apply and restore;
  • queue and pipeline baseline/final aggregates;
  • per-stage reconciliation counts;
  • redacted structural examples for a scan, an error/skip, and a finding;
  • JSONL/ledger reconciliation counts;
  • deviations, retries, quarantines, and unresolved questions;
  • final verdict with explicit tested and untested boundaries.