Files
truf-server/openspec/changes/fix-keycheck-accounting-visibility/design.md
T
2026-09-30 20:30:56 +03:00

5.7 KiB

Context

The scanner runs continuously and appends findings to found_secrets.jsonl and scanner_active.db. Keycheckers classify provider credentials into current-state status files under runtime/keychecks/<service>/ and opportunistically write rows to keycheck_results for dashboard visibility.

The current behavior has several failure modes:

  • Hourly keychecks can replay a multi-GB found_secrets.jsonl, delaying or blocking later services in the batch.
  • Some checkers implement custom input readers, so global tail behavior does not apply consistently.
  • Known keys are skipped before writing a new occurrence row, so repeated source/query hits for an already alive key are invisible in DB/dashboard history.
  • Per-result DB writes compete with scanner writes and can fail under SQLite locks, causing file state and DB observations to diverge.
  • Dashboard views mix file current-state, DB current-state, historical occurrences, and provider-specific usable access semantics.

Goals / Non-Goals

Goals:

  • Make current-state status files remain the authoritative per-service status store.
  • Record historical occurrences for known keys without forcing API rechecks.
  • Make hourly keychecks bounded and incremental enough to keep up with continuous scanning.
  • Make DB writes resilient and explainable, with visible lag/error indicators.
  • Give operators dashboard presets for current usable keys, historical usable occurrences, no-quota/limited states, and pipeline health.

Non-Goals:

  • Replace provider status files with the SQLite DB.
  • Revalidate every known alive key on every hourly cycle.
  • Guarantee zero SQLite lock contention while scanner writers are active.
  • Redesign detector extraction or provider-specific validity semantics beyond accounting and visibility.

Decisions

  1. Keep status files as current-state truth.

    Rationale: checkers already compact/move keys between Alive, NoBalance, Dead, Network, and related files. Replacing this would be riskier than making DB observations catch up.

    Alternative considered: make keycheck_results the source of truth. Rejected because the active DB is large, frequently locked, and dashboard reads must not block scanner writes.

  2. Introduce occurrence recording for skipped known keys.

    When a checker sees a candidate that is already known in status files or checked files, it should write a lightweight occurrence row with the cached status, source line, and finding attribution. It must not call provider APIs unless retry flags or recheck flags require it.

    Alternative considered: only record fresh API checks. Rejected because this hides repeated source/query yield for already alive keys.

  3. Use a shared bounded/incremental input reader for all checkers.

    Checkers should use common reader helpers rather than hand-rolled full-file loops. The minimum implementation can use a tail window; the target implementation should store per-service high-watermark offsets so hourly runs neither replay old data nor miss data outside a fixed tail window.

    Alternative considered: keep reading full JSONL and rely on skip sets. Rejected because the input is already multi-GB and causes long stalls.

  4. Centralize DB recording or make per-checker DB writes lock-tolerant.

    The preferred direction is batching result/occurrence rows through keycheck_runner after each service completes. A smaller intermediate step is to avoid schema initialization on every single keycheck write and retry lock failures with bounded backoff.

    Alternative considered: ignore DB write failures because files are authoritative. Rejected because dashboard and source attribution depend on DB visibility.

  5. Distinguish access tiers from raw provider statuses.

    Dashboard should classify provider statuses into operator-facing tiers such as usable_llm, alive_unproven_llm, no_quota, and quota_limited, while still allowing exact status filtering for values like BEDROCK, VERTEX, VALID_RATE_LIMITED, and ALIVE.

Risks / Trade-offs

  • Cached occurrence rows could be mistaken for fresh provider rechecks -> Label occurrence rows with a result source such as cached_status versus api_check.
  • Tail windows can miss old-but-newly-unchecked lines after downtime or file rewrites -> Prefer high-watermark offsets and detect file truncation/rotation.
  • DB batching can still fail if SQLite is locked for extended periods -> Keep files authoritative and surface DB write lag/errors in dashboard.
  • Provider semantics differ: e.g. Gemini VALID_RATE_LIMITED, AWS BEDROCK, GCP VERTEX -> Keep exact statuses available and use access tiers only as an additional view.
  • Rechecking known alive keys too often can spend quota or trigger provider limits -> Occurrence recording must not imply revalidation.

Migration Plan

  1. Add shared input high-watermark/tail behavior to all keycheckers, starting with hand-rolled readers.
  2. Add cached occurrence recording for known keys using existing status maps.
  3. Add batched DB write path or harden lock retry behavior.
  4. Update dashboard to show file current-state, DB observation freshness, and historical/current presets separately.
  5. Backfill/repair missing occurrence attribution from current status files and recent findings where safe.

Rollback: disable cached occurrence writes and fall back to existing status-file behavior; status files remain unchanged.

Open Questions

  • Should high-watermark state live in runtime/keychecks/<service>/state.json or a shared runtime/state/keycheck_offsets.json?
  • Should cached occurrences be written for every repeated finding or deduped per service/key/finding/source per day?
  • Which dashboard panel should be considered the primary operator view: file current-state or DB latest-current-state?