## ADDED Requirements ### Requirement: Current-state status files remain authoritative The system SHALL keep per-service keycheck status files as the authoritative current-state classification for keys. #### Scenario: Key status changes after recheck - **WHEN** a checker revalidates a key and receives a new status - **THEN** the key MUST be removed from other status files for that service and written to the status file for the new status #### Scenario: Dashboard compares files and DB - **WHEN** dashboard displays keycheck totals - **THEN** it MUST be clear whether each count comes from current-state status files or from DB observation rows ### Requirement: Known key occurrences are recorded The system SHALL record an occurrence when a checker sees a candidate that is already known in service status files or checked files. #### Scenario: Known alive key appears in a new finding - **WHEN** a key already classified as alive appears in a new scanner finding - **THEN** the system MUST record the new source/query/finding occurrence without requiring a provider API recheck #### Scenario: Known dead key appears in a new finding - **WHEN** a key already classified as dead appears in a new scanner finding - **THEN** the system MUST record the new source/query/finding occurrence with a cached dead status #### Scenario: Cached occurrence is distinguishable from API recheck - **WHEN** an occurrence row is written without calling the provider API - **THEN** the row MUST indicate that the status came from cached current-state classification ### Requirement: Keycheck input processing is bounded and consistent The system SHALL avoid replaying the full scanner JSONL input on every hourly keycheck run. #### Scenario: Hourly keychecks run on a multi-GB input file - **WHEN** `found_secrets.jsonl` is large - **THEN** each checker MUST process only a bounded recent range or an incremental range since its last processed offset #### Scenario: Checker has a custom input loop - **WHEN** a checker reads scanner findings - **THEN** it MUST use shared keycheck input-reading behavior or implement equivalent high-watermark/tail semantics #### Scenario: Input file rotates or shrinks - **WHEN** a stored high-watermark offset is larger than the current input file size - **THEN** the system MUST reset the offset safely and continue processing without crashing ### Requirement: Keycheck DB observation writes are resilient The system SHALL make keycheck DB observation writes resilient to active scanner DB contention. #### Scenario: SQLite database is temporarily locked - **WHEN** a keycheck result or occurrence is ready to record and SQLite is locked - **THEN** the system MUST retry with bounded backoff before reporting a DB write failure #### Scenario: DB write fails after retries - **WHEN** all DB write retries fail - **THEN** the status file write MUST remain intact and the failure MUST be visible in logs or dashboard health #### Scenario: Schema initialization would contend with active writers - **WHEN** a checker records a single result row - **THEN** it MUST NOT run schema initialization or migration DDL as part of that per-result write path ### Requirement: Dashboard exposes keycheck pipeline health The dashboard SHALL expose keycheck pipeline health and freshness separately from provider status counts. #### Scenario: DB observations lag behind status files - **WHEN** status files are newer than the latest DB keycheck row - **THEN** dashboard MUST show that DB observation data is stale relative to file current-state #### Scenario: Keycheck run is stuck on a service - **WHEN** the keychecks process has not advanced past a service for longer than expected - **THEN** dashboard or supervisor-visible status MUST make the stuck service and elapsed time visible #### Scenario: Operator wants current usable keys - **WHEN** an operator selects current usable key view - **THEN** dashboard MUST use provider-specific access tiers while retaining exact status filters such as `BEDROCK`, `VERTEX`, `ALIVE`, and `VALID_RATE_LIMITED` #### Scenario: Operator wants historical source yield - **WHEN** an operator selects historical yield view - **THEN** dashboard MUST include cached known-key occurrences so source/query yield is not lost after rechecks or known-key skips