## ADDED Requirements ### Requirement: Postman source registration The system SHALL provide a first-class `postman` source that can be configured, selected, supervised, queued, scanned, and displayed consistently with existing scanner sources. #### Scenario: Configured Postman source is selectable - **WHEN** a config contains `sources.postman.enabled: true` and the runner is invoked with `--source postman` - **THEN** the runner SHALL execute only the Postman source cycle using Postman source settings #### Scenario: Postman queue files are used - **WHEN** the Postman source prepares targets - **THEN** the system SHALL use `todo_postman.txt` and `checked_postman.txt` under the configured queue directory ### Requirement: GitHub code search discovery The Postman source SHALL discover public Postman artifacts from GitHub code search using configured queries and artifact kinds. #### Scenario: Collection files are discovered - **WHEN** `search_kinds` includes `collection` and the query is `openai` - **THEN** discovery SHALL search GitHub code for `filename:postman_collection.json openai` #### Scenario: Environment files are discovered - **WHEN** `search_kinds` includes `environment` and the query is `openai` - **THEN** discovery SHALL search GitHub code for `filename:postman_environment.json openai` #### Scenario: Search pagination is bounded - **WHEN** `pages` is `10` and `per_page` is `100` - **THEN** discovery SHALL request no more than 1000 search results per query and kind ### Requirement: GitHub auth pool rotation The Postman GitHub discovery flow SHALL use all available tokens from the configured GitHub auth pool before sleeping for rate limits. #### Scenario: Token rotates per GitHub request - **WHEN** multiple GitHub auth entries are available - **THEN** GitHub code search, commit lookup, and content download requests SHALL rotate across available tokens #### Scenario: One token is rate-limited - **WHEN** a GitHub request returns a primary or secondary rate limit for the current token - **THEN** the system SHALL mark only that token unavailable until its reset time or configured cooldown and continue with another available token #### Scenario: All tokens are unavailable - **WHEN** every configured GitHub token is rate-limited or temporarily unavailable - **THEN** the source SHALL sleep until the earliest known reset time, or for the configured fallback cooldown when no reset time is known #### Scenario: Token is invalid - **WHEN** a GitHub request returns an authentication-invalid response for a token - **THEN** the system SHALL exclude that token from the current cycle and report the authentication failure without marking other tokens invalid ### Requirement: Backfill freshness filtering The Postman source SHALL support a backfill mode that can scan deep code search pages while skipping GitHub artifacts older than a configured file age. #### Scenario: Recent file is queued - **WHEN** a GitHub code search result has a latest path commit within `max_file_age_days` - **THEN** the target SHALL be eligible for queueing #### Scenario: Old file is skipped - **WHEN** a GitHub code search result has a latest path commit older than `max_file_age_days` - **THEN** the target SHALL not be queued and SHALL be counted as skipped by freshness filtering #### Scenario: Freshness filtering is disabled - **WHEN** `max_file_age_days` is `0` - **THEN** the source SHALL not perform path commit age filtering ### Requirement: Tail mode early stop The Postman source SHALL support ongoing tail scans that stop pagination after consecutive known pages. #### Scenario: Known page increments stop counter - **WHEN** `stop_on_seen_pages` is enabled and every normalized target on a fetched page already exists in `todo_postman.txt` or `checked_postman.txt` - **THEN** the source SHALL count that page as known #### Scenario: Tail pagination stops - **WHEN** the known page count reaches `seen_page_threshold` after `min_pages_before_stop` - **THEN** the source SHALL stop fetching additional pages for that query and artifact kind ### Requirement: Postman target identity and deduplication The system SHALL normalize Postman targets so identical artifacts are not rescanned while changed artifacts are scanned again. #### Scenario: GitHub target identity includes SHA - **WHEN** a Postman target is discovered from GitHub code search - **THEN** its normalized target SHALL include source, repository, path, and file SHA #### Scenario: GitHub file changes - **WHEN** the same GitHub repository and path is discovered with a new SHA - **THEN** the system SHALL treat it as a new Postman target #### Scenario: Package target identity uses content hash - **WHEN** a Postman artifact is harvested from npm or PyPI - **THEN** its normalized target SHALL include the artifact content SHA-256 hash ### Requirement: Durable Postman artifact cache The system SHALL store discovered Postman artifact content in durable runtime cache before scanning. #### Scenario: GitHub content is cached - **WHEN** a GitHub code search target is queued for scanning - **THEN** the system SHALL download the artifact content and store it under the configured Postman cache directory #### Scenario: Package content is cached before cleanup - **WHEN** npm or PyPI extraction finds a Postman artifact - **THEN** the system SHALL copy the artifact into the durable Postman cache before the extraction directory is removed #### Scenario: Cache size is constrained - **WHEN** an artifact exceeds the configured maximum Postman artifact size - **THEN** the system SHALL skip the artifact and record a bounded error or skip reason ### Requirement: Postman artifact scanning The Postman source SHALL scan cached Postman collection and environment artifacts with TruffleHog filesystem scanning. #### Scenario: Cached artifact is scanned - **WHEN** a Postman target points to a cached JSON artifact - **THEN** the scanner SHALL run TruffleHog against a temporary filesystem directory containing that artifact #### Scenario: Findings are persisted - **WHEN** TruffleHog reports findings for a Postman target - **THEN** the system SHALL persist findings to existing JSONL outputs and scanner database tables with source `postman` #### Scenario: Scan finishes - **WHEN** a Postman target scan completes with findings, errors, skipped status, or clean status - **THEN** the target SHALL be moved from `todo_postman.txt` to `checked_postman.txt` ### Requirement: npm and PyPI Postman harvesting The npm and PyPI source flows SHALL harvest Postman artifacts discovered during existing package extraction and enqueue them for the Postman source. #### Scenario: npm package contains collection - **WHEN** an extracted npm package contains a file matching `*.postman_collection.json` - **THEN** the system SHALL cache the file and enqueue a Postman target with npm package origin metadata #### Scenario: PyPI package contains environment - **WHEN** an extracted PyPI artifact contains a file matching `*.postman_environment.json` - **THEN** the system SHALL cache the file and enqueue a Postman target with PyPI package origin metadata #### Scenario: Existing package scan continues - **WHEN** Postman harvesting fails for one package artifact - **THEN** the original npm or PyPI scan SHALL still complete and record the harvesting failure without failing unrelated package scanning ### Requirement: Postman-aware enrichment The system SHALL enrich Postman findings with contextual classification derived from Postman structure without replacing TruffleHog detection. #### Scenario: Header credential is classified - **WHEN** a finding appears in a Postman request header such as `Authorization` or `x-api-key` - **THEN** enrichment SHALL record the context location and infer credential kind from header type, value shape, and endpoint host when possible #### Scenario: Environment variable is classified - **WHEN** a finding appears in a Postman environment variable value - **THEN** enrichment SHALL record the variable name and classify provider or credential kind when supported by value shape or associated request endpoints #### Scenario: Placeholder is detected - **WHEN** a Postman value is a placeholder such as `{{API_KEY}}`, ``, `YOUR_API_KEY`, `example`, or `changeme` - **THEN** enrichment SHALL classify it as placeholder or low confidence rather than a live secret ### Requirement: Observability for Postman source The system SHALL expose Postman source activity through existing logs, queue counts, source cycle metrics, target scan records, findings, errors, and dashboard views. #### Scenario: Source cycle is recorded - **WHEN** a Postman source cycle runs - **THEN** the scanner database SHALL record source cycle metrics including fetched, queued, scanned, found, error, skipped, and queue counts #### Scenario: Dashboard shows Postman queues - **WHEN** Postman queue files exist - **THEN** the dashboard SHALL include Postman queue counts in the current queues view