Files
truf-server/openspec/changes/add-postman-source/specs/postman-source/spec.md
T
2026-09-30 20:30:56 +03:00

159 lines
8.8 KiB
Markdown

## ADDED Requirements
### Requirement: Postman source registration
The system SHALL provide a first-class `postman` source that can be configured, selected, supervised, queued, scanned, and displayed consistently with existing scanner sources.
#### Scenario: Configured Postman source is selectable
- **WHEN** a config contains `sources.postman.enabled: true` and the runner is invoked with `--source postman`
- **THEN** the runner SHALL execute only the Postman source cycle using Postman source settings
#### Scenario: Postman queue files are used
- **WHEN** the Postman source prepares targets
- **THEN** the system SHALL use `todo_postman.txt` and `checked_postman.txt` under the configured queue directory
### Requirement: GitHub code search discovery
The Postman source SHALL discover public Postman artifacts from GitHub code search using configured queries and artifact kinds.
#### Scenario: Collection files are discovered
- **WHEN** `search_kinds` includes `collection` and the query is `openai`
- **THEN** discovery SHALL search GitHub code for `filename:postman_collection.json openai`
#### Scenario: Environment files are discovered
- **WHEN** `search_kinds` includes `environment` and the query is `openai`
- **THEN** discovery SHALL search GitHub code for `filename:postman_environment.json openai`
#### Scenario: Search pagination is bounded
- **WHEN** `pages` is `10` and `per_page` is `100`
- **THEN** discovery SHALL request no more than 1000 search results per query and kind
### Requirement: GitHub auth pool rotation
The Postman GitHub discovery flow SHALL use all available tokens from the configured GitHub auth pool before sleeping for rate limits.
#### Scenario: Token rotates per GitHub request
- **WHEN** multiple GitHub auth entries are available
- **THEN** GitHub code search, commit lookup, and content download requests SHALL rotate across available tokens
#### Scenario: One token is rate-limited
- **WHEN** a GitHub request returns a primary or secondary rate limit for the current token
- **THEN** the system SHALL mark only that token unavailable until its reset time or configured cooldown and continue with another available token
#### Scenario: All tokens are unavailable
- **WHEN** every configured GitHub token is rate-limited or temporarily unavailable
- **THEN** the source SHALL sleep until the earliest known reset time, or for the configured fallback cooldown when no reset time is known
#### Scenario: Token is invalid
- **WHEN** a GitHub request returns an authentication-invalid response for a token
- **THEN** the system SHALL exclude that token from the current cycle and report the authentication failure without marking other tokens invalid
### Requirement: Backfill freshness filtering
The Postman source SHALL support a backfill mode that can scan deep code search pages while skipping GitHub artifacts older than a configured file age.
#### Scenario: Recent file is queued
- **WHEN** a GitHub code search result has a latest path commit within `max_file_age_days`
- **THEN** the target SHALL be eligible for queueing
#### Scenario: Old file is skipped
- **WHEN** a GitHub code search result has a latest path commit older than `max_file_age_days`
- **THEN** the target SHALL not be queued and SHALL be counted as skipped by freshness filtering
#### Scenario: Freshness filtering is disabled
- **WHEN** `max_file_age_days` is `0`
- **THEN** the source SHALL not perform path commit age filtering
### Requirement: Tail mode early stop
The Postman source SHALL support ongoing tail scans that stop pagination after consecutive known pages.
#### Scenario: Known page increments stop counter
- **WHEN** `stop_on_seen_pages` is enabled and every normalized target on a fetched page already exists in `todo_postman.txt` or `checked_postman.txt`
- **THEN** the source SHALL count that page as known
#### Scenario: Tail pagination stops
- **WHEN** the known page count reaches `seen_page_threshold` after `min_pages_before_stop`
- **THEN** the source SHALL stop fetching additional pages for that query and artifact kind
### Requirement: Postman target identity and deduplication
The system SHALL normalize Postman targets so identical artifacts are not rescanned while changed artifacts are scanned again.
#### Scenario: GitHub target identity includes SHA
- **WHEN** a Postman target is discovered from GitHub code search
- **THEN** its normalized target SHALL include source, repository, path, and file SHA
#### Scenario: GitHub file changes
- **WHEN** the same GitHub repository and path is discovered with a new SHA
- **THEN** the system SHALL treat it as a new Postman target
#### Scenario: Package target identity uses content hash
- **WHEN** a Postman artifact is harvested from npm or PyPI
- **THEN** its normalized target SHALL include the artifact content SHA-256 hash
### Requirement: Durable Postman artifact cache
The system SHALL store discovered Postman artifact content in durable runtime cache before scanning.
#### Scenario: GitHub content is cached
- **WHEN** a GitHub code search target is queued for scanning
- **THEN** the system SHALL download the artifact content and store it under the configured Postman cache directory
#### Scenario: Package content is cached before cleanup
- **WHEN** npm or PyPI extraction finds a Postman artifact
- **THEN** the system SHALL copy the artifact into the durable Postman cache before the extraction directory is removed
#### Scenario: Cache size is constrained
- **WHEN** an artifact exceeds the configured maximum Postman artifact size
- **THEN** the system SHALL skip the artifact and record a bounded error or skip reason
### Requirement: Postman artifact scanning
The Postman source SHALL scan cached Postman collection and environment artifacts with TruffleHog filesystem scanning.
#### Scenario: Cached artifact is scanned
- **WHEN** a Postman target points to a cached JSON artifact
- **THEN** the scanner SHALL run TruffleHog against a temporary filesystem directory containing that artifact
#### Scenario: Findings are persisted
- **WHEN** TruffleHog reports findings for a Postman target
- **THEN** the system SHALL persist findings to existing JSONL outputs and scanner database tables with source `postman`
#### Scenario: Scan finishes
- **WHEN** a Postman target scan completes with findings, errors, skipped status, or clean status
- **THEN** the target SHALL be moved from `todo_postman.txt` to `checked_postman.txt`
### Requirement: npm and PyPI Postman harvesting
The npm and PyPI source flows SHALL harvest Postman artifacts discovered during existing package extraction and enqueue them for the Postman source.
#### Scenario: npm package contains collection
- **WHEN** an extracted npm package contains a file matching `*.postman_collection.json`
- **THEN** the system SHALL cache the file and enqueue a Postman target with npm package origin metadata
#### Scenario: PyPI package contains environment
- **WHEN** an extracted PyPI artifact contains a file matching `*.postman_environment.json`
- **THEN** the system SHALL cache the file and enqueue a Postman target with PyPI package origin metadata
#### Scenario: Existing package scan continues
- **WHEN** Postman harvesting fails for one package artifact
- **THEN** the original npm or PyPI scan SHALL still complete and record the harvesting failure without failing unrelated package scanning
### Requirement: Postman-aware enrichment
The system SHALL enrich Postman findings with contextual classification derived from Postman structure without replacing TruffleHog detection.
#### Scenario: Header credential is classified
- **WHEN** a finding appears in a Postman request header such as `Authorization` or `x-api-key`
- **THEN** enrichment SHALL record the context location and infer credential kind from header type, value shape, and endpoint host when possible
#### Scenario: Environment variable is classified
- **WHEN** a finding appears in a Postman environment variable value
- **THEN** enrichment SHALL record the variable name and classify provider or credential kind when supported by value shape or associated request endpoints
#### Scenario: Placeholder is detected
- **WHEN** a Postman value is a placeholder such as `{{API_KEY}}`, `<api_key>`, `YOUR_API_KEY`, `example`, or `changeme`
- **THEN** enrichment SHALL classify it as placeholder or low confidence rather than a live secret
### Requirement: Observability for Postman source
The system SHALL expose Postman source activity through existing logs, queue counts, source cycle metrics, target scan records, findings, errors, and dashboard views.
#### Scenario: Source cycle is recorded
- **WHEN** a Postman source cycle runs
- **THEN** the scanner database SHALL record source cycle metrics including fetched, queued, scanned, found, error, skipped, and queue counts
#### Scenario: Dashboard shows Postman queues
- **WHEN** Postman queue files exist
- **THEN** the dashboard SHALL include Postman queue counts in the current queues view