Initial server source import
This commit is contained in:
@@ -0,0 +1,158 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Postman source registration
|
||||
The system SHALL provide a first-class `postman` source that can be configured, selected, supervised, queued, scanned, and displayed consistently with existing scanner sources.
|
||||
|
||||
#### Scenario: Configured Postman source is selectable
|
||||
- **WHEN** a config contains `sources.postman.enabled: true` and the runner is invoked with `--source postman`
|
||||
- **THEN** the runner SHALL execute only the Postman source cycle using Postman source settings
|
||||
|
||||
#### Scenario: Postman queue files are used
|
||||
- **WHEN** the Postman source prepares targets
|
||||
- **THEN** the system SHALL use `todo_postman.txt` and `checked_postman.txt` under the configured queue directory
|
||||
|
||||
### Requirement: GitHub code search discovery
|
||||
The Postman source SHALL discover public Postman artifacts from GitHub code search using configured queries and artifact kinds.
|
||||
|
||||
#### Scenario: Collection files are discovered
|
||||
- **WHEN** `search_kinds` includes `collection` and the query is `openai`
|
||||
- **THEN** discovery SHALL search GitHub code for `filename:postman_collection.json openai`
|
||||
|
||||
#### Scenario: Environment files are discovered
|
||||
- **WHEN** `search_kinds` includes `environment` and the query is `openai`
|
||||
- **THEN** discovery SHALL search GitHub code for `filename:postman_environment.json openai`
|
||||
|
||||
#### Scenario: Search pagination is bounded
|
||||
- **WHEN** `pages` is `10` and `per_page` is `100`
|
||||
- **THEN** discovery SHALL request no more than 1000 search results per query and kind
|
||||
|
||||
### Requirement: GitHub auth pool rotation
|
||||
The Postman GitHub discovery flow SHALL use all available tokens from the configured GitHub auth pool before sleeping for rate limits.
|
||||
|
||||
#### Scenario: Token rotates per GitHub request
|
||||
- **WHEN** multiple GitHub auth entries are available
|
||||
- **THEN** GitHub code search, commit lookup, and content download requests SHALL rotate across available tokens
|
||||
|
||||
#### Scenario: One token is rate-limited
|
||||
- **WHEN** a GitHub request returns a primary or secondary rate limit for the current token
|
||||
- **THEN** the system SHALL mark only that token unavailable until its reset time or configured cooldown and continue with another available token
|
||||
|
||||
#### Scenario: All tokens are unavailable
|
||||
- **WHEN** every configured GitHub token is rate-limited or temporarily unavailable
|
||||
- **THEN** the source SHALL sleep until the earliest known reset time, or for the configured fallback cooldown when no reset time is known
|
||||
|
||||
#### Scenario: Token is invalid
|
||||
- **WHEN** a GitHub request returns an authentication-invalid response for a token
|
||||
- **THEN** the system SHALL exclude that token from the current cycle and report the authentication failure without marking other tokens invalid
|
||||
|
||||
### Requirement: Backfill freshness filtering
|
||||
The Postman source SHALL support a backfill mode that can scan deep code search pages while skipping GitHub artifacts older than a configured file age.
|
||||
|
||||
#### Scenario: Recent file is queued
|
||||
- **WHEN** a GitHub code search result has a latest path commit within `max_file_age_days`
|
||||
- **THEN** the target SHALL be eligible for queueing
|
||||
|
||||
#### Scenario: Old file is skipped
|
||||
- **WHEN** a GitHub code search result has a latest path commit older than `max_file_age_days`
|
||||
- **THEN** the target SHALL not be queued and SHALL be counted as skipped by freshness filtering
|
||||
|
||||
#### Scenario: Freshness filtering is disabled
|
||||
- **WHEN** `max_file_age_days` is `0`
|
||||
- **THEN** the source SHALL not perform path commit age filtering
|
||||
|
||||
### Requirement: Tail mode early stop
|
||||
The Postman source SHALL support ongoing tail scans that stop pagination after consecutive known pages.
|
||||
|
||||
#### Scenario: Known page increments stop counter
|
||||
- **WHEN** `stop_on_seen_pages` is enabled and every normalized target on a fetched page already exists in `todo_postman.txt` or `checked_postman.txt`
|
||||
- **THEN** the source SHALL count that page as known
|
||||
|
||||
#### Scenario: Tail pagination stops
|
||||
- **WHEN** the known page count reaches `seen_page_threshold` after `min_pages_before_stop`
|
||||
- **THEN** the source SHALL stop fetching additional pages for that query and artifact kind
|
||||
|
||||
### Requirement: Postman target identity and deduplication
|
||||
The system SHALL normalize Postman targets so identical artifacts are not rescanned while changed artifacts are scanned again.
|
||||
|
||||
#### Scenario: GitHub target identity includes SHA
|
||||
- **WHEN** a Postman target is discovered from GitHub code search
|
||||
- **THEN** its normalized target SHALL include source, repository, path, and file SHA
|
||||
|
||||
#### Scenario: GitHub file changes
|
||||
- **WHEN** the same GitHub repository and path is discovered with a new SHA
|
||||
- **THEN** the system SHALL treat it as a new Postman target
|
||||
|
||||
#### Scenario: Package target identity uses content hash
|
||||
- **WHEN** a Postman artifact is harvested from npm or PyPI
|
||||
- **THEN** its normalized target SHALL include the artifact content SHA-256 hash
|
||||
|
||||
### Requirement: Durable Postman artifact cache
|
||||
The system SHALL store discovered Postman artifact content in durable runtime cache before scanning.
|
||||
|
||||
#### Scenario: GitHub content is cached
|
||||
- **WHEN** a GitHub code search target is queued for scanning
|
||||
- **THEN** the system SHALL download the artifact content and store it under the configured Postman cache directory
|
||||
|
||||
#### Scenario: Package content is cached before cleanup
|
||||
- **WHEN** npm or PyPI extraction finds a Postman artifact
|
||||
- **THEN** the system SHALL copy the artifact into the durable Postman cache before the extraction directory is removed
|
||||
|
||||
#### Scenario: Cache size is constrained
|
||||
- **WHEN** an artifact exceeds the configured maximum Postman artifact size
|
||||
- **THEN** the system SHALL skip the artifact and record a bounded error or skip reason
|
||||
|
||||
### Requirement: Postman artifact scanning
|
||||
The Postman source SHALL scan cached Postman collection and environment artifacts with TruffleHog filesystem scanning.
|
||||
|
||||
#### Scenario: Cached artifact is scanned
|
||||
- **WHEN** a Postman target points to a cached JSON artifact
|
||||
- **THEN** the scanner SHALL run TruffleHog against a temporary filesystem directory containing that artifact
|
||||
|
||||
#### Scenario: Findings are persisted
|
||||
- **WHEN** TruffleHog reports findings for a Postman target
|
||||
- **THEN** the system SHALL persist findings to existing JSONL outputs and scanner database tables with source `postman`
|
||||
|
||||
#### Scenario: Scan finishes
|
||||
- **WHEN** a Postman target scan completes with findings, errors, skipped status, or clean status
|
||||
- **THEN** the target SHALL be moved from `todo_postman.txt` to `checked_postman.txt`
|
||||
|
||||
### Requirement: npm and PyPI Postman harvesting
|
||||
The npm and PyPI source flows SHALL harvest Postman artifacts discovered during existing package extraction and enqueue them for the Postman source.
|
||||
|
||||
#### Scenario: npm package contains collection
|
||||
- **WHEN** an extracted npm package contains a file matching `*.postman_collection.json`
|
||||
- **THEN** the system SHALL cache the file and enqueue a Postman target with npm package origin metadata
|
||||
|
||||
#### Scenario: PyPI package contains environment
|
||||
- **WHEN** an extracted PyPI artifact contains a file matching `*.postman_environment.json`
|
||||
- **THEN** the system SHALL cache the file and enqueue a Postman target with PyPI package origin metadata
|
||||
|
||||
#### Scenario: Existing package scan continues
|
||||
- **WHEN** Postman harvesting fails for one package artifact
|
||||
- **THEN** the original npm or PyPI scan SHALL still complete and record the harvesting failure without failing unrelated package scanning
|
||||
|
||||
### Requirement: Postman-aware enrichment
|
||||
The system SHALL enrich Postman findings with contextual classification derived from Postman structure without replacing TruffleHog detection.
|
||||
|
||||
#### Scenario: Header credential is classified
|
||||
- **WHEN** a finding appears in a Postman request header such as `Authorization` or `x-api-key`
|
||||
- **THEN** enrichment SHALL record the context location and infer credential kind from header type, value shape, and endpoint host when possible
|
||||
|
||||
#### Scenario: Environment variable is classified
|
||||
- **WHEN** a finding appears in a Postman environment variable value
|
||||
- **THEN** enrichment SHALL record the variable name and classify provider or credential kind when supported by value shape or associated request endpoints
|
||||
|
||||
#### Scenario: Placeholder is detected
|
||||
- **WHEN** a Postman value is a placeholder such as `{{API_KEY}}`, `<api_key>`, `YOUR_API_KEY`, `example`, or `changeme`
|
||||
- **THEN** enrichment SHALL classify it as placeholder or low confidence rather than a live secret
|
||||
|
||||
### Requirement: Observability for Postman source
|
||||
The system SHALL expose Postman source activity through existing logs, queue counts, source cycle metrics, target scan records, findings, errors, and dashboard views.
|
||||
|
||||
#### Scenario: Source cycle is recorded
|
||||
- **WHEN** a Postman source cycle runs
|
||||
- **THEN** the scanner database SHALL record source cycle metrics including fetched, queued, scanned, found, error, skipped, and queue counts
|
||||
|
||||
#### Scenario: Dashboard shows Postman queues
|
||||
- **WHEN** Postman queue files exist
|
||||
- **THEN** the dashboard SHALL include Postman queue counts in the current queues view
|
||||
Reference in New Issue
Block a user