Files
2026-09-30 20:30:56 +03:00

8.8 KiB

ADDED Requirements

Requirement: Postman source registration

The system SHALL provide a first-class postman source that can be configured, selected, supervised, queued, scanned, and displayed consistently with existing scanner sources.

Scenario: Configured Postman source is selectable

  • WHEN a config contains sources.postman.enabled: true and the runner is invoked with --source postman
  • THEN the runner SHALL execute only the Postman source cycle using Postman source settings

Scenario: Postman queue files are used

  • WHEN the Postman source prepares targets
  • THEN the system SHALL use todo_postman.txt and checked_postman.txt under the configured queue directory

Requirement: GitHub code search discovery

The Postman source SHALL discover public Postman artifacts from GitHub code search using configured queries and artifact kinds.

Scenario: Collection files are discovered

  • WHEN search_kinds includes collection and the query is openai
  • THEN discovery SHALL search GitHub code for filename:postman_collection.json openai

Scenario: Environment files are discovered

  • WHEN search_kinds includes environment and the query is openai
  • THEN discovery SHALL search GitHub code for filename:postman_environment.json openai

Scenario: Search pagination is bounded

  • WHEN pages is 10 and per_page is 100
  • THEN discovery SHALL request no more than 1000 search results per query and kind

Requirement: GitHub auth pool rotation

The Postman GitHub discovery flow SHALL use all available tokens from the configured GitHub auth pool before sleeping for rate limits.

Scenario: Token rotates per GitHub request

  • WHEN multiple GitHub auth entries are available
  • THEN GitHub code search, commit lookup, and content download requests SHALL rotate across available tokens

Scenario: One token is rate-limited

  • WHEN a GitHub request returns a primary or secondary rate limit for the current token
  • THEN the system SHALL mark only that token unavailable until its reset time or configured cooldown and continue with another available token

Scenario: All tokens are unavailable

  • WHEN every configured GitHub token is rate-limited or temporarily unavailable
  • THEN the source SHALL sleep until the earliest known reset time, or for the configured fallback cooldown when no reset time is known

Scenario: Token is invalid

  • WHEN a GitHub request returns an authentication-invalid response for a token
  • THEN the system SHALL exclude that token from the current cycle and report the authentication failure without marking other tokens invalid

Requirement: Backfill freshness filtering

The Postman source SHALL support a backfill mode that can scan deep code search pages while skipping GitHub artifacts older than a configured file age.

Scenario: Recent file is queued

  • WHEN a GitHub code search result has a latest path commit within max_file_age_days
  • THEN the target SHALL be eligible for queueing

Scenario: Old file is skipped

  • WHEN a GitHub code search result has a latest path commit older than max_file_age_days
  • THEN the target SHALL not be queued and SHALL be counted as skipped by freshness filtering

Scenario: Freshness filtering is disabled

  • WHEN max_file_age_days is 0
  • THEN the source SHALL not perform path commit age filtering

Requirement: Tail mode early stop

The Postman source SHALL support ongoing tail scans that stop pagination after consecutive known pages.

Scenario: Known page increments stop counter

  • WHEN stop_on_seen_pages is enabled and every normalized target on a fetched page already exists in todo_postman.txt or checked_postman.txt
  • THEN the source SHALL count that page as known

Scenario: Tail pagination stops

  • WHEN the known page count reaches seen_page_threshold after min_pages_before_stop
  • THEN the source SHALL stop fetching additional pages for that query and artifact kind

Requirement: Postman target identity and deduplication

The system SHALL normalize Postman targets so identical artifacts are not rescanned while changed artifacts are scanned again.

Scenario: GitHub target identity includes SHA

  • WHEN a Postman target is discovered from GitHub code search
  • THEN its normalized target SHALL include source, repository, path, and file SHA

Scenario: GitHub file changes

  • WHEN the same GitHub repository and path is discovered with a new SHA
  • THEN the system SHALL treat it as a new Postman target

Scenario: Package target identity uses content hash

  • WHEN a Postman artifact is harvested from npm or PyPI
  • THEN its normalized target SHALL include the artifact content SHA-256 hash

Requirement: Durable Postman artifact cache

The system SHALL store discovered Postman artifact content in durable runtime cache before scanning.

Scenario: GitHub content is cached

  • WHEN a GitHub code search target is queued for scanning
  • THEN the system SHALL download the artifact content and store it under the configured Postman cache directory

Scenario: Package content is cached before cleanup

  • WHEN npm or PyPI extraction finds a Postman artifact
  • THEN the system SHALL copy the artifact into the durable Postman cache before the extraction directory is removed

Scenario: Cache size is constrained

  • WHEN an artifact exceeds the configured maximum Postman artifact size
  • THEN the system SHALL skip the artifact and record a bounded error or skip reason

Requirement: Postman artifact scanning

The Postman source SHALL scan cached Postman collection and environment artifacts with TruffleHog filesystem scanning.

Scenario: Cached artifact is scanned

  • WHEN a Postman target points to a cached JSON artifact
  • THEN the scanner SHALL run TruffleHog against a temporary filesystem directory containing that artifact

Scenario: Findings are persisted

  • WHEN TruffleHog reports findings for a Postman target
  • THEN the system SHALL persist findings to existing JSONL outputs and scanner database tables with source postman

Scenario: Scan finishes

  • WHEN a Postman target scan completes with findings, errors, skipped status, or clean status
  • THEN the target SHALL be moved from todo_postman.txt to checked_postman.txt

Requirement: npm and PyPI Postman harvesting

The npm and PyPI source flows SHALL harvest Postman artifacts discovered during existing package extraction and enqueue them for the Postman source.

Scenario: npm package contains collection

  • WHEN an extracted npm package contains a file matching *.postman_collection.json
  • THEN the system SHALL cache the file and enqueue a Postman target with npm package origin metadata

Scenario: PyPI package contains environment

  • WHEN an extracted PyPI artifact contains a file matching *.postman_environment.json
  • THEN the system SHALL cache the file and enqueue a Postman target with PyPI package origin metadata

Scenario: Existing package scan continues

  • WHEN Postman harvesting fails for one package artifact
  • THEN the original npm or PyPI scan SHALL still complete and record the harvesting failure without failing unrelated package scanning

Requirement: Postman-aware enrichment

The system SHALL enrich Postman findings with contextual classification derived from Postman structure without replacing TruffleHog detection.

Scenario: Header credential is classified

  • WHEN a finding appears in a Postman request header such as Authorization or x-api-key
  • THEN enrichment SHALL record the context location and infer credential kind from header type, value shape, and endpoint host when possible

Scenario: Environment variable is classified

  • WHEN a finding appears in a Postman environment variable value
  • THEN enrichment SHALL record the variable name and classify provider or credential kind when supported by value shape or associated request endpoints

Scenario: Placeholder is detected

  • WHEN a Postman value is a placeholder such as {{API_KEY}}, <api_key>, YOUR_API_KEY, example, or changeme
  • THEN enrichment SHALL classify it as placeholder or low confidence rather than a live secret

Requirement: Observability for Postman source

The system SHALL expose Postman source activity through existing logs, queue counts, source cycle metrics, target scan records, findings, errors, and dashboard views.

Scenario: Source cycle is recorded

  • WHEN a Postman source cycle runs
  • THEN the scanner database SHALL record source cycle metrics including fetched, queued, scanned, found, error, skipped, and queue counts

Scenario: Dashboard shows Postman queues

  • WHEN Postman queue files exist
  • THEN the dashboard SHALL include Postman queue counts in the current queues view