Initial server source import
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-06-02
|
||||
@@ -0,0 +1,80 @@
|
||||
## Context
|
||||
|
||||
The scanner is already organized around independent sources managed by `console_runner.py` and `supervisor.py`. Each source discovers targets, writes them to per-source queue files, scans targets through TruffleHog, records results in JSONL and SQLite, and can stop pagination when consecutive pages contain only known targets. Existing package sources already download and extract npm/PyPI artifacts before running TruffleHog filesystem scans.
|
||||
|
||||
Postman artifacts fit this model as filesystem scan targets, but they need separate discovery and enrichment. Public Postman web search is comparatively fragile, while GitHub code search exposes many real `*.postman_collection.json` and `*.postman_environment.json` files. npm and PyPI can also expose Postman artifacts during the package extraction windows that already exist.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Add a first-class `postman` source with standard queue, checked, supervisor, dashboard, and database behavior.
|
||||
- Discover Postman collection/environment JSON from GitHub code search using the existing GitHub auth pool.
|
||||
- Support initial backfill over up to the GitHub Search API result cap per query and ongoing tail scans over recently indexed pages.
|
||||
- Filter discovered GitHub code artifacts by last file commit age so stale files can be skipped during backfill.
|
||||
- Rotate across multiple GitHub tokens and pause only when all usable tokens are rate-limited.
|
||||
- Cache discovered Postman artifacts durably so package-derived artifacts survive temp directory cleanup.
|
||||
- Harvest Postman artifacts from npm and PyPI extraction flows without disrupting existing package scans.
|
||||
- Enrich findings with Postman-specific context derived from auth configuration, headers, query params, request bodies, environment variables, and endpoint hosts.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Scraping Postman's own web application in the first implementation.
|
||||
- Replacing TruffleHog detectors with custom regex-only detection.
|
||||
- Adding new keychecker services as part of this change.
|
||||
- Scanning private Postman workspaces through the Postman API.
|
||||
|
||||
## Decisions
|
||||
|
||||
1. Use GitHub code search as the primary discovery channel.
|
||||
|
||||
GitHub code search has authenticated API support, predictable pagination, and high-quality results for `filename:postman_collection.json <query>` and `filename:postman_environment.json <query>`. Public Postman web pages and `postman.com` links are lower-yield and more likely to change without notice. Direct Postman URL discovery can be added later as an additional provider without changing the scanner contract.
|
||||
|
||||
2. Treat Postman artifacts as durable filesystem scan targets.
|
||||
|
||||
The source will download or copy each discovered artifact into a runtime Postman cache and scan a temporary directory containing the cached JSON. This reuses the existing TruffleHog filesystem path and avoids keeping npm/PyPI extraction directories alive.
|
||||
|
||||
3. Identify GitHub-discovered targets by `repo:path:sha`.
|
||||
|
||||
The same file at the same SHA must not be rescanned, while a new SHA for the same path must be queued again. This matches the existing queue/checked model and makes `stop_on_seen_pages` useful for tail scans.
|
||||
|
||||
4. Identify package-harvested targets by content hash.
|
||||
|
||||
npm/PyPI packages can contain duplicate Postman artifacts across versions or package names. A SHA-256 content hash provides stable dedupe and allows different origins to point to the same cached artifact without rescanning identical content.
|
||||
|
||||
5. Filter GitHub code artifacts by path commit age after discovery.
|
||||
|
||||
The code search response does not include reliable file modification dates. The implementation will query the latest commit for each `repo:path` and skip artifacts older than `max_file_age_days`. This costs extra core API requests but keeps the code search query simple and reliable.
|
||||
|
||||
6. Use a source-local GitHub token pool for discovery.
|
||||
|
||||
Postman discovery may make many GitHub requests in one cycle. A per-request token pool can rotate across all configured GitHub auth entries, cool down only the token that failed, and sleep when no token remains available. This is more efficient than one token per source cycle.
|
||||
|
||||
7. Keep Postman enrichment separate from detection.
|
||||
|
||||
TruffleHog remains responsible for finding candidate secrets. Postman enrichment will add context and confidence by correlating findings with request auth, headers, variables, endpoints, and placeholder detection. This avoids increasing false positives from regex-only scans.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- GitHub code search is rate-limited to roughly 10 requests per minute per token -> throttle to a configurable safe RPM and rotate across the auth pool.
|
||||
- Commit-age filtering adds extra core API requests -> make `max_file_age_days` configurable and cache commit metadata per `repo:path:sha` within a cycle.
|
||||
- Broad queries such as `ai` can hit the 1000-result search cap and include noisy results -> use a Postman-specific query list and allow per-source query tuning.
|
||||
- Package harvesting adds small overhead during npm/PyPI scans -> limit file walking to reasonable extensions, max file size, and known Postman filename patterns.
|
||||
- Postman variables often contain placeholders rather than live secrets -> classify placeholders separately and keep TruffleHog verification/keycheckers as the authority for live/dead status.
|
||||
- All tokens may become unavailable -> sleep until the earliest known reset time, or a configured fallback such as 30 minutes when no reset is known.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add the `postman` source disabled by default in `config.yaml`.
|
||||
2. Add queue/dashboard/database support for `postman` without changing existing source behavior.
|
||||
3. Run a small verification cycle with `pages: 1`, `per_page: 10`, and `max_targets` set.
|
||||
4. Run the one-time backfill with `pages: 10`, `per_page: 100`, `stop_on_seen_pages: false`, and `max_file_age_days: 365`.
|
||||
5. Switch the source to tail mode with `pages: 1-3` and `stop_on_seen_pages: true`.
|
||||
6. Enable npm/PyPI harvesting after the base Postman source is verified.
|
||||
|
||||
Rollback is to disable `sources.postman.enabled`, leave its queues/cache intact, and continue running existing sources unchanged.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether the default backfill query list should include broad terms like `ai`, or keep only higher-intent terms such as `openai`, `anthropic`, `gemini`, `llm`, `rag`, and `agent`.
|
||||
- Whether package-harvested artifacts should always be enqueued, or only when they include auth/secret-related markers.
|
||||
@@ -0,0 +1,32 @@
|
||||
## Why
|
||||
|
||||
Public Postman collections and environments are a high-signal source for leaked API credentials because they often preserve request auth settings, headers, variables, and example payloads close to real API usage. The existing scanner already supports multi-source discovery, queues, TruffleHog filesystem scans, token rotation, and observability, so adding Postman can reuse the current architecture while expanding coverage beyond repositories, packages, containers, and HuggingFace Spaces.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a new `postman` source that discovers, queues, scans, and records Postman collection/environment artifacts.
|
||||
- Seed Postman targets from GitHub code search using public `*.postman_collection.json` and `*.postman_environment.json` files.
|
||||
- Support a one-time backfill mode that scans up to the GitHub Search API result limit per query while filtering out artifacts older than a configured age window.
|
||||
- Support a daily tail mode that fetches recently indexed pages and stops early when all targets on consecutive pages are already known.
|
||||
- Use the configured GitHub auth pool for Postman discovery, rotating across tokens and sleeping when all tokens are rate-limited.
|
||||
- Add durable Postman artifact caching so targets discovered from GitHub, npm, and PyPI can be scanned after temporary extraction directories are removed.
|
||||
- Harvest Postman artifacts from npm and PyPI packages during existing package extraction flows and enqueue them into the shared Postman queue.
|
||||
- Add Postman-aware result enrichment that classifies credentials using TruffleHog findings plus Postman auth/header/query/body/environment context.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `postman-source`: Discovery, queueing, scanning, caching, and enrichment for Postman collection and environment artifacts.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected scanner paths: `app/scanner.py`, `app/console_runner.py`, `app/scanner_db.py`, `app/dashboard.py`, and `app/config.yaml`.
|
||||
- Adds runtime files under `runtime/queues/` for `todo_postman.txt` and `checked_postman.txt`.
|
||||
- Adds durable artifact storage under a runtime Postman cache directory.
|
||||
- Uses existing GitHub auth pools from `secrets.yaml`; no new secret format is required for GitHub discovery.
|
||||
- Uses existing TruffleHog filesystem scanning and keychecker follow-up flows; no breaking changes to current sources are expected.
|
||||
@@ -0,0 +1,158 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Postman source registration
|
||||
The system SHALL provide a first-class `postman` source that can be configured, selected, supervised, queued, scanned, and displayed consistently with existing scanner sources.
|
||||
|
||||
#### Scenario: Configured Postman source is selectable
|
||||
- **WHEN** a config contains `sources.postman.enabled: true` and the runner is invoked with `--source postman`
|
||||
- **THEN** the runner SHALL execute only the Postman source cycle using Postman source settings
|
||||
|
||||
#### Scenario: Postman queue files are used
|
||||
- **WHEN** the Postman source prepares targets
|
||||
- **THEN** the system SHALL use `todo_postman.txt` and `checked_postman.txt` under the configured queue directory
|
||||
|
||||
### Requirement: GitHub code search discovery
|
||||
The Postman source SHALL discover public Postman artifacts from GitHub code search using configured queries and artifact kinds.
|
||||
|
||||
#### Scenario: Collection files are discovered
|
||||
- **WHEN** `search_kinds` includes `collection` and the query is `openai`
|
||||
- **THEN** discovery SHALL search GitHub code for `filename:postman_collection.json openai`
|
||||
|
||||
#### Scenario: Environment files are discovered
|
||||
- **WHEN** `search_kinds` includes `environment` and the query is `openai`
|
||||
- **THEN** discovery SHALL search GitHub code for `filename:postman_environment.json openai`
|
||||
|
||||
#### Scenario: Search pagination is bounded
|
||||
- **WHEN** `pages` is `10` and `per_page` is `100`
|
||||
- **THEN** discovery SHALL request no more than 1000 search results per query and kind
|
||||
|
||||
### Requirement: GitHub auth pool rotation
|
||||
The Postman GitHub discovery flow SHALL use all available tokens from the configured GitHub auth pool before sleeping for rate limits.
|
||||
|
||||
#### Scenario: Token rotates per GitHub request
|
||||
- **WHEN** multiple GitHub auth entries are available
|
||||
- **THEN** GitHub code search, commit lookup, and content download requests SHALL rotate across available tokens
|
||||
|
||||
#### Scenario: One token is rate-limited
|
||||
- **WHEN** a GitHub request returns a primary or secondary rate limit for the current token
|
||||
- **THEN** the system SHALL mark only that token unavailable until its reset time or configured cooldown and continue with another available token
|
||||
|
||||
#### Scenario: All tokens are unavailable
|
||||
- **WHEN** every configured GitHub token is rate-limited or temporarily unavailable
|
||||
- **THEN** the source SHALL sleep until the earliest known reset time, or for the configured fallback cooldown when no reset time is known
|
||||
|
||||
#### Scenario: Token is invalid
|
||||
- **WHEN** a GitHub request returns an authentication-invalid response for a token
|
||||
- **THEN** the system SHALL exclude that token from the current cycle and report the authentication failure without marking other tokens invalid
|
||||
|
||||
### Requirement: Backfill freshness filtering
|
||||
The Postman source SHALL support a backfill mode that can scan deep code search pages while skipping GitHub artifacts older than a configured file age.
|
||||
|
||||
#### Scenario: Recent file is queued
|
||||
- **WHEN** a GitHub code search result has a latest path commit within `max_file_age_days`
|
||||
- **THEN** the target SHALL be eligible for queueing
|
||||
|
||||
#### Scenario: Old file is skipped
|
||||
- **WHEN** a GitHub code search result has a latest path commit older than `max_file_age_days`
|
||||
- **THEN** the target SHALL not be queued and SHALL be counted as skipped by freshness filtering
|
||||
|
||||
#### Scenario: Freshness filtering is disabled
|
||||
- **WHEN** `max_file_age_days` is `0`
|
||||
- **THEN** the source SHALL not perform path commit age filtering
|
||||
|
||||
### Requirement: Tail mode early stop
|
||||
The Postman source SHALL support ongoing tail scans that stop pagination after consecutive known pages.
|
||||
|
||||
#### Scenario: Known page increments stop counter
|
||||
- **WHEN** `stop_on_seen_pages` is enabled and every normalized target on a fetched page already exists in `todo_postman.txt` or `checked_postman.txt`
|
||||
- **THEN** the source SHALL count that page as known
|
||||
|
||||
#### Scenario: Tail pagination stops
|
||||
- **WHEN** the known page count reaches `seen_page_threshold` after `min_pages_before_stop`
|
||||
- **THEN** the source SHALL stop fetching additional pages for that query and artifact kind
|
||||
|
||||
### Requirement: Postman target identity and deduplication
|
||||
The system SHALL normalize Postman targets so identical artifacts are not rescanned while changed artifacts are scanned again.
|
||||
|
||||
#### Scenario: GitHub target identity includes SHA
|
||||
- **WHEN** a Postman target is discovered from GitHub code search
|
||||
- **THEN** its normalized target SHALL include source, repository, path, and file SHA
|
||||
|
||||
#### Scenario: GitHub file changes
|
||||
- **WHEN** the same GitHub repository and path is discovered with a new SHA
|
||||
- **THEN** the system SHALL treat it as a new Postman target
|
||||
|
||||
#### Scenario: Package target identity uses content hash
|
||||
- **WHEN** a Postman artifact is harvested from npm or PyPI
|
||||
- **THEN** its normalized target SHALL include the artifact content SHA-256 hash
|
||||
|
||||
### Requirement: Durable Postman artifact cache
|
||||
The system SHALL store discovered Postman artifact content in durable runtime cache before scanning.
|
||||
|
||||
#### Scenario: GitHub content is cached
|
||||
- **WHEN** a GitHub code search target is queued for scanning
|
||||
- **THEN** the system SHALL download the artifact content and store it under the configured Postman cache directory
|
||||
|
||||
#### Scenario: Package content is cached before cleanup
|
||||
- **WHEN** npm or PyPI extraction finds a Postman artifact
|
||||
- **THEN** the system SHALL copy the artifact into the durable Postman cache before the extraction directory is removed
|
||||
|
||||
#### Scenario: Cache size is constrained
|
||||
- **WHEN** an artifact exceeds the configured maximum Postman artifact size
|
||||
- **THEN** the system SHALL skip the artifact and record a bounded error or skip reason
|
||||
|
||||
### Requirement: Postman artifact scanning
|
||||
The Postman source SHALL scan cached Postman collection and environment artifacts with TruffleHog filesystem scanning.
|
||||
|
||||
#### Scenario: Cached artifact is scanned
|
||||
- **WHEN** a Postman target points to a cached JSON artifact
|
||||
- **THEN** the scanner SHALL run TruffleHog against a temporary filesystem directory containing that artifact
|
||||
|
||||
#### Scenario: Findings are persisted
|
||||
- **WHEN** TruffleHog reports findings for a Postman target
|
||||
- **THEN** the system SHALL persist findings to existing JSONL outputs and scanner database tables with source `postman`
|
||||
|
||||
#### Scenario: Scan finishes
|
||||
- **WHEN** a Postman target scan completes with findings, errors, skipped status, or clean status
|
||||
- **THEN** the target SHALL be moved from `todo_postman.txt` to `checked_postman.txt`
|
||||
|
||||
### Requirement: npm and PyPI Postman harvesting
|
||||
The npm and PyPI source flows SHALL harvest Postman artifacts discovered during existing package extraction and enqueue them for the Postman source.
|
||||
|
||||
#### Scenario: npm package contains collection
|
||||
- **WHEN** an extracted npm package contains a file matching `*.postman_collection.json`
|
||||
- **THEN** the system SHALL cache the file and enqueue a Postman target with npm package origin metadata
|
||||
|
||||
#### Scenario: PyPI package contains environment
|
||||
- **WHEN** an extracted PyPI artifact contains a file matching `*.postman_environment.json`
|
||||
- **THEN** the system SHALL cache the file and enqueue a Postman target with PyPI package origin metadata
|
||||
|
||||
#### Scenario: Existing package scan continues
|
||||
- **WHEN** Postman harvesting fails for one package artifact
|
||||
- **THEN** the original npm or PyPI scan SHALL still complete and record the harvesting failure without failing unrelated package scanning
|
||||
|
||||
### Requirement: Postman-aware enrichment
|
||||
The system SHALL enrich Postman findings with contextual classification derived from Postman structure without replacing TruffleHog detection.
|
||||
|
||||
#### Scenario: Header credential is classified
|
||||
- **WHEN** a finding appears in a Postman request header such as `Authorization` or `x-api-key`
|
||||
- **THEN** enrichment SHALL record the context location and infer credential kind from header type, value shape, and endpoint host when possible
|
||||
|
||||
#### Scenario: Environment variable is classified
|
||||
- **WHEN** a finding appears in a Postman environment variable value
|
||||
- **THEN** enrichment SHALL record the variable name and classify provider or credential kind when supported by value shape or associated request endpoints
|
||||
|
||||
#### Scenario: Placeholder is detected
|
||||
- **WHEN** a Postman value is a placeholder such as `{{API_KEY}}`, `<api_key>`, `YOUR_API_KEY`, `example`, or `changeme`
|
||||
- **THEN** enrichment SHALL classify it as placeholder or low confidence rather than a live secret
|
||||
|
||||
### Requirement: Observability for Postman source
|
||||
The system SHALL expose Postman source activity through existing logs, queue counts, source cycle metrics, target scan records, findings, errors, and dashboard views.
|
||||
|
||||
#### Scenario: Source cycle is recorded
|
||||
- **WHEN** a Postman source cycle runs
|
||||
- **THEN** the scanner database SHALL record source cycle metrics including fetched, queued, scanned, found, error, skipped, and queue counts
|
||||
|
||||
#### Scenario: Dashboard shows Postman queues
|
||||
- **WHEN** Postman queue files exist
|
||||
- **THEN** the dashboard SHALL include Postman queue counts in the current queues view
|
||||
@@ -0,0 +1,69 @@
|
||||
## 1. Source Wiring
|
||||
|
||||
- [x] 1.1 Add `postman` to CLI `--source` and `--platform` choices and source-to-platform mapping.
|
||||
- [x] 1.2 Add `postman` queue file support through the existing `queue_files_for_args`, prepare, and mark-checked flow.
|
||||
- [x] 1.3 Add `postman` to supervisor source configuration and dashboard source lists.
|
||||
- [x] 1.4 Add disabled-by-default `sources.postman` configuration with GitHub auth pool, search kinds, backfill/tail controls, cache path, and rate-limit settings.
|
||||
|
||||
## 2. Target Model And Cache
|
||||
|
||||
- [x] 2.1 Define Postman target JSON formats for GitHub code search, npm package, PyPI package, local cache, and future URL targets.
|
||||
- [x] 2.2 Implement Postman target parsing and normalization in `console_runner.py` and `scanner_db.py`.
|
||||
- [x] 2.3 Implement durable Postman cache path resolution under the configured runtime directory.
|
||||
- [x] 2.4 Implement safe cache writes with SHA-256 content hashing, max artifact size checks, and origin metadata preservation.
|
||||
|
||||
## 3. GitHub Code Search Discovery
|
||||
|
||||
- [x] 3.1 Implement GitHub code search queries for `filename:postman_collection.json <query>` and `filename:postman_environment.json <query>` based on `search_kinds`.
|
||||
- [x] 3.2 Implement bounded pagination using configured `pages` and `per_page`, respecting the GitHub 1000-result search cap.
|
||||
- [x] 3.3 Convert GitHub code search items into Postman target JSON containing repository, path, SHA, kind, API URL, and HTML URL.
|
||||
- [x] 3.4 Implement latest path commit lookup for `max_file_age_days` filtering.
|
||||
- [x] 3.5 Integrate existing known-page early stop behavior for Postman tail scans.
|
||||
|
||||
## 4. GitHub Token Pool And Rate Limits
|
||||
|
||||
- [x] 4.1 Build a source-local GitHub token pool from configured `auth_pool` entries and fallback token settings.
|
||||
- [x] 4.2 Rotate tokens per GitHub code search, commit lookup, and content download request.
|
||||
- [x] 4.3 Mark only the failing token unavailable on primary rate limit, secondary rate limit, auth invalid, or auth forbidden responses.
|
||||
- [x] 4.4 Sleep until earliest known reset time, or configured fallback cooldown, when all GitHub tokens are unavailable.
|
||||
- [x] 4.5 Record token cooldown status in the source runtime state without exposing token values in logs or database snapshots.
|
||||
|
||||
## 5. Postman Artifact Scanning
|
||||
|
||||
- [x] 5.1 Implement Postman content download from GitHub Contents API and cache it before scanning.
|
||||
- [x] 5.2 Implement `scan_postman_target()` to stage cached JSON in a temporary directory and run TruffleHog filesystem scanning.
|
||||
- [x] 5.3 Add Postman branch to `scan_targets_batch()` and pass timeout, detectors, excluded detectors, and verification flags.
|
||||
- [x] 5.4 Preserve nearby file context and apply existing noisy finding filters to Postman scan results.
|
||||
- [x] 5.5 Ensure Postman findings, errors, skipped reasons, and clean scans are persisted through existing JSONL and scanner database writes.
|
||||
|
||||
## 6. npm And PyPI Harvesting
|
||||
|
||||
- [x] 6.1 Add a Postman artifact finder for extracted package directories that matches collection and environment filename patterns.
|
||||
- [x] 6.2 Cache npm package Postman artifacts before package temp directory cleanup and attach npm origin metadata.
|
||||
- [x] 6.3 Cache PyPI package Postman artifacts before package temp directory cleanup and attach PyPI origin metadata.
|
||||
- [x] 6.4 Enqueue harvested package artifacts into `todo_postman.txt` after package scan batches without failing the original package scan.
|
||||
- [x] 6.5 Deduplicate harvested package artifacts by content hash before enqueueing.
|
||||
|
||||
## 7. Postman-Aware Enrichment
|
||||
|
||||
- [x] 7.1 Parse Postman collection and environment JSON into request, auth, header, query, body, and variable context maps.
|
||||
- [x] 7.2 Correlate TruffleHog finding locations or nearby context with Postman context maps.
|
||||
- [x] 7.3 Classify credential kind and provider using DetectorName, value shape, auth/header type, variable name, and endpoint host.
|
||||
- [x] 7.4 Detect common placeholders and assign placeholder or low-confidence classification.
|
||||
- [x] 7.5 Persist enrichment fields using existing finding enrichment/database columns where possible.
|
||||
|
||||
## 8. Observability And Configuration
|
||||
|
||||
- [x] 8.1 Add Postman source cycle metrics, queue snapshots, target scan records, findings, and errors to existing database flows.
|
||||
- [x] 8.2 Add Postman queue counts to dashboard current queues and source health views.
|
||||
- [x] 8.3 Add redaction coverage for Postman/GitHub auth pool settings in config snapshots and logs.
|
||||
- [x] 8.4 Add backfill-friendly and tail-friendly config examples in `config.yaml` comments.
|
||||
|
||||
## 9. Verification
|
||||
|
||||
- [x] 9.1 Run a small Postman GitHub discovery cycle with `pages: 1`, `per_page: 10`, and `max_targets` set.
|
||||
- [x] 9.2 Re-run the same cycle and verify duplicate targets are skipped through `todo_postman.txt` and `checked_postman.txt`.
|
||||
- [x] 9.3 Verify all-token rate-limit fallback with a simulated or controlled token-unavailable state.
|
||||
- [x] 9.4 Verify npm and PyPI harvesting using a package fixture containing collection and environment JSON files.
|
||||
- [x] 9.5 Verify database and dashboard visibility for Postman source cycles, queues, target scans, findings, and errors.
|
||||
- [x] 9.6 Run `openspec status --change add-postman-source` and ensure all implementation tasks are complete before archive.
|
||||
Reference in New Issue
Block a user