Files
2026-09-30 20:30:56 +03:00

81 lines
6.3 KiB
Markdown

## Context
The scanner is already organized around independent sources managed by `console_runner.py` and `supervisor.py`. Each source discovers targets, writes them to per-source queue files, scans targets through TruffleHog, records results in JSONL and SQLite, and can stop pagination when consecutive pages contain only known targets. Existing package sources already download and extract npm/PyPI artifacts before running TruffleHog filesystem scans.
Postman artifacts fit this model as filesystem scan targets, but they need separate discovery and enrichment. Public Postman web search is comparatively fragile, while GitHub code search exposes many real `*.postman_collection.json` and `*.postman_environment.json` files. npm and PyPI can also expose Postman artifacts during the package extraction windows that already exist.
## Goals / Non-Goals
**Goals:**
- Add a first-class `postman` source with standard queue, checked, supervisor, dashboard, and database behavior.
- Discover Postman collection/environment JSON from GitHub code search using the existing GitHub auth pool.
- Support initial backfill over up to the GitHub Search API result cap per query and ongoing tail scans over recently indexed pages.
- Filter discovered GitHub code artifacts by last file commit age so stale files can be skipped during backfill.
- Rotate across multiple GitHub tokens and pause only when all usable tokens are rate-limited.
- Cache discovered Postman artifacts durably so package-derived artifacts survive temp directory cleanup.
- Harvest Postman artifacts from npm and PyPI extraction flows without disrupting existing package scans.
- Enrich findings with Postman-specific context derived from auth configuration, headers, query params, request bodies, environment variables, and endpoint hosts.
**Non-Goals:**
- Scraping Postman's own web application in the first implementation.
- Replacing TruffleHog detectors with custom regex-only detection.
- Adding new keychecker services as part of this change.
- Scanning private Postman workspaces through the Postman API.
## Decisions
1. Use GitHub code search as the primary discovery channel.
GitHub code search has authenticated API support, predictable pagination, and high-quality results for `filename:postman_collection.json <query>` and `filename:postman_environment.json <query>`. Public Postman web pages and `postman.com` links are lower-yield and more likely to change without notice. Direct Postman URL discovery can be added later as an additional provider without changing the scanner contract.
2. Treat Postman artifacts as durable filesystem scan targets.
The source will download or copy each discovered artifact into a runtime Postman cache and scan a temporary directory containing the cached JSON. This reuses the existing TruffleHog filesystem path and avoids keeping npm/PyPI extraction directories alive.
3. Identify GitHub-discovered targets by `repo:path:sha`.
The same file at the same SHA must not be rescanned, while a new SHA for the same path must be queued again. This matches the existing queue/checked model and makes `stop_on_seen_pages` useful for tail scans.
4. Identify package-harvested targets by content hash.
npm/PyPI packages can contain duplicate Postman artifacts across versions or package names. A SHA-256 content hash provides stable dedupe and allows different origins to point to the same cached artifact without rescanning identical content.
5. Filter GitHub code artifacts by path commit age after discovery.
The code search response does not include reliable file modification dates. The implementation will query the latest commit for each `repo:path` and skip artifacts older than `max_file_age_days`. This costs extra core API requests but keeps the code search query simple and reliable.
6. Use a source-local GitHub token pool for discovery.
Postman discovery may make many GitHub requests in one cycle. A per-request token pool can rotate across all configured GitHub auth entries, cool down only the token that failed, and sleep when no token remains available. This is more efficient than one token per source cycle.
7. Keep Postman enrichment separate from detection.
TruffleHog remains responsible for finding candidate secrets. Postman enrichment will add context and confidence by correlating findings with request auth, headers, variables, endpoints, and placeholder detection. This avoids increasing false positives from regex-only scans.
## Risks / Trade-offs
- GitHub code search is rate-limited to roughly 10 requests per minute per token -> throttle to a configurable safe RPM and rotate across the auth pool.
- Commit-age filtering adds extra core API requests -> make `max_file_age_days` configurable and cache commit metadata per `repo:path:sha` within a cycle.
- Broad queries such as `ai` can hit the 1000-result search cap and include noisy results -> use a Postman-specific query list and allow per-source query tuning.
- Package harvesting adds small overhead during npm/PyPI scans -> limit file walking to reasonable extensions, max file size, and known Postman filename patterns.
- Postman variables often contain placeholders rather than live secrets -> classify placeholders separately and keep TruffleHog verification/keycheckers as the authority for live/dead status.
- All tokens may become unavailable -> sleep until the earliest known reset time, or a configured fallback such as 30 minutes when no reset is known.
## Migration Plan
1. Add the `postman` source disabled by default in `config.yaml`.
2. Add queue/dashboard/database support for `postman` without changing existing source behavior.
3. Run a small verification cycle with `pages: 1`, `per_page: 10`, and `max_targets` set.
4. Run the one-time backfill with `pages: 10`, `per_page: 100`, `stop_on_seen_pages: false`, and `max_file_age_days: 365`.
5. Switch the source to tail mode with `pages: 1-3` and `stop_on_seen_pages: true`.
6. Enable npm/PyPI harvesting after the base Postman source is verified.
Rollback is to disable `sources.postman.enabled`, leave its queues/cache intact, and continue running existing sources unchanged.
## Open Questions
- Whether the default backfill query list should include broad terms like `ai`, or keep only higher-intent terms such as `openai`, `anthropic`, `gemini`, `llm`, `rag`, and `agent`.
- Whether package-harvested artifacts should always be enqueued, or only when they include auth/secret-related markers.