6.0 KiB
Context
Supersession, 2026-07-27: the approved cumulative-lag/memory/backpressure architecture supersedes this change's earlier no-new-queue and conservative-concurrency constraints. PostgreSQL is now sole authority, sources hand off S:-backed durable bundles, JSONL/status files are asynchronous projections, normal keychecks use PostgreSQL candidates, and the fair global scan limit is three. Process recycling, GC trimming, and reduced concurrency are not correctness mechanisms.
The scanner runs multiple supervised source processes that share runtime state, queue files, logs, and an observability database. Recent logs show source restarts caused by Windows PermissionError during state-file replacement, while dashboard and supervisor read the same state files for status display. Coverage is also limited by conservative artifact size caps and by discovery inputs that are broad but not always high-signal.
The original incremental constraints below remain historical context only where contradicted by the supersession above.
Goals / Non-Goals
Goals:
- Prevent transient Windows state-file read/write races from crashing source processes.
- Increase scan coverage through configurable size limit bumps while preserving conservative concurrency.
- Improve discovery signal by focusing metadata searches on provider, host, and framework terms.
- Strengthen
package_gitas a discovery backbone through better repository URL extraction and canonicalization. - Make GitHub Actions and GitLab CI scanning use better parsed seeds and modestly larger per-cycle target sets.
- Continue provider-specific detector and keychecker additions with safe validation and context routing.
Non-Goals:
- No deferred/deep queue in this change.
- No new checked/skipped classification model in this change.
- No scheduler rewrite or central scoring engine.
- No broad generic
api_key,secret, ortokenterms in normal repo/package metadata discovery by default. - No removal of existing command-line entry points or queue file formats.
Decisions
Use retrying unique-temp state writes instead of locking
State writes will use a unique temporary file name and retry os.replace on transient Windows permission failures. This keeps the existing JSON state model and avoids cross-process lock files or SQLite migration.
Alternatives considered:
- Lock files: rejected for now because they add another failure mode and require all readers/writers to cooperate.
- SQLite state: rejected as too large for this incremental reliability fix.
- Ignoring failed state writes: rejected because auth/query state must remain observable.
Increase limits through configuration first
Package, Postman, and CI artifact limits will be increased in config while keeping source worker counts conservative. The implementation will not introduce deferred queues or new outcome classes; oversized skips remain visible through existing logs and target scan records.
Alternatives considered:
- Remove limits entirely: rejected because large archives can exhaust disk, CPU, and scan slots.
- Add deep-lane queues: deferred to a separate design because it needs stronger classification semantics.
Keep metadata discovery provider-focused
Repository/package metadata search should use provider names, API hosts, framework terms, and ecosystem terms. Exact secret variable terms should be reserved for code/artifact-oriented searches where content is actually searched.
Alternatives considered:
- Add
api_keyglobally: rejected because most sources search names/descriptions/readmes and this produces low-signal security-tool/tutorial results.
Improve existing CI seed flow before increasing volume
CI sources should first parse more existing seed formats from DB records, package candidates, and findings. After parsing improves, ci_seed_scan_limit and per-cycle target counts can be raised modestly.
Alternatives considered:
- Increase CI volume immediately: rejected because current logs show many unparseable/known seeds, so raw volume would mostly amplify waste.
Add provider support as detector plus keychecker pairs
Provider expansion should follow the Qwen/DashScope pattern: contextual detector, safe keychecker, and routing safeguards for ambiguous sk-... formats.
Alternatives considered:
- Detector-only additions: rejected for providers where validation is feasible because they increase unverified noise.
Risks / Trade-offs
- Windows state retry may hide a persistent file access problem for a few hundred milliseconds -> surface the final error after bounded retries.
- Higher artifact size caps increase runtime and disk pressure -> keep workers conservative and rely on existing scan slot limits.
- Provider-focused queries may miss generic projects that leak keys -> use package/artifact/code-oriented paths for exact env var searches instead of metadata search.
- CI seed parsing improvements may still leave many stale/known targets -> postpone TTL/revisit policy until outcome semantics are revisited.
- Context routing can misclassify ambiguous keys when context is weak -> only route away from a provider when strong provider-specific context is present.
Migration Plan
- Apply state write retry first and monitor source restarts.
- Increase size caps in config and monitor disk usage, scan duration, and skipped/error counts.
- Adjust discovery query lists and verify target volume remains healthy.
- Improve
package_gitmetadata URL extraction/canonicalization. - Improve CI seed parsing, then raise CI limits modestly.
- Add the next provider detector/keychecker pair using the established pattern.
Rollback is straightforward for each step: revert the code change for state writes, restore previous config caps/query lists, or disable individual CI/provider changes.
Open Questions
- Should oversized artifacts remain checked under current semantics, or should that become a separate future change?
- Which provider detector/keychecker pair should be prioritized after Qwen/DashScope?
- What disk and scan-duration thresholds should trigger reducing size caps again?