Initial server source import

This commit is contained in:
sashatrask
2026-09-30 20:30:56 +03:00
commit 170dd941b9
498 changed files with 261563 additions and 0 deletions
@@ -0,0 +1,48 @@
## ADDED Requirements
### Requirement: PostgreSQL Candidate And Current-State Authority
Normal provider workers SHALL claim one fenced PostgreSQL candidate at a time and SHALL transactionally insert the result, update `keycheck_current_state`, complete that candidate, and create a projection job.
#### Scenario: Compatibility files lag or fail
- **WHEN** keycheck JSONL or status projection is delayed or quarantined
- **THEN** the committed PostgreSQL result and current state SHALL remain authoritative and unrelated candidates SHALL continue
#### Scenario: Explicit compatibility input
- **WHEN** an operator selects `--input-mode jsonl`
- **THEN** the bounded legacy JSONL reader MAY be used, but `--input` SHALL be rejected in normal PostgreSQL mode
### Requirement: Provider Support Uses Detector And Keychecker Pair
New provider API key support SHALL include both detection and validation when a safe validation endpoint is available.
#### Scenario: Provider has safe validation endpoint
- **WHEN** support is added for a provider with a non-generating authentication or model-list endpoint
- **THEN** the change SHALL include a detector or detector routing rule and a keychecker for that provider
#### Scenario: Provider has no safe validation endpoint
- **WHEN** a provider does not have a safe validation endpoint
- **THEN** the detector MAY be added only with explicit documentation that validation is unavailable
### Requirement: Ambiguous Key Formats Use Context Routing
The scanner SHALL use nearby provider-specific context to route ambiguous key formats to the correct keychecker when possible.
#### Scenario: Qwen context around generic key
- **WHEN** an `sk-...` key is found near Qwen or DashScope context
- **THEN** the key SHALL be routed away from unrelated generic `sk-...` checkers such as DeepSeek when the context does not also identify that provider
#### Scenario: Weak context around generic key
- **WHEN** an ambiguous key has no strong provider-specific context
- **THEN** the scanner SHALL avoid speculative rerouting that would suppress the existing detector result
### Requirement: Keycheckers Avoid Token-Generating Probes By Default
Provider keycheckers SHALL prefer safe non-generating validation endpoints where available.
#### Scenario: Provider exposes model-list endpoint
- **WHEN** a provider exposes an authenticated model-list or account-status endpoint
- **THEN** the keychecker SHALL use that endpoint before considering any generation-style probe
### Requirement: Keycheck Results Remain Linkable To Findings
Provider keycheckers SHALL emit detector names and result metadata that can be linked back to scanner findings.
#### Scenario: Custom detector result is validated
- **WHEN** a keychecker validates a key from a custom detector finding
- **THEN** the result SHALL include a stable detector name and source metadata sufficient for DB link repair
@@ -0,0 +1,33 @@
## ADDED Requirements
### Requirement: Configurable Larger Artifact Coverage
The scanner SHALL allow package, Postman, GitHub Actions, and GitLab CI artifact size limits to be raised through configuration without code changes.
#### Scenario: Package artifact limit is increased
- **WHEN** npm or PyPI `max_artifact_size_mb` is configured to a larger value
- **THEN** package scans SHALL use the configured size limit for download and scan decisions
#### Scenario: CI artifact limits are increased
- **WHEN** GitHub Actions or GitLab CI artifact archive and file limits are configured to larger values
- **THEN** CI scans SHALL use those configured limits for artifact download and extraction decisions
### Requirement: Conservative Concurrency Preserved
The scanner SHALL preserve per-source worker controls so larger artifact limits do not automatically increase concurrent heavy scans.
#### Scenario: Size limits increase with unchanged workers
- **WHEN** artifact size limits are raised in configuration and worker counts are unchanged
- **THEN** the scanner SHALL keep using the configured worker counts for the affected source
### Requirement: Oversized Artifact Visibility
The scanner SHALL record existing oversized-artifact skip reasons in logs and target scan records using the current result model.
#### Scenario: Artifact remains over configured limit
- **WHEN** an artifact exceeds the configured size limit
- **THEN** the scanner SHALL record a skipped result with the size-limit reason using existing logging and target scan recording paths
### Requirement: No New Queue Semantics
The scanner SHALL NOT introduce a deferred or deep artifact queue as part of this change.
#### Scenario: Artifact is too large for current run
- **WHEN** an artifact is skipped because it exceeds the configured limit
- **THEN** the scanner SHALL handle it through the current scan result and queue behavior without creating a new queue type
@@ -0,0 +1,40 @@
## ADDED Requirements
### Requirement: Durable Asynchronous Result Handoff Supersedes Source Publication
PostgreSQL SHALL be the sole result authority. A fair global permit count of three SHALL end after a pre-reserved result bundle is fsynced and atomically renamed on `S:`, without covering database ingest, JSONL projection, or keychecks.
#### Scenario: Projector or database is blocked after bundle handoff
- **WHEN** a source completes the durable ready rename
- **THEN** its scan permit SHALL be released while the bundle remains recoverable and downstream backlog applies capacity-based admission pressure
### Requirement: Correctness Does Not Depend On Recycling
The runtime SHALL NOT use GC trimming, source lifetime limits, private-memory restarts, periodic recycling, reduced concurrency, or time-based restarts to preserve correctness.
#### Scenario: Runtime memory grows after warmup
- **WHEN** a managed process reports increasing private memory
- **THEN** admission and durable queue capacity SHALL provide backpressure without recycling the process or reducing the configured three scan permits
### Requirement: Resilient Runner State Writes
The scanner SHALL persist runner state using a unique temporary file per write attempt and SHALL retry replacement when the operating system reports a transient file access error.
#### Scenario: Concurrent state read during write
- **WHEN** a supervised source writes `runner_state_<source>.json` while supervisor or dashboard reads the same state file
- **THEN** the source process SHALL retry the replacement and continue without crashing when the file becomes available within the retry window
#### Scenario: Persistent state write failure
- **WHEN** the state file cannot be replaced after the bounded retry window
- **THEN** the source process SHALL surface the final write error instead of silently discarding state changes
### Requirement: State Writes Avoid Shared Temp Path Contention
The scanner SHALL avoid using a single shared `.tmp` path for repeated state writes from supervised processes.
#### Scenario: Multiple state write attempts overlap
- **WHEN** two state write attempts occur close together for the same state file
- **THEN** each attempt SHALL use a distinct temporary file path before replacing the final state file
### Requirement: Restart Noise Reduction
The supervisor SHALL no longer restart sources due only to transient state-file replacement races that resolve within the retry window.
#### Scenario: GitLab state replacement race resolves
- **WHEN** the GitLab source hits a transient Windows file lock while saving state
- **THEN** the source SHALL complete the state save after retry and its `up` timer SHALL not reset because of that transient race
@@ -0,0 +1,48 @@
## ADDED Requirements
### Requirement: Metadata Discovery Uses High-Signal Queries
Repository, package, and image metadata discovery SHALL prefer provider names, API hosts, framework names, and ecosystem terms over generic secret-related words.
#### Scenario: Repo metadata query list is reviewed
- **WHEN** default repository or package metadata queries are configured
- **THEN** broad terms such as `api_key`, `secret`, and `token` SHALL NOT be added by default unless the source searches content rather than metadata
#### Scenario: Provider query is configured
- **WHEN** a provider such as Qwen, DashScope, Groq, or OpenRouter is targeted
- **THEN** provider names, API hostnames, and framework terms SHALL be eligible for metadata discovery queries
### Requirement: Exact Secret Terms Reserved For Content-Oriented Sources
Exact environment variable and API host searches SHALL be used for content-oriented sources such as Postman/API artifacts, code search, and CI artifacts rather than generic metadata searches.
#### Scenario: Env var search term is added
- **WHEN** an exact term such as `DASHSCOPE_API_KEY` is added to discovery
- **THEN** it SHALL be applied to a source that can inspect file or artifact content
### Requirement: Package Git Repository Canonicalization
`package_git` discovery SHALL canonicalize repository URLs from package metadata before queueing targets.
#### Scenario: Package metadata contains issue URL
- **WHEN** package metadata contains a repository-like URL ending in `/issues`
- **THEN** `package_git` discovery SHALL normalize it to the canonical repository URL when possible
#### Scenario: Package metadata contains repository URL variants
- **WHEN** package metadata contains `.git`, branch, tree, or homepage variants for the same repository
- **THEN** `package_git` discovery SHALL avoid queueing duplicate normalized repository targets
### Requirement: CI Seed Parsing From Existing Data
GitHub Actions and GitLab CI target discovery SHALL parse repository/project seeds from existing scanner DB records, package git candidates, findings, and target scan metadata where possible.
#### Scenario: Seed record contains package git target JSON
- **WHEN** a CI source examines a package git candidate or target scan record with repository metadata
- **THEN** it SHALL derive a GitHub repository or GitLab project seed when the URL provider matches the CI source
#### Scenario: Seed record cannot identify repository
- **WHEN** no repository or project can be derived from a seed record
- **THEN** the CI source SHALL count it as unparseable and continue processing other seeds
### Requirement: CI Scan Volume Is Configurable
CI source scan volume SHALL remain controlled by existing per-source configuration values.
#### Scenario: CI seed limit is raised
- **WHEN** `ci_seed_scan_limit` or `ci_max_repos_per_cycle` is increased in configuration
- **THEN** the CI source SHALL use the configured value without requiring code changes