Initial server source import

This commit is contained in:
sashatrask
2026-09-30 20:30:56 +03:00
commit 170dd941b9
498 changed files with 261563 additions and 0 deletions
@@ -0,0 +1,81 @@
## Context
The earlier global pruning pass removed 38 terms with zero strict-usable yield but deliberately preserved their existing backlog. The later cold-policy change added an audited reversible hold and applied it only to the retired Docker cohort. The current canonical configuration still has 324 exact source/query pairs across 95 query texts.
A fresh PostgreSQL dry-run attributed scans and candidates through immutable `target_scans.query` and `keycheck_candidates.query`, then linked candidates by `credential_id` to every historical `keycheck_results.status_group='alive'` result. After excluding dedicated source sentinels, 38 exact source/query pairs each have at least 1,000 successful scans, zero ever-alive credentials, and zero pending candidate checks. They account for 114,226 successful scans and 75,551 currently unfenced pending/deferred queue rows.
## Goals / Non-Goals
**Goals:**
- Retire source/query pairs only after substantial completed exposure and fully matured zero-alive evidence.
- Preserve a visible canonical rejection record with the evidence and decision reason.
- Stop future discovery and existing claimable work for the rejected pairs.
- Preserve every historical and queue authority record and make queue holds reversible.
- Apply the decision without racing workers or partially transitioning a reviewed cohort.
**Non-Goals:**
- Delete queue rows, scans, findings, candidates, credentials, keycheck results, reservations, or coverage.
- Treat source sentinels such as `gharchive`, `gharchive-files`, `gists`, `logs`, or `spaces` as discovery keywords.
- Retire a query globally because it failed in one source.
- Automatically reevaluate or reactivate rejected pairs during normal runtime.
- Claim that a rejected pair can never become productive in the future.
## Decisions
### Evaluate exact source/query pairs
The decision unit is the case-sensitive exact `(source, query)` pair. A keyword that produced an alive credential in DockerHub does not justify retaining a large npm backlog when the npm pair itself has substantial zero-alive evidence.
Alternative: retire only globally zero-alive query text. Rejected because the dry-run would hold only 3,626 rows and preserve most demonstrated source-specific waste.
### Require 1,000 successful scans and mature keychecks
A pair qualifies only when it has at least 1,000 ended scans with status `clean`, `found`, or `degraded`, no credential linked through any of its candidate occurrences has ever had a historical `status_group='alive'` result, and no candidate for the pair remains pending or leased. Historical alive status is intentionally used instead of only current status so a once-working credential permanently proves yield.
The 1,000-scan floor is deliberately stricter than a 100-scan cut. The latter selected 131 pairs and 187,917 rows but is too weak for rare useful credentials. Pair attribution uses immutable scan/candidate query snapshots for evidence; queue transition uses the row's current exact query because that is the durable admission attribution.
Alternative: use raw finding count, unique candidates, current status only, or the dashboard strict-usable tier. Rejected because the operator requested actual `alive`, findings do not prove provider utility, current-only status forgets historical success, and pending checks make zero yield unresolved.
### Preserve a canonical rejected registry
Each removed pair remains under canonical `query_policy.rejected` with status `rejected_zero_alive`, evidence cutoff, successful scan count, finding count, unique credential count, pending count, ever-alive count, and reviewed queue count. Active query lists and rejected entries must be disjoint. The registry is evidence and operator visibility; existing append-only queue policy events remain the state-transition audit.
Alternative: leave comments beside removed YAML entries. Rejected because comments are not machine-checkable and cannot fence future accidental reintroduction.
### Use the existing cold lifecycle
After configuration removal, the existing hash-fenced stopped-source manifest flow selects only unfenced `pending` or `deferred` rows whose exact query is absent from active policy. Each selected row becomes `cold`; no target value appears in review output, and no other row field or linked record changes. Active/fenced rows make apply fail closed. Reactivation remains possible only through the existing reviewed reverse action.
### Approve the exact cohort
- DockerHub: `OR`, `agent`.
- GitHub: `coding`, `memory`.
- npm: `OR`, `agent`, `agents`, `ai`, `assistant`, `benchmark`, `bot`, `chat`, `chats`, `completion`, `completions`, `conversation`, `gemini`, `groq`, `langchain`, `llm`, `mcp`, `open`, `openrouter`, `prompt`, `rag`, `semantic`, `studio`, `xai`.
- package-git: `agent`, `bot`, `completion`, `llm`, `open`, `semantic`, `studio`.
- Postman: `XAI_API_KEY`.
- PyPI: `langchain`, `open`.
GitLab has no pair meeting the 1,000-scan rule. Dedicated source sentinel queries remain untouched.
## Risks / Trade-offs
- [Rare future yield is lost] -> Preserve evidence, history, and reversible cold events; reactivation requires explicit review.
- [Current queue query may differ from an older scan query after rediscovery] -> Use immutable scan/candidate attribution for yield evidence, exact current queue attribution for holding, and preserve append-only reversal evidence.
- [A candidate becomes alive between dry-run and apply] -> Regenerate and verify evidence immediately before apply; fail if the approved zero-alive cohort drifts.
- [A worker owns a selected row] -> Stop sources and fail the entire manifest on any queue, resolver, reservation, or blob fence.
- [Configuration accidentally reintroduces a rejected pair] -> Validate active/rejected disjointness in configuration tests and policy loading.
## Migration Plan
1. Add canonical rejected-query evidence and exact configuration tests for the approved 38-pair cohort.
2. Verify runtime sources are stopped and cleanly stop the current maintenance database before changing authority-covered configuration.
3. Remove each rejected pair from its active source list and retain it in `query_policy.rejected` with reviewed evidence.
4. Restart maintenance PostgreSQL, rerun the zero-alive evidence query, and require the exact cohort and queue counts to match the reviewed decision.
5. Generate bounded private cold manifests per affected source/platform and apply them through the existing cluster lock and stopped-source action.
6. Verify 75,551 rows became cold, no selected row remains claimable, history counts are unchanged, and policy events reconcile exactly.
7. Cleanly stop maintenance PostgreSQL, start the canonical runtime, and verify retained queries, pipeline workers, sources, and queue movement.
8. Roll back through the existing reviewed reactivation manifests plus restoration of the active query entries.
## Open Questions
None.