Initial server source import

This commit is contained in:
sashatrask
2026-09-30 20:30:56 +03:00
commit 170dd941b9
498 changed files with 261563 additions and 0 deletions
+96
View File
@@ -0,0 +1,96 @@
# Runtime Reliability Parking Lot
Last reviewed: 2026-08-27
This file records observed follow-up work that is intentionally outside completed changes. Evidence must be refreshed before opening a new change.
## P1: GitLab discovery exits on one transient API timeout
- Status: implemented by `stabilize-gitlab-discovery-retries`; all 9 tasks complete and the change is ready for archive.
- Evidence: four GitLab source restarts during the current runtime window matched four `GET /api/v4/projects` read timeouts.
- Current behavior: direct API calls get one attempt; an exhausted discovery request escapes the source cycle and exits the supervised child.
- Impact: avoidable source churn and discovery delay. The saved query position is retained, so no confirmed target loss was observed.
- Candidate direction: bounded direct-request retries plus a source-cycle network disposition that does not terminate the child.
- Validation: inject timeout-then-success and timeout-exhaustion cases; verify query state, auth state, backoff, and no source restart.
## P1: GitLab scans have an incomplete RC=1 coverage gap
- Status: resolved by `stabilize-gitlab-trufflehog-lifecycle`; the production canary and bounded historical replay completed successfully.
- Evidence: 12 of 175 recent GitLab attempts exited RC=1 without a fatal diagnostic or `finished scanning` marker.
- Resolution evidence: 120 process attempts passed the rollout gate with 117 RC=0 completion markers and zero recurrence of the old signature. A three-target exact-signature replay completed cleanly on the first new attempt for every target; a further 30-minute soak reached 142 attempts and 139 RC=0 completion markers with the old signature still at zero. A later three-target sample produced two clean completions and one explicit retryable `trufflehog` deferral instead of the old terminal classification.
- Current behavior: all 12 were classified as non-retryable `command_exit` and became terminal failures after one attempt.
- Impact: these repositories have no confirmed complete scan.
- Candidate direction: run controlled Git `--local-dev` A/B tests, then consider extending explicit completion and bounded retry semantics beyond Docker.
- Validation: canary completion-marker rate, unexplained RC=1 rate, partial-finding preservation, retry exhaustion, and non-Git source isolation.
## P1: Runtime-wide external Windows termination has no autonomous recovery
- Evidence: at `2026-08-20 19:59:40 MSK`, PostgreSQL and unrelated source children were terminated together with Windows exception `0x40010004`; PostgreSQL shut down after its startup process also failed, and the supervisor was no longer alive to recover it.
- Impact: the pipeline remained offline until the next authority-checked `start_runtime.ps1` launch.
- Current state: PostgreSQL completed WAL recovery without corruption, all core sources and pipeline children restarted, and the post-recovery soak showed no child restarts or database failures.
- Candidate direction: run the authority-checked runtime under an external Windows Service or Task Scheduler watchdog that can restart a dead supervisor without weakening singleton or cluster-identity checks.
## Resolved: Updated core targets were permanently suppressed
- Status: resolved by `rescan-updated-core-targets`; the change is ready for archive.
- Evidence: the first production canary admitted three GitHub and three GitLab updated targets across six separate cycles, never exceeding the configured one-per-cycle cap. All six scans completed cleanly with exact remote revision snapshots, zero errors, zero findings, and applied queue completion.
- Isolation: DockerHub recorded no revision observations or updated admissions and continues to use digest identity. HuggingFace established its newest-modified revision baseline without an updated-target surge.
- Rollout decision: retain the cap at one per source cycle and the 24-hour per-target cooldown until longer-running yield data justifies a change.
## Resolved: Core discovery omitted an exact OpenAI query
- Status: resolved by `restore-openai-discovery-coverage`; the change is ready for archive.
- Evidence: the first bounded cycles fetched 44 GitHub and 42 GitLab results, admitted one changed GitLab target, and discovered 18 new DockerHub repository identities under the exact `openai` query. Source bounds remained one GitHub/GitLab page with at most five claims and two DockerHub pages with at most 20 claims.
- Funnel: all 18 DockerHub admissions reached terminal queue state (16 done, two explicit registry-access failures). Their scans produced 347 findings, including 53 OpenAI finding rows representing 24 distinct credentials. All 24 were genuinely new and received an authoritative API result: 20 quota/no-balance, four invalid/revoked, and zero alive.
- Durability: all 18 scan projections and 53 OpenAI keycheck projections completed with released capacity; the OpenAI candidate and publication backlogs drained to zero.
- Rollout decision: retain the bounded rotating query. It restored missing supply but did not produce a usable credential in the initial sample, so broader limits are not justified yet.
## Resolved: Bounded OpenAI ecosystem keyword wave
- Status: implemented by `expand-openai-ecosystem-discovery`; all nine source-specific queries completed one bounded production cycle.
- Git evidence: the three GitHub queries fetched 6, 0, and 3 results without new or updated admissions. GitLab `openai-api` and `openai-agents` admitted nothing; `librechat` admitted one bounded updated target whose scan completed cleanly without findings.
- Docker evidence: `librechat`, `lobechat`, and `openai-proxy` admitted 20, 19, and 16 immutable-digest targets. All 55 reached terminal queue state: 50 done and five explicit `auth_invalid` failures. Exact query scans produced 597 finding rows.
- Credential funnel: `librechat` produced five genuinely new credentials (GCP two, Gemini one, GitHub one, OpenAI one); `lobechat` produced two (AWS one, GitHub one); `openai-proxy` produced none. All seven received authoritative API outcomes and none was alive or usable. The OpenAI credential was explicitly invalid/revoked.
- Durability: every exact cohort candidate completed and all 147 cohort scan/keycheck projection jobs completed with released capacity.
- Rollout decision: retain the bounded wave provisionally without increasing any page, claim, worker, or global concurrency limit. Reassess after the next natural rotation; remove zero-yield terms if they remain empty rather than expanding this wave.
## P2: Heavy Docker images can exceed bounded resources
- Evidence: in the 197-attempt Docker canary, six scans reached the 600-second timeout, three hit explicit Windows `VirtualAlloc`/paging limits, and one had a mixed stream/timeout outcome.
- Current behavior: failures are explicit, partial findings are retained, and retryable outcomes use the bounded queue policy. The old unexplained RC=1 signature remained at zero.
- Impact: a small set of large images does not receive confirmed full coverage.
- Candidate direction: a separate heavy-image lane with lower concurrency and a larger target budget; do not raise global limits without host-level measurements.
- Validation: completion gain, host commit usage, scan-slot fairness, source throughput, and retry amplification.
## P2: Historical Docker replay remains intentionally gradual
- Evidence: the first exact-signature batch of 10 produced seven completed targets, one terminal registry-access failure, and two explicit deferred retries, with no old RC=1 recurrence. A second six-target batch produced four completed targets, one terminal registry-access failure, and one timeout deferral; all 41 emitted finding UIDs were globally unique and the timeout retained 35 partial findings.
- Current behavior: the historical terminal set has not been mass-requeued.
- Candidate direction: increase exact-signature batches conservatively while monitoring registry traffic and heavy-image failures.
## P2: GitHub manifests are not mined for Docker image references
- Evidence: 1,954 retained `github_archive_files` artifacts contained 158 unique image references, but only about 48 plausible namespaced Docker Hub repositories were incremental after filtering bases, local names, and existing queue identities.
- Current behavior: changed compose/workflow files are scanned for secrets only; `image:`, `FROM`, and `docker://` references do not feed Docker resolution. The bounded `github_archive_files` source is currently disabled.
- Impact: a small but potentially fresher source of Docker repositories is omitted. The measured one-time opportunity is much smaller than the existing unresolved Docker repository backlog.
- Candidate direction: after Docker resolver throughput is healthy, parse exact image references from changed compose, Kubernetes/Helm, workflow, and Dockerfile artifacts; resolve concrete tags directly to immutable platform digests and filter common base images.
- Validation: incremental repositories and digests, API calls per admitted target, source freshness, strict usable-key yield, duplicate rate, and GitHub quota cost.
## P2: Background supervisor launch nonce can parse as an option
- Evidence: one authority-checked launch on 2026-08-25 generated a URL-safe nonce beginning with an option-like prefix; `argparse` reported `argument --launch-nonce: expected one argument`. An immediate canonical retry succeeded.
- Impact: a rare transient startup refusal; no PostgreSQL or queue mutation occurred.
- Candidate direction: emit the hidden value as `--launch-nonce=<value>` and add a command-construction test with a leading-hyphen nonce.
## P3: GitLab removed/private repository churn
- Evidence: 10 recent clone-error scan events came from four targets; three targets exhausted all three retries.
- Current behavior: clone preparation errors are retryable even when a repository appears deleted, private, or deletion-scheduled.
- Impact: bounded but avoidable repeated work.
- Candidate direction: distinguish permanent not-found/access outcomes from transient clone transport failures before retrying.
## Needs Fresh Validation: Projector recovery/index path
- Earlier investigation suggested a rollback/missing-index weakness in projector recovery.
- Current state is healthy: ingester and projector report `ready`, with no active lease error.
- Before creating a change, reproduce or recover the original query/error evidence and determine whether the issue still exists.