Files
2026-09-30 20:30:56 +03:00

5.1 KiB

Context

The OpenAI checker, candidate router, scheduler, and projection path are healthy, but the fresh input funnel collapsed after old target backlogs drained. In the last six days, core sources produced only 27 OpenAI findings and 11 new OpenAI credentials. A read-only exact-query probe found one unseen GitHub repository, 27 safely rescan-eligible GitLab projects, and 20 unseen repositories in the first 20 DockerHub results.

The current source configuration accepts one plain string query per cycle and applies one source-wide page/target policy. Adding openai without a query-specific Docker bound could resolve and enqueue up to 200 repositories in one cycle. Runtime source state is authority-managed and must not be edited casually.

Goals / Non-Goals

Goals:

  • Restore explicit OpenAI-oriented discovery in the three query-driven core sources.
  • Bound only the exact openai query while preserving every other query's current limits.
  • Make the first post-deployment cycle deterministic without directly modifying runner state.
  • Preserve queue, revision, digest, pipeline, and keycheck authority guarantees.
  • Measure whether the added coverage yields new OpenAI credentials and usable checks.

Non-Goals:

  • Reclassifying provider API outcomes or weakening the successful-generation requirement.
  • Re-enabling package_git or other broad inactive sources.
  • Rechecking known credentials, changing HuggingFace discovery, or increasing global scan concurrency.
  • Guaranteeing that newly discovered credentials are valid or funded.

Decisions

Add an exact provider query to existing rotations

GitHub, GitLab, and DockerHub each receive one literal openai query. The term remains a normal persisted rotation entry, so completed cycles advance naturally and failures retain the query under existing semantics.

For rollout, each entry is inserted at that source's current persisted query index while the runtime is coordinately stopped. The first restarted cycle therefore exercises openai without mutating state files; the next successful cycle advances to the query that previously occupied that index.

Alternative: append the term and wait for a full rotation. Rejected because Docker cycles can be long and the canary would be delayed and difficult to attribute.

Apply a small allowlisted query override

build_args_from_source_config() merges only pages, per_page, and max_targets from query_overrides.<exact-query>. Overrides are exact string matches, non-mutating, and validated before use. Unknown keys or non-mapping override shapes fail closed.

Initial bounds:

  • GitHub openai: one page, five scan claims.
  • GitLab openai: one page, five scan claims; changed-target promotion remains capped at one.
  • DockerHub openai: two ten-result pages, twenty scan claims.

The Docker page limit bounds discovery admission to at most 20 repositories before tag resolution. max_targets separately bounds claims in that source cycle. Other queries continue using their source-wide page and target values.

Alternative: temporarily reduce source-wide pages. Rejected because it would silently reduce all-provider coverage after rollback mistakes. Alternative: add a dedicated Docker source alias. Rejected because it would duplicate lifecycle and queue ownership code.

Preserve downstream authority and deduplication

The change ends at target discovery arguments. Existing normalized identity deduplication, revision-aware completed-target promotion, Docker digest resolution, result-bundle fencing, credential deduplication, cached-known occurrence handling, and API keycheck rules remain authoritative.

Risks / Trade-offs

  • [Exact search still produces many known targets] -> Persist separate new/updated counts and evaluate the full candidate funnel, not fetched volume.
  • [Docker results could create expensive scans] -> Bound search to 20 repositories and claims to 20 while retaining the global three-slot limit.
  • [Query override could accidentally alter unrelated settings] -> Allow only three numeric discovery/claim keys and test that ordinary queries retain source defaults.
  • [Config edit triggers immutable-authority shutdown] -> Use coordinated stop/start scripts and verify PostgreSQL, pipeline workers, and every core source after deployment.
  • [No valid OpenAI keys appear] -> Treat zero usable outcomes as yield evidence, not as proof of pipeline failure, provided candidates complete with explicit API outcomes.

Migration Plan

  1. Add query-override parsing and focused tests.
  2. Add exact queries and bounded overrides at each source's current persisted index.
  3. Run focused and full regression suites plus strict OpenSpec validation.
  4. Coordinately restart the authority-managed runtime.
  5. Observe exactly the first openai source cycle for GitHub, GitLab, and DockerHub; verify limits, queue admission, scan completion, candidate checks, and pipeline drain.
  6. Keep the query in normal rotation if safety bounds hold. Roll back by removing the query and overrides, then coordinately restart; no data migration is required.

Open Questions

None. Further query weighting or expansion depends on measured canary yield.