Initial server source import
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-08-26
|
||||
@@ -0,0 +1,78 @@
|
||||
## Context
|
||||
|
||||
The exact `openai` rollout proved that bounded source queries can restore unseen credential supply without changing scanner or keycheck authority: 18 DockerHub targets produced 24 genuinely new OpenAI credentials, all with explicit terminal API outcomes. None were usable, so increasing that broad query is not justified. Historical production attribution instead shows usable OpenAI outcomes behind narrower agent, chatbot, and conversation ecosystems, while source semantics differ enough that one shared keyword list is inefficient.
|
||||
|
||||
The existing exact-query override mechanism already validates and applies `pages`, `per_page`, and `max_targets`. This change can therefore remain configuration-only plus focused contract tests. Runtime configuration is immutable-authority covered, so deployment must use coordinated stop/start and must not modify persisted query state directly.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Test nine high-signal OpenAI integration and deployable-application queries across the sources where their search semantics fit.
|
||||
- Bound discovery admission as well as scan claims for every new query.
|
||||
- Run each new query once promptly after deployment without directly editing query state.
|
||||
- Attribute the canary through durable scans, candidates, provider results, and projections.
|
||||
- Keep or remove each query based on its own measured useful yield and operational cost.
|
||||
|
||||
**Non-Goals:**
|
||||
- Expanding broad generic terms such as `gpt`, `llm`, `chatgpt`, `ai`, or `model`.
|
||||
- Increasing global scan concurrency, source worker counts, or updated-target promotion limits.
|
||||
- Changing detector routing, OpenAI checker classification, known-credential caching, or projection semantics.
|
||||
- Re-enabling inactive broad sources or guaranteeing a valid funded credential.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Use source-specific first-wave queries
|
||||
|
||||
The first wave is:
|
||||
- GitHub: `OPENAI_API_KEY`, `api.openai.com`, `openai-agents`.
|
||||
- GitLab: `openai-api`, `openai-agents`, `librechat`.
|
||||
- DockerHub: `librechat`, `lobechat`, `openai-proxy`.
|
||||
|
||||
GitHub can search README content, so direct environment and endpoint signatures are appropriate. GitLab project search is metadata-oriented, so branded slug terms are used. DockerHub searches repository metadata and then resolves immutable image digests, so deployable project and proxy names are used.
|
||||
|
||||
Alternative: add the same list to all sources. Rejected because it lengthens every rotation and ignores source search semantics. Alternative: expand historically broad terms. Rejected because those cohorts produced volume without usable OpenAI outcomes.
|
||||
|
||||
### Bound every query independently
|
||||
|
||||
Bounds are:
|
||||
- GitHub `OPENAI_API_KEY`: `pages=1`, `per_page=25`, `max_targets=5`.
|
||||
- GitHub `api.openai.com`: `1/25/5`.
|
||||
- GitHub `openai-agents`: `1/50/5`.
|
||||
- GitLab `openai-api`: `1/50/5`.
|
||||
- GitLab `openai-agents`: `1/50/5`.
|
||||
- GitLab `librechat`: `1/25/5`.
|
||||
- DockerHub `librechat`, `lobechat`, and `openai-proxy`: each `2/10/10`.
|
||||
|
||||
Page and page-size limits bound fetched/admitted identities; `max_targets` separately bounds claims in the active cycle. The existing global three-slot limit, GitLab one-updated-target-per-cycle cap, 24-hour update cooldown, and Docker digest requirement remain unchanged.
|
||||
|
||||
### Place the wave at each stopped source's current rotation index
|
||||
|
||||
After a coordinated runtime stop, capture each source's persisted `query_index` and insert that source's three-query block at the same index. Restarting then exercises the block naturally. Successful discovery advances through the block; failures and backlog-only work retain the current query under existing semantics. State files are never edited.
|
||||
|
||||
Alternative: append and wait for a full rotation. Rejected because DockerHub cycles can be long and attribution would be delayed. Alternative: edit persisted state. Rejected because state is runtime authority and direct edits would weaken recovery evidence.
|
||||
|
||||
### Evaluate individual query funnels
|
||||
|
||||
The canary reports fetched, new/updated admissions, durable scan outcomes, findings, genuinely new OpenAI credentials, `api_check` versus `cached_status`, explicit provider outcomes, first-alive/usable counts, projection drain, and runtime health. Aggregate volume alone cannot justify retention.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [README signatures produce placeholders] -> Count genuinely new credentials and explicit API outcomes; remove queries with only placeholder/dead yield.
|
||||
- [A query admits more work than its claim cap] -> Keep `pages × per_page` small and inspect the exact query-attributed queue until terminal.
|
||||
- [Docker scans are expensive or inaccessible] -> Cap each query at 20 repositories and 10 claims, retain digest authority, and classify registry failures separately.
|
||||
- [Three added terms lengthen source rotations] -> Keep only terms that add distinct credential or usable yield after the canary.
|
||||
- [Config deployment triggers authority fail-close] -> Stop coordinately before editing and restart only through `start_runtime.ps1`.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add focused tests for exact source membership, uniqueness, and all nine bounds.
|
||||
2. Coordinately stop the runtime and capture persisted source indices.
|
||||
3. Insert each three-query block at its source's current index; do not edit state files.
|
||||
4. Run focused and full regression suites plus strict OpenSpec validation with bytecode writes disabled.
|
||||
5. Start through the authoritative runtime script and verify PostgreSQL, pipeline, core sources, keychecks, and recorder.
|
||||
6. Observe each exact query cohort through terminal scans and keycheck projection, then retain or remove each query based on measured evidence.
|
||||
7. Roll back any low-value query by removing that query and override during a coordinated stop/start. No schema or data rollback is required.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None. A second wave remains gated on this canary's per-query useful-yield evidence.
|
||||
@@ -0,0 +1,24 @@
|
||||
## Why
|
||||
|
||||
The bounded exact `openai` canary restored fresh credential supply but produced no usable credentials, while historical production evidence shows that narrower ecosystem and integration terms can reach different cohorts. A small source-specific keyword wave can test those higher-signal surfaces without expanding broad generic discovery or scan concurrency.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add three source-specific OpenAI ecosystem queries to each of GitHub, GitLab, and DockerHub.
|
||||
- Apply exact per-query page, page-size, and claim bounds so every new query is independently constrained.
|
||||
- Preserve existing query rotation, target deduplication, revision-aware rescan limits, Docker digest authority, scan concurrency, and keycheck behavior.
|
||||
- Run one controlled production canary per new query and measure discovery, scans, new OpenAI credentials, explicit API outcomes, and usable yield.
|
||||
- Retain, revise, or remove individual queries based on measured bounded evidence rather than fetched volume.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `openai-ecosystem-discovery`: Source-specific bounded discovery for OpenAI integration signatures and deployable ecosystem projects, with per-query canary evidence.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None.
|
||||
|
||||
## Impact
|
||||
|
||||
The change affects `app/config.yaml`, focused query-configuration tests, authority-managed source rotation, and production canary operations. Existing allowlisted query override code is reused unchanged. There is no schema migration, new dependency, detector change, global concurrency increase, or credential recheck policy change.
|
||||
+73
@@ -0,0 +1,73 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Source-specific OpenAI ecosystem queries
|
||||
The system SHALL include the approved first-wave OpenAI ecosystem queries only in the source rotations whose search semantics match those queries.
|
||||
|
||||
#### Scenario: GitHub searches integration signatures
|
||||
- **WHEN** GitHub reaches the first-wave positions in its normal rotation
|
||||
- **THEN** it SHALL search `OPENAI_API_KEY`, `api.openai.com`, and `openai-agents`
|
||||
|
||||
#### Scenario: GitLab searches branded project metadata
|
||||
- **WHEN** GitLab reaches the first-wave positions in its normal rotation
|
||||
- **THEN** it SHALL search `openai-api`, `openai-agents`, and `librechat`
|
||||
|
||||
#### Scenario: DockerHub searches deployable ecosystems
|
||||
- **WHEN** DockerHub reaches the first-wave positions in its normal rotation
|
||||
- **THEN** it SHALL search `librechat`, `lobechat`, and `openai-proxy`
|
||||
|
||||
#### Scenario: Broad generic expansion is excluded
|
||||
- **WHEN** the first-wave configuration is evaluated
|
||||
- **THEN** it SHALL NOT add new broad variants of `gpt`, `llm`, `chatgpt`, `ai`, or `model`
|
||||
|
||||
### Requirement: Independent bounded query policies
|
||||
Each first-wave query SHALL have an exact allowlisted override that bounds both discovery volume and scan claims without changing source defaults or other query policies.
|
||||
|
||||
#### Scenario: GitHub signature queries are bounded
|
||||
- **WHEN** GitHub builds arguments for `OPENAI_API_KEY` or `api.openai.com`
|
||||
- **THEN** it SHALL use one page of 25 results and claim at most five targets
|
||||
|
||||
#### Scenario: GitHub agent query is bounded
|
||||
- **WHEN** GitHub builds arguments for `openai-agents`
|
||||
- **THEN** it SHALL use one page of 50 results and claim at most five targets
|
||||
|
||||
#### Scenario: GitLab queries are bounded
|
||||
- **WHEN** GitLab builds arguments for a first-wave query
|
||||
- **THEN** it SHALL use one page, the configured 25- or 50-result page size, and claim at most five targets
|
||||
|
||||
#### Scenario: DockerHub queries are bounded
|
||||
- **WHEN** DockerHub builds arguments for a first-wave query
|
||||
- **THEN** it SHALL use at most two pages of ten repositories and claim at most ten targets
|
||||
|
||||
#### Scenario: Existing authority limits remain unchanged
|
||||
- **WHEN** any first-wave query runs
|
||||
- **THEN** global scan concurrency, revision-aware promotion limits, cooldowns, and Docker digest requirements SHALL remain authoritative
|
||||
|
||||
### Requirement: Natural rotation and failure behavior
|
||||
The first-wave queries SHALL use the existing persisted source rotation without direct query-state modification.
|
||||
|
||||
#### Scenario: Successful query advances
|
||||
- **WHEN** a first-wave discovery cycle completes successfully
|
||||
- **THEN** the source SHALL advance through the existing persisted rotation semantics
|
||||
|
||||
#### Scenario: Failed query is retained
|
||||
- **WHEN** first-wave discovery fails before successful completion
|
||||
- **THEN** the source SHALL retain that query according to existing failure semantics
|
||||
|
||||
#### Scenario: Backlog work does not masquerade as discovery
|
||||
- **WHEN** a source drains existing backlog while a first-wave query is current
|
||||
- **THEN** canary attribution SHALL distinguish backlog-only cycles from the actual discovery cycle
|
||||
|
||||
### Requirement: Per-query end-to-end canary evidence
|
||||
Operators SHALL evaluate each first-wave query through durable source, scan, candidate, provider-result, and projection evidence without exposing targets or credential values.
|
||||
|
||||
#### Scenario: Query cohort reaches terminal accounting
|
||||
- **WHEN** a first-wave query admits new or updated targets
|
||||
- **THEN** operators SHALL verify queue dispositions, scan completion, candidate completion, and projection drain for that exact query cohort
|
||||
|
||||
#### Scenario: Useful yield is measured separately
|
||||
- **WHEN** a first-wave cohort creates OpenAI candidates
|
||||
- **THEN** genuinely new credentials and their explicit API outcomes SHALL be reported separately from cached-known occurrences
|
||||
|
||||
#### Scenario: Retention decision uses measured value
|
||||
- **WHEN** the bounded canary is complete
|
||||
- **THEN** each query SHALL be retained, revised, or removed using its distinct credential yield, usable outcomes, and operational error cost rather than fetched count alone
|
||||
@@ -0,0 +1,17 @@
|
||||
## 1. Source-specific configuration
|
||||
|
||||
- [x] 1.1 Coordinately stop the runtime, capture persisted source query indices, and insert each three-query block at its current index.
|
||||
- [x] 1.2 Add exact allowlisted bounds for all nine queries without changing source-wide defaults or existing safety policies.
|
||||
- [x] 1.3 Add focused configuration tests for source membership, uniqueness, exact bounds, and excluded broad expansion.
|
||||
|
||||
## 2. Verification
|
||||
|
||||
- [x] 2.1 Run focused and full regression suites with bytecode writes disabled.
|
||||
- [x] 2.2 Run strict OpenSpec validation and verify implementation against the artifacts.
|
||||
|
||||
## 3. Production canary
|
||||
|
||||
- [x] 3.1 Start the authority-managed runtime and verify PostgreSQL, pipeline, recorder, keychecks, and all core sources.
|
||||
- [x] 3.2 Observe one actual discovery cycle for each first-wave query and verify configured fetch and claim bounds.
|
||||
- [x] 3.3 Follow every exact query cohort through queue disposition, scan/candidate completion, provider outcomes, and projection drain.
|
||||
- [x] 3.4 Record per-query useful-yield evidence, retain or remove low-value terms, and update the parking lot.
|
||||
Reference in New Issue
Block a user