Files
2026-09-30 20:30:56 +03:00

88 lines
6.8 KiB
Markdown

## Context
GitHub, GitLab, and HuggingFace discovery return stable target identities together with remote update timestamps. The runner currently discards those timestamps and the PostgreSQL queue permanently deduplicates targets by `(source, normalized_target)`. A completed repository or Space is therefore never scanned again even when its content changes.
Production discovery is already saturated: less than one percent of fetched records are new target identities and the core queue is often empty. The change must restore changed-content coverage without turning every rediscovery into a rescan, growing one queue row per revision, or creating an unbounded backlog.
## Goals / Non-Goals
**Goals:**
- Preserve a bounded remote content-update signal for GitHub, GitLab, and HuggingFace targets.
- Rescan a completed target only when discovery observes content newer than the revision covered by its latest claim.
- Coalesce multiple remote updates into one mutable queue row and one pending follow-up.
- Bound changed-target admission and expose it separately from new-target admission.
- Roll out without treating every legacy row as changed.
**Non-Goals:**
- Requeue terminal failures or actively leased/deferred targets.
- Rescan every known target on a timer.
- Change DockerHub's digest-based identity and refresh behavior.
- Guarantee that provider activity timestamps always represent content changes.
- Enable inactive historical sources or enlarge global scan concurrency.
## Decisions
### Carry one normalized discovery record
The runner will represent eligible discovery results internally as a target plus an optional UTC `remote_modified_at`. GitHub uses `pushed_at`, GitLab uses `last_activity_at`, and HuggingFace uses `lastModified`. Missing, malformed, or non-monotonic timestamps remain valid discovery results but cannot trigger an updated-target rescan.
HuggingFace discovery will use a newest-modified feed so old Spaces changed recently are observable. Identity-only known-page stopping will be disabled when updated-target rescans are enabled; hard page and result limits remain the discovery bound.
Alternatives rejected:
- Repository `updated_at` on GitHub, because metadata-only edits are not content pushes.
- A HEAD-SHA request per repository, because it multiplies API traffic and rate-limit exposure.
- Revision-aware early stopping, because a known first page does not prove later pages contain no changed targets.
### Keep one queue row and two remote timestamps
`target_queue` will gain nullable `remote_modified_at` and `scan_remote_modified_at` columns. Discovery monotonically advances `remote_modified_at`. Claiming atomically copies the currently observed value into `scan_remote_modified_at`, recording what that scan covers.
The separate claim snapshot is required because a remote update can arrive while a scan is running. On completion, a newer observed timestamp remains ahead of the claimed timestamp and is eligible for exactly one later scan. Encoding revisions into `normalized_target` was rejected because it would grow queue rows and weaken queue authority.
For a legacy completed row with no claim snapshot, `completed_at` is the rollout baseline. It is eligible only when the first valid remote timestamp observed is newer than that completion. This prevents a migration surge while still admitting updates that occurred after the historical scan.
### Observe and requeue atomically under a hard budget
A PostgreSQL transaction will upsert discovery observations and requeue at most `updated_rescan_max_per_cycle` eligible rows for one source. Eligibility requires:
- `status='done'`;
- a strictly newer observed timestamp than `scan_remote_modified_at`, or than `completed_at` for a legacy row;
- elapsed `updated_rescan_cooldown_seconds` since completion;
- no lease, reservation, claim, or resolver authority.
Eligible rows are locked with `FOR UPDATE SKIP LOCKED`. Requeue resets only retry/completion scheduling fields required for a normal pending claim. Failed, pending, deferred, in-progress, unresolved, unchanged, and invalid-timestamp rows are never promoted by this path.
The initial production setting is one updated target per source cycle. Existing backlog-first behavior remains enabled, so a source drains its admitted work before discovery can admit more; `refresh_registry` is not enabled by this change.
Alternatives rejected:
- Global `requeue_done=True`, because it requeues unchanged rows on every cycle.
- Comparing only remote time with local completion time forever, because provider and host clocks differ and a mid-scan update can be lost.
- A separate maintenance queue, because it duplicates existing lease, reservation, and completion authority.
### Account for updated targets separately
`source_cycles` will gain `queued_updated_count`. `queued_new_count` keeps its current meaning. Source logs and dashboard aggregation will report changed-target admissions separately so rollout volume and yield can be audited.
## Risks / Trade-offs
- [GitLab activity can change without a repository push] -> Use a one-target-per-cycle cap and cooldown; report updated admissions separately.
- [Provider clock skew] -> Require strict monotonicity and use the claimed remote timestamp after the first revision-aware scan.
- [More discovery API traffic after disabling identity-only early stop] -> Keep existing hard page/per-page limits and source intervals.
- [Changed targets consume capacity without useful findings] -> Start at one per cycle and compare updated-target yield before increasing the cap.
- [Schema rollout while runtime is active] -> Stop the authority-managed runtime, apply schema through the normal initialization path, run tests, then restart through `start_runtime.ps1`.
- [Rollback leaves nullable columns] -> Disable the feature in source configuration; nullable columns and metrics are backward-compatible and can remain.
## Migration Plan
1. Add nullable queue columns, the source-cycle counter, and a partial eligibility index through idempotent schema initialization.
2. Deploy code and tests with updated-target rescans disabled by default.
3. Enable GitHub, GitLab, and HuggingFace with a cap of one and a conservative cooldown; leave DockerHub unchanged.
4. Restart the authority-managed runtime and verify discovery, queue authority, projection, and source health.
5. Observe changed-target admission, completion, findings, and worker occupancy before changing any cap.
Rollback is configuration-first: disable updated-target rescans and restart the managed runtime. Existing pending work completes under normal queue semantics; no destructive data migration is required.
## Open Questions
- Whether production evidence supports different cooldowns per source after the initial canary.
- Whether a later change should add provider-specific immutable revisions when APIs can supply them without extra requests.