56 lines
3.3 KiB
Markdown
56 lines
3.3 KiB
Markdown
## Context
|
|
|
|
GitLab project discovery calls `api_request` without an explicit retry budget. Direct requests therefore receive one attempt, and `ApiRequestError` escapes `run_configured_source`, which terminates the supervised child. The supervisor restarts the child and the persisted query position prevents confirmed target loss, but every transient 30-second read timeout creates avoidable churn and delay.
|
|
|
|
## Goals / Non-Goals
|
|
|
|
**Goals:**
|
|
|
|
- Retry idempotent GitLab project discovery requests with a small bounded budget.
|
|
- Keep the GitLab child alive when the bounded network budget is exhausted.
|
|
- Persist a failed source cycle without advancing the query or damaging auth state.
|
|
- Preserve existing rate-limit and authentication handling.
|
|
|
|
**Non-Goals:**
|
|
|
|
- Change TruffleHog scan commands or target retry policy.
|
|
- Retry non-idempotent requests.
|
|
- Hide persistent GitLab outages or loop without delay.
|
|
- Change other source APIs in this change.
|
|
|
|
## Decisions
|
|
|
|
### Configure attempts at the GitLab source boundary
|
|
|
|
GitLab source configuration will provide `discovery_request_attempts: 3` and `discovery_retry_delay: 5`. These values flow only into project discovery calls and use the existing `api_request` retry implementation.
|
|
|
|
Alternative considered: change the direct-request default globally. Rejected because it would silently alter every source and API call without source-specific evidence.
|
|
|
|
### Convert exhausted discovery transport errors into failed cycles
|
|
|
|
GitLab project discovery will wrap exhausted request transport failures in a dedicated `GitLabDiscoveryTransportError`. `run_configured_source` will catch only that error for GitLab, roll back any open database transaction, finish the source cycle as failed, retain the current query, and return to the normal configured-source cooldown. It will not mark the token invalid or terminate the process. Payload and programming errors retain fail-fast behavior.
|
|
|
|
Alternative considered: rely on supervisor restart. Rejected because process restart is expensive error handling for an ordinary transient network condition.
|
|
|
|
### Preserve rate-limit and auth paths
|
|
|
|
HTTP 429, 401, and 403 responses continue through the existing GitLab API and auth-pool policy. Only transport failures and the existing retryable HTTP status set use the discovery retry budget.
|
|
|
|
## Risks / Trade-offs
|
|
|
|
- [A persistent outage lengthens a failed cycle] -> Bound attempts to three and delay to five seconds.
|
|
- [A failed page could cause partial discovery ambiguity] -> Discard the cycle's fetched list on failure and retain the current query for the next cycle.
|
|
- [A catch could hide programming errors] -> Catch only `ApiRequestError` for the GitLab source; all other exceptions retain fail-fast behavior.
|
|
- [Auth state could be corrupted] -> Do not invoke token cooldown or invalidation for transport-only failures.
|
|
|
|
## Migration Plan
|
|
|
|
1. Add retry and failed-cycle tests using injected timeout responses.
|
|
2. Deploy the GitLab-only configuration and error handling.
|
|
3. Restart the managed runtime and observe at least one full query rotation or an injected exhaustion test.
|
|
4. Roll back by removing the GitLab source settings and narrow exception handler.
|
|
|
|
## Open Questions
|
|
|
|
- Whether the same direct-request policy should later be adopted by other sources requires separate evidence.
|