Initial server source import

This commit is contained in:
sashatrask
2026-09-30 20:30:56 +03:00
commit 170dd941b9
498 changed files with 261563 additions and 0 deletions
@@ -0,0 +1,71 @@
## ADDED Requirements
### Requirement: Separate assignment and scan outcomes
The worker administration list SHALL display assignment transport outcome, scan outcome, and diagnostic count as separate fields and SHALL not label `last_error_code` as the complete scan error category.
#### Scenario: Accepted scan has an error outcome
- **WHEN** a reservation has a durable accepted bundle whose target scan status is `error`
- **THEN** the list SHALL show assignment `accepted`, scan `error`, and the diagnostic count/categories
#### Scenario: Assignment expires before scan ingestion
- **WHEN** a reservation expires without an accepted bundle
- **THEN** the list SHALL show assignment `expired`, scan outcome unavailable, and the expiry diagnostic
### Requirement: Assignment detail timeline
Each worker assignment row SHALL link to a detail page that reconstructs issued, phase-progress, bundle/terminal report, receipt, ingestion, queue settlement, and projection timestamps that exist for that assignment.
#### Scenario: Administrator opens an active assignment
- **WHEN** progress events exist for an unresolved assignment
- **THEN** the page SHALL show current phase, phase age, last progress age, scan deadline, assignment deadline, and ordered prior phases
#### Scenario: Administrator opens a settled assignment
- **WHEN** the assignment has been accepted and projected
- **THEN** the timeline SHALL distinguish acceptance, ingestion, queue settlement, and projection completion rather than collapsing them into one completion time
### Requirement: Clickable diagnostic detail
The assignment detail page SHALL list diagnostics and SHALL provide human summary, canonical envelope JSON, raw body view, process log view, transformation metadata, and copy/download actions for each diagnostic.
#### Scenario: Diagnostic body is complete
- **WHEN** an HTTP diagnostic contains an untruncated body
- **THEN** the raw-body view SHALL identify it as complete and display the captured content and metadata
#### Scenario: Diagnostic material is truncated
- **WHEN** body or log material was size-truncated
- **THEN** the view SHALL prominently display original/stored sizes, hash, and truncation state
#### Scenario: Legacy scan error has no diagnostic envelope
- **WHEN** an older scan has only existing `errors.raw_error` or result metadata
- **THEN** the detail page SHALL display those fields as legacy evidence and SHALL not invent a new envelope
### Requirement: Diagnostic filtering and grouping
The admin UI SHALL filter independently by source, time, worker/device, assignment outcome, scan outcome, phase, category, stable code, and retryability and SHALL group repeated diagnostic fingerprints without hiding individual occurrences.
#### Scenario: Administrator filters rate-limit errors
- **WHEN** category `rate_limit` and a time window are selected
- **THEN** results SHALL include matching diagnostics regardless of whether their assignments were accepted or prebundle-failed
#### Scenario: Repeated diagnostics are grouped
- **WHEN** multiple diagnostics share a fingerprint
- **THEN** the UI SHALL show aggregate count and affected assignments while retaining links to each occurrence
### Requirement: Worker fleet status
The admin UI SHALL display each worker's configured cap, active slots, current phases, package identity, latest contact/progress ages, pending local-recovery indication when reported, and known idle/backoff reason.
#### Scenario: Worker is scanning without recent API contact
- **WHEN** a worker has an active assignment and recent progress events but its authentication contact timestamp is old
- **THEN** fleet status SHALL show active progress rather than classifying the worker as idle solely from contact age
#### Scenario: Worker cannot claim due to capacity
- **WHEN** the server rejects claims because a pipeline capacity axis is closed
- **THEN** fleet status SHALL show the capacity reason instead of a generic offline/idle state
### Requirement: Deadline and duration administration
The runtime editor and worker observability pages SHALL explain effective scan, upload, and assignment deadlines and SHALL show p50/p95/p99 duration metrics with sample counts by source and phase.
#### Scenario: Administrator edits a source assignment deadline
- **WHEN** a per-source TTL candidate is previewed
- **THEN** the editor SHALL show the effective policy, validation relationship to scan/upload bounds, and that only future assignments are affected
#### Scenario: Administrator compares policy to observations
- **WHEN** sufficient phase-duration samples exist
- **THEN** the page SHALL show the configured deadline alongside source-specific percentile values without automatically changing configuration
@@ -0,0 +1,86 @@
## ADDED Requirements
### Requirement: Unified versioned diagnostic envelope
Worker scan errors, provider failures, process failures, local exceptions, timeouts, prebundle failures, and assignment expiry context SHALL use one versioned diagnostic envelope with independent phase, kind, category, stable code, summary, retryability, attempt, timestamps, and optional HTTP/process/exception material.
#### Scenario: Provider returns an HTTP error
- **WHEN** a provider operation receives an unsuccessful HTTP response
- **THEN** the diagnostic SHALL identify the phase, provider operation, HTTP status/content type, stable category/code, retryability, and captured response material
#### Scenario: Scanner process fails
- **WHEN** a scanner process exits unsuccessfully or is terminated at its deadline
- **THEN** the diagnostic SHALL identify its process result, timeout/signal state, phase, stable category/code, and captured log material
#### Scenario: Assignment expires without a worker result
- **WHEN** the server expires an unresolved assignment
- **THEN** it SHALL create or expose a diagnostic describing assignment expiry and the last accepted progress phase without claiming a scanner error occurred
### Requirement: Diagnostic fidelity and explicit transformation
Captured body and log bytes SHALL be preserved without silent semantic rewriting. Every size limit, encoding conversion, or truncation SHALL record original bytes, stored bytes, content hash, encoding, and truncation state.
#### Scenario: Text body fits the bound
- **WHEN** a captured provider response body fits the configured diagnostic body bound
- **THEN** the transmitted diagnostic SHALL contain the complete captured text and SHALL mark it untruncated
#### Scenario: Body exceeds the bound
- **WHEN** captured body bytes exceed the transmitted bound
- **THEN** the diagnostic SHALL carry the bounded material plus original/stored sizes, full captured-content hash when available, and `truncated=true`
#### Scenario: Body is not text
- **WHEN** captured diagnostic body bytes are not valid text in the declared encoding
- **THEN** the envelope SHALL use an explicit binary encoding representation and SHALL preserve the same transformation metadata
### Requirement: Bounded diagnostic transport
The protocol SHALL enforce deterministic per-body, per-log, per-envelope, diagnostic-count, and aggregate diagnostic bounds while rejecting envelopes whose declared and actual sizes disagree.
#### Scenario: Accepted bundle contains diagnostics
- **WHEN** a worker uploads a result bundle with diagnostic frames within all bounds
- **THEN** bundle acceptance and ingestion SHALL validate and persist each diagnostic idempotently with the scan
#### Scenario: Diagnostic aggregate exceeds its limit
- **WHEN** a bundle or terminal report exceeds a diagnostic count or byte limit
- **THEN** the API SHALL reject it with a stable protocol error and SHALL NOT partially persist diagnostics
### Requirement: Prebundle and accepted-result parity
The same diagnostic envelope SHALL be usable in prebundle terminal reports and accepted scan-result bundles, with only transport-size profiles differing.
#### Scenario: Worker storage fails before bundle creation
- **WHEN** the worker cannot create a result bundle
- **THEN** its terminal report SHALL include a diagnostic envelope rather than replacing the exception with one generic fixed detail string
#### Scenario: Scan returns errors in a valid bundle
- **WHEN** scanning completes with structured errors and a valid bundle
- **THEN** those errors SHALL be represented as diagnostics attached to the ingested target scan and SHALL remain distinct from assignment transport outcome
### Requirement: Deterministic diagnostic identity
Each diagnostic SHALL have a deterministic UID derived from its canonical identity and content so retries and replay cannot create duplicates.
#### Scenario: Accepted upload is replayed
- **WHEN** an identical accepted result bundle is uploaded again
- **THEN** the server SHALL return the durable receipt and SHALL NOT insert duplicate diagnostic rows
#### Scenario: Same code occurs twice in one assignment
- **WHEN** two distinct occurrences share category and code but differ in occurrence identity or content
- **THEN** both SHALL be retained as distinct diagnostics with stable UIDs
### Requirement: Local diagnostic archive
The worker SHALL retain a queryable local JSON diagnostic envelope and optional body/log artifacts per assignment, with configurable age/byte rotation and explicit artifact-availability state in history.
#### Scenario: Operator opens a local diagnostic
- **WHEN** `truf-worker history` or `logs` selects a retained diagnostic
- **THEN** the worker SHALL present the canonical envelope and exact paths/availability of its body and log artifacts
#### Scenario: Artifact rotates out
- **WHEN** a body or log artifact is removed by configured local rotation
- **THEN** terminal history SHALL remain and SHALL state that the artifact is no longer locally retained
### Requirement: Orthogonal error taxonomy
The diagnostic model SHALL keep phase, kind, broad category, stable code, retryability, assignment outcome, and scan outcome as separate dimensions.
#### Scenario: Accepted scan has provider errors
- **WHEN** a result bundle is durably accepted but the scan outcome is `error`
- **THEN** the assignment outcome SHALL remain `accepted`, scan outcome SHALL be `error`, and provider diagnostics SHALL retain their own categories/codes
#### Scenario: Assignment expires
- **WHEN** an assignment expires before bundle acceptance
- **THEN** assignment outcome SHALL be `expired`, scan outcome SHALL be unavailable, and the expiry diagnostic SHALL not be categorized as a provider scan failure
@@ -0,0 +1,74 @@
## ADDED Requirements
### Requirement: Unified worker lifecycle CLI
The worker package SHALL provide `install`, `run`, `start`, `stop`, `status`, `attach`, `logs`, `history`, and `doctor` commands with equivalent lifecycle semantics on supported Windows and Linux/WSL platforms.
#### Scenario: Operator starts a detached worker
- **WHEN** an installed operator invokes `truf-worker start`
- **THEN** the command SHALL launch the worker supervisor, wait for its startup handshake, and return the verified instance identity and current state
#### Scenario: Operator runs in the foreground
- **WHEN** an operator invokes `truf-worker run`
- **THEN** the same supervisor implementation SHALL run in the foreground and SHALL begin graceful drain on the first interrupt
### Requirement: Verified detached supervisor lifecycle
The supervisor SHALL publish a versioned instance record, status projection, local control endpoint, startup result, and shutdown receipt tied to the exact running process identity.
#### Scenario: Status finds a stale instance record
- **WHEN** the recorded process no longer matches the recorded executable, creation identity, or live control handshake
- **THEN** `status` SHALL report the instance as stale and SHALL NOT represent it as a running worker
#### Scenario: Graceful stop has active slots
- **WHEN** `stop` is requested while one or more slots own assignments
- **THEN** the supervisor SHALL stop new claims, display the draining slots, and wait for terminal local reconciliation up to the requested stop deadline
### Requirement: Attachable live operator view
The supervisor SHALL provide an `attach` session that renders current worker and per-slot state and follows new events without making attachment own the worker lifetime.
#### Scenario: Operator detaches
- **WHEN** the operator presses `q`, sends EOF, or interrupts the attach client
- **THEN** only the attach session SHALL end and the worker supervisor SHALL continue running
#### Scenario: Concurrent slots update
- **WHEN** multiple slots emit interleaved phase events
- **THEN** attach SHALL render one coherent row per slot and SHALL preserve event ordering by local sequence
### Requirement: Honest per-slot status
Status and attach SHALL display source, phase, phase elapsed time, scan deadline, assignment time remaining, last progress age, and only counters measured by the execution path. They SHALL NOT synthesize percentage completion without a reliable denominator.
#### Scenario: Long scanner execution
- **WHEN** a slot remains in `scanning` with a live runner process
- **THEN** status SHALL continue updating elapsed time, deadline remaining, and last-progress age rather than appearing frozen
#### Scenario: Worker has no assignment
- **WHEN** a slot is idle because of server backoff, cap, paused dispatch, capacity, or an empty queue
- **THEN** status SHALL report the known idle/backoff reason and next claim time when supplied by the server
### Requirement: Human and machine output contracts
Every non-interactive inspection command SHALL support versioned JSON output, and every follow command SHALL support versioned NDJSON output containing no human decoration.
#### Scenario: Automation requests status
- **WHEN** `truf-worker status --json` is invoked
- **THEN** stdout SHALL contain exactly one parseable versioned status object representing the same state as the human view
#### Scenario: Automation follows events
- **WHEN** `truf-worker logs --follow --json` is invoked
- **THEN** stdout SHALL contain one complete versioned event object per line in sequence order
### Requirement: Local history and operational diagnosis
The supervisor SHALL retain terminal assignment history, rotating worker logs, event history, diagnostic references, and local storage usage, and SHALL expose them through `history`, `logs`, and `doctor`.
#### Scenario: Operator investigates a completed assignment
- **WHEN** the operator requests history for a terminal reservation
- **THEN** the worker SHALL show its terminal local/receipt outcome, durations, phase timeline, and available diagnostic artifact references
#### Scenario: Operator runs doctor
- **WHEN** `truf-worker doctor` is invoked
- **THEN** it SHALL inspect package identity, singleton/process state, local state readability, disk usage, server reachability, and retained work without claiming an assignment
### Requirement: From-zero operator documentation
The release SHALL include one canonical guide from package acquisition through installation, first start, attach/status interpretation, graceful stop, recovery, update, diagnostics, and removal.
#### Scenario: New operator follows the guide
- **WHEN** an operator starts with a supported worker package and issued server enrollment data
- **THEN** the documented commands SHALL lead to a running verified worker and explain every state visible before the first assignment
@@ -0,0 +1,79 @@
## ADDED Requirements
### Requirement: Canonical assignment phase model
The worker SHALL represent assignment execution with one versioned phase/event model shared by local status, local history, server progress, and administrative views.
#### Scenario: Assignment completes normally
- **WHEN** a slot claims, executes, stages, uploads, and receives acceptance for an assignment
- **THEN** it SHALL emit monotonic phase events sufficient to reconstruct the time spent from `assigned` through `awaiting_receipt`
#### Scenario: Process restarts during an assignment
- **WHEN** a worker restarts with persisted slot state
- **THEN** recovered events SHALL continue from the persisted sequence and SHALL record recovery without rewriting the prior timeline
### Requirement: Complete scan-stage watchdog
The worker SHALL enforce one hard scan-stage deadline across permit acquisition, local preparation/resolution, data acquisition, scanner execution, filtering, cleanup, and result staging by supervising the complete execution unit outside the controller process.
#### Scenario: Scanner child exceeds the deadline
- **WHEN** the assignment runner remains active at the scan-stage deadline
- **THEN** the controller SHALL terminate its complete process tree, release/detach local resources, and produce a timeout result identifying the final phase
#### Scenario: Cleanup blocks after scanner exit
- **WHEN** scanner execution has ended but cleanup or staging remains blocked at the deadline
- **THEN** the same hard deadline SHALL terminate the runner and SHALL prevent the slot from remaining occupied until assignment expiry
#### Scenario: Permit acquisition consumes the budget
- **WHEN** no scan permit is acquired before the scan-stage deadline
- **THEN** the worker SHALL produce a phase-specific timeout result without starting the scanner
### Requirement: Non-renewing server progress
The Worker API SHALL accept idempotent monotonic progress events for the current reservation while preserving the original immutable assignment and queue deadlines.
#### Scenario: Progress is accepted
- **WHEN** the assigned device submits the next valid event sequence for its unresolved reservation
- **THEN** the server SHALL persist the event/latest phase and SHALL NOT alter assignment expiry or ownership
#### Scenario: Duplicate progress is retried
- **WHEN** an already accepted event sequence is submitted again
- **THEN** the server SHALL return the prior acceptance without creating a duplicate timeline event
#### Scenario: Progress cannot reach the server
- **WHEN** local phase transitions occur during a temporary connection failure
- **THEN** execution SHALL continue under the fixed deadline and events SHALL remain available locally for ordered retry
### Requirement: Observable deadline semantics
Assignments SHALL carry distinct effective target-scan, result-upload, and end-to-end assignment deadlines, and every operator/admin view SHALL label them by those meanings.
#### Scenario: Operator inspects active work
- **WHEN** status or admin renders an active reservation
- **THEN** it SHALL show the effective scan deadline, assignment deadline, time remaining, and current phase without conflating them
#### Scenario: Assignment expires
- **WHEN** the immutable assignment deadline passes without an accepted terminal result
- **THEN** expiry evidence SHALL include the last accepted phase and last-progress timestamp when available
### Requirement: Global and per-source assignment policy
The managed runtime configuration SHALL provide a global assignment TTL fallback and optional explicit overrides for GitLab, DockerHub, and HuggingFace, selected by the server at issuance.
#### Scenario: Source override exists
- **WHEN** a DockerHub assignment is issued and a DockerHub assignment TTL override is configured
- **THEN** its immutable expiry SHALL use the override and the assignment SHALL report that effective policy
#### Scenario: Source override is absent
- **WHEN** an assignment is issued for a source without an override
- **THEN** the global assignment TTL SHALL be used
#### Scenario: Invalid deadline policy is previewed
- **WHEN** an effective assignment deadline cannot cover its source scan timeout, upload deadline, and required handoff margin
- **THEN** managed configuration preview SHALL reject the candidate with a field-specific explanation
### Requirement: Phase duration percentiles
The server SHALL expose p50, p95, and p99 durations by source, phase, outcome, and selected time window, based only on completed observations appropriate to that metric.
#### Scenario: Administrator reviews DockerHub latency
- **WHEN** duration metrics are requested for DockerHub
- **THEN** the result SHALL separate end-to-end, scanning, cleanup, bundling, and upload percentiles and SHALL report sample counts
#### Scenario: Insufficient samples exist
- **WHEN** a percentile does not have the configured minimum sample count
- **THEN** the UI/API SHALL label it insufficient rather than presenting it as a stable policy recommendation