Collectors and durable delivery
Daemon collector and durable delivery
The sole private logger sink at bootstrap composes the plugin ring and ProblemService independently.
Typed plugin identity owns ring entries. Warn/error events reach the collector regardless of the local
log threshold, while console, files and ring retain that threshold. No prefix/scope parser remains.
Core Logger.warn/error require an explicit third metadata argument: a declared problem code, or
reporting: 'local-only' for a specific operational diagnostic. Merely naming a plugin is not a reporting
decision. External provider, process and transport error prose uses reporting: 'code-only'; the local log
keeps the diagnostic while the outbox carries only the typed code and count. Normal progress and successful
fallbacks use info. No message-text classification or collector suppression is involved.
The collector projects a problem code, level, real typed plugin names, validated inventory versions, counts and
at most three distinct redacted message samples per aggregate. The code is the call site's declared
PROBLEM_CODES entry. Live metadata validates the code against that catalog and permits only the declared
local-only or code-only reporting policy. An untyped producer with absent or invalid classification is
counted as runtime.reporting_invalid or plugin.reporting_invalid, without a sample, so metadata faults
stay visible without exporting unknown prose. Codes are never derived from scope or message text. Restored immutable batches keep their original
validated codes, including older scope catch-alls and unclassified, so already queued reports remain deliverable. Samples pass src/problems/redact.ts inside ProblemService.project(), which observe() calls, before anything reaches
the outbox, and are single lines of at most 500 characters. code-only lines are counted without a sample.
All plugin names are included, including private, unlicensed and non-bundled plugins. There is no synthetic
local bucket. The shared strict entry schema validates the same name/version rule for live and restored
batches: names follow the canonical 2–64 character grammar; versions are numeric triplets or null.
No-owner daemon events use none with a null version. Plugins actually named local or none remain ordinary
validated names. Plugin version inventory is a boot snapshot. Observation performs no credential,
filesystem or database reads; only the independent sender resolves the current credential before transport.
Reportability is decided at the producer: a diagnostic is sent when it can drive an Elowen bug fix or refactor,
not merely because it is useful to the local operator. Awaited boot-step completion timings and Sites
application HTTP 5xx responses are local-only; synchronous blocking steps, event-loop saturation, unexpected
provider-cache drops inside the TTL and missing hosted-search replay remain reported. Boot steps still waiting
at 30 seconds remain reported. Successful impersonation transitions are info-level audit facts. Asynchronous write-lock waits and holds of 250 ms or more are local-only (src/store/writeLock.ts:116-118), while synchronous slow steps are reported as runtime.slow_step (src/shared/slowStep.ts:21).
The event-loop watchdog retains its 150 ms sustained p99 and 2-second severe-stall budget; an unattributed
multi-second stall remains actionable even without a known producer. The same monitor exposes gcMs in /health.eventLoop and warnings: native GC durations clipped to the older staggered window, accumulated into two scalar sums by a PerformanceObserver. Pending records are drained before reads and window rolls; stop() disconnects the observer. Asynchronous delivery can lag a stall, and activity is not a causal attribution. No collector message filtering is used.
Provider retries retain one elapsed timer across attempts and reset it at settlement or a new turn. guardProviderStreamIdle aborts silent bodies at its existing 120 s deadline and logs the attempt locally at INFO. Its stream wrapper retains timeout evidence on the actual failed assistant object in a WeakSet without parsing provider prose or changing the wire transcript. The terminal spawnEventReducer reads that evidence and reports provider.stream_idle_timeout at WARN only when agent_end.willRetry is false; a successful retry produces no timeout problem.
ProviderRequestRecorder emits transport-attempt diagnostics at info/local-only; SpawnEventReducer owns
terminal provider reporting. Retry exhaustion comes from typed auto_retry_start attempt/maxAttempts,
cleared on successful recovery, settlement and a newly admitted turn. Provider prose never determines
retry exhaustion, timeout or image-rejection notices. Retry scheduling, permitted model fallback, cache
eviction, tool activation and fan-out misses are info; retry exhaustion emits one terminal classification,
and unexpected prefix mutation remains reportable. The per-retry notice is logged at info with the provider.retry code (src/brain/service/spawnEventReducer.ts:201), and the collector observes only warn and error lines (src/problems/service.ts:100), so the notice never enters the outbox.
In-session compaction measures normalized system, tools and history plus its appended instruction and
output allowance before sending. A warm request that cannot fit selects the standalone span, then PI's
native compaction when needed. Fallback selection is info; terminal summarization failure is reportable.
Hosted replay loss is reportable only after hosted content occurred, including content observed after
capture abandonment. Reference filtering is info. Unsafe capture, refusal, successful repair, unsuccessful
repair and wrapper failure declare distinct codes; their external prose is code-only.
Delegated progress returns SubagentProgressOutcome: accepted, already-terminal, or rejected with a
validation, ownership or persistence reason. The durable lifecycle owns finality. Only accepted writes
publish persisted progress and update live claims. Terminal updates are consumed without publication
or claim changes. Background completion commits its durable result and terminal run projection in one
write-lock transaction and returns the same typed outcome. Accepted completion publishes the persisted
terminal row, releases its progress claim and starts the existing result drain. Already-terminal callbacks
are consumed silently; validation and ownership refusals report their explicit completion codes.
Persistence failure leaves the run open for restart recovery. Delegate and DelegateContinue issue no
independent terminal progress update after background completion. Live results enter through
completeSubagentRun; restart and late-result recovery use the existing recovery ingress. There is
no public non-atomic child-result enqueue path.
Runner startup carries a typed SpawnFailureCode through validated IPC and RunnerStartupError. The pool owns its one report and health classification; diagnostic prose never determines the category. Dispatch may fall back locally because a startup failure occurs before delegated turn execution.
Pending-delivery recovery snapshots raw pending result IDs, including malformed payloads. Wake-accounting
and abandonment affect only captured IDs that remain pending for that parent, under the existing write
lock and counters. Empty batches do no work; later arrivals keep their own budget. Recovery joins the
existing delivery drain outside turn locks after continuation, including a continuation that throws.
slowStep accepts the live waitingOnProvider lifecycle predicate. Active provider compaction logs
waiting and elapsed duration as local information; ordinary stuck-step diagnostics resume when it ends.
The provider stream watchdog continues to own detection of a silent stream.
Dashboard digest eligibility is checked from bounded user messages and conversation summaries before claiming a daily generation. Usage counters, conversation titles and long-term memories alone do not trigger inference. Empty material does not consume an attempt; newly arriving material can make the next request eligible.
Maintenance shutdown cancels its timers and catalog/embedding workers, then awaits owned asynchronous sweeps before dependencies close. Plugin interval shutdown stops all host timers and drains returned tick promises before stopping contributed services, independent of registration order. Plugins must return their asynchronous work for the host to own it. The daemon checkpoints resumable turns first and keeps its bounded exit guard; wedged cleanup exits without closing dependencies underneath live work. Forwarded plugin metadata is retained while the host stamps the actual plugin owner. Runtime capability refusals and operator configuration diagnostics remain local-only.
Hook registry failures use plugin.registry_unavailable with code-only reporting, except the licensing
report endpoint, which stays local-only to prevent reporting transport recursion. Project provider
failures carry explicit plugin ownership and operation-specific codes. Conversation titling and activity
event persistence retain local diagnostics while reporting only their code and count. Core-owned logger
dependencies and the brainStatusFor warning callback use canonical Logger signatures.
Reporter and report-ingress diagnostics are typed local-only. Hook failures (published pages, webhooks)
are typed hooks.handler_failed or hooks.stream_failed and are code-only by default: a line whose text can carry message content or business data. Names may stay. A plugin HTTP route can declare problemReporting: 'local-only', which applies to that route's failures (src/api/routes/hooks.ts:86, src/api/routes/hooks.ts:93, src/plugins/api.ts:588). Batches pack entries under both the 50-entry
and the 16 KiB body limit.
This is intentionally daemon-only: early construction, locked boot, runners and web are outside coverage.
Every alert producer supplies an explicit reporting decision: local-only, already-reported, or problem
with a code-defined category. AlertService validates it and reports created, escalated or reopened
warning/critical transitions after any hold, once per condition rather than per recipient. Dynamic alert
keys, rendered templates and business parameters never enter reports; category reports are code-only.
An alert accompanying an existing error emitter uses already-reported. Operator quota, disk and Sites
serving conditions stay local-only. Source core has no plugin owner; plugin sources retain the host-stamped
identity and inventory version. Unchanged, dismissed, info and clear transitions do not report. Problem categories map to four codes: alert.environment_disk, alert.gateway_unavailable, alert.worker_unavailable and alert.scheduled_job_failed (src/alerts/alertService.ts:129-134). The alert.raised code is declared in the catalog (src/licence/contract.ts:93), but no producer in src/ emits it.
update.rolled_back and update.needs_operator project the updater's existing validated journal at boot,
on enable and each existing 60-second flush. The follow-up read is necessary: rollback writes its terminal
state only after proving the restored daemon's boot. A single problem_update_marker row records the last
transaction handed off, in the same SQLite transaction as its immutable redacted error batch. Restart,
successful delivery and consent Off do not erase that marker. No journal mutations or journald parsing
are involved; the journal's core from/to versions and persisted rollback note supply the sample. The updater
replaces the one retained journal with the next transaction, so both the journal and marker stay bounded.
While Off, neither journal reads nor marker writes run. Enabling later can hand off only the latest retained
failure, not a historical flood. A failed storage transaction advances neither batch nor marker. Existing
outbox limits, expiry, permanent rejection and explicit Off still apply after handoff: exactly-once here
means durable handoff, not guaranteed network delivery. A daemon that cannot boot cannot report until it
runs again; replacing a failure journal before the next successful flush also removes its evidence.
ProblemService holds at most 200 aggregate keys and 1,000,000 counts per key. A 60-second flush freezes immutable UUID batches in problem_outbox. Pending storage is bounded to 200 rows, 4 MiB and seven days. Pruning runs one flush interval early so normal scheduling stays under the age maximum, and never evicts the one bounded request in flight. Its expiry may overshoot by at most the 15-second request deadline, an accepted limit. Drop counters change only after their transaction commits. Every restored row is strictly validated before transport; malformed, oversized or mismatched batches are discarded with a fixed local-only warning and one lost unit when their original count is unknowable. Boot replays leftovers. One request runs at a time through licence/reports.ts, using the existing 15-second deadline and Bearer credential without redirects, token renewal, locking or update coupling. Jittered retries run from 30 seconds to 30 minutes and honor bounded Retry-After. Permanent payload failures drop the batch; 401/403 pause the current credential, and 404 retries slowly. Reporter-owned metadata in settings preserves the config revision. ConfigStore carries it opaquely during ordinary config reads and writes; only problemState/setProblemState validate it. Invalid metadata emits a fixed local-only warning and disables the reporter with storage status, without resetting other settings or fabricating successful status. The stored consent choice remains intact. Explicit off followed by confirmed on resets metadata and recovers reporting. Disable atomically clears pending rows and last-sent metadata, clears memory and aborts transport. An explicit off write clears leftovers even if already off; an off boot also purges restored leftovers without a config revision or consent event. Failed off-boot cleanup remains mandatory before any later enable can send. A generation check discards stale replies. Shutdown synchronously freezes memory into the outbox, aborts transport and awaits its owned cleanup before maintenance dependencies close. It does not wait for successful delivery. Repeated stop calls join the same settlement; bootstrap awaits this stop before pausing the privileged worker. The daemon exit guard still bounds a wedged cleanup. Expected setup and licence states and successful orphan-process cleanup use info; boot, stdio and snapshot failures retain typed code-only reports. A hard crash can lose the unfinished memory window.
ConfigStore is the sole setting-write boundary. An explicit authenticated admin or explicit local-admin
actor must accompany boolean changes, with confirmation on enable. The local-admin actor kind exists in the type (src/problems/state.ts:22), but no code in src/ constructs it; the config route passes kind: 'admin' (src/api/routes/config.ts:116), and a change without an actor is refused (src/store/configStore.ts:1415-1417). Enabling without problemReportingConfirmation and expectedRevision answers 409 confirmation_required (src/api/routes/config.ts:75-76). It depends only on ProblemConsentWriter,
implemented by EventStore, rather than importing the event store back into configuration. EventStore writes one nonaggregated
admin-setting consent record in the same transaction, including actor, timestamp and config revision.
These records are admin-only and use the existing runtime.limits.eventRetentionDays retention.
The general plugin events-read seam excludes them; only authenticated admin activity reads expose them.
The existing CLI API door uses authenticated daemon routes, not an inferred local-admin actor.