Restart and recovery
Restart and recovery
SIGTERM and SIGINT trigger a bounded pause. The daemon stops admitting new ordinary turns, checkpoints queued messages and resumable live work, records what cannot be resumed, and exits. A requested restart uses the same pause path and lets its supervisor start a new daemon process.
Boot recovery has two phases. Before platform traffic starts, recovery providers synchronously claim durable work left by the previous process. After platform startup, the coordinator resumes claimed work in dependency order and waves. This prevents inbound traffic from observing stale running rows and lets recovery turns use the normal channel and brain path.
BrainBootRecovery in src/brain/recovery/boot.ts implements BootRecoveryHost and owns the shared claim state: claimed runs and calls, owner result parents, and workflows still owed to the engine. It composes three internal recovery owners from the same live dependency object. The type-only recovery/deps.ts owns BootRecoveryDeps for that shared object, keeping each owner independent of the boot orchestrator. BrainPauseRecovery in pause.ts owns PauseSummary, the synchronous checkpoint, the single bounded unparkable wait, and paused-queue replay. WorkflowResumeRecovery in workflowResume.ts owns the existing attempt cap, current-authority journal check, output persistence and durable completion, including terminalization when the workflow control is absent or declines recovery. ParkedConversationRecovery in parkedConversations.ts owns result wakes and silent owner continuations, their marker/activity compare-and-set and the failure budget for the original result IDs. An owner continuation injects no generic restart message; an already-settled transcript needs no provider call. A drain that returns landed-unanswered has already delivered those original rows into context. Recovery checks the current durable park through claimParkedResultContinuation, without spending a generic park attempt. If restart or crash left that marker standing, the existing subagent-result-resume custom message continues the parent from its result, and delivered rows are acknowledged only after a settled answer. Recovery joins any continuation-triggered drain before checking the original batch. User Stop and turn admission clear the marker: a cleared park or a result-only wake is released without a continuation or failed attempt, retaining unanswered rows for the next user turn. Failed continuations retain the original result budget and report recovery.result_continuation_failed. No additional flag or storage is introduced. Undelivered original rows still consume their own bounded budget; late arrivals do not. Result wake and park resume failures log WARN before the third attempt and ERROR at exhaustion; timed result delivery uses the same rule with its five-attempt budget. Its outer finally releases result delivery and replays checkpointed user input last, including on failed or released recovery. These internal boundaries add no storage, provider calls, retries or timeout budgets. BrainService remains the real consumer through its unchanged public verbs, and createBootRecovery is still handed the brain. Claims stay synchronous before platforms start, resumes run afterwards in coordinator waves, and publishToSession remains the brain's injected fan-out callback.
The shared slowStep seam measures awaited boot phases and turn-start work; slowSync measures synchronous work, logging local INFO from 250 ms to below one second, reporting WARN at one second and above, and retaining the last 32 slow intervals in memory. recordSlowSync(label, startedAt, endedAt, caller?) feeds already-timed phases into that same threshold/history without allocating entries below the threshold; the optional caller function is invoked only for a slow interval, and its answer is stored as the entry's caller and appended to the diagnostic as in <caller>. withWriteLock uses it to distinguish database write lock wait, including SQLite busy timeout and synchronous retry backoff, from database write lock hold, including callback execution and commit/rollback, and names the two stack frames below the lock (the method that took it and its caller, so a plugin behind PluginDb.transaction is named too; for a lock nested inside another, the second frame is the outer transaction's callback), e.g. slow sync step: database write lock hold took 5268 ms in MemoryStore.setStatus (…/memoryStore.js:760:12) < MemoryStore.softDelete (…). The stack is captured only on that slow path. A hold longer than the 5 s busy_timeout turns every other writer's wait into SQLITE_BUSY, which is what this attribution is for; withWriteLock retries only a busy answer that came back before a full busy_timeout, so such a wait blocks the event loop for one timeout, not five. Statements outside a transaction hold the lock too but are not timed. Exhausted acquisition and failed callbacks are recorded too. Both intervals are captured before warning I/O and recorded after the transaction wrapper commits or rolls back, so logging does not extend the measured hold or run inside that wrapper's transaction. Timing adds only clock reads and preserves transaction/retry behavior. Nested transactions retain SQLite's savepoint behavior, so their recorded holds can overlap the outer hold. The activity snapshot refresh uses the ordinary wrapper. The loop-lag warning names recorded synchronous steps overlapping its window, or unattributed when none matches. startDaemon times licence evaluation, app construction and bind, and the database open/migrations step is timed inside buildBrainCore (src/daemon/brainCore.ts); startLoops times claims, sweeps, awaited plugin/platform startup and synchronous recovery resume dispatch. The background resumeAll promise includes respawned child turns and is not an awaited boot step: readiness and the boot announcement precede it, so it has no stuck-boot watchdog. Dependency waves, item ordering and provider outcome summaries still await each provider's own resume work; a delegation's resumed count means its continuation completed and queued the result, not merely that dispatch began. No entries persist across process exits, and uninstrumented synchronous work remains unattributed. Awaited slowStep latency is not treated as synchronous blockage.
After the claim pass and plugin registry reconciliation, adapter connects run concurrently with plugin boot reconciles. Each adapter has a 30-second start deadline and each plugin reconcile has a 20-second budget; an expired task continues in the background and late failures are logged. Plugin services start once both startup waits have settled, then the boot announcement and recovery resume follow. onPlatformsStarted and /health.platformsReady mean that the platform startup phase has settled or exhausted its deadlines, not that every gateway is connected; a timed-out gateway can connect later. The sub-agent runner still starts only its own adapter and no plugin services. A process with no adapters resolves immediately.
The durable platform turn envelope and the authority re-derived from it are pure functions in src/brain/platformTurnEnvelope.ts (PlatformTurnResumeEnvelope, normalizePlatformTurnEnvelope, provePlatformSenderBinding, resolvePlatformTurnAuthority, and the serializer and history seed that share its bytes). ChannelSessionService.send writes the envelope; platformTurnRecovery.ts reads it back. The serialized bytes are persisted and compared on resume, so they stay identical.
PlatformOrchestrator.pendingConnection(platform) exposes the actual adapter connect settlement separately from its boot wait deadline. BrainBootRecovery.resumeParkedPlatformTurn waits on only the affected adapter before entering the existing recovery seam. A pending connection leaves park markers, resume envelopes and already-computed delivery rows intact without consuming attempts; readiness and independent recovery items do not wait for it. After success or rejection, recovery re-reads durable state and current authority, so a room that spoke or aborted during the wait is not resurrected. A connection that never settles keeps its work durable and its recovery item pending, including completion of that provider's recovery wave.
The reconcile budget also releases only the host's wait, not a plugin's work. Sandbox coalesces overlapping reconcile calls onto one sweep. Its automatic recovery re-reads the environment inside the existing store transaction before both exhaustion and enqueue writes, matching the inspected generation and lifecycle epoch and rejecting a stopped/deleted intent or active operation. A late observation cannot restore an obsolete generation or overwrite a newer lifecycle decision. Cronjob's boot reconcile is synchronous and has no post-budget continuation.
Local calls that were in flight when the daemon stopped are settled with an interrupted tool result saying that the effect may or may not have happened and must be verified before repeating it. Unanswered delegation calls are handled differently: the child is recovered and its result is delivered, so the parent must not delegate the same task again. A partial assistant stream that is not provably final is removed so the model can regenerate that step. A final assistant answer is retained.
Recovery is fail-closed for malformed or unverifiable durable envelopes. It does not invent authority or replay an unknown side effect. Completed child results remain durable until their parent delivery is acknowledged.
Interrupted cron and scheduled runs are terminalized without another model turn. Boot recovery attempts their owner notice once through PlatformOrchestrator.notify; a rejected delivery is logged as platform.notice_delivery_failed with code-only reporting, and the scheduled work remains terminalized.
A plugin change is a separate lifecycle operation: the daemon restarts and rebuilds the plugin registry generation from disk, while SQLite state and conversation history remain. All requested restarts use ShutdownControl.requestRestart: a 2.5-second quiet window measured from the last request joins a burst into one restart. It does not wait for parkable turns. The existing pause checkpoints resumable work immediately after the window and gives only unparkable turns their unchanged, single 20-second settleUnparkable budget. Joined HTTP callers retain bounded response-flush protection. Signals bypass the request window and still use the same pause. /system/restart, the CLI, core updates and plugin marketplace mutations reach this one path through applyChange; there is no per-caller bypass. RestartRequest.quietMs replaces the old turn-count/ceiling response.
Owning code, minimal use and limits
Owning code:
src/daemon/shutdown.ts: signal pause, restart quiet window (RESTART_QUIET_MS, 2.5 s), checkpoint guard (PAUSE_CHECKPOINT_GUARD_MS, 5 s).src/daemon/applyChange.ts: the one restart and reload path (createApplyChange).src/brain/recovery/:boot.ts(BrainBootRecovery),deps.ts,pause.ts(unparkable wait, 20 s),workflowResume.ts,parkedConversations.ts,providers.ts(createBootRecovery).src/brain/platformTurnEnvelope.tsandsrc/brain/platformTurnRecovery.ts: the durable platform turn envelope and its resume.src/shared/slowStep.ts:slowSync(logs at 250 ms, keeps the last 32 slow intervals) andslowStep(warns after 1.5 s and every 30 s while a step is still pending).src/store/writeLock.ts: write lock wait and hold attribution.- The CLI verb
elowen restart <daemon|web|all>is insrc/cli/commands.ts.
Minimal use, from the system route:
const { request } = await applyChange({ mode: 'restart', byUserId, reason: '/system/restart', flushed });
Limits and absence:
- A process without an
applyChangequeues no restart. The plugin marketplace route returns no request in that case (src/api/routes/plugins/index.ts). - The pause is bounded: the checkpoint is synchronous SQLite work that no timer can preempt, and the unparkable wait is capped at 20 s inside a unit whose stop timeout is 30 s.
- Platform adapter startup is capped at 30 s and each plugin reconcile at 20 s; expired work continues in the background.
- Slow-step history is process-local and not persisted.
Real consumers: BrainService.settleUnparkable (src/brain/brainService.ts), the /system/restart route (src/api/routes/system.ts) and the platform /restart command (src/brain/platforms.ts).