NAVIGATION
ELOWEN DOCUMENTATION

Last updated: 10 October 2026

Browse documentation · Infrastructure and process topology
Developer reference

Infrastructure and process topology

Owning code and consumers

Owning code, by area:

  • Ports and CLI endpoint resolution: src/shared/servicePorts.ts, src/cli/launcher.ts (localDaemonUrl, daemonUrl, start lock).
  • Process locks: scripts/kernelLock.mjs (descriptor lock), scripts/privilegedLocks.mjs (lock table, ranks, wait classes).
  • Root helper operations: scripts/elowen-site-gateway.mjs (operation table), scripts/siteGatewayContract.mjs (installed contract).
  • Privileged client and installation: src/privileged/workerClient.ts, src/cli/install/rootFiles.ts, src/cli/install/privilegedCutover.ts, src/privileged/privilegedFiles.ts.
  • Deployment records: src/cli/provision/deployment.ts.
  • Atomic JSON writes: packages/plugin-shared/atomicJson.mjs.
  • Process topology, sub-agents and managed branch reads: src/daemon/brainCore.ts, src/subagent/pool.ts, src/subagent/sizing.ts, src/brain/service/sessionProcesses.ts, src/brain/processRegistry.ts, src/integrations/managedProjectBranch.ts.
  • Host alerts: src/daemon/resourceAlerts.ts, src/update/report.ts, plugins/sandbox/lib/environmentAlerts.mjs.

Minimal use of the CLI endpoint and a kernel lock:

import { daemonUrl } from '../cli/launcher.js';
import { acquireKernelLock } from '../../scripts/kernelLock.mjs';

const base = daemonUrl(process.env);                       // ELOWEN_URL wins; otherwise the local daemon
const release = await acquireKernelLock('<dataDir>/example.lock', { timeoutMs: 5_000 });
try { /* exclusive section */ } finally { await release(); }

Limits: lock waits are bounded by a named wait class and by the request deadline, and a lock is never released by a stale PID file. Without an installation record, local lifecycle commands select the port from the environment or the tracked run state. Remote ELOWEN_URL targets are for HTTP clients only.

PrivilegedWorkerClosedError in src/privileged/workerClient.ts is the client lifecycle failure for intentional close or pause, with the single readonly code closed. Existing connection, execution and request shutdown branches instantiate it without changing their messages or cleanup. A socket close during pause also uses it because pending requests may settle from that event before pause's final cleanup; unexpected socket closure remains an ordinary transport error. createPublishedSitesGatewayControl recognizes the typed error and exposes its code in the existing gateway status. No consumer hook is required, other errors receive no lifecycle classification, and the conversion adds no transport work. The registry Sites gateway and worker monitor use the resulting status to distinguish daemon shutdown from gateway failure.

Infrastructure lifecycle and deployment facts

src/shared/servicePorts.ts owns validated daemon/web bind ports. CLI HTTP clients resolve through daemonUrl (src/cli/launcher.ts): an explicit ELOWEN_URL may target a remote instance; otherwise the installed daemon port wins even when an operator shell exports a different ELOWEN_PORT, because installed systemd units still bind their recorded port. Without an installation record, only the supplied environment selects the port, followed by its tracked port and shared default; an explicit environment object never borrows missing endpoint values from process.env. Tests model installation facts and run state in fixtures and clear inherited endpoint variables rather than reading host installation state. localDaemonUrl ignores remote overrides for updater health and offline lifecycle checks. Auto-start refuses unreachable remote targets and preserves the systemd guard. It calls the launcher's daemon-only mode, which starts and tracks the existing daemonSupervisor without requiring the web bundle; a later full start adopts that supervisor and adds the web.

scripts/kernelLock.mjs is the shared Linux process-lock seam for launcher start, update/recovery, explicit deployment reconfiguration and privileged mutations. It opens an owned regular file without following symlinks and passes its open-file description to /usr/bin/flock for acquisition. The caller keeps its own descriptor after that subprocess exits; release or caller death closes it. There is no separate holder process whose death could unlock a still-running mutation. The lock pathname is never unlinked and PID text is never authority. Namespaces remain separate; startup waits five seconds and updates fail fast. A lock may be taken shared or exclusive; the privileged operation locks below use both. The installed root helper imports a root-owned copy beside it; privileged installation and the worker/client digest include that copy. Refresh the privileged files before restarting an upgraded daemon. Gateway state JSON and publication identities are not migrated or rewritten by this convergence.

Every operation of the root helper (sites, nspawn and plugin-service) is one entry of a single definition table in scripts/elowen-site-gateway.mjs. An entry declares its effect (read, mutation or execution), its lock scope, the wait class that bounds acquiring that scope, its worker admission class (control or bulk), what it needs read after the locks are held (owner proof, deployment, storage roots), any phases and its handler. The scope is mandatory and explicit: lockScope.none says so deliberately, lockScope.host(name, mode) names a host resource, lockScope.environments(resolver) names the environments a request touches, and lockScope.all(...) combines them. A definition or phase with no declared scope, a scope that contradicts its wait class, or a phase that breaks the lock order makes the module fail to load, and tests/contract/privilegedOperationScopes.test.ts pins the declared scope of every operation. The worker handoff policy (only mutations are handed to a later caller), the capability list the broker advertises, whether a deployment record is read and the admission class are all derived from that table; there is no second list of lock-free or record-free operations.

scripts/privilegedLocks.mjs owns the lock table, scope identities, claim resolution, operation-definition validation and the descriptor executor built on kernelLock.mjs. The root operation table and all request-to-environment validators stay in elowen-site-gateway.mjs. Trusted maintenance callers import lockScope and runWithDeclaredLocks directly from the lock module; handleRequest uses its claim executor, and so do the two trusted maintenance callers (explicit deployment reconfiguration, which holds the gateway lock, and the proof-host teardown, which holds the uid-registry lock). No handler takes a lock itself. Locks have a fixed rank and are acquired in rank order, then in key order within a rank, and released in reverse: the provisioning barrier (10, shared by work that needs a provisioned host and exclusive for provisioning), environments (20), the Sites gateway and plugin-service ledger (30) and the machine uid registry (40). A nested phase may only take a lock ranked above everything already held, which is what makes a cycle between two requests impossible; there is no upgrade from shared to exclusive. Every wait is bounded by a named class (none, uid-registry 5 s, environment-control 30 s, environment-bulk 15 min, plugin-services 120 s, host-transaction 9 min) and by the request's own deadline, and a request logs how long it waited for and held its locks. The maintenance-only cutover lock ranks below all of them (rank 5); no operation may declare it.

An environment lock is named by a validated identity, never by a request string. A Project request names its resource by kind and resource (and its machine, which must agree), a tree request names the owner of each tree it touches (owner, or sourceOwner and targetOwner, each exactly {kind, resource}), and the helper requires every tree path to lie inside that owner's storage root. A copy between two environments takes both locks in canonical order. Lock files live under /var/lib/elowen/locks (the gateway lock keeps its historical path), are never unlinked, and are refused unless their directories are not symlinks, are owned by the helper and are closed to group and world (the gateway lock keeps the instance state directory it always had, so only its file checks apply). A bind source of a machine envelope must lie in the storage of the environment the envelope belongs to. The machine uid range is reserved in a declared uid phase under the short registry lock, which is released before extraction or an ownership pass starts. Command execution, freezing and thawing a machine (which touch no helper data, so a recovery thaw never waits behind a long tree operation), status and the other reads take no lock, and the worker keeps one operation slot for control work so bulk operations cannot occupy every slot. Replacing the installed helper is a cutover, not an in-place swap, and installPrivilegedFiles is the only writer of those files: elowen privileged refresh and elowen install both go through src/cli/install/privilegedCutover.ts, which serializes against another replacement (the maintenance-only cutover lock), stops admission and drains the worker, holds every host lock while the files are written, and only then reopens what was open, so an old and a new lock policy never mutate the same host at once.

scripts/siteGatewayContract.mjs is the pure installed contract for canonical hostnames, binding count and byte limits, token/email/slug grammars, fixed owner/deployment/storage receipt paths and plugin-service socket derivation. The root helper, daemon gateway and service-unit renderer import it directly; CLI proxy readiness checks use the same socket function. Boundary-specific binding diagnostics and installed routing semantics stay with their callers. Neither shared module reads host configuration on import: undeclared scopes and invalid socket names fail explicitly, and an unlocked scope touches no lock files. Both modules and their declarations ship in the package; the JavaScript files are installed root-owned before the helper and worker through PRIVILEGED_FILES. The original inventory order is retained with these dependencies inserted before the helper. Worker and client SHA-256 digests retain the original file order and NUL separators and append both dependencies in the same order, so missing or mismatched bytes refuse the handshake. Cutover, owner checks and historical lock paths are unchanged. The proof-host inventory copies both modules and rewrites only the lock module's paths into its isolated lock namespace.

On the client, PhaseDeadlines in src/privileged/workerClient.ts is the one lifecycle for a request: handshake, preparation, attachment and the overall deadline are phases with budgets defined together. The five-second attachment timer covers only the wait from writing exec-start until the worker reports exec-started (admission, the three stream attachments and the launcher handoff). Time spent connecting or preparing (which may wait for an environment lock) never counts against it, and a missing attachment still times out. The caller's overall deadline is preserved. The phase budgets are handshake 5 s, preparation 120 s and attachment 5 s.

Fresh install and explicit setup use recordGatewayDeployment for the gateway's independently validated host/port facts. reconfigureDeployment holds the declared gateway lock, updates those facts and the canonical install URL together, preserves the install inventory, and restores both records and the app proxy on a failed operation. A hostname change with existing Site bindings or pending removals is refused before any proxy mutation: those identities require a separately coordinated DNS/publication migration. Existing bindings also forbid an HTTPS downgrade; failed TLS provisioning restores the prior deployment. Neither upgrade nor privileged refresh invokes this explicit reconfiguration; its only caller is the setup step (src/cli/setup/steps/deployment.ts).

Core catalog caches use packages/plugin-shared/atomicJson.mjs with its TypeScript declaration, private mode (the callers pass mode: 0o600) and non-durable replacement (durable defaults to false). Shared staging uses exclusive creation and removes only a staging file it created, including when a colliding name prevents acquisition.

Process topology

The normal topology has one daemon process and one web process:

browser
  via web application
    daemon HTTP/SSE
CLI
  daemon HTTP/SSE
platform adapter
  daemon
    brain and SQLite
    plugin services and platform gateways
    optional delegated runner processes

The daemon creates the Hono HTTP server, registers core and plugin routes, and dispatches authenticated plugin root mounts without core domain fallbacks. Unknown mounts return 404, while a known unavailable plugin returns 503. It attaches plugin WebSocket upgrades to that same server. Ticketed routes retain live account/grant resolution; public integration routes require a registered authorize({params,headers}) function before upgrade and expose no Elowen identity or ticket payload. False answers HTTP 401 without a 101, verifier failures answer 500, and the same-host Origin gate remains first. Registry Twilio consumes this generic transport. The registry preserves each declared route pattern for refusal, transport and handler logs so path credentials are not logged; the existing bounded-frame, paused-read, backpressure and generation teardown paths are shared. Without registration there is no public verifier work. The daemon starts platform adapters and plugin services, and owns maintenance, shutdown, recovery, and process registries. Plugin load, boot-reconciliation, and service-start failures produce warning alerts for administrators and clear when the producer recovers. Maintenance produces core alerts for provider usage at 80% warning and 95% critical, for host disk space (a warning when 20% or less is free, critical when at most 2 GiB and at most 5% are free), and for updates: a rolled-back or failed run raises update-failed (critical when the previous set could not even be restored), a run that deliberately left plugins behind raises update-skipped, and a clean commit clears both. The update producer reads the updater's durable last-run report (report.json in InstancePaths.updateDir, written by every decided elowen update path), because the CLI updater runs in its own process without an AlertService; one stable key per condition makes a repeated hourly failure upsert instead of spam. Provider alerts clear below 75% in every usage window. A critical host-disk alert drops back to a warning once 3 GiB or 7.5% is free, and the host-disk alert clears only when less than 75% of the disk is used. The disk dial in Settings → System (GET /system diagnostics) reads the same figure through hostDiskUsage in src/api/systemDiagnostics.ts, the filesystem holding the database, so the dial and the alert always agree. The Sandbox plugin raises project-scoped alerts: a warning at 80% of a managed environment's disk ceiling, a critical alert when usage exceeds the ceiling, a warning when a start is refused for disk, and a critical alert when automatic recovery is exhausted.

The shared buildBrainCore factory never creates a privileged connection. Process entry points explicitly supply privilegedTransport: new PrivilegedWorkerClient(): the daemon passes it through buildApp, and the delegated runner supplies it directly. Embedded and test cores omit it by default. Machine gateway calls then reject locally with no privileged transport in this process; daemon-mode Sites gateway calls report unavailable with the same detail. The runner still has no Sites gateway. A supplied client is reused for both domain-stamped gateways and the daemon's pause hook; omission creates no connection or pause work. Sandbox is the machine-runtime consumer, and Sites is the published-site consumer. Loading a registry without Sites, including a disabled or failed plugin, sends no gateway request and leaves the existing deployment untouched.

Background shell telemetry has one daemon-owned conversation projection in SessionProcessService.snapshot. GET /brain/status, the first SSE snapshot and live process events carry its processes field. It contains the focused conversation and its direct durable delegated children, filtered by the existing operator/account ownership rule. processHandleOwnedByAccount is the canonical identity predicate for local handles, accountless runner snapshots and teardown sweeps: a non-null explicit account wins; otherwise the originating durable session owner decides. Snapshots expose no explicit account field. Foreground controls describe the focused session only. Terminal registers both explicit background launches and foreground runs in ProcessRegistry, and re-registers the same handle as a job after manual or deadline detach. Runner IPC supplies metadata to this read model and clears it when the runner exits; output and kill still reach the owning registry through their existing gates. No process cards or client account-wide polling remain. Empty registries produce empty telemetry; this adds no persisted state, changes no restart lifetime, and costs only a scan of registered handles. CLI and web rails render running non-foreground entries from this field.

Direct-host descendant discovery has one scanner in packages/plugin-shared/processTokens.mjs. Terminal stamps DIRECT_PROCESS_TOKEN_ENV into each run and calls tokenPids([token]) while confirming cleanup; src/brain/processTokens.ts retains only core's best-effort killTokenProcesses(tokens) wrapper for runner-exit cleanup. One synchronous Linux /proc pass handles the whole token set, including setsid descendants; the scanner excludes itself and unreadable or exited processes and returns descending pids. Empty tokens or no readable matches return no pids and send no signals. Scan cost scales with process environments and token comparisons; no hook, persistence or provider request is added. Terminal keeps its kill-error diagnostics and waits for confirmed exit before releasing a lease. ProcessRegistry retains session-scoped output/kill and account-scoped listing; the unused account-specific output/kill methods are removed after checking core, bundled, registry and installed consumers. The account-wide kill, killAccount (src/brain/processRegistry.ts), stays: it is the live path behind killAccountProcesses from the users route (src/api/routes/users.ts).

The Next.js web process owns browser rendering, route composition, and the same-origin browser-to-daemon proxy. The catch-all proxy forwards REST and streaming responses to the daemon and keeps browser authentication at the web boundary. Web modules display daemon state but do not make authorization or persistence decisions.

When the delegated runner pool is enabled and a delegated turn is dispatched, the daemon can fork runner processes. A runner rebuilds the same brain core and plugin registry needed for its child session, but does not start an HTTP server, ordinary platform gateways, scheduler, migrations, or maintenance loops. The daemon remains responsible for dispatch, recovery, abort fencing, reverse workflow calls, and delivery.

SubagentRunnerPool in src/subagent/pool.ts owns one channel route map whose values hold both the runner entry and its touchedAt activity time. Dispatch and accepted nested edges establish or refresh that record; tapSessionSnapshot, used by delegated-child drill-in, refreshes the same time before asking the owning runner for history and live events. Missing routes send no remote request and let ordinary dispatch choose a runner. The maintenance sweep retires routes idle for 30 minutes only after confirmed release. Busy or failed release keeps ownership and refreshes activity to the sweep's start; each candidate reads current activity because an earlier release can await while another route is used. Successful release removes a route only while that runner still owns it. Runner exit removes all of its routes, and reset clears the map; terminal nested-edge retractions after reset notify the sink without recreating routes. Stats and reap decisions count this map directly, with no separate activity index, persistence, or added IPC. Route lookup is constant time; sweep and session counts are linear in the current routes.

Runner memory admission uses a bounded asynchronous Linux process-tree sample in src/subagent/sizing.ts, triggered by the existing two-second heartbeat. It follows every thread's children file, sums VmRSS once per PID, and stops at 256 processes or 1,024 threads. This includes language-server wrappers, semantic/syntax servers, and typings installers without scanning unrelated host processes. At most one sample per runner is in flight. An unreadable or truncated sample blocks growth until a complete sample arrives; it never becomes a success-shaped zero. A discovered descendant that exits before its status read contributes no resident memory.

The sizing denominator R is the high-water maximum of reported runner RSS and sampled whole-tree RSS, or the existing 512 MiB estimate before a measurement. The memory cap is floor(0.5 * totalHostMemory / R). Automatic capacity is min(max(1, CPUs - 1), memoryCap); an explicit operator maximum replaces the CPU ceiling but still gives min(operatorMax, memoryCap). Spawning also needs available host memory of at least 1.5 * R, the existing growth cooldown and saturation rules. Pending or incomplete samples prevent growth, while admitted work continues. For example, on an 8 GiB host a 256 MiB runner with 1,792 MiB of descendants has a 2 GiB denominator and memory cap 2, rather than the root-only cap 16. This is a conservative RSS budget, not a hard process or cgroup limit: shared resident pages can be counted more than once, swap is not RSS, and descendants reparented after runner death belong to the existing containment cleanup.

/health.subagentPool.runners keeps rssBytes for the runner itself and adds treeRssBytes, treeProcesses, and treeMemoryComplete for the latest sample. Those fields are null before sampling. The pool-wide runnerRssBytes exposes the sizing denominator, which deliberately retains the high-water mark even after descendants exit. nestedActiveTurns projects the runner's existing live child edges, validated against the owning route and deduplicated by child session. activeTurns remains the outer dispatch count. Both cover call settlement as well as execution. The effective load for placement, least-loaded selection, saturation and growth is activeTurns + nestedActiveTurns, derived from that same validated edge projection without another counter. A new directly dispatched turn fits only below load eight; saturation starts at load four or sustained loop lag. Nested work is also a growth demand signal, reconsidered on edge changes and existing heartbeats after the fifteen-second cooldown.

Nested children stay in their parent's runner and start without waiting for a pool slot. Their parent keeps its outer count, including while waiting for children. No parent/child slot dependency is created; a locally nested fan-out may therefore exceed eight, at which point new outer placements wait or use another runner. Memory and runner-count ceilings still apply to growth, and existing work is never killed to meet a reduced memory cap. This changes load accounting only: dispatcher predictions, provider prompt bytes, workflow DAG ownership and abort/recovery routing keep their existing behavior. Spreading work means placing new outer calls on other runners; existing nested trees are never moved.

Managed Project execution is another boundary, not another daemon. The Sandbox plugin and the Project execution contracts prepare and control a persistent managed environment. Core integrations reach it through the typed environment control rather than switching to host paths implicitly. prepareExecution() returns either a local or confined launch shape or a ManagedPreparedExecution; managed consumers call start() for a ManagedExecutionSession with guest stdin, stdout, stderr, and a settled closed promise. The persistent privileged worker carries those sessions over protocol 2 and direct socket-stdio attachments, with interruption recorded before cancellation settles.

The checked-out branch of where a conversation works has one source per execution kind, and every client reads it from one field, project.branch of GET /brain/status. A host directory is a memoized read of .git/HEAD (src/brain/service/gitBranch.ts). A managed Project has no host path, so its branch is read inside its environment by src/integrations/managedProjectBranch.ts: the environment state read, then git rev-parse in the guest through runManagedProjectCommand, the same seam as /projects/:id/git. GET /projects/summary (the register cards) uses the same module, so the CLI rail, the web telemetry foot and the project card cannot disagree. The read never starts an environment: a stopped or unprovisioned environment, a refused state read, a failed guest command and a timeout (2 s) all yield no branch, which clients draw as nothing (web) or a quiet "unknown" (CLI rail). Answers are remembered per project id, 5 s for the status (so a git checkout made during a turn shows when the turn ends) and 30 s for the register, a failed read, including an environment-state failure, for 5 s so a wedged guest cannot stall every poll. Pending execution cleanup and the 2 s telemetry timeout are local INFO diagnostics; other failures retain project.branch_read_failed. runManagedProjectCommand raises the standard DOMException with name TimeoutError on its deadline, after verified cancellation, so consumers classify the outcome without parsing its message. BrainStatusService.status() stays synchronous and only peeks that memory; the /brain/status route first awaits BrainStatusService.refreshProjectBranch(), which does the (at most one) guest read when the memory is stale and re-checks the caller's access to the Project. Example: a conversation in a running container project on main: the first status poll costs one short guest command, the polls of the next 5 s cost none. With no Sandbox plugin or no running environment the field is null.