NAVIGATION
ELOWEN DOCUMENTATION

Last updated: 10 October 2026

Browse documentation · Isolated fixtures and CLI evidence
Developer reference

Isolated fixtures and CLI evidence

Isolated daemon test preparation

The API, brain, migration, licence and web E2E harnesses share test-only preparation in scripts/tests/lib/daemonHarness.mjs: call isolatedChildEnv(dataDir, scenarioValues), allocate freePort(), then await waitForHealth(baseUrl, timeoutMs, readinessPredicate). Inherited production ELOWEN_* and agent configuration paths are stripped; HOME, DB, project state and logs point into the temporary scenario directory. Explicit scenario values remain caller-owned. The sole inherited Elowen key is ELOWEN_E2E_DIST, made absolute before children change working directory. resolveTestEntry.mjs checks the selected build's daemon/main.js for test-licence boot, or daemon/index.js for the licence scenario's real locked boot. Without a selector it uses the checkout's own dist/; a missing selected entry throws before spawn without switching builds. Brain also imports its restart exit code from that same selected build.

Minimal use, as brain-e2e/spawn-daemon.mjs does it:

const port = await freePort();
const baseUrl = `http://127.0.0.1:${port}`;
const env = isolatedChildEnv(dataDir, { ELOWEN_PORT: String(port) });
const entry = resolveTestDaemonEntry('main.js', env); // throws before spawn if the selected build is missing
// Spawn `entry` with `env`, then wait for readiness:
await waitForHealth(baseUrl, 30_000, (body) => body?.ok === true && body.platformsReady === true);

Ports are released ephemeral IPv4 loopback ports, not reserved leases. Readiness polls every 100 ms and bounds requests and response-body reads by the total deadline. Brain requires platformsReady: true; the locked licence scenario accepts ok: true alone. Spawn, licence minting, restart, logs and teardown remain scenario-local. These helpers never ship or run on normal boot. Preparation regression tests mock child processes and build artifacts, so they do not execute a daemon or touch live state; full E2E proof still requires a separately built, isolated test environment. The preparation behavior is tested in tests/scripts/daemonHarness.test.ts, and the entry selection rules have their own test in tests/scripts/resolveTestEntry.test.ts.

The transport helpers in scripts/tests/lib/sseStream.mjs and scriptedOpenAi.mjs are also test-only. openSseStream(baseUrl, path, token, observers) opens authenticated HTTP and incrementally decodes UTF-8 and LF-separated SSE frames. Scenario adapters own their counters, readiness promises, predicates and timeout messages: for example, delegate-e2e/run.mjs observes : connected through onComment, while hooks counts JSON events through onEvent. The parity and lazy-MCP watchers use onLine to retain raw-line idle counting and their historical lack of an HTTP status assertion. Readers discard incomplete/non-JSON frames and tolerate transport teardown as before; abort() cancels immediately and close(deadlineMs) bounds reader teardown for the API lifecycle assertion. Memory follows the current incomplete frame or line, while scenario event collections remain scenario-owned.

Minimal stream use, as continuity-e2e/run.mjs does it:

const transport = await openSseStream(baseUrl, `/brain/stream?session=${encodeURIComponent(session)}`, token, {
  onEvent(evt) { if (evt.type === 'idle') idles += 1; },
});
// ...when the scenario is finished:
transport.abort();

Scripted model servers import readJson, chunkFrame, contentText, startCompletionStream and endCompletionStream directly. Empty or malformed request bodies record null; valid JSON scalars are retained. Content matching preserves spaces for non-text parts, and response headers, data bytes and the terminal [DONE] marker are shared. Scenario modes, chunk ids, usage, tool calls, held-response timing and server teardown remain local, for example in recovery-e2e/model.mjs. These helpers only run when a test explicitly opens a stream or starts a fixture; they add no daemon startup work or external services. Continuity waits for both native idle and the host's non-streaming turn control, so automatic compaction finishes before the next send. Its prompt probe selects the first chat completion after admission, skipping independent model-catalog requests and excluding later summary requests; it still reads only that request's final user message. The helper contracts are tested in tests/scripts/e2eTransport.test.ts, and the probe in tests/scripts/continuityPrompt.test.ts. Run the continuity suite with npm run test:e2e:continuity. Each turn gets 40 seconds to settle, and the suite adds at most 14 filler turns while it waits for a compaction (scripts/tests/continuity-e2e/run.mjs:123 and :160).

Recovery E2E keeps scripts/tests/recovery-e2e/run.mjs as the sequential package and CI entrypoint. It runs all ten named modules under recovery-e2e/scenarios/ by default; RECOVERY_E2E_ONLY=1,7 selects their existing one-based positions. Each scenario owns its scripted model, real daemon, restart sequence and finally teardown. The test-only recovery-e2e/lib.mjs owns shared read-only DB observations, the store's display projection, provider-wire checks, silent-resume assertions and the attached client's real transcript fold. setCurrentDaemon(daemon) updates the timeout log source as soon as a scenario boots. Failure counts and restart timing evidence are shared with the runner; polling deadlines, log tails and the three-second pause bound are unchanged. ELOWEN_SUBAGENT_RUNNER=1 still enables and proves the forked-runner path. Daemon preparation continues through spawnRealDaemon, including its selected scratch build; these modules add no production entrypoints, persistence or background work. Run it with npm run test:e2e:recovery. CI runs the same entrypoint in the recovery-e2e job (.github/workflows/ci.yml:290). Polling waits default to a 60 second deadline (scripts/tests/recovery-e2e/lib.mjs:14), and the pause check allows 3 seconds for a clean exit (lib.mjs:17).

CLI tmux driver and evidence ownership

The eight CLI tmux scenarios share test-only process and input primitives in scripts/tests/lib/tmuxDriver.mjs. Bind createTmuxInput(server, session) to the private server returned by createTmuxServer(label); use shellQuote for launch arguments and startMock(fixture, spawnOptions) for the fixture's child and output snapshots. Readiness parsing, environment, teardown and scenario assertions remain caller-owned. Deadline polling keeps each scenario's existing interval, strict or retrying predicate errors and timeout capture. Rail polling separately preserves its 200 attempts and final 50 ms sleep. No helper starts a daemon or builds code. The default deadline interval is 30 ms (tmuxDriver.mjs:33). The driver's tests are in scripts/tests/lib/tmuxDriver.test.mjs.

Evidence has six direct owners under scripts/tests/tmux-evidence/: artifacts creates fresh scenario directories, cleans process-owned temporary directories and writes reports; frames reads strict or live JSONL and validates diagnosed frames; capture binds plain/ANSI/state triplets within one completed frame and owns pane analysis; diagnostics summarizes timing and viewport work; metadata binds the Git commit, complete dist hash and run identity; aggregate revalidates all raw evidence against those owners. Capture retries remain capped at six, ordinary frames remain below 50 ms, and required labels and report fields are unchanged. Scrollbar visibility and drag coordinates use one isolated-cell finder restricted to transcript rows. That finder is scrollbarThumb in tmux-evidence/capture.mjs:121, and the drag in cli-tmux.mjs:553 takes its coordinates from it.

For example, cli-tmux.mjs imports capture analysis and the input driver directly; cli-tmux-built.mjs runs support and driver tests before its four evidence scenarios and validates the completed root through aggregateTmuxReports. Missing, stale, malformed or inconsistent evidence still fails the gate. These modules run only in the test harness. Support tests use independent expected capture labels, analysis, performance and build-identity fixtures, not the analyzers' own output as their oracle. Full CLI proof requires a separately built test checkout with tmux. CI runs the gate with npm run test:cli-tmux:built (.github/workflows/ci.yml:131). When ELOWEN_TMUX_RUN_ID is unset, the gate generates a local- run ID once and passes it to every scenario. The gate also refuses a configured ELOWEN_TMUX_ARTIFACT_ROOT that already holds files (scripts/tests/cli-tmux-built.mjs:12-21).

Brain service scenario fakes share the tool loadout through fakePiToolState() in tests/helpers/fakePiSession.ts. Spread its result into the fake session; the existing createSession callback fills __tools, __active and model for each spawn. Each call creates independent empty tool and active-name arrays, an empty rendered system prompt, an unset model slot and a tool-selection spy. It acts only in tests that explicitly include it and installs no runtime hooks. Without it, a scenario supplies its own tool surface. It does not model prompts, events, persistence, queues, thinking, compaction or billing. The shared BrainService fixture in tests/brain/brainService/fixtures.ts is the current consumer.

Complete API and plugin disk fixtures reuse tests/helpers/fixturePlugin.ts writeFixturePlugins(root, specs); manifestOverrides preserves each fixture's declarations and moduleSource copies the complete module verbatim. Discovery-only manifests and malformed inputs stay local, and each suite still owns its roots, ancillary assets, loading options and cleanup.

For process-route authorization tests, setup({ ownerRoutes: true }) loads the terminal and subagent plugins, grants terminal to the ordinary account, and returns the in-memory UserStore as users so a test can create and grant a second administrator explicitly. Without ownerRoutes, no owner-operation plugins or extra terminal grant are installed. These are test-only fixtures with fresh stores per invocation; they do not change production permissions or perform provider requests. p34OperatorBoundary.test.ts exercises the ordinary-account denial and second-administrator success through the real plugin dispatcher and terminal gate.

Memory and plugin unit-test fixtures

tests/helpers/memoryEmbedding.ts owns FakeMemoryEmbeddings, a table-backed fake used by the five memory ranking and recall suites. Known texts return fresh float vectors; unknown texts return three zero dimensions, or two when requested by the vitality suite. Its failFor option rejects only single-text embedding with embed boom; batches still use the table. Snapshot and memory-space identity come from fakeSnapshotConfig and fakeMemorySpace in the same file (memoryEmbedding.ts:6 and :10), which eight embedding and brain test files import. It makes no network requests and has no production caller.

directEmbedBatch(embeddings) in the same file adapts a fake service's embedBatch to the worker's batch function, answered in-process, for suites that do not fork the background worker (memoryEmbedding.ts:40). A failing request fails the whole batch, as an outage would. Three test files use it: api/memoryRoutes.test.ts, api/memoryVitality.test.ts and brain/memoryMaintenanceService.test.ts.

tests/helpers/fixturePlugin.ts owns disk plugin preparation for loader, controls, alerts, daemon boot, update preflight and UI-route tests, and also config, system readiness, plugin skill manifest, licence snapshot, publisher gate, apply-change and API transition tests. writeFixturePlugins(root, specs) creates each plugin directory, writes the default manifest and then applies manifestOverrides after declaration blocks. The default manifest is name, version 1.0.0, apiVersion 2, a fixture plugin <name> description and entry index.mjs (fixturePlugin.ts:40-46). Each spec supplies either a register body, wrapped in the existing function template, or moduleSource, copied byte-for-byte without source inspection. The optional provides, configSchema and userConfigSchema fields are written into the manifest verbatim (fixturePlugin.ts:8-18). For example, loader tests provide complete modules and set the manifest version, description and apiVersion through manifestOverrides (tests/plugins/loader.test.ts:23-34). Without overrides, existing defaults stay unchanged; deliberately malformed JSON is still written directly by its test. fixturePlugins loads an authoritative test context with in-process delegation, matching the existing context defaults. Extra assets, temporary roots and cleanup remain owned by callers; these fixtures never load installed plugins or run on normal daemon boot.

Minimal use, with cleanup in finally:

const fixture = fixturePlugins([{ name: 'ledger', register: 'ctx.registerSystemPromptFragment("raw");' }]);
try {
  // Exercise the code under test with fixture.provider.
} finally {
  fixture.cleanup();
}

writeFixturePlugins(dir, specs) is the variant for a plugin root that already exists, such as a folder that appears while the daemon is running. It is used by daemon/daemonBootWiring.test.ts and update/apiTransition.test.ts. Use fixturePlugins when the test only needs a fresh root.

  • The test-only tmux driver exposes stopMock(child) for the seven CLI scenarios. It ignores absent or already settled children, sends SIGTERM, waits up to one second, then sends SIGKILL if necessary. It clears its timer and exit listener. Session/server cleanup and owned temporary-directory removal stay in each scenario; cli-tmux-signals.mjs is a consumer.
  • daemonHarness.readDaemonLog(logDir) joins sorted daemon-* files and queryDaemonPid() returns the trimmed systemd MainPID output. Both propagate errors; each scenario retains its current unavailable-log/service and zero-PID policy. These helpers do not own boot, restart or teardown. Integration consumers are recovery/delegate diagnostics and the API/web PID safety checks; their migration must land with the helpers.
  • daemonHarness.configureInstanceMcp(baseUrl, token, name, args) creates an enabled instance-scoped stdio server using process.execPath, then reads the instance list and requires connected status. It checks both HTTP responses and exposes HTTP failures. It acts during parity/lazy-MCP preparation, never production boot. No server is configured unless a scenario calls it; lazy MCP is a current consumer.
  • sseStream.openEventCollector(baseUrl, path, token) collects parsed events in order, exposes connected after the subscription comment and retains deadline-bound predicate waiting with event evidence. abort is immediate; close waits for reader settlement. Scenario callers select their existing close policy and bound readiness waits. No stream is opened without a call; events are retained for the stream lifetime. Integration consumers are the API, brain and delegate harnesses, whose adapters must migrate together.
  • sseStream.watchIdle(baseUrl, token, session) preserves the raw-line idle counting shared by parity and lazy MCP. stop uses the existing bounded stream close. It does not parse JSON or change each suite's turn deadline; without a call there is no subscription. Both MCP suites consume it.
  • scriptedOpenAi.completionFrames({ id, created, model }) returns delta(value, finishReason) and usage(prompt, completion) SSE strings with the existing OpenAI completion envelopes. Metadata and token counts remain supplied by the scenario; stream headers, terminal markers, routing, held responses and socket lifetime remain separate. It has no network side effects. Parity and lazy MCP consume it; brain/hooks/workflow model migrations use the same seam.
  • Tmux evidence uses one ORDINARY_FRAME_LIMIT_MS policy of 50 ms for both raw frames and aggregate summaries. analyzeFrameDiagnostics(frames) and validateFrame(frame, prefix) have no budget override. aggregateTmuxReports accepts expectedRounds, repo, expectedCommit and expectedDistHash; it snapshots Date.now once and uses a fixed one-hour evidence window plus one-minute future skew. Forced frames retain their existing treatment. No evidence is checked until an analyzer runs; cli-tmux-support.test.mjs and the aggregate runner are consumers.