Provider catalogs and model roles
Provider model catalog
src/brain/providerModelCatalog.ts is the shared OpenAI-compatible /models fetch and parser. It requests
output_modalities=all and carries validated architecture.output_modalities through ReportedModel.outputs;
endpoints that omit the field leave it absent. The provider probe returns { models: [{ id, outputs? }] }, and
GET /brain/models (src/api/routes/brainProviders.ts:26) exposes the same optional outputs array on BrainModelOption. A provider's explicit
manual model list stays authoritative, including non-chat kinds. With no manual list, discovery keeps only models
whose outputs are unknown or include text.
src/shared/modelCatalog.ts owns modelSupportsOutput and chatModelCatalog. Core discovery, model-role
validation and CLI chat picking, direct switching and autocomplete use this output-kind seam. The browser
build-root projection in web/lib/modelCatalog.ts is byte-pinned by webMirrorContracts.test.ts; browser
callers import these filters directly. web/lib/modelProvider.ts owns brain model identity and presentation,
plus the shared embeddingModelCatalog consumed by core role and plugin configuration editors.
BrainModelOption, the complete BrainModelCatalogOption REST row and
ProviderProbeResponse live in the importless wire contract. The CLI preserves all catalog metadata,
validates provider probe object rows, turns their ids into prompt labels and distinguishes HTTP or malformed
replies from an empty chat catalog. General catalogs and explicit provider allowlists remain all-kinds.
Chat pickers use the latter helper, while embedding settings and plugin embeddingModel fields use the same
output predicate for embeddings, retaining unknown-kind API-key endpoints. OAuth embedding choices require
positively reported embeddings outputs, and the shared embeddingModelCatalog includes connected OpenRouter routes. The provider dialog passes reported kinds to
the shared ManageSelectionModal filter options; if the endpoint reports no kinds, the additional filter is absent.
For example, an OpenRouter catalog can expose an embeddings model in the provider selection while omitting it from
chat menus and offering it to embedding-role selection.
Configured endpoint identity belongs to src/brain/providerEndpoint.ts. Its pure,
browser-safe providerEndpointKind({ type, baseUrl }) returns openai, anthropic,
azure-openai-responses or custom; web/CLI provenance can import it without the registry
or credentials. The WHATWG URL parser retains host, non-default port and path, excludes non-HTTPS
URLs and userinfo/query/fragment, and uses vendor defaults for empty custom-provider URLs.
An HTTPS default port is official; proxy paths, lookalike hosts and non-default ports are not.
OAuth accounts remain a separate credential-origin decision, not an endpoint classification.
openAiApiFor lives beside this classifier and uses it only for automatic wire selection; an explicit API wins.
CLI provider controls import this pure resolver; web provider drafts use its byte-pinned build-root projection
in web/lib/providerEndpoint.ts, including the shared trailing-slash primitive. Runtime registration, automatic
API descriptions, compatibility controls and the stateful-responses gate therefore resolve the same endpoint.
supportsResponsesStateful in src/brain/responsesStateful.ts is that gate: an openai entry on the
Responses wire or the xAI account. Config persistence keeps responsesStatefulEnabled only where it holds, the
responses-stateful status route lists exactly those providers, and the spawner opts a session into chaining
under the model's PI provider; web settings read the switch from that route and hold no type list of their own.
A request continues the previous response only when every field except input, store and the per-request
max_output_tokens ceiling matches it (responsesChaining.ts); the ceiling is sized from the room left in the
window, so including it would break every continuation.
The browser never imports provider registry or credential code.
Fast routing shares it. registerCustomProvider (src/brain/providers.ts:375) owns normal and ad-hoc catalog projection,
including configured/discovered IDs, earlier explicit additions, pins and native descriptor provenance;
adding a hand-typed model cannot drop discovered models or make a relay inherit native input limits.
These are local parser/catalog reads, with no fetch or extra credential access.
Vision fallback reads session.model.input, the resolved PI descriptor used to send or strip images,
so endpoint reports and explicit capability overrides apply before decideVisionHop (src/brain/visionFallback.ts:11).
Instance conversation model role
brain.defaultModel is the persisted atomic { providerId, model } pair in ConfigStore, or null before a provider is connected. brainDefaultSelection (src/brain/providers.ts:589) is the sole runtime resolver for an empty personal chat selection, catalog default flags and serverDefaultRoute. Config provider IDs are never replaced with registry IDs. Settings' Model roles row writes the pair through admin-only PUT /config; the API validates the complete pair against the offered text-chat catalog with isOfferableExec (src/shared/execs.ts:152). A provider list and role in one patch use the sanitized draft catalog, preserving omitted credentials. Config revision checks and the store write lock still govern persistence.
The existing ConfigStore.migrateRetiredPluginConfig (src/store/configStore.ts:1086) boot cleanup removes inert root customModels, hiddenPresets, and modelNotes in one write-locked transaction. Unknown installation keys, nested plugin settings, executor allow-lists, and current brain/provider fields are preserved. A changed settings row advances its revision once; rerunning the cleanup performs no write. There is no replacement store or network lookup, and runner construction with migrate: false never runs it.
At daemon boot, initializeBrainDefaultModel (src/brain/config.ts:69) migrates only settings that lack the field, preserving the formerly effective first-provider model, including catalog-backed OAuth defaults. Runners with migrate: false do not migrate. A fresh or providerless install stays null. Provider saves, OAuth connection completion and background catalog snapshot changes initialize a pending null role once a model is available. There are no network requests on reads or turn startup; initialization uses local PI descriptors and persisted provider snapshots. Empty custom OpenAI-compatible providers register the same discovered text catalog that the picker offers, filtered once by discoveredChatModels (src/brain/providerModelCatalog.ts:31); this lets delayed setup and stored discovered defaults resolve without ad-hoc positional routing. BackgroundWorkerDeps.onModelsChanged is an optional awaited hook after changed snapshots are restored, used by maintenance for this pending initialization; absent, refresh behavior is unchanged.
Removing a provider or model retains the stored pair. The Settings picker pins it as unavailable rather than rewriting it. Spawns retain the existing account authorization chain: the permitted instance role, then account allow-list pairs, then permitted configured models, including local registry descriptors for OAuth and automatically discovered providers with empty manual lists. ServerDeps.effectiveChatRoute projects the existing daemon accountChatRoute authority into GET/PATCH personal CLI settings, separately from the global serverDefaultRoute; absent wiring reports null, never a guessed route. Account's model label, reasoning controls and vision capabilities consume this projection without new routing logic. Provider connection tests name a concrete offered chat model, so testing a non-default provider never relies on a positional conversation default. Personal chat and compaction choices remain independent; without a distinct valid compaction choice, compaction uses the resolved chat model. For example, an explicit Codex role wins even if Claude/Fable is first in provider order.
PI model catalog overlay and provider model lists
buildBrainRegistry in src/brain/providers.ts is the sole descriptor preparation seam for the shared
ModelRuntime. Spawning, inference, status and delegated-child identity all call it. It retains one read
facade per runtime and compares provider settings and model overrides by value against a private snapshot,
so fresh config objects and in-place edits have the same semantics. An unchanged generation only compares
that small configuration and looks up descriptors; it never registers providers, refreshes availability or
reads credentials. No providers means no custom endpoint registrations; the same native catalog and models.dev supplements remain readable by account pickers.
modelRegistryGeneration.ts carries explicit publication revisions. A successful PI catalog restore,
a changed endpoint snapshot or retained-endpoint pruning, a changed file credential, and completed OAuth
connect/disconnect invalidate preparation. Snapshot timestamps alone, failed reads, timers and PI's own
registration-generated refreshes do not. Revisions and prepared facades are process-local, not shared
between the daemon and runners. After a successful PI restore or effective endpoint-snapshot change,
createBackgroundWorker.onModelsChanged flows through maintenance to SubagentRunnerPool.modelsChanged.
The pool sends one typed modelsChanged IPC frame to each existing process, including a booting runner;
with no pool or no runners it does nothing and forks nothing. RunnerModelCatalog serializes offline
runtime.refresh, local generation invalidation and endpoint-snapshot reload. Later turn frames await
the publication tail before resolving their models; a failed PI restore is logged by the runner and its
last good catalog stays in use, as in the daemon's own worker pass, until a later publication succeeds.
Unreadable endpoint snapshots retain the last good copy and report the read problem.
A runner booted afterwards restores current files as usual. No polling or per-progress work is added:
publication costs one offline refresh per runner, only when descriptors change. IPC backpressure queues
frames and is not mistaken for delivery failure. Provider additions, edits, deletes, connected-account
changes and
overrides are detected on the next call by their effective config values. Preparation removes deleted custom
providers and departed OAuth overlays before applying current descriptors. A missing OAuth credential removes
an account from the runnable set, but configuredBrainProviders still retains its explicit stored identity:
revocation is not provider deletion and must not erase preferences or permissions. Request-time authentication
still uses the fresh credential store, never a cached token.
An ad-hoc route registers through the same custom-provider projection, records only explicitly added ids,
and advances the prepared generation immediately only when no earlier publication is pending. An ad-hoc
registration updates one provider and must not consume unapplied changes to others. Later config/catalog
preparation reapplies those ids with current metadata; discovered ids are not retained as ad-hoc additions, and provider deletion forgets them.
For example, non-live delegated progress reads only the selected descriptor from this prepared source.
Its thinking fallback uses BrainStore.journalContextState and PI's default off, then the existing model
clamp, rather than rebuilding a session. Its cost is independent of transcript size, including branch moves
and compaction. Real descriptor changes still incur PI registration/catalog work once per generation.
Built-in providers (Claude and ChatGPT accounts, GitHub Copilot, Kimi Code, xAI, Meta and OpenRouter accounts and the rest of PI's catalog) list the models
of the pinned PI release. PI also publishes its catalog to https://pi.dev/api/models/providers/<id>, and
ModelRuntime layers that list over the built-in one, so a model PI adds appears without an Elowen release. Elowen
keeps PI's descriptors authoritative. The worker also supplements new official Anthropic, xAI and Meta
chat releases from models.dev while PI's curated remote catalog catches up. Codex is excluded: its
subscription API differs from the OpenAI API catalog, and models.dev has no matching Codex section.
catalogSupplement.ts validates the models.dev boundary with TypeBox and persists only model identity,
name, all token prices including context-length tiers, limits and a PI sibling id in
catalog-supplement.json. buildBrainRegistry composes these through the existing OAuth descriptor overlay used by model overrides and OpenRouter
additions. There is no second registry or transport definition. The sibling comes from the same provider,
API and canonical family: Claude haiku/sonnet/opus/fable, Grok, or Muse Spark. Numeric versions are compared
deterministically; only releases newer than the newest native sibling are added, and an earlier release
date than that sibling's models.dev date is excluded when available. Dated snapshots, aliases,
deprecated and non-text entries are excluded. A missing same-family sibling skips the model and logs
its identity once per daemon lifetime instead of guessing transport settings.
For example, claude-haiku-5-5 inherits Anthropic request compatibility, thinking map, prompt cache,
reasoning, modalities and input limits from claude-haiku-4-5; its prices and context/output limits
come from models.dev. Every base and tier price must explicitly include input, output, cache read and
cache write rates. The fetch validates the top-level provider/models envelope; each model row is
validated independently by the supplement reducer. Raw rows may carry non-chat limits such as an output
limit of zero; only supplement candidates require positive integer context and output limits. Invalid
rows and incomplete or unknown prices skip only that model. Price objects and context tiers accept only
explicitly supported keys; an unknown surcharge such as context_over_200k is not silently dropped.
Skipped identities are logged at most once each (src/brain/modelCatalog.ts:324), capped at 1000 identities per daemon lifetime to bound
both memory and log volume. Missing rates never become zero. Context tiers are validated at the row boundary and persisted in PI's cost.tiers
shape with inputTokensAbove. PI selects the highest threshold exceeded by the full input usage,
including cache reads and writes, and applies that tier to the entire request. Spend admission and
settled usage continue through the existing PI pricing path.
A native PI entry always wins by id, including while models.dev is down. The worker retains takeover
descriptors in the supplement and returns their identities in private IPC retired; only the daemon
removes them after its offline PI store restore succeeds. A failure of a different provider does not
prevent restoring successful PI publications. Failed restores retain the takeover descriptors and the
last restored store revision, so even an unchanged store is retried on the next pass.
Failed fetches or invalid top-level documents preserve the last supplement and PI catalog and use the existing
one-warning-per-outage provider-catalog reporting. Failed disk reads retain the process's last good
copy; a corrupt supplement is reported independently and does not abort PI or provider refreshes.
A valid models.dev reduction can atomically replace that corrupt file.
The fetch is shared with the Fast overlay: one multi-megabyte download on boot, every four hours and a
forced refresh, bounded by the existing thirty-second fetch and sixty-second worker deadlines.
PI_OFFLINE prevents it. Only the small reduction remains on disk and in memory; registry preparation
does no network work. New supplement ids appear in the refresh result's added list. The daemon reloads
the reduction and publishes effective changes through the existing model generation and runner
notification seam. Its last announced supplement content is tracked independently of picker reloads,
so a picker cannot consume a runner notification. Warm runners reload it after their offline PI restore.
Fresh picker runtimes read the same instance file, including the credential-less OAuth catalog used in Settings. Existing account
allowlists and model overrides still apply; no new picker marker is introduced.
A configured OpenAI-compatible provider is the other half of the same job: its own GET /models answers what it serves
and with which limits (context window, output tokens, vision, output kinds, prices). src/brain/providerModelCatalog.ts
owns that read model. No daemon path calls /models, and no conversation start waits on a provider's server.
- Persistence.
models-store.jsonin the brain directory, besideauth.json, never PI's default agent directory, and next to itfast-models.json(the Fast-mode overlay, see "Fast mode: capability and choice") andprovider-models.json: one entry per endpoint, keyed by the endpoint's normalized base URL plus a SHA-256 of its API key (providerModelsKey) and never the key itself, holding the models that endpoint last reported and when. The worker replaces that file atomically (temp file + rename).buildBrainCore(src/daemon/brainCore.ts:308) points every runtime of the process at the overlay (useCatalogStore) and every reader at the snapshot (useProviderModelsSnapshot), and hands the daemon's maintenance loop every path. The daemon, every forked sub-agent runner and the credential-less runtime behindlistBrainModels,oauthBuiltinCatalogand the account default (inMemoryModelRuntime) restore the same last good copies, so a picker offers what a session can run. A missing snapshot (a fresh install, the:memory:test database) is an EMPTY one: manual model lists and catalog limits answer, exactly as the previous in-memory cache did when it held nothing, and a read path still never calls the network. - Who fetches. A dedicated WORKER PROCESS (
src/brain/backgroundWorker.ts), forked by the daemon for one request and nothing else. That is the point of the split: the network calls and PI's rebuild of the model runtime stay off the daemon's event loop, next to live turns.catalogRuntimenever enables network, so creating a runtime restores the file and returns without waiting on pi.dev, and a runner, a picker or a test cannot reach the network. One pass does ALL its fetches — PI's catalog, every configured endpoint's/modelslist and, on the catalog's cadence, models.dev for Fast support and supplements (fast-models.json,catalog-supplement.json) — side by side. Supplement reduction waits for the refreshed PI descriptors so new PI entries take precedence. The daemon'screateMaintenanceLoops(src/daemon/maintenance.ts:85) forks the worker in the background right after boot, then everyWORKER_PASS_INTERVAL_MS(15 minutes,src/brain/modelCatalog.ts:66); it never fetches itself. PI's own four-hour freshness gate decides inside a pass when its catalog is actually refetched, which is what keeps that cadence from hammering pi.dev — the cadence exists for the provider lists. A provider saved in Settings does not wait for a tick either: the live worker is asked for a pass (retainProviderModels, called by the config store (src/store/configStore.ts:1436) on every provider change). Successful OAuth connect/disconnect also requests a pass. The worker registration carries its livebackgroundWorkerTargetsgetter, andretainProviderModelsuses that same target set so a synthetic connected OpenRouter snapshot survives unrelated provider edits and is removed on disconnect. Stopping the registered worker releases its pass requester and target getter; processes without a worker retain only the endpoints supplied by the caller. The OpenRouter account target fetcheshttps://openrouter.ai/api/v1/models?output_modalities=allwithout an authorization key; other OAuth accounts add no endpoint target. One pass at a time, and a request that lands while a pass runs is remembered and runs once after it. A pass that has not answered within 60 s is given up on and logged, a pause or shutdown gives up the same way, and the child gets SIGTERM first so it can abort its fetches and release PI's models-store lock, with SIGKILL only after a five-second grace. An admin can callPOST /brain/models/refreshfrom Settings → Models: the same worker queues one forced pass behind any current pass, joins repeated manual requests, asks PI withforce: true(bypassing PI's stored freshness gate), refreshes the provider/modelslists and the Fast-mode overlay, and resets the daemon's four-hour gate. It reports added models and per-provider failures, then restores the shared runtime, provider snapshot and catalog supplement before replying when successful. It never overrides PI'sPI_OFFLINEnetwork switch; a request before worker startup returns 503 (src/api/routes/brainProviders.ts:41-44).:memory:has no paths and never forks. - Module ownership.
modelCatalog.tsowns instance catalog paths, offline runtime restoration, the pass cadence and serialized probes.backgroundWorkerProtocol.tsowns request/reply types and their boundary parsers; it has no process or persistence side effects.backgroundWorkerFork.tsresolves the packaged/source worker entry and owns IPC send, reply/close/timeout settlement, cancellation and TERM/KILL escalation.embedBatchWorker.tsowns all live daemon-side embedding children:createBackgroundWorkerresumes batch admission at startup and stops every active batch at pause/shutdown, refusing new batches until startup.buildBrainCoresupplies the queue and owner reindex withembedBatchThroughWorker, which admits the account, resolves its endpoint and commits reported usage before judging vector count or shape. Without worker paths, maintenance creates no worker; missing embedding configuration leaves queue bodies pending. The pass, probe and batch bounds remain 60, 10 and 45 seconds (src/brain/modelCatalog.ts:75,80,src/brain/embedBatchWorker.ts:14), followed by the same five-second kill grace (src/brain/backgroundWorkerFork.ts:10).embeddings/embeddingResponse.tsis the pure response owner shared by direct requests and worker replies: it ownsEmbeddingRequestUsage, parses reported usage and validates row ordering, finite nonempty vectors and configured dimensions. Service and IPC consumers import the usage type from this response boundary.EmbeddingService(src/embeddings/embeddingService.ts) retains HTTP, credentials, admission and accounting; parsing never performs I/O. - How it is asked. The request travels over the forked child's IPC CHANNEL, never as an argument: a pass carries the
providers' API keys, and
/proc/<pid>/cmdlineis world-readable. A PROBE — the provider dialog testing an endpoint nobody has saved yet — is the same worker and the same parse code on a shorter leash, as a one-off request that stores nothing (probeProviderModels, reached throughPOST /brain/providers/probe). Probes are SERIALIZED (src/brain/modelCatalog.ts:421-426) — one child per daemon at a time, each caller waiting its turn for its own endpoint's answer — and the HTTP caller's own signal cancels its probe, so closing the dialog stops the fetch instead of leaving it running. The memory drain's embed batch is the same pattern on the same process: one request, one answer, nothing persisted by the child. - Memory embeddings. This is the worker's third job, and its only input that is not a file. The drain
(
EmbeddingQueue,src/embeddings/embedQueue.ts:71) picks the memories that still need a vector out of the store, inside a body count AND a character budget, and the daemon hands the batch to one child (embedBatchInWorker), which posts it to the configured/v1/embeddingsendpoint in a SINGLE request and answers ONE RESULT PER BODY over the IPC channel. The credential travels with the request — the child resolves none of its own — and the daemon validates every vector itself (width, finiteness, emptiness) before storing it, so a reply from a stray process cannot become a memory's embedding. Rows are placed by theindexthe endpoint sends when it sends one, so a provider that answers out of order cannot hand one body's vector to another. A request the provider REFUSES is narrowed to one body at a time inside the child, which is what keeps a memory over the model's limit from costing the memories batched beside it: it loses its own vector, is retried once per tick and never blocks the backlog behind it. Only a failure of the endpoint as a whole — a transport error, a timeout, a rate limit, a 5xx, a dead child — costs the tick, with its reason logged and the same bodies retried later, and every body of it counted against the budget. The narrowing is bounded too: a provider that refuses body after body is refusing the REQUEST, so after a handful of individually refused bodies (MAX_INDIVIDUAL_REFUSALS, 5, insrc/brain/backgroundWorker.ts:125) the rest of the batch is reported with the batch's own reason instead of one paid request each, and the batch is retried as a whole on a later tick at one attempt per body. A pause or shutdown gives up on a running batch instead of waiting for its own timeout. Keeping this out of the daemon is what stops a slow or unreachable embeddings provider from holding the event loop a live turn runs on. The QUERY embedding of a turn (turn-start recall and live recall) stays a DIRECT call: a turn cannot wait for a forked child. The drain sends ONE BATCH PER USER (the caps still bound the whole tick), because the usage of a request can only be attributed to one account, and the child reports what each HTTP request it made consumed (usagein its reply, one entry per request, refused requests excluded); see "Embedding usage" under the usage tables below for where that goes. The owner reindex (MemoryMaintenanceService, behindPOST /memory/maintenance/reindex) sends its snapshot through the sameEmbeddingBatch, cut by the drain's body count and character budget, so a few hundred memories cost a handful of requests; a batch that failed as a whole counts against its own bodies and the job moves on to the next one. Its batches can be in flight beside the drain's, and a pause or shutdown gives up all of them. Maintenance progress and results live inmemory_maintenance_runs, exposed throughMemoryStore.maintenance. The existing additive schema migration creates this table on upgrade without advancing the numbered migration ladder.MemoryMaintenanceServicecheckpoints each completed embedding batch or categorized memory and publishes terminal counts. Only daemon bootstrap callsrecoverInterruptedRuns(): attached runners share the same database but must not interrupt live daemon work. Unfinished runs becomeinterrupted, retaining observed counts, and are never resumed or reported as successful.GET /memory/maintenancereturns the latest two operation slots plusruns, a bounded twenty-row projection of the caller's UI-started running work and unacknowledged results from the last 24 hours. Full human browser credentials mark API starts as UI work; tools and scoped credentials do not.POST /memory/maintenance/:id/acknowledgeacknowledges only a settled receipt owned by the caller, including earlier receipts after another start. No maintenance service means the existing routes report unavailable. Account removal deletes the receipts through the existing memory cleanup transaction. Reads make no model calls; starts keep the existing single-running-operation-per-owner lock and inference costs. - Read-only credentials. The worker opens
auth.jsonthroughReadOnlyCredentialStore(src/brain/credentialStore.ts:373):read/listdelegate,modifyhands back the credential on disk without running its callback, anddeleteis refused. Token rotation therefore stays with the daemon's own refresh loop. It is enough because PI consults a credential only to decide which providers are refreshable (resolveRefreshCredential) and the catalog request carries none, so an expired token still lets the fetch run. - Picking the file up. A pass that rewrote the store is handed to the daemon's own runtime
(
brainRuntime.refresh({ allowNetwork: false })) — a file read and a provider recompose, never a request to pi.dev, and nothing at all for a pass that wrote nothing (a gated boot pass, an air-gapped host, a provider without a credential). Which one it was is judged by the store file's identity around the pass, not by the added-models list: a descriptor that changed under an existing id never appears in that list, and a304that only moves the freshness marker is still a write to the file. Without the restore the shared runtime keeps the overlay it restored at boot, so a model the worker added would reach the settings pickers at once but a live session only after the next restart. A runner keeps the overlay it restored when it was forked, so it learns a model added later on its next start (idle runners are reaped after two minutes,src/subagent/sizing.ts:142). The provider snapshot is re-read the same way, by the file's identity around the pass, and reloaded even when the catalog fetch failed: the two jobs of one pass are independent. - Failure. A failed pass keeps the last overlay and logs one warning per outage, not one per tick; a pass that adds
models logs them at info, and providers that failed are named while the rest of the pass still lands. An endpoint
whose
/modelsfetch failed keeps the list the last good pass stored, and the warning names the SET of endpoints that are failing, so a second one going down is logged while the first is still down and a pass that never answered (a timeout, a dead child) leaves that state alone. A snapshot the worker cannot WRITE is reported as one more endpoint failure rather than costing the pass its catalog result; the daemon's own re-read of the file keeps the lists it is answering from when the file cannot be read or parsed, and reports that once. An endpoint that is no longer configured leaves the snapshot on the next pass, and an unreadable provider list is that pass's failure rather than a rejected tick. - No cache to invalidate.
buildBrainRegistryruns per spawn and per listing and reads the runtime's current models; the OAuth override pass unregisters its own earlier registration first, so a pinned account also sees new models on the next build. A running conversation keeps the descriptor it started with. - Air-gapped installs.
PI_OFFLINEset to any value in the daemon's environment turns PI'S catalog fetch and the models.dev Fast fetch off entirely — the worker inherits that environment, PI reads the variable itself and the worker checks it before fetching models.dev — and the built-in catalog, the generated Fast baseline and whatever overlays are already persisted stay in use. It does not stop the provider lists: those endpoints are the ones the operator configured for their own conversations, so a host that can reach them still gets their model lists. Seedocs/DEPLOYMENT.md.
Custom openai-type endpoints still price and size a model from PI's bundled models.dev data (nativeListPrice,
catalogLimits in providers.ts), not from the overlay, so a relay serving a model newer than the pinned release keeps
the operator's own price and limit overrides as its source. An anthropic-type entry and every OAuth account read the
overlay descriptor. For example, claude-sonnet-5-5 comes from the overlay with its own name, price and limits.
Provider temperature remains configured on the provider entry. buildBrainRegistry publishes a
runtime-local policy keyed by PI registry provider id; wrapModelStream passes the actual request
model's value through PI's typed temperature stream option. Nested wrappers share policy updates
without mutating PI runtimes or descriptors. Owner sessions, channels, delegated runners and
connectivity probes use this seam. Unset temperature sends no provider-specific value. PI owns
transport-specific temperature guards and thinking serialization; Elowen has no provider request-profile
payload hook or Alibaba/Qwen budget rewrite. Policy lookup adds no provider request or model tokens.
Custom OpenAI-compatible Chat Completions endpoints retain their explicit compatibility baseline. Base-URL catalog lookup supplies model limits, not native provider compatibility. A custom DashScope endpoint therefore does not inherit native Qwen thinking serialization or its token-budget field merely by sharing a native provider's URL.
OAuth model overrides preserve every PI model kind, keyed by kind and model id. Chat settings apply only to chat descriptors; image and classifier descriptors survive unchanged, including ids shared with chat models. OpenRouter snapshot additions remain chat-only.
catalogRuntime reads models.json only beside the instance's models-store.json, never from PI's global agent
directory. Without a configured catalog store it uses modelsPath: null and keeps the runtime in memory. Missing
instance configuration leaves PI's native catalog intact; runtime creation performs no network requests.
The independent CLI update coordinator uses the CoreHalf seam in src/update/transaction.ts: prepare downloads
and verifies the candidate while services run. Bootstrap, update and the licensing release proof share
scripts/releaseInstall.mjs: extract the release as a root package and run npm ci --omit=dev --engine-strict
with native install scripts enabled, preparing Python, make and g++ first. install.sh downloads and verifies
that release before provisioning Node, reads its engines.node with Python's JSON parser, and derives both
the minimum runtime and NodeSource major from that manifest. The bootstrap contract accepts one canonical
>=major.minor.patch requirement and fails closed on any other range; it has no hand-kept version copy and
needs no checkout manifest when invoked directly from the repository. Missing Python is provisioned only
after the licence and artifact gates. The activation key and installation identity are passed only in
the provisioner's child environment, not exported to later release commit or recovery processes.
Pack-time artifacts derive a production
manifest without development dependencies or workspace discovery, and its production root lock, from the committed
manifest and lock. The shipped shared workspace remains a locked local link; the unshipped UI workspace is excluded.
apply moves the candidate under <prefix>/lib/elowen-releases/<version>-<transaction> and atomically replaces
<prefix>/lib/node_modules/elowen. Stable command, systemd and sudoers paths use that link.
src/shared/installResolution.ts owns the same install-kind resolution for the updater, provisioner and
plugin host. src/plugins/hostModules.ts publishes the stable package's node_modules and package link
for installed plugins, including the shared user-plugin node_modules/elowen link. Those links follow
core-only swaps while marketplace housekeeping is locked, and survive retained-release pruning;
developer checkouts keep their dependency-derived module layout. The existing core journal
records intent before the switch, with a null previous version only for fresh bootstrap. Directory changes are synced;
converge refreshes root-owned host files before start where the updater has root access. commit marks proven
bytes before the journal becomes terminal and prunes only marked owned releases, retaining current and previous.
Each retained backup keeps the existing package inventory beside it; cleanup checks those bytes before removal
and leaves unrecorded backups untouched. Fresh-bootstrap recovery removes only its recorded package and command
links, refuses unfamiliar replacements, and uses the candidate recovery module even before its live link exists.
restore atomically selects the retained release and is followed by the same converge before restarting it.
Without systemd, the updater resolves and imports the currently selected launcher's canonical path through
that stable installation on every start, including rollback; it never starts from the coordinator's
already-imported retained launcher. Developer checkouts continue using their own launcher. clean removes staging only when the
core swap record is no longer applied. src/cli/updateTransactionIO.tsupdateCoreHalf, lines 158-211) supplies the writable-package and root-owned implementations, src/cli/update.ts decides writability (line 484) and runs the root-owned unit convergence (convergeRootOwnedUnits, line 205); for a plugin-only transaction there is no core apply, converge, commit, or restore. For example, the root-owned implementation sends a token only with the staging request and re-executes the installed release's privileged refresh before its daemon boots. For a writable service-user installation, convergehas no privileged write to perform and the existing drift report runs after restart. Seedocs/DEPLOYMENT.md` for operator recovery commands and the worker's
drain limit.
Plugin candidate validation uses MarketplaceServiceOptions.candidateManifest (src/plugins/marketplace.ts:233): the CLI (src/cli/updateTransactionIO.ts:244-265) binds the
parseManifest export from the verified staged target's dist/plugins/manifest.js after core preparation,
before extracting or validating plugin payloads. A plugin-only run and the daemon use their own parser.
There is no API-version fallback or compatibility mode. A selected payload reported as cannot-stage
aborts the entire transaction before stopping services; consent refusals remain explicit skipped candidates.