Voice calls and account spending
Voice and account spend seams
The voice foundation defines service interfaces in src/voice/types.ts for pricing, usage settlement,
speech, calls and rings. Optional VoiceServices slots pass through BrainCore and ServerDeps;
RouteContext exposes only the call and dictation engines, ring coordinator, pricing and provider predicate
used by routes. Speech, connection, usage settlement and account spend remain on ServerDeps. The call engine
is assembled in daemon bootstrap from the injected voice connection,
usage sink, pricing and spend gate, all required assembly inputs. Only an absent brain leaves the engine
unassembled. POST /voice/calls requires a full human-login credential and reserves one account slot before
provider admission. Expected refusals use VoiceCallRefusedError.code, exhaustively mapped to 403/404/409/503;
provider faults return 502 without inspecting exception text. Account spend refusals use sessionRouteError
and preserve its 402 reason and reset date. /ws/voice/ redeems a ticket bound to that account and call;
attaching clears its 30-second reservation deadline. Transport faults are reported separately from invalid
browser frames. Both plugin and voice upgrades flush the shared bodyless refuseWebSocketUpgrade response.
RingCoordinator.initialize binds core collaborators after construction; bindWebPresence binds the API's
existing presence registry through the same interface, without requiring the concrete coordinator class.
A stale, expired, declined or already claimed ring refuses a claim with callback_unavailable.
The upgrade dispatcher shares the existing
same-origin check and one-shot webSocketTickets store; voice uses the reserved core:voice namespace.
PluginHostVoice carries an explicit SessionSource and declares the reads:['voice'] capability;
the host method and its core speech wiring are required. Missing capability throws; disabled voice,
an unlinked sender, a missing grant or a reached spend limit returns a typed refusal.
GPT-Live calls
VoicePeer in src/voice/types.ts is the one far-end transport for calls and dictation.
browserSocket.ts binds a browser socket to PCM audio, optional display controls, graceful close and abort.
It preserves readiness and backpressure checks. sendPeer routes audio separately from display frames;
without a peer, nothing is sent. Shared teardown settles usage before sending the ending frame and closing,
and aborts the peer if either send or close fails. Plugins supply CallMedia, PCM16LE mono 24 kHz audio
and a close callback, without a display channel; conversion belongs to the plugin.
src/voice/live/session.ts is the only call engine. It connects with the shared VoiceConnect
from src/voice/realtime/openaiSocket.ts to /v1/live/sessions, without query parameters.
session.start specifies gpt-live-1, store: false, PCM16LE mono at 24 kHz, the chosen voice,
and delegation: { type: 'client' }. It waits for session.started before accepting the browser.
The browser still uses an account-bound one-shot ticket, bounded binary PCM and strict hangup/played controls.
Call and dictation ticket redemption re-resolves the owner from the live user store; a deleted owner
is refused with 401 before WebSocket upgrade or engine attachment.
Microphone frames become session.input_audio.append; session.output_audio.delta.delta supplies
ordered PCM. Primary OpenAI WebSocket output has no audio timestamps or output-done event. Live itself
controls simultaneous listening and speaking. There is no Realtime response cancellation, item truncation,
VAD configuration or browser speech-start flush. One audio_started identifies the continuous stream
as live:<callId>. Monotonic, bounded played.ms tracks every consumed sample, including exact-zero PCM
frames that continue throughout silence. The shared pcm16HasSignal helper in shared/voice.ts
separates playback activity from transport activity, with a byte-pinned browser mirror. Only a non-silent
frame advances the last audible stream position and restarts the one-second goodbye gap timer.
Closing requires new non-silent output after its directive and playback through that stable position;
later silence does not postpone closing. The existing bounded closing deadline remains. This is playback
tracking, never caller VAD, turn detection or interruption control. The browser plays queued PCM in
arrival order and does not flush it when the caller speaks. Live can change its ongoing output while
listening, but previously buffered audio may continue playing; interruption is not an instant playback
cut. The worklet's thirty-second queue limit fails the call on overflow rather than trimming audio.
The continuous stream has no interruption boundary with which a worklet-only flush could distinguish
obsolete output from the next response.
Transcript deltas have delta, start_ms, end_ms on the session timeline that input and output share.
Caller words are transcribed after the model has started answering them, so arrival order can put a reply
before the words it answers, or split one reply around them. callState.ts therefore keeps the fragments in
start_ms order, and transcriptEntries builds the record from them, ordered by where each entry starts:
a speaker's fragments less than 1.2 seconds apart form one entry, which never defines an authoritative caller
turn. The other speaker closes that entry only by starting at or after its end; talking over it does not split
it, because on a phone the model often starts its reply before a short answer reaches it, and splitting at the
overlap broke the reply mid-word. A short answer spoken during the reply is therefore listed after it. Interpreter input, the stored call record, the third-party result
and the ended frame share this ordered transcript, preserving the readback before a caller confirmation.
Browser live captions keep arrival order with the same grouping until the ended frame replaces them.
The record retains the newest context within VOICE_CALL_RECORD_BOUNDS: entries retain their last
16,000 UTF-16 code units, the oldest fragments are removed beyond 200,000 total code units and the
oldest entries beyond 1,000. These record limits never end the call. Browser deltas and exact new caller
words continue unchanged; interpretation receives the retained transcript. Every fragment carries its
arrival number, and a delegation's mark is the first arrival it did not see, so a caller word transcribed
after one delegation reaches the next one exactly once, in its spoken place.
session.start.input contains the last voice.contextMessages settled written messages, without tools.
src/voice/live/context.ts selects newest messages before serialization, reserves the pending question
as an indivisible native text message, and conservatively budgets 7,600 UTF-8 bytes including per-message
overhead below the provider's 8,192-token and 128-message bounds. An oversized question refuses startup. The oldest message that still fits is kept only as its leading part, and older messages are dropped.
A zero history count still preserves a callback question. Question callbacks actively append their
opening question on browser attachment, before the caller speaks. Done callbacks likewise speak their result.
session.delegation.created provides an opaque client delegation ID, not task text.
The engine captures caller fragment boundaries and serializes interpretation and admissions.
src/voice/live/delegation.ts uses the conversation's permitted model through voiceInference,
the existing account-owned piInferenceClient and withOriginUsage, without tools.
Strict JSON decisions select work, a confirmed question answer, clarification, call end or callback.
A work decision also carries summary, one clean sentence in the caller's language. Work is admitted
only when new, exactly received caller words exist; concurrent delegations compute the remainder after
earlier admissions. The model-facing text is the dialogue alone (Voice call transcript: then
role: text lines): the first request carries the retained transcript, and every later request in the
same call only the dialogue after the previous consumed delegation's mark, so no words reach the agent
twice. The summary is only the display text, the bubble label and the conversation title source; the
agent works from the transcript. The immutable caller, session and origin supply
authority. Teardown also waits for dispatched admissions within the settlement budget, so an accepted
brief cannot be reported as unsubmitted when hangup races admission. Every work request carries
surface: 'voice', trusted provenance and a bounded VoiceCallRecord holding exactly its own transcript
segment, which the web renders as a call card. Status and memory queries use
this same BrainService path. There is no Live function-tool catalog or OpenAI Responses backend.
Question answers reuse brainAnswerSchema and BrainService.answerQuestion, rechecking the currently
pending, claimed non-approval question and bound account after inference. Approvals remain in written chat.
Delegation IDs are deduplicated; another ID with no new caller words cannot repeat a successful submission.
Speech interruption and hanging up leave admitted backend work running. Ending a call cancels any
in-flight interpretation through its own abort signal, without closing the voice transport before final
usage settlement. No interpretation result can start work after hangup. Invalid interpretation fails
explicitly and does not mutate the conversation.
Holding uses quiet session.thinking.append; visible progress and verified results use session.commentary.append;
trusted steering and goodbyes use session.instructions.append. Plain content is split at Unicode
boundaries into at most 480 UTF-8 bytes, conservatively below 500 tokens. Task progress and the consolidated
conversation result use the most recently accepted work or confirmed-answer delegation_id; ignored,
clarifying and refused later events cannot replace it. Accepted follow-ups steer the same conversation,
whose result is delivered once after every accepted operation completes and the conversation is idle.
Direct clarification uses its own event ID; session-wide steering uses null. Append acknowledgments describe
context injection, not speech completion; the engine never uses them as evidence of playback.
Third-party media and external costs
The same engine has typed owner and third-party call states. Owner-only helpers accept OwnerCall;
third-party state contains no conversation identity, callback, ring or backend admission state.
startThirdParty({ userId, origin, language, brief }) reserves the account synchronously and
checks human voice authority, enabled voice, spend, provider/voice eligibility before dialing. Failure releases the reservation and never opens an upstream. A 120-second
attach deadline prevents a lost bridge from retaining the account slot. attach(media) opens the
same Live connection and session configuration with empty history, plugin-supplied instructions and
the account's voice, then instructs Live to start the call as the brief says. hear and played use the existing PCM
bounds and cumulative playback accounting; transport conversion remains plugin-owned.
| Path | Owner | Third-party |
|---|---|---|
| History and identity | Owned conversation | Brief only, empty history |
| Delegation | BrainService admission | Commentary refusal with the delegation ID |
| Ring and subscriptions | Existing callbacks and progress | None |
| Closing directives | Owner templates | Neutral third-party goodbye |
| Result | Existing summary | Same summary: callId, reason, seconds and the bounded transcript |
PluginHostVoice.thirdPartyCall requires reads:['voice'], a userGrantable plugin and the paying
human account's explicit plugin grant. Core stamps platform:<plugin>; no plugin-supplied origin or
conversation authority crosses the boundary. Bootstrap late-binds the engine and its existing usage
sink through setPluginHostCallEngine; a forked runner has neither and rejects daemon-only work.
Registry Twilio consumes the handle, bridging PCM16LE mono 24 kHz. With no consumer nothing runs.
After media and usage teardown, the result promise returns callId, the end reason, seconds
and the bounded transcript; it never rejects. Restart takes priority during settlement and reports
daemon_restart. Core performs no outcome inference. finishedCalls keeps the record for the placing plugin and its control callers for six hours or until restart, so a later conversations.send with that call.callId can attach the call card. Live duration bills the account and plugin origin;
the plugin owns any interpretation and can use the existing conversation seam to supply the transcript
to its agent.
PluginHostUsage.reportCost requires reads:['usage'] and the same explicit human-account plugin grant,
without requiring a voice grant for generic external work. Core stamps a namespaced event ID, plugin
provider, item model, current time, plugin origin and zero tokens. VoiceUsageSink.recordCost passes
plugin-computed micro-USD through VoiceUsageStore.record, including its nonnegative safe-integer check,
idempotent receipt, write lock, atomic existing provider/day and origin rollups, and failure health guard.
There is no independent usage authority. Registry Twilio reports once after a billed call; the next
account admission sees it. Without a report, no cost is added. There is no mid-call cost estimation.
Holding the line, and calling back
Admitted work sets browser state on_hold while microphone and output audio continue.
All admitted briefs must settle and the conversation must be idle before the newest written result
is spoken through commentary. Delivery atomically takes the durable callback request.
While work is pending, the call subscribes with BrainService.tapSession(userId, sessionId, listener)
to the existing fixed-session BrainEvent stream, the same seam used by bound chat clients.
Ownership is checked by core; the tap follows respawns and its unsubscribe joins the call's disposal list.
Only new visible assistant text deltas are buffered, bounded to 240 Unicode characters with
clipCodePoints. A following tool_authoring, tool or step consumes that buffer and proves it
is intermediate work; tool details, output, reasoning and empty text never become progress.
Terminal idle, errors, questions, new user turns and leaving hold discard buffered text, leaving
latestReply/resultItem as the unchanged authority for the final written result.
The engine sends at most three updates per call, at least 20 seconds apart and no earlier than
20 seconds after the brief is admitted, so the agent's opening acknowledgement never repeats what
the call already said (PROGRESS_THROTTLE_MS, PROGRESS_MAX_CHARS in voice/live/session.ts).
Skipped updates are consumed rather than queued: no timer, history scan, setting or extra inference runs.
With no visible text there is no update. progressItem frames text as untrusted data through
wrapUntrusted and asks Live to speak one short sentence in the caller's language only when the caller
is not speaking, without claiming completion or delegating more work. Progress preserves hold state and
uses the accepted delegation ID; it stops while a question is held, after result delivery, during closing
and after hangup. GPT-Live remains billed by call duration; progress has no separate token accounting.
A non-approval question pauses holding, is read on the call, and resumes holding after confirmed answers.
The limit is voice.holdSeconds, bounded by the wrap-up point of maxCallSeconds; zero says goodbye
immediately. Expiry, an early callback request or hangup during unfinished work marks callback_when_done
only when the account permits callbacks. The result stays in written chat otherwise.
A goodbye closes after new audio, a one-second output gap and consumed playback, or a ten-second deadline.
Caller speech can cancel ordinary completion, but never the spend or maximum-duration limit.
GPT-Live has no way to end a session itself and often answers a goodbye without delegating it, so the
caller's goodbye reaches the backend only when Live delegates it (prompts/voice/call.md asks for that).
As a backstop, IDLE_END_MS (twenty seconds) without caller transcript, audible output or a returned
result begins a silence goodbye (prompts/voice/silence.md) that closes like any other goodbye.
Pending work never ends a call this way, because holding has its own window, and caller speech cancels
the idle goodbye.
Idle has one definition, BrainStatusService.idle (BrainService.conversationIdle): no running turn (the durable activity
claim), nothing queued behind one, no parked question, no running child or workflow, no child result waiting to be delivered
into the conversation. Background shell jobs do not count. BrainService.subscribeConversationIdle fires from the same
edges that publish the owner-scoped conversation event (activityChanged: a turn settling, the last child ending, boot
recovery), only when the conversation is idle afterwards, possibly more than once per idle moment, so listeners are idempotent. Nothing
polls.
Done callback. RingCoordinator listens for idle moments and rings marked conversations through the existing SSE
and push path with a question-less kind: 'done' ring (VoiceRing.kind is question or done; push carries ringKind).
The durable request is authoritative even if a typed follow-up cleared voice_briefed, which gates only question rings.
An owned user conversation, the voice grant, voice enabled, and the callback preference are required; permanent
ineligibility consumes the request and leaves the result as text. Busy work, an active call, transient spend refusal,
and an offline user without push retain it. The request also survives an open ring and daemon restart; only acceptance,
decline or timeout consumes it through takeCallbackWhenDone. Duplicate idle signals cannot open another done ring
for the same conversation. sweepDone runs after boot recovery has marked parked, running and undelivered work busy.
The same sweep, scoped to an account, runs when its call slot frees or WebPresence.subscribeOnline reports its first
visible authenticated stream. That subscription is removed on rebind and shutdown; without presence, push remains
the reachability path. Accepting starts a call with the done-callback task template. Once its browser attaches, the engine sends
voice/result through the same latestReply/resultItem path as a hold result as spoken commentary before user speech. The result is independent of contextMessages and history trimming; the call permits
a new brief. Result delivery in a call clears the request, so an earlier mark does not ring for the same result.
Prompts and call summaries
Templates live in prompts/voice/. call uses the Live prompting guide's Backchannel, Interruption
and Delegation policy sections; brief, callback and done-callback select the opening purpose.
question and result frame returned data; wrap-up, spend-limit, silence, goodbye-chat, goodbye-callback and third-party-goodbye
steer the existing session. The obsolete hold tool-result prompt is removed.
Clock and authenticated identity use the chat's shared runtimeClock, runtimeIdentity,
createTimezoneResolver and BrainService.ownerChatIdentity. Language names use Intl.DisplayNames;
Czech and Slovak pronunciation guides remain. Metadata is substituted last and history is native input.
CallEndSummary records reason, duration, submission, callback, operation kinds and failures,
provider close/error codes, observed usageSeconds and finalUsageConfirmed, without user content.
Once audio flowed it also records audio: milliseconds of PCM received from the provider against the
wall time between the first and last frame, and the browser's last played position against the wall
time from the first frame to that report. Received audio spanning more wall time than its length means
the provider or daemon delivered slower than real time; played audio falling further behind than the
playback reserve means the browser waited for audio on the way to the caller.
Duration accounting and verified protocol
gpt-live-1 costs USD 0.05 per minute, billed per second, including silence and holding.
Elowen's own backend and delegation interpretation bill independently through existing brain/inference seams.
session.usage.updated and final session.closed supply cumulative usage.seconds.
VoiceUsage.duration is a cumulative snapshot; VoicePricing.costMicrousd rounds the cumulative price,
never fabricates token charges or locally estimated billable seconds.
voice_duration_usage checkpoints request/account/provider/model, seconds and cumulative micro-USD.
One SQLite write transaction inserts the durable receipt, advances the checkpoint and writes only the
price delta to existing provider and origin rollups. Duplicate or older snapshots do nothing, including
after sink recreation; any writer failure rolls back all writes. The first snapshot counts one logical
provider request and origin turn per call; later snapshots charge only the cumulative difference with
zero request/turn increments. UsageOriginStore.addTurn accepts an explicit cost-only continuation,
while every existing caller keeps its default settled-turn behavior. Checkpoints and receipts survive
usage reset, so replay cannot restore cleared spend or counts; only new price differences accrue afterward.
usage_by_origin remains the only origin ledger.
The engine checks account headroom each second, subtracts not-yet-reported elapsed time only to bound
admission, and begins a goodbye within five seconds of the remaining budget. Exhaustion closes immediately,
including during another goodbye; provider seconds alone settle
charges. It warns before maximum duration and enforces a hard maximum. Graceful teardown sends
session.close and drains final events and settlement for at most fifteen seconds before releasing.
Transport closure without session.closed settles observed usage and logs finalUsageConfirmed: false.
Settlement failure or an unconfirmed final usage tail from a configured paid call latches the existing
account usage health guard before releasing the call and blocks further paid work until restart.
Only provider-reported seconds are charged; unpaid, unconfigured admission failures do not latch it.
On 2026-10-06, real OpenAI API-key sessions with CoreSpeechService-generated PCM confirmed spoken
requests producing session.delegation.created { offset_ms, delegation: { id, type: 'delegation', target: 'client' }, event_id }
without task text. Thinking and commentary appends with that ID received timed *.appended acknowledgments.
Both transcript directions carry delta/start_ms/end_ms/event_id. Output can precede the commentary
acknowledgment, so result handling does not wait for it. Audio carried only type/delta, no timestamp,
item ID or completion event, and 100-ms PCM frames continued with exact zeros during silence.
A spoken interruption stopped a long answer before the caller finished speaking, without a cancel,
flush or interruption event; continuous silent output followed. An unanswered delegation produced a
short initial acknowledgment, then silence; after thinking progress it briefly said the task remained
pending and went quiet again. Thinking is accepted context, not a guarantee of silence or completion.
Observed cumulative usage was 14, 29, 44, 60 seconds at roughly 15-second intervals, with
context_window: { usage_ratio } and event_id; final closures included the resolved session snapshot,
client_event_id, reason: 'close_requested' and usage.seconds of 68 and 24 respectively.
The parser accepts these extra fields and does not assume an update cadence. Fake-server fixtures now
include resolved session/acknowledgment metadata, rich usage, output-before-ack ordering and continuous
silent PCM. Azure and every selectable voice were not individually tested against a live deployment.
Primary protocol references: Live sessions,
client delegation,
WebSockets,
prompting.
Dictation and voice-note services
Live dictation is a second engine on the same call machinery (src/voice/realtime/dictation.ts). POST /voice/dictations
(session-only like /voice/calls, no body) checks the voice grant, voice enabled, the provider predicate, the spend gate
and the priced voice.dictationModel (a liveTranscription entry of priceSupplement.ts, the only kind the Settings
picker offers for it), opens the upstream session through the one VoiceConnect ({ providerId, intent: 'transcription' }
selects wss://api.openai.com/v1/realtime?intent=transcription; Realtime speech names its model; calls select { api: 'live' }), sends
session.update with the transcription model, the account's language hint (languages) and turn_detection: null (the live model has no
server VAD), waits for session.updated and answers { dictationId, ticket }. The ticket carries
{ dictationId } instead of { callId } in the core:voice namespace; /ws/voice/<id> stays the only voice socket mount
and picks the engine from that payload. The browser never reaches the provider.
Calls and dictation share what is not specific to either: SessionRegistry (one active session per account, reserved
synchronously; each engine owns one registry, so an account may hold one call and one dictation at once, which cannot
collide because they are separate upstream sessions and separate browser sockets). A second dictation reservation for
that account throws dictation_active, mapped to HTTP 409 and an already-active notice in the browser. VoiceSession (upstream, browser,
timers, disposers, pending usage), bindBrowser (frame size bounds, fifteen seconds of browser silence ends the session,
a malformed or unknown frame closes it with 1008; the engine supplies its control schema and handlers), configureUpstream
(the ten-second protocol acknowledgment wait), finishSession (Live allows fifteen seconds for final duration settlement, dictation two; the one teardown: settle usage, clear timers, close both sockets,
free the slot, summary line). Its optional drainSignal shortens an already running drain without closing
the provider early: Live shutdown aborts settlement after VOICE_CALL_STOP_BUDGET.shutdownDrainMs,
three seconds, including pending writes. Missing final usage or timed-out writes latch account health.
The daemon checkpoint/beforePause guard remains five seconds and its later stop guard remains twenty-four,
so bounded voice settlement precedes plugin teardown and total guarded stop stays below thirty seconds.
Ordinary hangup keeps its fifteen-second budget. VOICE_CALL_STOP_BUDGET owns that sessionDrainMs,
shutdownDrainMs and browserWaitMs of sixteen seconds, including a one-second transport margin; the web mirror
is byte-pinned. Hangup immediately stops browser media but retains the control socket for authoritative
ended/submitted. A sticky submitted frame is emitted on admission and retained across transport loss.
accountUsage (record once per item id for the account and origin, then check headroom).
Dictation closes the upstream socket as soon as it ends, so a normal end does not wait for the provider.
Browser to daemon, besides binary PCM16 24 kHz frames: {type:'commit'} finishes the current segment (ignored below
100 ms of buffered audio, the provider's minimum), {type:'stop'} finishes the open segment only when the provider has
already produced text for it, waits at most three seconds for the final text of every committed segment and ends with
hangup, {type:'cancel'} ends at once and discards; audio and commits after stop are dropped. Daemon to browser:
{type:'delta', itemId, text} (interim), {type:'final', itemId, text} (replaces that item's interim text; a segment whose
conversation.item.input_audio_transcription.failed arrived gets an empty final, so no new frame exists) and
{type:'ended', reason}; provider events are never forwarded. Each completed event's duration usage is priced by the voice
pricing resolver and recorded once per item through the usage sink; an exhausted account ends the dictation with
spend_limit. The maximum length is the instance maxCallSeconds (a graceful stop ending with max_duration), a
daemon shutdown ends it with daemon_restart, and every exit closes both sockets and frees the slot.
Language hints: GPT-Transcribe takes languages (an array of language subtags), not language. transcriptionLanguages
(src/voice/transcriptionLanguages.ts) is the one place that reduces the account's tag (cs-CZ to ['cs']); dictation and the voice-note upload (languages[] form field) all use it, and an empty preference
omits the field so the model detects the language.
Dictation and Realtime speech retain their existing typed recoverable/fatal provider-error handling
in src/voice/realtime/providerError.ts. Live command errors instead end the call explicitly and log
only the provider code. Dictation transcription failure drops its segment; it is never a completed call turn.
The browser side (web/modules/voice/dictationSession.ts, useDictation) captures with the call's AudioWorklet (capture
only, no playback), decides segments from frame energy (700 ms of silence after speech sends commit, audio-time based)
and stops by itself after the account's dictationPauseSeconds of continuous silence (3, 5, 10, 20 or never). The old
upload path POST /voice/transcriptions is gone; CoreSpeechService.transcribe stays for the chat platforms.
web/modules/voice/dictationDraft.ts owns the pure DictationDraft text model, imported directly by
useDictation and its text characterization tests. It tracks the inserted range and spacing from
the previous draft and the native beforeinput selection, shifts ownership for edits before it and
relinquishes intersecting edits. Late updates for detached item IDs are ignored; withdrawal removes
only the current owned text. It performs no I/O or React work and processes only the current draft
and session items. The hook retains microphone/session lifetime, state, timers and target writes;
ChatComposer remains the owner of the editable draft and waits for final text before sending.
An idle hook makes no draft writes.
The voice config block is sanitized and merged per field; dictationModel is validated against the catalog like the
other model fields. Account voice preferences (voice, language, callback, dictationPauseSeconds) are a revisioned
user_settings resource. brain_sessions.voice_briefed has explicit get/set/clear accessors; callback_when_done has set, get, atomic take and list accessors.
The shared wire contract carries bounded call records, WebSocket controls, ring events and REST DTOs;
src/shared/voice.ts owns runtime bounds. TurnRequest.provenance is core-only and HTTP rejects it.
StoredUserMessageMetadata stays an importless wire-contract interface; the shared voice metadata validator is pinned to it with satisfies. It carries voiceCall
and trusted provenance through admission and the steering mirror; only voiceCall is saved in the PI journal.
Voice admissions store the brief as display text and preserve the full
model-facing brief plus transcript in the native PI user entry's providerText; turn framing persists separately as anchored elowen.turn-context records. Only a trusted accepted voice submission sets the
voice_briefed mark. The additive brain_paused_queue.metadata column keeps provenance, transcript and
brief display through pause checkpoint, boot replay and queue promotion. Invalid stored metadata fails
closed inside the queue transaction instead of dropping the message. Transcript rendering validates
voice metadata per row with safeParse and omits invalid call metadata while preserving the row's text
and the rest of the conversation. HTTP brain sends accept only web and cli surfaces; the daemon's own
voice engine submits the voice surface directly through BrainService.
Account spend admission uses PayerResolver and AccountSpendGate in
src/brain/session/spendAdmission.ts; accountSpend.ts implements the monthly gate and runtime guard. The active request pin wins, then the durable session owner,
then an explicit prospective owner for a new room. admitPayer accepts that prospective owner separately
from the shared spend wiring. Admission and billSettledTurn share that resolver.
The nullable account micro-USD limit applies to administrators too. createAccountSpendGate receives a
Pick<UsageProviderStore, 'spentBetween'> reader, not a database connection. spentBetween(userId, fromDay, toDayExclusive) sums reported USD costs across providers for that account over an inclusive-start,
exclusive-end UTC day range, reading only usage_by_provider_day through its (user_id, day) index.
Empty and entirely unpriced ranges return zero; unpriced-model admission remains a separate check.
The gate supplies the current UTC month boundaries and rounds the total once to micro-USD. Database
errors propagate. The daemon passes its existing usageProviders store; subscription tokens use list prices.
Unknown requests remain visible separately. Origins and providers are never summed together.
Every run is admitted before preparation or cold compaction, including delegated runs without origins.
An admitted run may finish above its limit; each actual request model must still be priced for a limited
account. The factory runtime guard checks compaction's actual model too, using the provider pricing seam.
Account-owned sessionless inference and direct embeddings carry an explicit payer and check admission per request. WorkAccount and validateWorkAccount distinguish an explicit null instance payer from a missing account. Instance plugin inference and embeddings run without account admission or per-account usage rows; account paths retain their payer. Palette Ask AI passes the authenticated account into this same inference seam. Worker embedding batches check admission before the daemon dispatches them.
InferenceClient.decide accepts an optional sessionId and forwards it to the provider runtime. Conversation titling supplies the conversation id so session-bound transports such as OpenCode receive their required session header; standalone background consumers omit it without creating a conversation. This metadata does not change billing or persist a transcript.
Sessionless inference writes provider usage exactly once, including paid failures; origin accounting remains
in withOriginUsage. ImageService resolves the current tool session through the same payer resolver, admits before dispatch, and uses recordInferenceProviderUsage once per successful generation or edit and for provider-reported charges on failures. Image prices come only from a provider-reported charge; missing prices remain unknown. Unknown embedding prices remain unknown and do not bypass exhausted limits.
The shared createAccountUsageHealth counter wrapper logs usage.provider_counter_failed and throws
AccountUsageUnavailableError, preserving the saved transcript rather than claiming a transcript failure.
It blocks later paid work for that account in the current process, including background requests and voice.
HTTP chat, compaction, goal, headless-run and search routes preserve its specific wording with status 503.
A later successful write does not clear the failure: it cannot restore the missing charge. An administrator
must reconcile usage before restarting the daemon, which clears this process-local health state.
Embeddings use the same recordInferenceProviderUsage sink as secondary inference and images.
InferenceAccount names secondary inference admission and accounting wiring. Absent wiring runs instance work
without account counters. ConversationTitler requires the shared payer resolver, with the facade supplying
the owner-only resolver when spend wiring is absent.
Spend refusals carry monthly_limit or unpriced_model and the next UTC month start.
AccountSpendLimitError owns the reset sentence reused by admission, provider streams and goal pauses. Before acceptance
chat, compaction, goal and headless-run HTTP routes return 402 through sessionRouteError;
after acceptance a durable chat error carries the same typed notice. Goal turns
pause with the reason. The core alert key spend-limit:<YYYY-MM> coalesces the monthly account notice.
GET /users/:id/spend is admin-only and returns UserSpendSummary through the shared projection.
Owner turn admission checks spend before busy refusal and cold setup, then rechecks after awaited setup. Admin permission patches use setSpendAndVoicePermissions for both the nullable
limit and the explicit voice grant. POST /users/:id/usage/reset is administrator-only and atomically resets the target account's provider and origin counters and monthly spend, retaining conversations. Its shared UserUsageResetResult reports ok, providersCleared and originsCleared; advancing the conversation usage epoch does not report a cleared-message count. BrainStore.resetUsage returns only originsCleared, and BrainUsageStore.advanceUsageEpoch returns no value.
Voice provider validation uses the shared voiceProviderPredicate in src/voice/provider.ts:
keyed official OpenAI and Azure https://<resource>.openai.azure.com/openai/v1 entries qualify;
proxies, OAuth and other Azure endpoints do not. resolveVoiceEndpoint owns all credential and URL resolution.
Live uses wss://api.openai.com/v1/live/sessions with Bearer auth or
wss://<resource>.openai.azure.com/openai/v1/live/sessions with api-key, without query parameters.
Realtime speech uses /realtime?model=<model>, dictation /realtime?intent=transcription.
Azure model IDs remain exact Global Standard deployment names; no mapping or fallback is applied.
Batch transcription retains the deployment-specific audio/transcriptions?api-version=2024-10-21
route on Azure and the official /v1/audio/transcriptions endpoint on OpenAI.
src/voice/priceSupplement.ts is the one curated price/model catalog: Live for calls, Realtime for speech,
transcription for voice notes and liveTranscription for dictation. Live's duration row declares
unit: 'minute', billedBy: 'second' and source date 2026-09-10; unsupported dimensions are zero
and rejected by billing. Selection rechecks provider eligibility; admitted usage pricing survives provider removal.
Realtime speech/dictation keep src/voice/realtime/usage.ts for their unchanged modality/token/duration parsing.
One VoiceConnect implementation and APP_IDENTITY_HEADERS serve both protocols; no credentials cross it.
ConfigStore.migrateLiveCalls runs in authoritative daemon assembly inside the write lock.
It changes only persisted voice.callModel equal to gpt-realtime-2.1-mini, gpt-realtime-2.1
or gpt-realtime-2 to gpt-live-1 and increments settings revision once.
Other fields, including speech and dictation models, are preserved; repeating migration is a no-op.
Fresh defaults and validation offer only gpt-live-1 for calls, even when voice is disabled.
src/store/voiceUsageStore.ts shares receipt, checkpoint and provider/origin writers in one transaction.
Receipts and duration checkpoints survive usage resets so duplicates cannot resurrect cleared spend.
GET /voice/catalog offers the priced model groups and eligible providers without secrets.
Its active selection uses activeVoiceVoices: the account voice is a call voice, so it offers the Live
call model's voices (the twelve GPT-Live-1 voices; Realtime models reject them). Voice notes offer no
choice and always speak VOICE_NOTE_VOICE (marin) on the Realtime speech model, so a speech model
without that voice leaves voice unavailable. Migration v54 moved stored older voices to the default quartz.
GET/PATCH /auth/me/voice is revisioned and account-bound; admission revalidates stored voices and
BCP 47 language choices rather than silently replacing them.
The active catalog is null while voice is disabled; unavailable selections return no voices. Personal settings require expectedRevision for PATCH and return the current snapshot on conflict. Empty language inherits the account locale; supplied tags use Intl.Locale for speech. BrainCore constructs pricing and settlement once and passes VoiceServices slots to HTTP routes. Delegation answers reuse brainAnswerSchema and briefs use VOICE_CALL_RECORD_BOUNDS.maxEntryChars. Text owner conversations and voice assembly use the same memoryRecallScope selector, passing the selected identity validated by the execution boundary before host-cwd inference. Voice never chooses locally between project and cwd authority. Guest working directories are never interpreted as host project roots when a project is selected. Each live pass rebuilds private/global and currently authorized shared-pool categories, so project switches and share-list revocation apply immediately. Ordinary shared platform rooms retain global-only recall.
Voice callbacks
BrainService.subscribeElicitations and pendingQuestion delegate to the existing private
ElicitationRegistry. Asked events follow parking; settlement events follow removal and precede promise
settlement on answer, timeout, session cancellation, global cancellation and replacement. Listener
failures are isolated and logged. Queued approvals that never park do not emit lifecycle events.
src/voice/ring.ts owns generation-local rings. It checks a pending non-approval question in an
owned, non-delegated user conversation, its voice-briefed mark, the human owner's explicit voice grant,
instance enablement, callback preference, spend admission and absence of an active call. The account
must have visible web presence or a push subscription. Bootstrap creates the shared coordinator slot before building BrainCore, so the call engine and REST
routes receive the same object. It initializes observation and delivery after the core services and
push transport exist. API route assembly binds the existing RouteContext.webPresence registry to
the coordinator; no second presence tracker exists.
The configured ring time sets expiry. Read, decline and claim revalidate ownership, question and expiry. Claim reserves the ring synchronously for the exact account and conversation and returns a VoiceRingClaim with bound question ids and reservation-specific commit/release methods. The call engine commits after upstream session.started confirms configuration; only then is the ring removed and accepted published. Failed starts release only their own reservation, allowing another accept before expiry. Opaque reservation identity prevents a stale claim from releasing or committing a newer one. Commit revalidates eligibility, the pending question and expiry. The engine's already reserved account slot is permitted. Decline, timeout, settlement and restart invalidate outstanding claims, and a later release cannot revive the ring. Decline, timeout and restart leave the text question pending. Shutdown unsubscribes and closes remaining rings before pausing conversations.
GET /voice/rings/:id and POST /voice/rings/:id/decline require a full human login and never
widen access for admins. Foreign and stale ids both return 404. Private voice_ring and
voice_ring_end SSE frames reach only that owner's full human-login streams, with routing
userId stripped. Setup, advisor, chatbot and impersonation streams receive no voice frames.
EventBus.subscribeVoice is a separate account lane: generic event subscribers, including plugin row
resolvers and the activity recorder, never receive ring payloads. Push builders in src/push/messages.ts
share a ring tag, expiry and encoded conversation URL;
localized Accept opens the page and requires a further explicit in-page microphone/playback tap.
Owner-turn completion uses the same push payload owner: bootstrap resolves UserSettingStore.locale
for the actual recipient at delivery and passes it to buildTurnDone(input, locale). Only an empty
conversation title or an answer preview emptied by the existing Markdown flattener uses the cs/sk/en
defaults. The fallback title uses the current product name. Supplied titles and answer previews retain
the existing 60/180-code-point limits; notification actions and the /chat tap target are unchanged.
The service worker displays the completed payload without translating it. No push subscription means
no delivery, and locale lookup uses the store's existing English default for absent or unknown values.
At owner-turn admission, an explicit external submission clears voice_briefed. Core voice
provenance, interrupt replay, unattended automation and internal turns leave it alone. Surface names
are never accepted as proof of voice. Queue and recovery paths must preserve trusted provenance.