Voice and account surfaces
Owning code
- Settings voice section:
web/modules/settings/VoiceSettingsSection.tsx. - Account voice section and the Users permission toggle:
web/modules/account/VoiceAccountSection.tsxandweb/modules/users/VoicePermission.tsx. - Call, ring and dictation state:
web/modules/voice/(callController.tsx,useVoiceCall.tsx,useDictation.ts,dictationSession.ts,audioWorklets.tsx,socket.ts,IncomingCallPrompt.tsx,CallPanel.tsx). - Composer entry points:
web/modules/advisor/ComposerActions.tsx,web/modules/advisor/ChatComposer.tsxandweb/modules/dashboard/HomeComposer.tsx. - Service worker notification handling:
web/public/sw.js. - Daemon routes:
src/api/routes/voiceCalls.ts,src/api/routes/voiceRings.tsandsrc/api/routes/selfSettings.ts. Shared budgets and bounds:src/shared/voice.ts, which the browser mirrors inweb/lib/voice.ts.
Voice and account spending
Settings/Voice configures the provider, models, bounded call/ring durations and the hold limit (voice.holdSeconds, 0 to 300 s in steps of 5, 0 reads "No holding") through the shared config
mutation. Account/Voice edits personal voice, language, callbacks and the dictation pause through useSaveMyVoicePreferences;
it reads the cached revision at each save, adopts acknowledged revisions and offers AutoSaveStatus's
reload action on conflict. Language and callbacks remain editable when no voices are available.
Config voice/provider saves invalidate the voice catalog (web/lib/mutations.ts). The personal picker uses catalog.active.voices,
computed by the daemon, and ChoiceField pins the saved
value. All cs/sk/en labels and hints live in the voice namespace. Both sections register stable row ids
in rowAnchors, shared with siteSearch, so palette links reveal the actual control.
SliderWithReading composes the scalar shadcn Slider with a fixed-width visible reading and
aria-valuetext. Pass the usual Slider props and reading; readingWidth and action support a compact
reading with a trailing reset button. Voice durations, Terminal and Account appearance share it;
without an action only the track and reading render. It adds no requests or timers.
SpendLimit shares MICROUSD_PER_USD from lib/format with token-limit formatting, pins its wire bounds
to the daemon schema and displays usage with four decimals. The editor offers whole cents and derives
its bounds and translated validation copy from the same constants.
Shell mounts VoiceCallProvider once inside BrainChatProvider. The composer's single Options menu
(ComposerActions, trigger chat-actions, menu chat-actions-menu) sits left of send/stop and replaces the
former phone button and work-mode pill. Its trigger shows one general sliders glyph, tinted outside Build, with
the mode in its title and aria-label; the visible "Options" word joins it from md: up. The menu has a
"Work mode" radio group (the daemon's kind:'mode' commands, run through runSlash) and, only when the account
has the voice grant and the instance enables voice, a "Voice" group with "Call" and "Dictate". While a call
is active the textarea and attachment/menu controls yield to a single-line CallStatus inside the composer:
phone, tabular elapsed time and translated state. The primary action is a shared danger Button with PhoneOff,
disabled during ending. The Options menu passes the composer's callActionRef to its close-focus handler,
so selecting Call moves keyboard focus to hang-up after the menu releases focus. If a turn runs, the separate
existing accent square STOP remains immediately before hang-up and only aborts that turn. The draft remains
owned by BrainChatProvider. After ending, the textarea is editable immediately; the reason occupies a wrapping
status row inside the composer. Normal reasons disappear after five seconds once transcript restoration
settles; close, expiry and starting or accepting another call wait for pending restoration so delayed failures
cannot erase the copyable transcript.
Errors remain as dismissible alerts until acknowledged or a new call starts. Only failed transcript recovery
opens CallPanel, a small shared Modal containing the copyable transcript. Its existing closeDisabled contract blocks header, Escape and backdrop dismissal; the explicit footer Close action acknowledges the only remaining copy. A ready incoming ring explains unavailable acceptance in the body while keeping Decline available.
The persistent controller exposes the snapshot, recovering, hangup and dismiss through useVoiceCallController,
so compact and full composers share one call with no extra transport. Dashboard HomeComposer uses its startNew()
action from a shared ghost/icon Phone button immediately before Send, with the same instance-enabled/account-granted
availability and existing cs/sk/en call label. VoiceCallSession.start accepts either an existing session id or an
async target factory receiving the preparation's AbortSignal: capture/playback still starts in the click stack,
then the factory resolves before the paid
call request. startNew(target?) creates the conversation through BrainChatProvider.startNewConversation({ target })
in the destination chosen in the dashboard, never raises the project question or opens the chat dock, and navigates
to /chat while the persistent call connects. Ordinary
chat Call and incoming rings still pass their existing ids. HomeComposer shows the shared CallStatus and hang-up
while pending, disables sending during a call/recovery, and retains microphone/creation errors inline. Cancelling
while creation is pending prevents a paid call and late navigation. Without voice availability the phone is absent;
this adds no transport, polling, or provider request beyond the selected call. Call and dictation capture share one microphone request that asks the browser for echo cancellation, noise suppression, automatic gain control and voice isolation; GPT-Live has no noise setting of its own, so the browser is the only filter before the model and prompts/voice/call.md tells it to treat background sound as not addressed to it. Dictation is live:
ChatComposer owns useDictation, which runs a DictationSession (microphone first through createVoiceCapture, the
call's capture worklet without playback, then POST /voice/dictations and a same-origin /ws/voice/<dictationId>?ticket=
socket) and writes the daemon's interim text into the textarea at the caret where dictation started while the person
speaks; a final text replaces the item's interim text, and items are separated by single spaces. DictationDraft in
web/modules/voice/dictationDraft.ts is the pure text model imported by useDictation; it owns an exact range, including
inserted spacing, against its previous draft snapshot. The composer's native beforeinput
listener calls useDictation.beforeEdit(start, end) with the selection before the change, and every textarea onChange
calls useDictation.edit(value, caret). This anchors edits even beside identical text: edits before the range shift it,
edits after it leave it intact, and intersecting edits detach its existing item IDs. Delta, final and withdrawal also
reconcile the current snapshot; ownership never comes from searching for matching words elsewhere. Detached IDs' late
updates are ignored. New item IDs insert at the current caret; cancelling removes only the new owned span and preserves manual edits. The pill in the trigger's
slot (dot and clock, the clock hidden below md:, tick, cancel) stays compact so a 320 px composer remains readable; the
send button stays visible and enabled beside it. Send (button or Enter) while dictating stops the dictation and waits
for the daemon's ended before sending the composed draft. The shared voice contract's VOICE_DICTATION_STOP_BUDGET grants three seconds for committed
finals and two seconds for upstream closure and usage settlement; the browser's six-second fallback includes a one-second
transport margin and keeps receiving text throughout that wait. The tick stops and keeps the text; the
cross or Escape cancels and removes only the text this dictation inserted. Holding the send button for 400 ms (HOLD_TO_DICTATE_MS) starts the same dictation for as long as it is held and releasing it is the tick: the text stays and nothing is sent, and the click that ends a hold is swallowed. It applies only when voice is available and no call or dictation runs; a short press still sends. While voice is available an empty-draft send button is aria-disabled and dimmed rather than disabled (a disabled button gets no pointer events), its click is ignored, and it carries touch-none select-none with the context menu suppressed so a touch long-press neither selects nor opens a menu. Keyboard is unchanged (Enter sends). 700 ms of silence after speech finishes a
segment; the account's dictationPauseSeconds (Account → Voice: 3, 5, 10 or 20 seconds, or "Never"; default 5; null is "Never",
src/shared/voice.ts lines 18-20, VoiceAccountSection.tsx lines 86-89) of continuous silence stops it. While connecting, the existing three-second pre-roll queue retains audio and commit boundaries in capture order;
trimming drops a boundary only after all preceding audio for that segment has gone. Opening flushes the queue before any
new controls. The microphone is released on every exit, and failures show translated toasts without touching the draft.
An account already dictating in another tab receives HTTP 409 dictation_active; the toast asks the user to end that
existing dictation instead of reporting a lost connection. Calls and dictation both use createVoiceSocket in
web/modules/voice/socket.ts: it encodes the session ID, selects WS/WSS from the page origin and attaches the one-shot
ticket to /ws/voice/<id>. The hooks provide it directly to their session dependency objects.
useVoiceCall owns the call state machine: POST /voice/calls through elowenClient, then a same-origin
/ws/voice/src/shared/voice.ts lines 21-27; useVoiceCall.tsx line 173). The daemon's own live shutdown uses
shutdownDrainMs (3 s), which must settle before the five-second pause barrier (src/voice/live/session.ts line 536). Sticky submitted
admission frames prevent accepted work from returning to the draft on a later transport loss. Calls without an admitted submission append their transcript to the existing composer draft in the
bound conversation; admitted work remains in the daemon conversation. The daemon's state frames are
connecting, listening, on_hold and ending; actual worklet playback adds the local speaking indicator
while listening. Live also streams exact-zero PCM during silence, so queue occupancy does not imply speech.
The shared pcm16HasSignal helper (src/shared/voice.ts, mirrored byte for byte in web/lib/voice.ts because
the web root cannot import outside itself) marks each queued frame; the worklet reports whether the frame being
played has signal and clears the indicator on silence or an empty queue. The internal played callback
carries speaking, while wire played.ms still counts all samples continuously. No samples are dropped,
no caller VAD is added, and dictation has no playback callback. The elapsed
clock stays outside the live region so it does not interrupt announcements every second.
on_hold shows "Elowen is working" in the composer's status line, with hang-up unchanged
and no tone (the daemon holds the line while the work it started runs, see "Holding the line, and calling back" in docs/ARCHITECTURE.md).
During that wait, Elowen can speak up to three short progress updates, at least twenty seconds apart
and the first one twenty seconds after the work starts (src/voice/live/session.ts lines 72, 186-188, 327). Each
comes from assistant text that a following tool or step event closes, not from every written message. She is instructed to wait while the caller speaks. Updates keep
on_hold active; the final answer returns the call to listening as before.
The persistent controller validates /chat?voiceRing=kind: question or done) and renders malformed-ring errors; a done ring takes the title "Elowen has the result". IncomingCallPrompt
uses the shared Modal and overlayStack with no feature-local portal. Acceptance always requires an
in-page tap, which activates microphone/playback before asynchronous navigation and starts the call
with ringId. Decline uses elowenClient.declineVoiceRing. Missing or expired declines become stale; transport failures
retain an explicit retry. Settlement, expiry and account unavailability disable acceptance. Call refusals
show their specific translated code. Visible call state follows daemon echoes; useNow supplies elapsed
time. A call requests microphone and playback in the click stack. When the Permissions API already
reports the microphone as granted, it requests the daemon reservation at the same time, so the provider
session opens while the browser prepares audio; otherwise it waits for capture first, so a denied
microphone never makes a paid request. The call socket opens once both succeed. When either fails or the
attempt is cancelled, the other is undone: prepared audio stops, and a reservation the attempt no longer
uses is consumed and hung up instead of waiting out its attach deadline. Cleanup also owns late-reservation sockets and their deadlines.
The service worker handles ring/ring_end push variants, restricts ring navigation to the app origin,
and opens the app when an authenticated decline fails. A ring notification vibrates in a call pattern
where the device supports it; browsers give notifications no custom sound, so its sound stays the
system default. While IncomingCallPrompt can be accepted, useRingtone (web/modules/voice/ringtone.ts)
loops a synthesized ringtone through its own AudioContext and closes it on accept, decline, expiry or
unmount. A page without earlier user interaction keeps that context suspended and stays silent.
The fake daemon's voice handlers script REST, private ring/core SSE and ticketed WebSocket audio.
POST /__test/voice controls ring, ring_end, audio, emit, drop, admissionError, catalogMode, reset, callScript and
dictationScript (web/tests/e2e/fake-daemon/handlers/voice.ts lines 280-300); GET returns
call frame/control observations without tickets. Both the automatic POST /__test/reset and the voice reset control
share one voice-state reset: close live transports and clear calls, dictations, rings, admission errors and scripts.
Connecting-dictation assertions pin the admitted session ID so an earlier commit cannot satisfy the new session's wait.
The audio script emits audio_started before PCM.
The voice Playwright fixture forwards only /ws/voice/ to the local fake daemon because the test Next
server has no nginx upstream. Product sockets remain same-origin. Playwright launches with fake
microphone permissions and devices. voice.settings.e2e.ts covers instance settings, account choices,
and the user voice grant in cs/sk/en at 320 px with a mobile context and 1440 px, including loading,
empty and recoverable catalog errors and the per-second call-price help. voice.call.e2e.ts covers
full-duplex playback, hold, goodbye, failures and transcript recovery in cs/sk/en at both widths;
caller transcript deltas do not flush playback. In-process route tests share createVoiceSpendApp; the fake user
PATCH uses the real userPermissionsSchema. Browser verification belongs to integration.
Voice preferences and call transcripts
Settings → Voice mode uses the existing config query and config mutation. The provider and four
model pickers (call, transcription, dictation and speech) consume GET /voice/catalog; only GPT-Live 1 is offered for calls, with shared help explaining
USD 0.05 per minute billed per second, paid waiting and separately billed work. More history does not
change that call rate. No provider eligibility or model/voice list is inferred
in the browser. Call and ring sliders mirror the shared protocol bounds using literal contract types.
An empty provider catalog prevents enabling voice but leaves disabling and duration edits reachable.
Read failures replace controls with the shared retry state; invalid enabled configuration cannot save.
Account → Voice uses the shared voice-preferences query in modules/voice/voicePreferences.ts.
Catalog and revisioned account responses are validated at the HTTP boundary. Autosave patches only
changed preferences through PATCH /auth/me/voice with the last acknowledged expectedRevision.
A conflict remains an explicit failed save with the draft retained. Reloading the saved preferences
is an explicit action that discards the draft; there is no automatic overwrite.
The catalog's active selection identifies the configured call and speech models for every eligible
account without granting access to instance config. The voice picker offers the intersection of those
models' voice lists, or the single configured model's list when only one is set. Inactive models never
restrict the selection. A null active selection, missing provider/model or empty intersection makes
voice selection unavailable while language and callback preferences remain editable. A previously
saved unsupported voice stays visible as the current value but is not a selectable option.
The Users pane independently autosaves only can_use_voice, including for administrator accounts.
Both user-bubble layouts render a turn with typed voiceCall metadata through the shared
VoiceCallMessage (web/modules/voice/VoiceCallTranscript.tsx) in place of the plain text: a phone
header with the call duration, the message text (the interpreter's one-sentence summary), and that
request's own transcript segment collapsed under it as plain-text speaker entries. History and live-user projections
retain the same record. VoiceTranscriptEntries renders the plain-text speaker list once for this disclosure and CallPanel, preserving line breaks and using the shared empty state when there are no entries. Rendering is linear in the supplied transcript and performs no reads. Brief headings are never parsed to recover a transcript.