Voice Devices
Physical smart-speaker devices bound to one user account: registry, pairing codes, device tokens, heartbeat, conversation modes (push-to-talk or live), and the chat/live-voice endpoints they may call.
A voice device is a physical smart speaker that is bound to exactly one Clapilot user account on exactly one Clapilot instance. After pairing, the device authenticates with its own bearer token and talks to the existing personal chat (push-to-talk / walkie-talkie) and live-voice relay endpoints on behalf of its owner. The first hardware target is the M5Stack Atom VoiceS3R (ESP32-S3).
The server side (device registry, pairing flow, device authentication) and the firmware for the Atom VoiceS3R ship together. The firmware lives in firmware/clapilot-voice/ (ESP-IDF, see its README for build, flash and setup) and supports both conversation modes, push-to-talk and live voice; a wake word is a later step. The API is designed so that any device that can perform HTTPS requests can pair and talk to Clapilot.
UI name: Sprachgeräte (DE), Voice devices (EN), Dispositivi vocali (IT). The settings surface exists in the web app (Settings -> Sprachgeräte, route /settings/voice-devices) and in the iOS/macOS apps under the profile settings section.
Supported hardware
The registry identifies hardware by the device_kind string the firmware sends in the pairing request. Known kinds are defined once in src/lib/voice-devices/kinds.ts (VOICE_DEVICE_KINDS, with per-kind capability metadata) and each has a translated label (voiceDevices.kind.<device_kind>) that the web and Apple settings lists show instead of the raw value. An unknown or missing device_kind (older firmware sends none) pairs normally and is stored as the default kind m5stack_atom_voices3r; it never rejects the pairing.
device_kind | Hardware | What differs |
|---|---|---|
m5stack_atom_voices3r (default) | M5Stack Atom VoiceS3R (ESP32-S3, PSRAM, ES8311 codec) | No display or keyboard; Wi-Fi and pairing code are entered through the captive-portal setup assistant. Full-duplex microphone and speaker. |
m5stack_cardputer | M5Stack Cardputer (ESP32-S3 StampS3, keyboard, 1.14" display, PDM microphone, small speaker, no PSRAM) | Wi-Fi credentials and the pairing code can be typed directly on the device keyboard and the display shows setup and turn state. Microphone and speaker are half-duplex (record or play, not both). Without PSRAM, push-to-talk recordings are shorter than on the other boards. |
m5stack_cardputer_adv | M5Stack Cardputer-Adv (StampS3A, ESP32-S3FN8, no PSRAM; TCA8418 keyboard controller, same display, ES8311 codec) | Same on-device keyboard/display setup as the Cardputer; the ES8311 codec gives full-duplex audio like the Atom. Verified on hardware. |
Meta Ray-Ban Display web client
The sibling ClapilotRayban client pairs using device_kind: meta_rayban_display. The web, iOS and macOS settings lists show Meta Ray-Ban Display. Pairing, revocation, owner identity, target routing and chat storage use the existing device contracts.
The glasses web client uses Meta's system text composer (focus a text field, then tap to dictate or write) and sends reviewed text turns to POST /api/chat. It resumes the device chat on reopen, sends visible-page heartbeats, renders streamed replies and fetches reply audio with bearer authentication for explicit playback where supported. It polls reply-audio for asynchronous specialist replies. Main/team history remains available in Clapilot; the client shows replies to its new turns. Stop is supported for the exposed device/specialist session, not main/team targets whose session IDs are not exposed by heartbeat.
This client does not capture raw microphone audio or implement continuous live relay sessions. Its settings explain that continuous live audio requires a native companion integration. The system composer is the voice-entry surface even if the owner has selected live mode. The browser stores a revocable device credential locally; forgetting the browser is not server-side unpairing. No new management or business-data tools are exposed.
Capability flags per kind (hasDisplay, hasKeyboard, hasPsram, fullDuplexAudio) live next to the kind list so server and UI code can branch on hardware features without string comparisons.
Related, but separate: Home Assistant assist satellites and media-player speakers are output-only announcement targets managed by the Home Assistant module (Settings -> Module -> Home Assistant, tool home_assistant_announce). They receive text that Home Assistant renders with its own TTS; they never pair with Clapilot, hold no device token, do not record speech, and do not appear in this registry. Everything below applies only to paired Clapilot voice devices.
Binding semantics
- One device belongs to one user (
voice_devices.user_id). A device row is the binding, not a visibility filter: voice devices are a personal-first module, like personal chat sessions and push-device registrations. Owner-scoped queries here are intentional and are not the workspace-global visibility bug described in the repository guidelines. - One device belongs to one instance. The pairing response includes the
instance_urlthe device must use for all later calls; a device that should move to another instance or another user is unpaired and paired again. - Each device gets its own personal chat session (
voice_devices.chat_session_id, titled after the device). By default the device talks in that session, so conversations from the device show up in the owner's normal chat list; the owner can read and continue them in the web or Apple chat UI. The owner can redirect a device's spoken turns elsewhere with its turn target. - Unpairing deletes the device row and revokes its token immediately. The device's chat session stays in the owner's chat list.
Pairing flow
- The owner opens
Settings -> Sprachgeräteand presses Gerät koppeln. The app callsPOST /api/voice-devices/pairing-codesand shows the returned code asXXXX-XXXXin a monospace block, with a countdown until expiry and a copy button. Generating a new code invalidates the user's previous unconsumed codes. - The owner enters the code in the device's setup assistant (captive-portal or companion-app flow, defined by the firmware step). The device sends the code to
POST /api/voice-devices/pairtogether with itshardware_id(for example its MAC address),device_kind, andfirmware_version. - The server validates the code, creates the
voice_devicesrow bound to the code's user, creates the device chat session, and returns the device token once plus everything the device needs to operate:instance_url,chat_session_id, the owner's display name, and the relative paths of the endpoints it may call. - While a code is active the settings panel polls
GET /api/voice-devicesevery three seconds; as soon as the new device appears in the list the panel shows a success line and stops polling. - The device stores the token in its own secure storage and starts sending heartbeats.
Pairing codes are eight characters from the alphabet ABCDEFGHJKLMNPQRSTUVWXYZ23456789 (no 0, O, 1, I), displayed as XXXX-XXXX, valid for ten minutes, and single-use. Input is normalised before lookup: uppercase, then strip separators and whitespace, so the dash and lowercase input are accepted; any character outside the alphabet makes the code invalid instead of being remapped.
Device token and allowlisted endpoints
The device token has the format clpd_<base64url 32 bytes>. It is returned once in the pairing response and stored as a SHA-256 hash (voice_devices.token_hash); only the first twelve characters (token_prefix, for example clpd_AbCdEf1) are kept in clear text so the owner can recognise a device in the settings list.
Devices send Authorization: Bearer clpd_.... Device bearer auth is accepted only on this allowlist:
| Path | Purpose |
|---|---|
GET /api/voice-devices/me | read own identity and chat session (no writes) |
POST /api/voice-devices/heartbeat | report firmware/status and refresh last_seen_* |
GET /api/voice-devices/reply-audio | look up the newest spoken reply for the device's turn target (fallback when the stream carried no audio frame) |
POST /api/chat | push-to-talk turn (audio attachment) |
POST /api/chat/stop | stop the running assistant turn |
/api/chat/audio/* | stream stored user audio and assistant TTS replies |
/api/chat/sessions/*/messages | read the device session history |
/api/chat/live/relay/** | live-voice relay (create session, send input, read events, close) |
POST /api/chat/live/tools | live-voice tool execution |
Every other route returns 401 for a device token, including all user-facing /api/voice-devices/* management routes, /api/auth/*, and the rest of the business APIs. A device token never grants admin rights, never accesses other users' data, and cannot pair or unpair devices.
Heartbeat
The device calls POST /api/voice-devices/heartbeat periodically (the Atom firmware sends one every 60 s; the settings UI treats a device as online when last_seen_at is less than two minutes old). The body carries an optional firmware_version and a free-form status object (for example rssi, ip, uptime_s). The server updates last_seen_at, last_seen_ip, firmware_version, and last_status, and answers with the device identity, its chat_session_id, the owner's current device settings (voice_mode, live_voice, turn_target_kind), and server_time so the device can align its clock. The heartbeat is how settings changes reach the device: the firmware applies the returned voice_mode and live_voice to its next button press, so a mode switch made in Clapilot takes effect within about a minute. GET /api/voice-devices/me returns the same shape without writing anything.
Push-to-talk (walkie talkie) through POST /api/chat
For a push-to-talk turn the device sends one utterance as a normal personal-chat turn to POST /api/chat with a type: "audio" attachment (base64 data, mimeType, name, optional durationMs), exactly like the Apple chat composer's press-and-hold mic. The firmware streams the microphone while the button is held: a capture task fills a ~0.5 s ring buffer, the request goes out with chunked transfer encoding as raw audio/pcm (PCM16 mono 24 kHz, little-endian) in 32 ms base64 chunks, and the server wraps the bytes into a WAV and computes durationMs itself. So no board has to hold the whole utterance in memory (the Cardputer's free heap would allow only a second or two), there is no length cap, and the upload is finished a few hundred milliseconds after release. A tap shorter than 300 ms opens no request. The heartbeat and the voice turns share one kept-alive HTTPS connection: a TLS handshake costs about 3 s on a PSRAM-less board (longer than a short utterance) and Cloudflare closes the connection after every streamed reply, so the heartbeat keeps it warm, a heartbeat issued right after each reply rebuilds it in the background, and a button press only writes the request head (measured: 12 ms instead of 2.8 s). The spoken reply is streamed as well: a reader task fills a 32 KB ring from /api/chat/audio/[id]?format=pcm16 while the codec drains it, so replies of any length play on every board. If the chat stream ends before a slow agent has answered, the device keeps asking GET /api/voice-devices/reply-audio while the server reports the reply as pending, for up to five minutes, and shows the elapsed time on a display. The backend stores the audio as a user-owned chat asset, runs speech-to-text, injects the transcript into the agent turn, and marks the turn for an assistant audio reply. For device-authenticated turns the TTS reply is synthesized before the stream ends and announced in-stream as a clapilot.audio SSE frame carrying the attachment (id, url, mimeType, durationMs, transcript) right before [DONE]; the device then fetches /api/chat/audio/[id]?format=pcm16&sample_rate=<rate> (the Atom firmware asks for 24000, its codec rate; the route resamples to any supported rate) and streams the raw PCM16 mono bytes into its I2S output. If the frame is missing (older server, asynchronous specialist reply), the device falls back to GET /api/voice-devices/reply-audio?since=<ISO>, which returns the newest spoken reply for the device's configured target regardless of where the turn was delivered (see Turn targets), or, on servers without that route, to polling the session history for the newest assistant audio attachment.
A device-authenticated POST /api/chat without a sessionId uses the device's own chat session, so the conversation appears in the owner's chat list under the device name, unless the device's turn target says otherwise. The device may pass its chat_session_id explicitly; it cannot address sessions that do not belong to its owner. Routing fields in the device request body (roomId, group flags) are ignored: the stored turn target alone decides where a device turn goes.
Turn targets
Each paired device has a turn target (voice_devices.turn_target, JSON) that decides where its spoken turns are delivered. The owner chooses it per device; the device itself has no say and cannot change it. Four kinds exist:
| Kind | Where the turn goes | What the owner sees |
|---|---|---|
device_session (default) | the device's own personal chat session | a chat titled after the device in the owner's chat list |
main_session | the owner's main personal chat (Hauptchat) | the spoken turn and the reply appear in the Hauptchat, mixed with the owner's typed conversation |
team_chat (roomId) | a team chat room (channel or group the owner is a member of; public channels are always allowed, human DMs and archived rooms are not) | the transcript is posted as the owner into the room. The main agent answers a speaker there even in rooms where it normally replies only when mentioned (provided the main agent is enabled for the room), unless the utterance explicitly addresses a specialist (@handle); in that case the specialist answers as in any other team-chat turn |
specialized_agent (agentId) | a direct conversation with one enabled specialist | the device session is pinned to that specialist (model = agent:<handle>), so the device chat becomes a direct conversation with the specialist; switching back to device_session unpins it again |
Spoken replies exist for all four kinds. For the personal targets and for the main agent in a room the TTS reply is synthesized before the chat stream ends and announced with the clapilot.audio frame as described above; room replies are stored as audio assets on the team-chat message. Specialist replies run asynchronously: their audio is synthesized when the specialist turn completes, and the device picks it up through GET /api/voice-devices/reply-audio?since=<ISO>, which reports the newest reply for the target (audio: null, pending: true while the specialist is still working).
The owner changes the target under Settings -> Sprachgeräte in the device row (Ziel / Target / Destinazione), or in the same voice-device settings screen of the iOS/macOS app. The list of selectable targets comes from GET /api/voice-devices/options (targets: device session, main chat, the owner's team rooms, enabled specialists; labels are server-localized; the older GET /api/voice-devices/targets still returns the same list on its own), and the choice is saved with PATCH /api/voice-devices/{id} and { turn_target }. A target the owner may not use (room they are not a member of, disabled specialist, agent DM room) is rejected with 400. A malformed stored value (for example a team_chat target without roomId) is read as device_session.
The turn target governs push-to-talk turns. A device in live mode talks through the relay, which is a personal-session concept: live conversations are not posted into team rooms and are not pinned to a specialist. A room or specialist target therefore only takes effect while the device is in push-to-talk mode.
Conversation modes
Each device has a conversation mode (voice_devices.voice_mode) that the owner sets per device; the device reads it from the heartbeat response and cannot change it itself.
| Mode | Button | What happens |
|---|---|---|
push_to_talk (default) | hold to talk, release to send | one utterance is recorded, sent as a chat audio turn to the device's turn target, transcribed, answered, and the spoken reply is played (see Push-to-talk) |
live | press to start, press again to stop | the device opens a continuous live conversation through the GPT-Live relay; the user speaks freely, the assistant answers in real time, and tool calls run during the conversation (see Live voice through the relay) |
Live mode additionally has a voice (voice_devices.live_voice): one of the GPT-Live relay voices (marin, cedar, quartz, ripple, vesper, willow, stone, gleam, meridian, bossa, tempo, beacon, delta, cinder), or null for the server default (marin). The device passes it as voice when it creates the relay session.
Live mode is billed per minute of open relay session by the provider (the relay keeps the provider Realtime socket open for the whole conversation, including pauses), whereas push-to-talk is billed per turn (speech-to-text, chat completion, text-to-speech). The firmware ends idle live sessions on its own, see below.
The owner sets mode and voice under Settings -> Sprachgeräte in the device row (Modus / Mode / Modalità next to the target; Stimme / Voice / Voce appears only while the mode is live, with Standard for the server default), or in the same voice-device settings screen of the iOS/macOS app. Mode labels, descriptions and the voice list come from GET /api/voice-devices/options, and the choice is saved with PATCH /api/voice-devices/{id} and { voice_mode } / { live_voice } (null resets the voice). Invalid values are rejected with 400. The device picks the change up with its next heartbeat (about a minute); a live conversation that is already running keeps its voice until it ends.
Live voice through the relay
For a live conversation the device uses the same server-held relay that the Apple Watch and other native clients use. The flow implemented by the Atom firmware (firmware/clapilot-voice/main/live_session.c):
- On the starting button press the device calls
POST /api/chat/live/relay/sessionwith its device token,voice(when the owner set one) andclientContext: { routePath: "/voice-device", pageContext: { scope: "voice-device", platform: "esp32", deviceKind } }. The server creates the provider Realtime session under the owner's identity and returnstransport: "clapilot_relay",relay_session_id, the model and voice in use, and PCM 24 kHz audio metadata. - The response also carries
ws_path(/api/chat/live/relay/ws) and a one-timews_ticket. The device openswss://<instance>/api/chat/live/relay/ws?session=<relay_session_id>&ticket=<ws_ticket>; that single socket carries the conversation in both directions. Next.js route handlers cannot answer an HTTP upgrade, so the handler lives inscripts/standalone-ws-server.mjs, the container entry that attaches anupgradelistener to the standalone server and then runs it unchanged. It authenticates with the ticket (constant-time compare) of a relay session that an authenticated request created, so no auth logic is duplicated, and it forwards frames verbatim between device and relay. When the server offers no ticket (older instance) or the socket cannot be opened, the device falls back to the POST/SSE transport described below. - The device streams its microphone as mono 24 kHz PCM16 in 320 ms blocks. Over the WebSocket it sends them as binary frames and the bridge wraps each one into an
input_audio_buffer.appendevent with base64 on the server; that removes a third of the uplink bytes and the encoding cost from the device, which is what lets it keep up with real time. Text frames stay available and are forwarded verbatim, and the fallback path base64-encodes on the device and posts toPOST /api/chat/live/relay/[sessionId]/input. Capture and transport run in separate tasks with a bounded queue, so a slow network never stalls recording; the oldest block is dropped when the queue is full. Turn detection runs on the server (provider VAD); the device sends no commit events. - It reads relay events from the same WebSocket (fallback: the SSE stream
GET /api/chat/live/relay/[sessionId]/events) and playsresponse.output_audio.deltaframes (base64 PCM16 at 24 kHz) through a playback buffer that primes with 300 ms before a burst so short gaps between deltas do not click.input_audio_buffer.speech_startedclears whatever is still queued so a reply does not keep playing over a new question; transcript events are only logged. GPT-Live keeps its output channel open and streams silence between replies, so silent deltas (peak below about -36 dBFS) are dropped unless a reply is still queued; otherwise the idle silence would count as "assistant speaking" and keep the microphone muted. - Function calls (
response.function_call_arguments.doneor afunction_callitem inresponse.output_item.done) are executed by the device throughPOST /api/chat/live/toolswithtoolName,arguments, the device'schat_session_idassessionId, and the sameclientContextas above. The tool output is answered back into the relay as aconversation.item.create/function_call_outputevent followed byresponse.create, exactly as the Apple clients do; the existing live tool contract applies unchanged, see Agent Tool Contracts. - The conversation ends on the next button press, after 60 s without audio in either direction (only counted while the assistant is not speaking), after a hard cap of 10 minutes, or when the relay closes the session (
*.closedevent, or404/409on the input route). The device then callsDELETE /api/chat/live/relay/[sessionId]and returns to idle.
Relay sessions created by a device token use a stricter server VAD on the GA path (threshold: 0.65, silence_duration_ms: 800) because a room microphone without echo cancellation otherwise starts a turn on breathing, and each such turn mutes the microphone while the model answers it. Browser and Apple clients keep the default sensitivity. The GPT-Live path (gpt-live-1) does its own turn detection and takes no threshold.
A reconnecting WebSocket client deliberately does not get buffered audio replayed (the SSE transport still does): a device that drops and reconnects mid-conversation would otherwise play seconds of stale speech over the live reply, and keep its own microphone muted while doing so. Non-audio events are replayed so session and tool state stay consistent.
The bridge pings a connected device every 20 s but treats any inbound frame as a sign of life (audio, text, the client's own pings, pongs); only a connection that stays completely silent for three intervals is dropped. A device whose pong arrives late is therefore not disconnected mid-reply. The web container log ([live-ws]) records one line per conversational milestone for device sessions (speech started, user: …, assistant: …, response completed, upstream errors) and, on disconnect, the close code plus how many frames and seconds of audio went up, how many pings/pongs the device sent, and a type×count summary of the events that went down. On the GPT-Live path the transcript lines come from clapilot.live.transcript.segment and tool delegations from response.function_call_arguments.done; that path has no speech-started/response-done events. This is enough to tell "the provider never heard speech" from "the reply was generated but not played" without a serial console.
Loudness on the Atom is set in firmware/clapilot-voice/main/app_config.h (AUDIO_OUTPUT_VOLUME_PERCENT, AUDIO_PLAYBACK_GAIN through a soft limiter). The level is bounded by the USB power budget, not by the codec: raising the codec to 0 dB with a 1.4x-2.5x software gain made the NS4150B pull the supply rail down, beeps and replies stuttered, the button read spurious presses and the chip reset. The codec therefore stays at the 85 % that runs stable from a hub port, and loudness comes from a compressor in audio_hal_play: samples are multiplied by AUDIO_PLAYBACK_GAIN (2.0) and everything above AUDIO_COMPRESSOR_KNEE (8000) is squeezed 4:1, so speech comes out about 6 dB louder while a full-scale sample still ends below the old peak level and the supply sees no higher current spikes. The boot banner and the ready: line log the last reset reason, so a BROWNOUT is visible in the console even though the USB console re-enumerates on a reset.
The ES8311 codec path on the Atom has no acoustic echo cancellation, so the firmware keeps the microphone muted while the speaker plays and for 250 ms afterwards; otherwise the server VAD would hear the reply and answer itself. Barge-in by voice is therefore not possible yet: the user waits for the reply to finish (or presses the button to end the conversation). This is a device-side limitation, not a relay one; devices with echo cancellation can leave the microphone open.
See GPT-Live voice transport for the relay routes and event types.
Firmware builds per board
The firmware in firmware/clapilot-voice/ is one code base with a board layer (main/board.h, main/boards/*.h) selected through idf.py menuconfig → Clapilot Voice → Target board, or through the per-board sdkconfig.defaults.* overlay. Each board has its own build directory and generated sdkconfig, so the three targets can be built side by side:
| Board | Build | Audio path | Setup |
|---|---|---|---|
| Atom VoiceS3R | idf.py build (default) | ES8311 codec, full duplex | captive portal |
| Cardputer | idf.py -B build-cardputer -DSDKCONFIG=sdkconfig.cardputer -DSDKCONFIG_DEFAULTS="sdkconfig.defaults;sdkconfig.defaults.cardputer" build | PDM microphone + NS4168 amplifier, half duplex (both need GPIO 43 as clock; the HAL switches direction, playback wins) | typed on the device (74HC138 matrix; untested so far) |
| Cardputer-Adv | idf.py -B build-cardputer-adv -DSDKCONFIG=sdkconfig.cardputer_adv -DSDKCONFIG_DEFAULTS="sdkconfig.defaults;sdkconfig.defaults.cardputer_adv" build | ES8311 codec (MCLK from BCLK), full duplex | typed on the device (TCA8418 keyboard controller on the internal I2C bus) |
Flash with idf.py -B <build dir> -p <port> flash. Pairing data in NVS survives re-flashing.
Cardputer behaviour
- Display (LVGL on the ST7789): header with device name, mode (
talk/live) and Wi-Fi signal; a scrolling conversation area withYou:andAgent:lines (push-to-talk turns show(voice message)for the user side, live sessions show both transcripts as the relay delivers them); the typed line above the status; a status line (Listening…,Thinking…,Speaking…,Ready). - Keys: hold Space to talk (push-to-talk) or press it to start/stop a live conversation, exactly like the Atom's button; G0 on the side does the same. Any other printable key starts a typed message; Enter sends it as a text turn to the same target (
POST /api/chatwithpageContext.input = "keyboard"), the reply is shown and spoken; Backspace edits, Esc (Fn+`) or Del (Fn+Backspace) clears. - Setup: with nothing configured (or after
wifi_failed,invalid_code,unauthorized), the device shows a five-step form instead of opening the portal: Wi-Fi SSID, password, instance URL (prefilled withhttps://app.clapilot.com), pairing code, device name. Enter moves on, Esc goes back, the last Enter saves and restarts into the normal pairing flow. The captive portal is not started on keyboard boards. - Memory: neither Cardputer has PSRAM (the Adv is a StampS3A with an ESP32-S3FN8 too). Push-to-talk streams the microphone as it records (see above), the live playback ring holds ~1.3 s and the WebSocket frame buffer is 8 KB; both modes run within the internal RAM.
- Loudness:
AUDIO_PLAYBACK_GAINis 4.0 on the Cardputer (the NS4168 path is quiet at unity) and the PDM microphone is scaled byAUDIO_MIC_SOFT_GAIN(8); both live inmain/boards/cardputer.hand are the first thing to tune on hardware. - Audio on the Adv: the ES8311's left I2S slot carries the microphone (the right slot showed a constant ±15480 waveform on this unit, so the slot is fixed rather than picked by level), the microphone gain is 18 dB (30 dB clipped at 0 dBFS), and the DAC volume register is raised by +6 dB on top of the 100 % codec level because the speaker path is far quieter than the Atom's; a boot-time acoustic self-test logs the microphone level with the speaker silent versus during a 1 kHz tone.
- Telling the variants apart: both look alike; the Adv has an I2C bus on GPIO 8/9 with the ES8311 (0x18), the TCA8418 (0x34) and an IMU (0x69), the classic drives its matrix through those pins. If keys do nothing after flashing, the other build is the right one. Verified on hardware 2026-09-21: Cardputer-Adv (display, keyboard, ES8311).
- Display strings are English only for now (device UI, not covered by the app localization system).
Unpairing
DELETE /api/voice-devices/{id} (owner, cookie auth) hard-deletes the device row. The device receives 401 on its next request and must be paired again to be used. Renaming (PATCH /api/voice-devices/{id} with { name }, 1–80 characters) only changes the display name and the title shown in the owner's chat list from then on; the token and binding are unchanged. The same route accepts { turn_target }, { voice_mode } and { live_voice } (any combination, optionally together with name) to change where the device's turns go and how it converses, see Turn targets and Conversation modes.
Security notes
- The device token is shown once at pairing and stored only as a SHA-256 hash. Clapilot cannot display it again; a lost token means unpair and pair again.
- Pairing codes are hashed at rest (
code_hash), expire after ten minutes, are single-use, and are invalidated when the user generates a new code. POST /api/voice-devices/pairis public (no cookie) and rate limited; throttled calls receive429withRetry-After. Unknown, expired, and already-used codes all return the same400withcode: "invalid_code"so nothing about code state is leaked.- Device tokens are rejected outside the allowlist above. The management routes (
GET /api/voice-devices, pairing-code creation, options/target list, rename, target/mode/voice change, unpair) require the owner's normalclapilot_sessioncookie. - Turn targets are validated against the owner's own access (room membership, enabled specialists) when set and are never taken from the device request; a device token cannot post into rooms its owner cannot see.
last_seen_ipandlast_statusare visible only to the owning user in their device list.- There is deliberately no chat or live-voice agent tool for pairing or unpairing devices; issuing a device credential is a settings-only workflow (same policy as API keys).
Data model
Migration db/migrations/319_voice_devices.sql adds:
voice_device_pairing_codes:id,user_id(FKusers),code_hash(unique),expires_at,consumed_at,device_id(nullable),created_atvoice_devices:id,user_id(FKusers, owner/binding),name,device_kind(defaultm5stack_atom_voices3r; see Supported hardware),hardware_id,firmware_version,token_hash(unique),token_prefix,chat_session_id(FKchat_sessions, nullable),last_seen_at,last_seen_ip,last_status(jsonb),created_at,updated_at
Migration db/migrations/320_voice_device_turn_targets.sql adds:
voice_devices.turn_target(jsonb, not null, default{"kind":"device_session"}), see Turn targetschat_audio_assets.group_message_id(FKchat_group_messages, nullable, cascade delete) so spoken replies for room turns can belong to a team-chat message;chat_message_idbecomes nullable and a check constraint requires at least one of the two references to be set
Migration db/migrations/321_voice_device_voice_mode.sql adds:
voice_devices.voice_mode(text, not null, default'push_to_talk', checkpush_to_talk | live)voice_devices.live_voice(text, nullable;null= server default voice), see Conversation modes
API
The full request/response contract for all user-facing and device-facing routes is in the API Reference.
When Live Voice uses a connected Codex/ChatGPT subscription provider, device options offer arbor, breeze, cove, ember, juniper, maple, sol, spruce, and vale. The same PCM relay runs gpt-live-1-codex through server-side WebRTC; the device keeps its existing owner-scoped token and tool permissions. Stored voices from another provider normalize to the selected model default. Subscription login failure does not switch to API billing.
