WebSocket audio endpoint
A websocket endpoint streams raw PCM audio to/from an external peer over a WebSocket, instead of negotiating RTP/SRTP via SDP. Internally it behaves like a bridge endpoint (L16, payload type 127) but at the wire sample_rate rather than a fixed 48 kHz, so it participates in routing, transcoding, and multi-party mixing exactly like any other endpoint. Running at the wire rate means the session's per-edge resampling converts directly to each peer (e.g. an 8 kHz PCMU leg goes 16 kHz↔8 kHz, not 16→48→8) — no unnecessary 48 kHz detour.
Lifecycle
- Create (control plane). Call
endpoint.create_websocketon a session. The endpoint is created in theconnectingstate (not yet routed) and a single-useconnect_tokenis returned. - Dial in (audio plane). The audio peer opens a WebSocket to
ws://<host>:<port>/audio/<connect_token>on the same port as the control/HTTP server, orwss://when rtpbridge TLS is enabled. On success the endpoint transitions toconnected, joins the routing table, and anendpoint.ws.connectedevent is emitted. - Stream. Binary WebSocket frames carry raw audio (see Wire format). Audio routes to/from other endpoints in the session per the endpoint's direction.
- Disconnect. When the socket closes, the endpoint moves to
disconnected, leaves the routing table, andendpoint.ws.disconnectedis emitted. There is no automatic reconnect — create a new endpoint to reconnect.
endpoint.create_websocket
Request params:
| field | type | default | notes |
|---|---|---|---|
direction | string | sendrecv | sendrecv / recvonly / sendonly / inactive (SDP sense) |
sample_rate | number | 8000 | wire PCM rate in Hz: 8000, 16000, or 48000 |
flush_ms | number | 0 | outbound coalescing window, multiple of 20 (0 = passthrough) |
Result:
{ "endpoint_id": "<uuid>", "connect_token": "<uuid>" }The client constructs the audio URL itself as /audio/<connect_token>.
Direction uses the same peer/SDP perspective as plain RTP and WebRTC endpoints:
sendonly— the endpoint is a source only: the peer's audio enters rtpbridge (other endpoints hear the WS peer).recvonly— the endpoint is a destination only: rtpbridge sends other endpoints' audio to the WS peer.sendrecv— both.inactive— neither source nor destination; the endpoint stays attached but is isolated from routing.
Wire format
Binary WebSocket frames only. Each frame is raw little-endian 16-bit mono PCM at the negotiated sample_rate. Frames may vary in length within the limits below; rtpbridge reframes the inbound stream to 20 ms internally (a trailing partial sample is buffered until the next frame). Text frames are ignored; Ping is answered with Pong; Close disconnects the audio endpoint.
- Inbound (peer → rtpbridge): carried as L16 at
sample_rate(no resampling at the socket) and routed to destinations (resampled/transcoded to their codecs as needed). rtpbridge synthesizes a monotonic RTP timeline so downstream RTP/WebRTC peers see advancing timestamps. - Outbound (rtpbridge → peer): each source is transcoded to the WS endpoint's L16 stream at
sample_rateand written as binary frames. Output is source-clocked: a frame is produced for each 20 ms of routed audio (no synthetic silence is sent when sources are idle). Withflush_ms > 0, that many milliseconds of audio are coalesced into a single WebSocket message (e.g.flush_ms: 100at 8 kHz → 1600-byte messages).
Events
| event | data | when |
|---|---|---|
endpoint.ws.connected | { endpoint_id, connected_at_epoch_ms } | audio socket attached; media-host routing boundary |
endpoint.ws.disconnected | { endpoint_id } | audio socket closed / errored |
endpoint.ws.connect_timeout | { endpoint_id } | no audio socket dialed in within 30 s; the endpoint is auto-removed |
Notes & limits
The
connect_tokenis a single-use secret; reusing it (or presenting an unknown or malformed token) closes the audio socket with a 1008 (policy) close.A configured control-plane HMAC secret does not apply to this route. The single-use server-minted token is the audio capability; audio consumers must not receive the control HMAC signing key.
A created endpoint that is never dialed into is auto-removed after 30 s (
endpoint.ws.connect_timeout), reclaiming its endpoint slot and token.WebSocket endpoints cannot be transferred between sessions (
endpoint.transfer).Backpressure: if the peer can't keep up, the newest outbound frame is dropped rather than blocking the session (same policy as bridge endpoints).
The inbound byte ring buffers at most 256 KiB of PCM (about 16.4 seconds at 8 kHz or 2.7 seconds at 48 kHz). Accepted bytes are paced into the session at 20 ms per frame; trailing odd bytes are carried across messages. Sending beyond the budget closes the connection. Closing the endpoint discards remaining buffered input.
Both transports close after a socket write exceeds five seconds or a matching Pong is absent for ten seconds after a Ping. Audio Pings are sent every 30 seconds; control uses
ws_ping_interval_secs. Reads and heartbeat processing run independently of media output and control handlers.Outbound packet queueing remains nonblocking and may drop the newest packet under backpressure. The socket writer also has a finite message/byte budget. Disconnect completion is reserved at attach time so a full session command channel cannot lose cleanup.
The 256 KiB limit applies to the IO ring. Also budget for the currently decoded WebSocket message, the session’s bounded 256-packet input queue, its playout frames, and transport buffers; those are separate fixed bounds. The ring drains at most one 20 ms frame per tick, so accepted TTS bursts are paced instead of filling the downstream queue at once.