Troubleshooting
Common Failure Modes
Session creation fails with SHUTTING_DOWN
The server is in graceful shutdown (received SIGINT/SIGTERM). New sessions are rejected while existing sessions drain. Wait for the restart to complete or redirect traffic to another instance.
Endpoint creation fails with SDP errors
- "SDP too large" — The offer exceeds
max_sdp_size_kb(default: 64 KB). If legitimate, increase the limit in your config. - "No supported codec" — The SDP offer doesn't include any of the supported codecs (PCMU, G.722, Opus). Ensure the remote side offers at least one.
- "Port allocation failed" — All ports in
rtp_port_rangeare in use. Either increase the range or reduce the number of concurrent endpoints.
WebSocket connection drops
When a control WebSocket disconnects, the session enters an orphaned state for disconnect_timeout_secs (default: 30s). During this window:
- Media continues flowing between endpoints
- The session can be reclaimed via
session.attachon a new connection - After the timeout, the session and all its endpoints are destroyed
If you see frequent orphaning, check network stability between your application and rtpbridge.
No audio / one-way audio
- Check endpoint directions — Verify
sendrecv/sendonly/recvonly/inactiveare set correctly. Arecvonlyendpoint won't forward media, and aninactiveendpoint is isolated. - Check routing — Use
stats.subscribeto monitor packet counts. Ifinbound.packetsincrements butoutbound.packetsdoesn't, the routing table may not include the expected path. - Check codec mismatch — If two endpoints negotiate different codecs, transcoding kicks in automatically. Monitor
rtpbridge_transcode_errors_totalfor failures. - Check NAT/firewall — For plain RTP endpoints, ensure the remote can reach
media_ipand the allocated even/odd ports fromrtp_port_range. For WebRTC endpoints, ensure peers can reach the advertisedmedia_ip:portICE host candidate; WebRTC uses OS-assigned UDP ports, notrtp_port_range.
SRTP errors
A rising rtpbridge_srtp_errors_total counter indicates authentication or replay failures. Common causes:
- Key mismatch — The SRTP keys in the SDP don't match what the remote is using.
- Replay attack detection — Duplicate or reordered packets beyond the replay window. Usually indicates network issues.
- SSRC collision — Two sources using the same SSRC. Rare but possible.
Recording failures
- "Cannot create recording file" — The
recording_dirdoesn't exist, isn't writable, or disk is full. - "Maximum concurrent recordings reached" — Hit the
max_recordings_per_sessionlimit (default: 100). Stop unused recordings or increase the limit. - Channel full warnings — The recording writer can't keep up with packet rate. This usually means disk I/O is saturated. Packets are dropped from the recording (not from the media path).
Events dropped
rtpbridge_events_dropped_total incrementing means the WebSocket client isn't consuming events fast enough. The event channel has a fixed buffer; when full, new events are dropped. Ensure your client reads messages promptly without blocking.
Session idle timeout
If session_idle_timeout_secs is configured (non-zero), sessions with no media packets and no commands for that duration are automatically destroyed. Check this setting if sessions are disappearing unexpectedly.
Debugging
Enable debug logging
rtpbridge --log-level debugOr in your config file:
log_level = "debug"Use trace for maximum verbosity (includes per-packet logging).
Inspect live sessions
Query the HTTP API on the control plane port:
# List all active sessions
curl http://localhost:9100/sessions
# Check server health
curl http://localhost:9100/health
# Dump Prometheus metrics
curl http://localhost:9100/metricsInspect recordings with Wireshark
Recording files are standard PCAP format with synthetic Ethernet/IPv4/UDP headers. Open them directly in Wireshark:
wireshark /var/lib/rtpbridge/recordings/session-abc/recording.pcapWireshark will automatically dissect RTP/RTCP protocols. Endpoints are assigned deterministic synthetic IP addresses (10.x.0.y:10000) for easy filtering. The bridge side of outbound packets uses 10.255.0.1.
Use stats for live diagnosis
Subscribe to per-endpoint statistics over the WebSocket. For normal polling, use the default compact payload. For receive-path diagnosis, opt into diagnostics:
{"id":"1","method":"stats.subscribe","params":{"interval_ms":1000}}{"id":"2","method":"stats.subscribe","params":{"interval_ms":1000,"include_diagnostics":true}}Key fields to watch:
inbound.packets/outbound.packets— Are packets flowing?inbound.raw_packets— Are socket-level datagrams arriving? If raw packets keep rising whileinbound.packetsis flat, the path is alive but no media is being decoded. If both are flat, the remote path is likely dead. Present only withinclude_diagnostics: true.inbound.raw_rtp_packets_lost/inbound.raw_rtp_sequence_gaps— RTP sequence loss before bridge media processing. For WebRTC, this separates browser/TURN or network loss from loss added after socket ingress. Present only withinclude_diagnostics: true.inbound.packets_lost— Media-plane loss after endpoint processing. If this rises while raw RTP sequence counters stay clean, investigate bridge-side WebRTC/RTP processing.inbound.jitter_ms— Network jitterinbound.recv_loop_gap_ms,inbound.dequeue_delay_ms, andinbound.channel_overflows— Receive-loop stalls or session-channel backpressure. Present only withinclude_diagnostics: true.inbound.last_received_ms_ago— Time since last packet (high values indicate stalled media)rtt_ms— Round-trip time from RTCP for plain RTP/SRTP endpoints
Prometheus also exposes send-result counters for socket-backed endpoints: rtpbridge_udp_send_ok_total{endpoint_type,packet_type,family} and rtpbridge_udp_send_errors_total{endpoint_type,packet_type,family}. A rising error counter means packets reached the endpoint writer but failed at the UDP socket send step; compare it with endpoint outbound.packets to separate media routing from OS/network egress problems.
Check resource limits
If sessions or endpoints fail to create, verify your limits:
max_sessions = 10000 # 0 = unlimited
max_endpoints_per_session = 20 # 0 = unlimited
max_recordings_per_session = 100Also check OS-level limits:
- Open file descriptors — Each UDP socket and recording file consumes an fd. With 10,000 sessions and multiple endpoints each, you may need
ulimit -nin the hundreds of thousands. - UDP port range — The
rtp_port_rangemust have enough ports for all concurrent plain RTP endpoints. WebRTC endpoints use OS-assigned UDP ports outside this range.