Security and performance remediation plan
Prepared September 5, 2026, against review baseline 7614703. Status: runtime remediation implemented and reviewed again September 8; validation and release follow-ups are tracked below and in the implementation report. This plan covers all 16 numbered review findings and the additional WebSocket buffering, allocation, PCAP conversion, deployment, dependency, and documentation issues. Finding numbers below refer to that review.
Complete the security and ownership fixes before changing the shared media pipeline. Keep each step independently reviewable, with its regression checks in the same change. Steps 1–7 close all seven P1 findings; step 0 addresses unsafe deployment defaults immediately. The later performance work must preserve the corrected security and lifecycle behavior.
The original review was static. Tests, release builds, benchmarks, a 50-call soak, and candidate-image scanning are now part of implementation validation. Checked items describe delivered changes; acceptance paragraphs remain the intended coverage envelope, not a claim that every listed fault combination has been exercised. The implementation report records actual commands, results, and remaining coverage limits.
Implementation sequence and coverage
| Step | Change | Review coverage | Depends on | Primary files |
|---|---|---|---|---|
| 0 | Safe control defaults and maintained PEM parsing | Deployment and dependency items | None | src/config.rs, src/control/server.rs, Cargo.toml, deployment examples |
| 1 | Validate secure negotiation and preserve key epochs | 1, 3, 10 | None | src/media/sdp.rs, src/media/srtp.rs, src/session/endpoint_rtp.rs |
| 2 | Enforce RTP and RTCP source policy | 2 | 1 | src/session/endpoint_rtp.rs, src/session/media_session.rs, SDP/configuration |
| 3 | Bound transfer cancellation and make attach atomic | 4, 12 | None | src/session/mod.rs, src/session/endpoint_rtp.rs, src/session/media_session.rs, src/control/handler.rs |
| 4 | Own cache references and bound download admission | 5, 9; playback download memory | 3 lifecycle conventions | src/playback/file_cache.rs, src/session/media_session.rs, src/config.rs |
| 5 | Enforce playback destination policy and redact secrets | 8, 15 | 4 | src/playback/file_cache.rs, src/control/logging.rs, configuration |
| 6 | Isolate storage work and fix shared-playback ownership | 6, 16 | 4 | src/recording/recorder.rs, src/session/endpoint_file.rs, src/playback/shared_playback.rs, session integration |
| 7 | Stream recording HTTP responses | 7 | 6 bounded storage facilities | src/control/server.rs, configuration |
| 8 | Supervise WebSocket IO and preserve accepted audio bursts | 11; WebSocket reframing and Synth overflow | 3 lifecycle conventions | src/control/connection.rs, src/control/ws_audio.rs, src/session/endpoint_websocket.rs, src/session/playout.rs |
| 9 | Preserve sample duration through transcoding and mixing | 14 | 1–3, 8 | src/media/transcode.rs, src/session/playout.rs, src/session/mixer.rs, session routing |
| 10 | Share source decoding and remove measured packet allocations | 13; allocation items | 6, 9 | src/session/media_session.rs, src/session/mixer.rs, src/session/audio_analysis.rs, media helpers |
| 11 | Bound and stream offline PCAP conversion | PCAP conversion item | 9 framing semantics | src/bin/pcap2audio.rs, tests/pcap2audio_test.rs |
| 12 | Verify integration, update sizing guidance, and stage release | All findings and supporting documentation | 0–11 | Existing tests/benches, e2e/soak50, docs, CI |
These are implementation units, not a proposal to combine unrelated changes into one large pull request. Step 6 can separate recording workers from playback workers; step 4 should establish ownership before adding streaming. Each split must leave a buildable, internally consistent state.
Shared design rules
- Keep one media task per session and keep each str0m instance owned by that task. Move blocking work out through bounded messages; do not introduce locks around live WebRTC state.
- Make resource ownership explicit across asynchronous completion. An endpoint, transfer, download request, cache file, playback instance, or connection has one defined cleanup path. Late results carry ownership that can be safely dropped.
- Bound both active work and waiting work. Account for bytes and retained buffers as well as task counts; acquire admission before spawning. A timeout starts at admission and includes queueing. Never hold a map lock across network, file, or channel waits.
- Use synchronous, short ownership release where possible. Do not depend on spawning an asynchronous task from
Dropto repair reference counts. - Keep diagnostic counters for rejected or dropped work separate from accepted-media statistics. Use bounded metric labels without URLs, tokens, arbitrary SSRCs, or session IDs.
- Validate new limits and checked size/time arithmetic before listening. New security/resource limits have finite defaults; zero must not silently mean unlimited.
- Preserve existing authentication, single-use audio tokens, recording containment and exclusive creation, message/SDP limits, SRTP authentication before receive-state allocation, and the receive-SSRC cap.
- Follow the repository async rule: name any awaited result before using it in another operation.
Step 0 — Safe control defaults and dependency maintenance
- [x] Change both
Cli.listenandConfig::default().listento127.0.0.1:9100; verify CLI/file precedence still works. Explicit IPv6 loopback remains supported. - [x] Validate every configured listener. A non-loopback listener requires HMAC and TLS by default. Add separate explicit
allow_plaintext_controlandallow_unauthenticated_controlexceptions, both false by default. A TLS-terminating proxy uses only the plaintext exception and retains HMAC. Development deployments requesting both exceptions receive a clear startup warning. - [x] Document that loopback access trusts local processes, HMAC is administrative authorization across all sessions, and a trusted controller remains responsible for tenant isolation. Preserve the documented health/metrics and single-use audio-token route policies.
- [x] Update container, systemd, Kubernetes, configuration, and getting-started examples in the same change, including container port publishing and proxy reachability. Do not silently keep wildcard insecure defaults in an example.
- [x] Replace the vulnerable automatically vendored OpenSSL 3.6.3 with a checksum-verified 3.6.4 build shared by development, CI and Docker; validate the linked version at startup and record it in startup logs. Compatible
openssl-srccrates do not yet contain the patch. - [x] Replace the direct
rustls-pemfiledependency with PEM parsing fromrustls-pki-types, updating certificate-chain and supported private-key loading. Restrict lockfile changes to the replacement and required resolution changes. The reported advisory is an unmaintained-package notice, not an exploit: RUSTSEC-2025-0134.
Acceptance: configuration coverage includes mixed loopback/non-loopback listeners, IPv4/IPv6, TLS only, HMAC only, explicit proxy mode, and explicit insecure development mode. TLS tests cover valid chains/keys, supported PEM key encodings, malformed material, and mismatched keys. Existing control authorization behavior remains covered.
Step 1 — Secure negotiation and key lifetime
- [x] Parse the transport profile and crypto attributes into a validated negotiation result before constructing or mutating an RTP endpoint. Secure profiles with missing, malformed, or unsupported crypto fail with a stable protocol error. Validate every offer, answer, and renegotiation entry point. Plaintext fallback is available only for the explicitly supported opportunistic mode.
- [x] Make renegotiation transactional: invalid SDP leaves the previous transport, codecs, remote addresses, keys, replay history, and endpoint resources intact. Failure during initial creation returns socket/endpoint capacity.
- [x] Remove same-key SRTP/SRTCP receive-state resets. Preserve replay windows and rollover counters for each SSRC for the key lifetime. A new SSRC uses its own bounded state; a peer restarting the same SSRC and packet index must negotiate a fresh key.
- [x] Represent rekey retirement explicitly for both SRTP and SRTCP. Keep the current five-second transition as the proposed default. Successful new-key RTP cannot cancel retirement of the old SRTCP key. Enforce expiration on the endpoint timer and in both receive paths, including RTCP-only and completely idle transitions.
- [x] Retire old contexts idempotently, preserve independent transmit/receive keys, and release obsolete key material through the existing zeroization mechanisms where applicable.
Acceptance: rejected secure offers retain no endpoint/port allocation; invalid renegotiation leaves a working call unchanged. Previously accepted RTP and SRTCP remain rejected after a same-key update. Continuous traffic survives rollover; a new SSRC works. Old-key RTP and SRTCP fail after grace expiry even when only one protocol used the new key. Exercise replay and rollover on both sides of rekey with deterministic clocks/packet fixtures.
Step 2 — Source validation before media state changes
- [x] Carry the UDP source address through RTP and RTCP handling. Classify RTP/RTCP mux safely and validate packet structure, negotiated payload types, header extensions/padding lengths, and RTCP structure before applying accepted-media state or recording data.
- [x] Default plain RTP to the SDP peer IP with symmetric port learning. An IP different from the SDP address requires an explicitly configured peer network or controller-supplied expected source. Once established, enforce the selected tuple. Receiving a new SSRC alone must not reopen address learning.
- [x] Use a bounded, explicit learning state and a deliberate renegotiation/reset transition. Allow calls to wait for their first media during ringing; do not expire legitimate initial learning merely because media is delayed. Any packet probation is a compatibility heuristic, not authentication.
- [x] Enforce RTCP independently: support the negotiated RTCP destination and constrained same-IP symmetric learning when non-mux NAT mappings differ. Do not assume every NAT preserves the RTP-plus-one port relationship. Add parsing/validation for the necessary SDP RTCP attributes.
- [x] For SRTP/SRTCP, authenticate before accepting a candidate address migration, and still apply the configured address policy. Preserve address-family binding and the existing prohibition on plain RTP renegotiation across families.
- [x] Reject disallowed sources before SSRC changes, latching, accepted-media stats, routing, DTMF/BYE events, or recording. Record only aggregate rejection diagnostics.
Acceptance: a second socket cannot inject media/control or relatch an established endpoint. Test first-packet injection from a disallowed IP, locked-port changes, malformed/unknown-PT packets, valid same-IP NAT learning, approved alternate peer addresses, RTP/RTCP mux and separate ports, IPv4/IPv6, and authenticated SRTP migration. Actual source-address spoofing remains outside the protection offered by plaintext tuple checks; deployments needing that guarantee use SRTP.
Step 3 — Transfer and attach ownership
- [x] Make RTP and RTCP receive enqueue nonblocking with counted drops, matching the existing WebRTC approach. Cancellation and joining must finish even when the session packet queue is full.
- [x] Give receive-task joins a deadline and an explicit fallback for asynchronous tasks that fail to exit. Confirm socket/port ownership remains in the endpoint bundle until handoff completes.
- [x] Handle the whole transfer exchange: extraction, destination admission, commit, source restoration, dropped reply receivers, controller disconnect, destination destruction, and command deadlines. Keep the bundle in a supervised operation through completion. Return it to the source on pre-commit failure; after destination commit, reconcile completion without inserting a duplicate at the source. A caller timeout must not strand or duplicate the endpoint.
- [x] Use a bounded pending-transfer record only if needed for the existing command protocol; remove it on terminal completion. No persistent transaction history is required.
- [x] Reserve command-channel capacity and commit attachment/state-generation changes atomically with respect to orphan expiry. Cancel the old orphan timer only after attachment can be delivered. Failed attachment keeps the original deadline. Expiry must check and destroy the same orphan generation atomically.
Acceptance: force a full packet queue and transfer while unrelated endpoints exchange media. Force destination queue failure, dropped replies, controller disconnect, and destruction around commit; assert one owner and no leaked ports/tasks. Force a full attach command queue and verify cleanup at the original orphan deadline. Race successful attachment against expiry using barriers or injected clocks, not probabilistic sleeps.
Step 4 — Owned cache leases and bounded downloads
- [x] Return a
CacheLeasecarrying the exact URL-plus-headers key and a unique cache-entry instance. Move the lease intoFileReadyand then into the endpoint/decoder only if that endpoint still exists. Discarding late completion releases it automatically. Remove URL-only release bookkeeping. - [x] Separate pending demand from a completed file lease. Acquire a bounded request-owner slot before creating waiting tasks, including waiters on an identical URL. Bound queued distinct downloads separately from active transfers. Reject overload promptly using a documented resource-limit error.
- [x] Start each owner's deadline at the request. Cancellation removes that owner's demand; the final owner cancels queued/active work. A surviving owner keeps a shared transfer alive, and one owner's short timeout must not fail every other owner. Give the shared transfer its own finite maximum lifetime.
- [x] Stream response chunks to unique exclusively created temporary files instead of collecting whole bodies. Enforce the byte limit while reading, independently of
Content-Length. Include temporary files, reservations, and completed files in an aggregate disk budget; count every retained response/write buffer in a memory budget. - [x] Publish a completed file atomically under an instance-specific name. Cache cleanup claims only unleased entries and deletes that exact file instance outside the lock. Old cleanup cannot remove a replacement download of the same key. Handle write errors, cancellation, disk full, initialization failure, and shutdown through the same ownership path.
- [x] Make the entry and byte limits hard admission limits when all existing files are pinned. Do not evict a leased file to meet the target. Recover only files owned by this cache implementation, using a dedicated cache directory and recognizable names.
Acceptance: repeat authenticated and unauthenticated variants of the same URL; remove endpoints during queueing, download, completion delivery, and decoder startup. Cache references, pending owners, permits, temporary files, and bytes return to baseline after owners finish and eligible eviction runs. A live shared owner keeps its file. Saturated create/remove loops have a fixed task/byte ceiling. Chunked, missing-length, oversized, and slowly delivered bodies respect limits and the end-to-end deadline.
Step 5 — Playback URL policy and credential-safe diagnostics
- [x] Require configured HTTP(S) media origins for URL playback; an empty allowlist disables remote playback. Explicitly support private media services through approved origins and destination networks. Reject URL userinfo, unsupported schemes, and unapproved ports. Header credentials remain scoped to the approved original origin.
- [x] Enforce the policy on actual connections. Validate all candidate IPv4/IPv6 addresses, normalize mapped addresses, and pin connection selection to allowed results while preserving the original HTTP Host and TLS server name. Ensure connection pooling cannot reuse a connection outside that policy. DNS preflight followed by an independent unrestricted resolution is insufficient.
- [x] Handle redirects explicitly, preserving the current maximum of five hops. Revalidate origin and destination at every hop; block HTTPS downgrade and strip caller credentials/custom sensitive headers on cross-origin redirects. Disable environment-derived proxies by default; any supported explicit proxy must enforce equivalent destination policy.
- [x] Sanitize download errors at their source. Logs/events contain an opaque request identifier, approved host, size/status, and error category. Exclude URL paths as well as userinfo/query/fragment, header values, raw reqwest error URLs, and audio/control tokens. Reuse the same sanitizer in the control summary layer.
- [x] Document allowed-origin migration and provide a working private-media example. Retain network egress restrictions as defense in depth, consistent with OWASP SSRF guidance.
Acceptance: approved public/private media works; forbidden loopback, link-local, metadata, private and mapped-address destinations fail unless explicitly allowed. Cover DNS changes between validation and connection, redirects, cross-origin headers, and proxy environment variables using local fixtures. Capture all log levels and error outputs with distinctive test credentials in URL path/query/userinfo and headers; none may appear.
Step 6 — Storage isolation and shared-playback instances
- [x] Put synchronous recording writes/flushes on a bounded dedicated worker facility. Add a process-wide active recording limit and a byte limit for queued recording packets, in addition to existing per-session limits. The media task enqueues without waiting; saturation increments recording-drop counters and follows the documented recording error policy.
- [x] Move local and downloaded file open/probe/read/decode/seek/rewind onto bounded workers. Keep decoder state there and send ready PCM through bounded prefetch queues. Session commands request work and process completion later; they do not await slow storage on the sole session task. Apply this to shared and nonshared playback, initial opens, loop rewinds, and seeks.
- [x] Bound the work inside the actual decode loop: packets skipped, consecutive errors, decoded duration, output bytes, and total work per dispatch. Yield between bounded batches and reject malformed streams that exceed the terminal error/work policy. Accept only supported regular-file inputs and validate decoded channel/rate/allocation sizes before use.
- [x] Give each shared playback a unique instance identity. Subscribers and worker completion can alter/remove only that instance. Every terminal path, including open failure and cancellation, closes or completes subscribers with an explicit EOF/error result. Subscriber cleanup is idempotent and does not require a live Tokio runtime.
- [x] Keep the cache lease alive in the decoder worker for the entire period the file can be accessed. A removed endpoint cannot make cleanup delete a file still in use by a shared or winding-down decoder.
- [x] Define stop/flush results honestly. A running blocking filesystem call cannot be aborted by cancelling its future. Hold worker/admission capacity until the actual operation exits; do not spawn replacement workers beyond the cap. Use dedicated threads whose stalled jobs cannot indefinitely block Tokio runtime shutdown. Drain within the configured shutdown deadline and report unfinished recordings; hard per-job termination would require a separate worker process.
Acceptance: delayed reads/writes/flushes do not block unrelated session commands or media scheduling; saturated workers cause bounded queueing/rejection. Verify recording drop/error reporting and finite shutdown. Malformed files reach bounded terminal failure. Finish playback A, start B for the same source, then remove A; B continues. Repeat cancelled-old-worker completion, immediate unsubscribe/resubscribe, invalid-file startup, seek/loop changes, and final-subscriber removal. File leases release only after actual worker use ends.
Step 7 — Recording HTTP streaming
- [x] Replace the whole-response vector with a response representation that can send small headers followed by a bounded file stream. Keep existing small JSON responses simple.
- [x] Open the authorized contained file safely and inspect metadata from the handle. Preserve path/symlink defenses and require a regular file. For an active recording, serve a snapshot of the length observed at open: advertise that length, cap reads to it, and exclude subsequent growth. A concurrent truncation/early EOF terminates the incomplete response.
- [x] Use fixed chunks, initially 64 KiB, with a separately bounded read-ahead queue. Add an independent recording-download concurrency cap, initially four, and read/write idle plus overall response deadlines. Use the storage admission facilities so slow file reads cannot create unlimited blocking work. Release the connection permit when the network operation ends; retain any worker permit until its real IO ends.
- [x] Preserve the existing recording-size admission limit, HTTP authorization, content type, deletion behavior, and recording path containment. Document that an active snapshot can end with a partial PCAP record; clients needing a complete file should stop recording first.
Acceptance: four slow downloads near the file-size cap retain memory proportional to configured chunks/read-ahead, never file length. A fifth download is rejected or waits only in an explicitly bounded queue. Growing files send no more than the snapshot length. Stalled readers and storage hit their respective limits; unrelated calls continue. Traversal, symlink, delete/open races, and authentication regressions remain covered.
Step 8 — WebSocket lifecycle and burst handling
- [x] Give each control/audio connection one supervisor owning its connection permit and cancellation token. Separate reading from writing with bounded channels and one ordered writer. Keep control command execution from preventing the reader from observing liveness frames; retain response/event ordering and critical-event priority.
- [x] Track outstanding Ping/Pong liveness and apply deadlines to send, flush, and close. Add audio keepalive. Outbound-only audio with no inbound samples remains valid when the transport responds. Cancellation must interrupt channel waits and pending network writes.
- [x] Ensure endpoint/session disconnect state is delivered exactly once through reserved lifecycle capacity or a supervisor-observed state transition. Do not replace an unbounded final channel wait with a best-effort dropped notification. The supervisor joins its asynchronous children within a deadline and releases the permit.
- [x] Replace front-draining sample vectors with a byte cursor/ring and retain a trailing odd byte across messages. Reframe without repeated copying of the remaining burst; bound conversion work per dispatch.
- [x] Separate WebSocket burst storage from the real-time Synth jitter queue. Implemented bound: a 256 KiB inbound IO ring, plus the separately bounded current wire message, session packet queue and playout frames documented in the WebSocket protocol. This retains one currently legal maximum-size message: about 16.4 seconds at 8 kHz, or 2.7 seconds at 48 kHz. Pace accepted data into playout without the current 14-frame truncation.
- [x] Preserve Ping/Pong processing while audio is buffered. A producer exceeding the configured aggregate burst budget receives an explicit overload close/error; do not silently discard accepted TTS content. Keep Bridge real-time overflow policy separate. Document the memory/maximum-latency tradeoff and retain the configured wire-message cap for deployments needing smaller individual messages. The IO ring limit is fixed in this implementation.
Acceptance: stop reading, stop answering Ping, disconnect during writes, remove endpoints with full queues, and stall command handling. Cleanup and permit release stay within deadlines. Outbound-only audio stays connected. At 8/16/48 kHz, one maximum allowed burst and arbitrarily split odd-byte messages preserve sample order and duration; oversized aggregate bursts follow the explicit policy. Copy/allocation work grows linearly with input bytes, and peak retained bytes stay within the advertised budget.
Step 9 — Duration-correct audio framing
- [x] Introduce a source-owned decoded-audio frame representation carrying sample rate, source/codec epoch, media position, and exact decoded sample count. Replace the one-input-packet/one-output-frame assumption with bounded PCM accumulation and zero-or-many output frames.
- [x] Reorder encoded packets before stateful decoding where playout requires it. Convert timestamps using the codec's RTP clock separately from its PCM rate, especially G.722. Drain/publish media according to duration, supporting multiple short input packets or multiple output frames from a long packet.
- [x] Preserve remainder samples across calls; remove unconditional 20 ms truncation/padding from transcoding, mixing, and WebSocket output. Use 20 ms generated output frames initially. At a declared stream end, pad only the final incomplete frame if required by the encoder; do not inject padding between consecutive short packets.
- [x] Accept supported legal Opus durations up to the decoder's 120 ms packet limit, including sub-20-ms input; cover 10/20/40/60 ms packetization for other applicable paths. Validate packet/decode sizes before allocation and keep PCM/reorder budgets bounded by bytes and media duration. On real-time overflow use an explicit discontinuity/drop policy with counters, rather than silently altering every packet.
- [x] Preserve transparent same-codec single-source forwarding where decoding is unnecessary. Keep DTMF outside audio decode. Define codec/SSRC changes, loss/PLC, hold/resume, direction changes, transfer, and synthetic WS/Bridge clocks so stale buffered samples cannot enter a new source epoch.
Acceptance: identifiable sample sequences and duration/timestamp checks survive transcode, mix, WebSocket output, and Bridge paths at variable packetization. Compare exact sample counts before lossy encoding and duration/signal tolerances afterward; account for resampler delay explicitly. Cover sequence/timestamp wrap, reordering, loss, dynamic codec changes, VAD/fax taps, and a source that sends eight 2.5 ms packets per output tick. Queues remain bounded with no accumulating latency under valid real-time input.
Step 10 — Shared decode and packet-path efficiency
- [x] Decode each real source packet once whenever any mixer, transcode destination, VAD, or fax consumer needs PCM. Share immutable decoded frames within the session. Share stateful resampling per source epoch and distinct required PCM rate; retain destination-specific encoders and mix timing.
- [x] Remove source decoders/resamplers from every destination mixer. Rebuild only changed routing/output state. Preserve all direction, bridge-loop, source-exclusion, transfer, and one-source passthrough behavior.
- [x] Sum into a wide accumulator and clamp once at output. A full-conference mix-minus optimization is optional only after proving routing equivalence; the required initial improvement is shared decoding/resampling.
- [x] Bound media state through validated finite session and endpoint admission limits: at most one source decoder, three resampling rates per source, and one mixer/output encoder per destination. Require the transcode cache to cover the endpoint cap and prune obsolete edges, preventing eviction of active pipelines. These are resource bounds, not a CPU reservation or scheduling SLO.
- [x] Gate recording descriptor construction before allocating strings and reuse the recording manager’s descriptor identities. Starting a recording mid-call must still obtain current descriptors before its first packet. Retain clear buffer ownership. Destination-info allocation, payload cloning, and metric micro-optimizations remain subject to profiling; no speculative pooling was added.
- [ ] Conference benchmarks now cover 2/3/10/20-party PCMU, G.722, Opus and mixed codecs, plus the former per-destination decode path. Extend the measured reference-host matrix with analysis taps, recording off/on, churn, allocation profiling, and media scheduling p95/p99 before publishing production capacity claims.
Acceptance: a fully active 20-party conference with 20 ms input performs approximately 1,000 real-packet decodes per second, instead of the reviewed 19,000 per-destination decodes; distinguish PLC/recovery operations in the instrumentation. Decoder state scales with sources, resamplers with source/rate pairs, and encoders with outputs. Passthrough remains decode-free unless a tap needs PCM. No-recording traffic creates no recording-descriptor strings per packet. Publish measured end-to-end results rather than inferring CPU speedup directly from decode counts.
Step 11 — Bounded PCAP conversion
- [x] Validate maximum duration, channels, sample rate, output bytes, and temporary-disk bytes before large allocations or silence expansion. Use checked arithmetic for timestamps, offsets, sample counts, and WAV/RIFF size fields. Reject impossible or over-limit capture spans with a useful error.
- [x] Apply the duration bound to both RTP-derived and capture-derived time. The existing fallback from implausible RTP time must not turn an unbounded capture timestamp into a huge silence allocation.
- [x] Replace whole-capture packet/PCM/interleaved vectors with bounded processing. Preflight seekable input for channel/timeline metadata and limits; spool per-channel data/PCM within a disk quota, then interleave fixed-size blocks to a temporary output and publish it on success. Use bounded reorder windows and explicitly reject or report captures beyond the supported disorder window.
- [x] Preserve codec epochs, RTP rollover, duplicate handling, channel alignment, mono/concatenated/multichannel modes, and required metadata. Remove temporary artifacts and incomplete final output on failure or cancellation.
Acceptance: a tiny capture with an enormous timestamp gap fails before large allocation or silence writes. Long ordinary captures have bounded RSS independent of decoded duration. Channel/rate/RIFF limits fail cleanly; valid fixtures retain existing timing and codec behavior. Disk usage never exceeds the configured spool/output budget without a reported error.
Step 12 — Integration gates and release
- [x] Add focused regressions to the existing suites: SRTP/security and IPv6 media tests; transfer/reconnect/concurrent mutations; file playback/resource exhaustion; recording/transport security; WS audio; transcoding/mixing/VAD/fax/Bridge; and PCAP conversion. Reuse existing helpers and add deterministic saturation/failure hooks where needed. Do not create tests that merely repeat the implementation.
- [x] After writing the implementation and regression code, run affected tests first, then the existing CI-required formatting, Clippy, unit/integration and release checks. Preserve the repository's serial integration-test and timeout conventions. Run benchmarks when relevant changes are ready; record their actual execution in the implementation report.
- [x] Compare the implemented conference workload with an isolated checkout of
7614703using the same toolchain/native libraries. The checkout was created after implementation, so it provides a retrospective comparison. Record host details and raw timings; allocation and scheduling measurements remain a separate production-sizing exercise. - [ ] Exercise combined faults: slow recording disk plus conference media; slow HTTP clients plus URL downloads; full packet/command queues during transfer/attach; repeated playback create/remove plus shared decoding; and maximum WS bursts during bidirectional audio. Verify the documented task, byte, disk, and worker ceilings, not only successful API responses.
- [ ] Run the existing opt-in soak harness on a dedicated environment, extending coverage for SRTP rekey, IPv6, recording, variable packetization, and conference workloads. Include a sustained resource-churn run long enough to observe multiple cache cleanup/orphan cycles. After teardown, live sessions, ports, owners, tasks, and files return to their expected baseline.
- [ ] Define the reference-host latency/CPU acceptance envelope from the baseline before performance changes land. Require all targeted resource/correctness invariants and no unexplained regression outside repeated-run variability for established passthrough/transcode workloads. Report maximum supported load at the chosen scheduling SLO rather than promising unmeasured calls-per-core capacity.
- [x] Retain existing cargo-audit, cargo-deny, SBOM, and container scanning. Recheck the locked dependency set at release time and record the actual deployed native libopus/OpenSSL and OS packages. Scan the exact candidate image before promotion; current scanning after a push alone is not a pre-release gate. Fix applicable findings or explicitly document any unresolved release blocker.
- [x] Rewrite performance/deployment guidance using measured results. Correct same-codec conference costs, supported WebRTC codec advice, 20 ms generated output versus variable input packetization, dual-stack socket counts, recording queue/flush behavior, and unsupported per-session/per-recording memory and capacity claims. Show sizing from configured queues, active workers, cache bytes, and observed codec costs.
- [x] Update protocol/configuration docs alongside their owning steps: negotiation errors, source learning and RTCP, URL allowlists, resource-limit failures, WS keepalive/burst policy, recording snapshots, and lifecycle timeout outcomes. Describe the shared HMAC trust boundary without adding a speculative tenant database/API.
- [ ] Stage deployment on a canary with explicit source/URL policy, TLS/HMAC or documented proxy mode, and finite budgets. Watch rejection rates, queue drops, scheduling delay, cache/worker counts, and recording failures. Drain sessions before changing binaries; in-memory state is not migrated across restart. Keep a previous binary/config for operational recovery, with exposure restricted if recovery would reintroduce a reviewed security defect.
All 16 numbered runtime findings have implementation changes and targeted regression coverage. Expanded combined-fault/load experiments and deployment remain open release work; the standard soak does not establish the proposed extended soak matrix. Keep unperformed validation items open. The dependency advisory, optional deeper mix-minus optimization, and deployment assumptions must remain distinguishable from confirmed runtime defects when reporting progress.