- expand auth.token entry with canPublishData effect on datachannel
- note in compatibility matrix and transport table
- add moderator-grant screenshots and ru sync
- record future auto-grant idea as a docs note
allow supplying a wb account access token in config so the session
joins as that account instead of an anonymous guest. when auth.token
is empty the provider falls back to the guest-register flow.
threads auth.token through config -> session -> server/client ->
transport -> engine builtin -> auth provider.
- add package and exported-method doc comments
- use errors.New for sentinel, drop unused nolint directives
- wrap AcceptTCP error
- fix errcheck/noctx in tests
- extract newSettingEngine to keep negotiatePC under complexity limit
vp8channel runs kcp with block=nil (no fec, no checksum), so the
fake-vp8 carrier has no integrity protection. an sfu that perturbs a
payload byte lets a corrupt-but-parseable kcp segment ride through as
valid in-order data and break the muxconn aead above it, surfacing as
chacha20poly1305: message authentication failed under sustained
transfers (issue #109).
append a crc32 trailer to every kcp packet at the kcpconn seam and
verify+strip it on deliver; a mismatch drops the packet so kcp
retransmits via sack, restoring udp-equivalent semantics below kcp.
the pacer (52aea2d) and its per-tick byte budget capped the wire well
below the sfu's real ceiling: telemost ran ~10mbit before it landed and
2-3mbit after. it was added to fix the #95 collapse, but that collapse
was keyframe starvation (sfu dropped the track with no decodable
keyframe after ~40s), fixed separately by forceKeepalive in b20554b. the
pacer was treating a symptom that no longer exists.
drop perTickBytes and the vp8.max_bytes_per_sec knob, and restore the
kcp send window (512->4096) and tick (20ms->5ms) that were shrunk
alongside the pacer. control liveness no longer depends on a small data
window: ping/pong runs on its own isolated kcp session drained with
priority by writerLoop, so the original starvation rationale is gone.
batchSample still bounds one sample to defaultMaxPayloadSize.
refs #107#95
the aimd pacer read kcp's GetSRTT as its congestion signal, but the data
kcp runs with nc=1 and a fixed 512-segment (~716kb) window it keeps full
on purpose (kcp.go: data queues ~2-3s). that self-inflicted bufferbloat
dominates srtt no matter where the rate sits relative to the sfu policer
knee, so every tick tripped the high-delay mark and the rate decayed to
the 120kb/s floor. the result was slower, not faster.
the srtt signal is unusable as a knee detector on this carrier, so drop
adaptive control and keep the simple static pacer. raise the default to
the 1mb/s target from #107 (the old ~42s collapse at 1.2mib/s was the
keyframe-starvation drop fixed in #95, not a raw rate ceiling), still
tunable via vp8.max_bytes_per_sec.
refs #107
the hardcoded 400kb/s was a stability compromise and the previous knob
just moved the guesswork to operators: they had to hand-measure each
sfu's policer knee by ramping until stalls. replace both with a
delay-based aimd controller that finds the knee at runtime.
the sfu policer queues before it drops, so kcp's smoothed rtt inflates
ahead of the throughput collapse from #95/#107. the pacer folds one
srtt sample per writer tick: probe the rate up additively while delay
sits near the path baseline, back off multiplicatively once it
inflates. it settles just under the per-sfu knee on its own, both on
the client->server writerLoop and each server->client peer pump.
vp8.max_bytes_per_sec is now just the probe ceiling (default 1mb/s);
a fresh session still starts from the old 400kb/s operating point so a
healthy path only ramps up. no per-service hand-tuning left.
refs #107
the pacer was pinned at a hardcoded 400kb/s stability compromise, so
operators could not reach their sfu's real throughput ceiling without
recompiling. the options.maxbytespersec knob already existed but was
never wired from yaml. plumb it through config -> session -> transport,
keeping 400kb/s as the conservative default.
refs #107
detect a server restart when a frame from a new epoch arrives after the
latched peer has gone silent past a grace window, then drive the full
carrier rebuild (stream.Reconnect) instead of a bare re-handshake over
the stale carrier. the restarted server rejoins the sfu as a fresh
participant, so re-handshaking on the old media path only times out;
rebuilding the carrier re-establishes a path the new server answers on
and recovers in seconds instead of waiting out the ~70s liveness window
A restarted server rejoins the SFU with a new vp8channel epoch. The
client stayed latched to the old epoch and dropped every frame from the
new one, so recovery only happened once the relaxed control-liveness
window expired (~70s).
Watch for a frame from a different epoch after the latched peer has been
silent past a short grace window and fire the carrier reconnect callback,
driving a re-handshake against the new epoch in seconds. A live peer keeps
the latch fresh via its ~2s keepalives, so the watchdog never trips under
load or during SFU renegotiation.
- swap readme.md to english as default entry point
- move russian readme to readme.ru.md
- use RU / EN language switch buttons
- drop ffmpeg remark for videochannel
WaitForPeer makes the client block before it sends its SYN, so the
server cannot yet know the client epoch and its first welcome arrives as
a broadcast (receiverEpoch=0). The same happens on every reconnect:
peerEpoch is reset to 0 and the peer re-announces via a broadcast
Send(nil). The strict requireTargetedPeer guard dropped those broadcasts
while unlatched, so peerEpoch never latched and WaitForPeer hung for its
whole timeout - the link connected and pinged but no data ever flowed
("ping works, no connection"), most visibly when a peer left and
reconnected. Accept a broadcast while unlatched (peerEpoch==0) and keep
rejecting third-party broadcasts once latched.
peerWriterPump emitted every KCP frame the instant it was queued. With
KCP congestion control off (nc=1) and a BDP-sized window that overran the
SFU policer on the server->client direction. Pace it on the same frame
ticker and per-tick byte budget as writerLoop, sharing batchSample via
batchSampleFrom.
Server data-plane peer sessions used the conservative 30s smux keepalive
while the client used the relaxed window (SmuxConfigFor) for ControlPlane
transports like vp8channel. When the carrier went legitimately silent for
tens of seconds (publisher-PC reconnect / SFU renegotiation), the server
tore down its peer data session first at 30s, surfacing as 'closed pipe'
on the client and triggering a reconnect storm. Mirror the client and
apply SmuxConfigFor on the server data plane too.
pion TrackLocalStaticSample.WriteSample packetizes under its lock but
emits RTP after releasing it, so concurrent callers interleave sequence
numbers on the wire. The server runs a per-peer writer pump for bulk data
alongside writerLoop for control/keepalive, both hitting the shared track;
the remote VP8 reassembler enforces strict sequence contiguity and drops
the interleaved frames, stalling server->client traffic. Funnel every
WriteSample through one mutex so each frame's packetize+emit is atomic.
- bound WaitForPeer in bringUpLink/tryReopenSession to handshake timeout
so a missing peer fails fast instead of hanging before the SOCKS listener
- call waitForPeer on the reconnect path too: resetLinkPeer clears the peer
epoch, so without it the post-reconnect SYN re-races the server bridge
- restore strict requireTargetedPeer guard: the server always replies
targeted (it latches the client epoch from the SYN before smux replies),
so untargeted broadcasts before latch must be dropped
- remove per-frame debug logging on the bridge send/recv hot path
- extract buildSmuxClient / establishPeerSession helpers and wrap interface
errors to satisfy cyclop/nestif/wrapcheck
add a silent dead-link mode to the memory carrier (link stays up,
frames vanish) and a flag-gated TestDeadLinkDetectionWindow that
measures how long the datachannel client takes to notice a dead link
and reconnect. on the conservative liveness window it reconnects in
~58s; with the buggy 120s smux keepalive it fails to detect within 75s.
also drop dead restartControlKCP (both callers moved to
restartControlKCPWithHeader)