I shipped audio-relay v0.1.0
this week. It captures whatever my laptop is playing and streams it to my
phone over the local network, so the earbuds already paired to my phone can
play my laptop's audio without re-pairing anything.
For a while it worked perfectly, in the sense that every number I could measure was correct, and it was completely unlistenable.
The measurement that lied
The pipeline is not complicated. Capture 10ms of audio, stamp it with a sequence number, send it over UDP, buffer it briefly on the phone, play it.
I measured throughput. 48kHz stereo, 16-bit — about 1.5 Mbps, arriving intact. I measured loss: negligible on my router. I measured the jitter buffer's average depth: sitting right around its target. Every average was where it should be.
And the audio cut out roughly twice a second.
The thing I had not measured was when the audio arrived. Here is the cadence of 200 consecutive capture reads on Linux:
194 reads returned in < 1ms
6 reads returned after ~341msThe mean gap across those 200 reads is about 10ms, which is exactly right. The distribution is a disaster.
Where the 341ms came from
The Linux capture backend records from a PulseAudio monitor source — the
same "record what you hear" trick WASAPI loopback does on Windows. You open
it with pa_simple_new, which takes a BufferAttr describing how you want
audio delivered.
I passed None. The documentation for what happens then is short, and I had
read it, and I had not understood what it meant:
the server will pick something like 2s
I had read that as a buffer capacity — a ceiling, the most it would ever
hold. It is a fragment size: the quantum in which the server hands you
audio. Passing None did not mean "give me audio promptly, with room for up
to 2 seconds." It meant "wake me up when you have collected a large block."
So the server collected. My reads drained an already-full buffer instantly, 194 times, and then blocked for a third of a second while it filled again.
The fix is one argument — an explicit fragment size — and it took the median gap to 10.65ms with a worst case of 11.11ms. Finding it took two days, almost all of it spent looking at the receiver, because the receiver was where the symptom was.
Why a burst is worse than latency
Here is the part I want to keep.
A jitter buffer's whole job is to absorb variation in arrival time. Mine targeted 20–40ms of depth, which is a normal figure — comfortably more than ordinary Wi-Fi jitter.
A 341ms gap in delivery cannot be absorbed by a 40ms buffer. It cannot be absorbed by any buffer this app would be willing to have, because a buffer deep enough to hide it is 341ms of latency, permanently, which is the thing the entire project exists to avoid.
So a burst is not just delay. It is unhideable delay. The receiver underran on every single burst and played its concealment silence, which is what "it cuts out constantly" actually was — not lost audio, not late audio, just audio that arrived in the wrong shape.
It got worse on a phone hotspot. A burst does not only stall the receiver, it dumps tens of UDP packets into the network at once, and a hotspot answers a sudden burst by dropping most of it. A pacing bug had quietly become a loss bug too, which is why the loss numbers I trusted on my router meant nothing about the case I actually cared about.
Two more bugs that were the same bug
Once I had the shape of it, two other defects turned out to be the same mistake in different clothing: something in the path was accumulating, and nothing was giving it back.
Every packet was being fragmented. 10ms of 48kHz stereo 16-bit audio is 1920 bytes. A 1500-byte MTU carries about 1472 bytes of UDP payload. So every audio packet — not some, every one — was split into two IP fragments.
IP fragments are all-or-nothing. Lose either half and the whole packet is gone. Fragmenting roughly doubles the effective loss rate, and it does so on exactly the marginal links where loss is already what limits you. The sender now splits chunks itself, keeping each packet under 1200 bytes. (1200, not 1472, to leave headroom for IPv6 headers and any tunnel in the path — the same budget QUIC and WebRTC use.)
Backlog never drained. If the buffer's depth ever rose above target and stayed there, that was not jitter tolerance. That was delay the listener paid on every packet for the rest of the session.
I had a clock-drift correction loop that moves one PCM frame per ten packets. That is the right rate for cancelling crystal drift between two devices — tens of parts per million, accumulating over hours. It is useless for shedding 200ms of backlog, which at that rate takes about ten minutes. So one transient stall left playback running seconds behind live, forever.
The buffer now skips forward when depth exceeds target by a wide margin, dropping the backlog in a single step. It costs an audible discontinuity, and for live audio that is obviously the right trade — you would rather hear one click than be permanently two seconds late. The two mechanisms stay separate on purpose: frame-level correction for continuous drift, skip-ahead for accumulated backlog. Trying to make one mechanism do both is how you get something that does neither.
The rule I actually took away
Real-time systems are judged by their tail. Averages actively hide the failures that matter.
Every bug above had a healthy mean. Mean throughput: correct. Mean gap between reads: 10ms, correct. Mean buffer depth: on target. Mean loss rate: fine. The distribution behind each of those numbers was the entire problem, and no amount of staring at the averages was going to show me that.
There is a version of this I already knew — p99 latency is the number that matters, everyone says so — but I knew it as advice about serving HTTP, where a slow tail costs some users some patience. In an audio pipeline the tail is not a degraded experience for a few requests. It is the only thing the listener hears, because the ear notices the one gap and not the thousands of packets that were on time.
The corollary is that pacing is something you measure rather than assume.
There is a probe in the repository now, capture_delivery_cadence, whose
whole job is to report the distribution of gaps between capture reads. It
exists because I want the next person who touches a capture backend — very
possibly me — to find out from a test rather than from their earbuds.
What I'm still not claiming
The honest latency budget lives in the case study, and the top line is this: Bluetooth A2DP adds 100–200ms all by itself, which is more than everything my code does combined, and no software on either end can remove it.
So the claim is not "low latency." It is "adds under 100ms on top of a floor somebody else set." Realistic end to end is 150–290ms — fine for video, music, and meetings; not for competitive gaming.
And the parts that need physical hardware to verify — WASAPI capture on real
Windows, AudioTrack routing to real A2DP earbuds, mDNS over an actual
Android hotspot — are written against documented APIs and have never run on a
device. They are open checkboxes in the roadmap rather than quiet
assumptions, because I have now had a very concrete lesson in what an
assumption that looks verified costs.