Skip to content
ALL WRITING

// build log

The audio was on time. It just arrived all at once.

Shipping audio-relay meant learning that a real-time pipeline is judged by its tail, not its average. The throughput was correct the whole time it was unlistenable — here is the 341ms stall that caused it, and the two other bugs that were the same bug wearing different clothes.

min read1,290 words

I shipped audio-relay v0.1.0 this week. It captures whatever my laptop is playing and streams it to my phone over the local network, so the earbuds already paired to my phone can play my laptop's audio without re-pairing anything.

For a while it worked perfectly, in the sense that every number I could measure was correct, and it was completely unlistenable.

The measurement that lied

The pipeline is not complicated. Capture 10ms of audio, stamp it with a sequence number, send it over UDP, buffer it briefly on the phone, play it.

I measured throughput. 48kHz stereo, 16-bit — about 1.5 Mbps, arriving intact. I measured loss: negligible on my router. I measured the jitter buffer's average depth: sitting right around its target. Every average was where it should be.

And the audio cut out roughly twice a second.

The thing I had not measured was when the audio arrived. Here is the cadence of 200 consecutive capture reads on Linux:

code
194 reads returned in  < 1ms
  6 reads returned after ~341ms

The mean gap across those 200 reads is about 10ms, which is exactly right. The distribution is a disaster.

Where the 341ms came from

The Linux capture backend records from a PulseAudio monitor source — the same "record what you hear" trick WASAPI loopback does on Windows. You open it with pa_simple_new, which takes a BufferAttr describing how you want audio delivered.

I passed None. The documentation for what happens then is short, and I had read it, and I had not understood what it meant:

the server will pick something like 2s

I had read that as a buffer capacity — a ceiling, the most it would ever hold. It is a fragment size: the quantum in which the server hands you audio. Passing None did not mean "give me audio promptly, with room for up to 2 seconds." It meant "wake me up when you have collected a large block."

So the server collected. My reads drained an already-full buffer instantly, 194 times, and then blocked for a third of a second while it filled again.

The fix is one argument — an explicit fragment size — and it took the median gap to 10.65ms with a worst case of 11.11ms. Finding it took two days, almost all of it spent looking at the receiver, because the receiver was where the symptom was.

Why a burst is worse than latency

Here is the part I want to keep.

A jitter buffer's whole job is to absorb variation in arrival time. Mine targeted 20–40ms of depth, which is a normal figure — comfortably more than ordinary Wi-Fi jitter.

A 341ms gap in delivery cannot be absorbed by a 40ms buffer. It cannot be absorbed by any buffer this app would be willing to have, because a buffer deep enough to hide it is 341ms of latency, permanently, which is the thing the entire project exists to avoid.

So a burst is not just delay. It is unhideable delay. The receiver underran on every single burst and played its concealment silence, which is what "it cuts out constantly" actually was — not lost audio, not late audio, just audio that arrived in the wrong shape.

It got worse on a phone hotspot. A burst does not only stall the receiver, it dumps tens of UDP packets into the network at once, and a hotspot answers a sudden burst by dropping most of it. A pacing bug had quietly become a loss bug too, which is why the loss numbers I trusted on my router meant nothing about the case I actually cared about.

Two more bugs that were the same bug

Once I had the shape of it, two other defects turned out to be the same mistake in different clothing: something in the path was accumulating, and nothing was giving it back.

Every packet was being fragmented. 10ms of 48kHz stereo 16-bit audio is 1920 bytes. A 1500-byte MTU carries about 1472 bytes of UDP payload. So every audio packet — not some, every one — was split into two IP fragments.

IP fragments are all-or-nothing. Lose either half and the whole packet is gone. Fragmenting roughly doubles the effective loss rate, and it does so on exactly the marginal links where loss is already what limits you. The sender now splits chunks itself, keeping each packet under 1200 bytes. (1200, not 1472, to leave headroom for IPv6 headers and any tunnel in the path — the same budget QUIC and WebRTC use.)

Backlog never drained. If the buffer's depth ever rose above target and stayed there, that was not jitter tolerance. That was delay the listener paid on every packet for the rest of the session.

I had a clock-drift correction loop that moves one PCM frame per ten packets. That is the right rate for cancelling crystal drift between two devices — tens of parts per million, accumulating over hours. It is useless for shedding 200ms of backlog, which at that rate takes about ten minutes. So one transient stall left playback running seconds behind live, forever.

The buffer now skips forward when depth exceeds target by a wide margin, dropping the backlog in a single step. It costs an audible discontinuity, and for live audio that is obviously the right trade — you would rather hear one click than be permanently two seconds late. The two mechanisms stay separate on purpose: frame-level correction for continuous drift, skip-ahead for accumulated backlog. Trying to make one mechanism do both is how you get something that does neither.

The rule I actually took away

Real-time systems are judged by their tail. Averages actively hide the failures that matter.

Every bug above had a healthy mean. Mean throughput: correct. Mean gap between reads: 10ms, correct. Mean buffer depth: on target. Mean loss rate: fine. The distribution behind each of those numbers was the entire problem, and no amount of staring at the averages was going to show me that.

There is a version of this I already knew — p99 latency is the number that matters, everyone says so — but I knew it as advice about serving HTTP, where a slow tail costs some users some patience. In an audio pipeline the tail is not a degraded experience for a few requests. It is the only thing the listener hears, because the ear notices the one gap and not the thousands of packets that were on time.

The corollary is that pacing is something you measure rather than assume. There is a probe in the repository now, capture_delivery_cadence, whose whole job is to report the distribution of gaps between capture reads. It exists because I want the next person who touches a capture backend — very possibly me — to find out from a test rather than from their earbuds.

What I'm still not claiming

The honest latency budget lives in the case study, and the top line is this: Bluetooth A2DP adds 100–200ms all by itself, which is more than everything my code does combined, and no software on either end can remove it.

So the claim is not "low latency." It is "adds under 100ms on top of a floor somebody else set." Realistic end to end is 150–290ms — fine for video, music, and meetings; not for competitive gaming.

And the parts that need physical hardware to verify — WASAPI capture on real Windows, AudioTrack routing to real A2DP earbuds, mDNS over an actual Android hotspot — are written against documented APIs and have never run on a device. They are open checkboxes in the roadmap rather than quiet assumptions, because I have now had a very concrete lesson in what an assumption that looks verified costs.