Skip to content
JEEL.GAJERA/ CASE STUDY · AUDIO-RELAY

// case study 01 · real-time audio

audio-relay

Low-latency desktop-to-Android audio relay, played back through whatever Bluetooth device your phone already has connected. No admin rights, no virtual audio driver, no cloud service.

TYPE
REAL-TIME AUDIO
STATUS
EARLY ACTIVE DEVELOPMENT
READING
8 MIN READ
SURFACES
V0.1.0 ON GITHUB

The problem

My laptop's audio is stuck on my laptop. My earbuds are paired to my phone. Re-pairing them every time I want to watch something on the bigger screen is a thirty-second ritual that fails about a third of the time, and it means only one device can play at once.

The usual answers all ask for something I did not want to give. Virtual audio drivers need admin rights and a kernel-mode install. Cloud relays send my desktop audio to somebody else's server to get it back to a device two feet away. Bluetooth-from-the-laptop means unpairing from the phone, which is the problem I started with.

There is a version of this that asks for none of that: capture what the laptop is already playing, send it over the local network, and let the phone play it as ordinary media. Android's normal audio routing then does the last hop, sending it to whatever the phone is already connected to — exactly the way it would for Spotify. The phone stays the Bluetooth endpoint. Nothing in the path needs privileges.

A side effect turned out to be the feature I use most: because loopback capture does not interrupt local playback, the laptop keeps playing while it relays. Headphones on the laptop and earbuds on the phone hear the same audio, with nothing to set up.

The approach

Two apps, both required. A Rust desktop app captures and sends; a Kotlin Android app receives and plays.

code
Laptop (Windows/Linux)                    Phone (Android)
┌──────────────────────────┐               ┌────────────────────────────┐
│ Loopback capture         │   UDP (PCM)   │ UDP receiver → jitter buf  │
│  → framer/sequencer      │──────────────▶│  → AudioTrack (USAGE_MEDIA)│
│ TCP control (pairing,    │◀─────────────▶│  → routed to your BT       │
│  heartbeat, reconnect)   │   TCP + mDNS  │    device by Android       │
└──────────────────────────┘               └────────────────────────────┘

Every hard constraint maps onto an API that already exists for this purpose:

ConstraintSatisfied by
No admin installPortable binary — a .exe, or a .tar.gz/.deb
No virtual audio driverWASAPI loopback capture (user-mode, in Windows since Vista)
Same trick on LinuxPulseAudio monitor-source capture, which covers PipeWire through pipewire-pulse
Phone stays the BT endpointAudioTrack with USAGE_MEDIA; Android routes it like any other media
Works on a phone hotspotPlain LAN sockets plus mDNS — no cloud, no relay server

Audio goes over UDP because it is loss-tolerant and latency-sensitive. Pairing, heartbeat, and reconnection go over TCP because they are none of those things. Pairing is a six-digit code shown on the laptop, which seeds an HKDF-derived key that encrypts the audio payloads with ChaCha20-Poly1305.

Raw PCM, not Opus. On a LAN, bandwidth is not the constraint — 48kHz stereo is about 1.5 Mbps, trivial for any Wi-Fi — and a codec would buy that non-problem at the cost of 5–20ms of encode and decode. The packet header reserves a codec_id byte anyway, so adding Opus later is a version bump rather than a redesign.

Capture has to be paced, not just correct

The worst defect this project has had was not a wrong calculation. It was audio arriving in bursts.

Both capture APIs will hand you audio in large infrequent blocks, and both do it by default. On Linux, pa_simple_new lets the server pick the fragment size when you pass no BufferAttr, which the documentation describes as defaulting to "something like 2s". Measured against a real PipeWire server, that meant 194 of every 200 reads returned instantly and then the stream stalled for 341 milliseconds. Passing an explicit fragment size took the median gap to 10.65ms, worst case 11.11ms.

What makes this worse than it sounds is that a burst is unhideable latency. A 341ms gap in delivery needs a jitter buffer deeper than 341ms to conceal, which is far more than this app targets — so the receiver underran on every burst and played concealment silence instead. That is what "it cuts out constantly" actually was. Bursts also dump tens of packets into the network at once, which a phone hotspot answers by dropping them, turning a pacing bug into a loss bug as well.

The lesson generalised: pacing is a thing to measure, not to assume. There is a capture_delivery_cadence probe in the repository that measures it, and any change to a capture backend has to keep it honest.

One packet, one MTU

Ten milliseconds of 48kHz stereo 16-bit audio is 1920 bytes. A 1500-byte MTU allows about 1472 bytes of UDP payload. So every single audio packet was being IP-fragmented into two.

IP fragments are all-or-nothing: lose either half and the whole packet is gone. Fragmenting therefore roughly doubles the effective loss rate, on precisely the marginal Wi-Fi and hotspot links where loss is already the limiting factor. It showed up as constant brief dropouts.

The sender now splits a captured chunk across as many packets as it takes to keep each one under 1200 bytes — 1200 rather than 1472, leaving headroom for IPv6's larger header and any tunnel in the path, the same conservative budget QUIC and WebRTC use. Because each split packet carries its own sequence number and timestamp, the receiver cannot tell a split chunk from a natively small one, so this needed no protocol change at all.

It did impose one rule on the receiver: never assume a packet size. The jitter buffer learns the real one at runtime, because concealing a lost packet with the wrong amount of silence injects drift that the correction loop then has to fight.

Backlog is latency, and only skipping sheds it

A jitter buffer sitting persistently above its target depth is not holding jitter tolerance. It is holding delay that the listener pays on every packet for the rest of the session, and nothing in the steady-state design gives it back.

The clock-drift loop moves one PCM frame per ten packets. That is correct for cancelling tens of parts-per-million of crystal drift between two devices, and useless here: shedding 200ms of backlog one frame at a time takes about ten minutes. So a single transient stall — a descheduled receive loop, a burst of Wi-Fi retransmits — used to leave playback running seconds behind live, permanently.

So the buffer skips forward when depth exceeds target by a wide margin, discarding the backlog in one step. That costs a brief discontinuity and buys back correct latency, which for live audio is overwhelmingly the right trade. The two mechanisms stay deliberately separate: frame-level correction for continuous drift, skip-ahead for accumulated backlog.

Two related rules fell out of the same thinking. Queues between real-time stages are bounded and drop when full, because an unbounded queue between a fast producer and a slow consumer does not buffer, it accumulates. And the buffer has to be able to resynchronise: if the play position ever runs past the sender, every subsequent packet looks late, so a buffer that only discards late packets goes silent forever while the sender transmits normally. Sustained starvation is now treated as a signal to re-prebuffer from wherever the sender actually is.

Pairing state is two one-sided facts

"Remembering each other" is not shared state. It is each side's own store, and treating it as shared produced a class of bug that needed an app restart to escape.

Forgetting a laptop on the phone only touches the phone's storage. The laptop's HELLO_ACK.paired flag is therefore a hint, never a guarantee the phone can actually re-pair. The Android service now treats "the laptop says we're paired" and "we hold a usable local key" as two independent facts, and falls back to the pairing-code flow whenever the local key is missing, whatever the laptop claims.

Symmetrically, forgetting has to end the live session, not just erase the key. Erasing only the key left audio flowing on a credential that had just been revoked, with the laptop still in Streaming — so it never showed a pairing code, and neither side could recover. And the laptop now shows a pairing code whenever nothing is actively streaming, not only before the first pairing, because a phone can need a fresh code at any point.

One more, in the same family: an explicit user action must preempt an automatic one. connectTo used to decline to start while an attempt was in flight, so tapping Connect during an automatic retry did nothing at all — which a user experiences as the app looping on its own instead of asking for a code.

The latency budget, honestly

StageTypical
Loopback capture buffer3–10ms
Packetization5–10ms
Network (router Wi-Fi)1–5ms
Network (phone hotspot)5–15ms
Jitter buffer20–40ms
AudioTrack low-latency buffer10–20ms
Pipeline subtotal~45–90ms
Bluetooth A2DP (SBC/AAC)100–200ms
Realistic end-to-end~150–290ms

The part worth stating plainly: Bluetooth A2DP contributes more latency than everything this project does, combined, and no software on either end can remove it. That is true of every audio relay tool that exists. The honest claim is not "low latency" in the abstract — it is that the pipeline adds under 100ms on top of a floor somebody else set. If your earbuds negotiate aptX-LL or LE Audio the floor drops to 40–80ms automatically, and this project gets no credit for that either.

That makes it a watch-videos, listen-to-music, sit-in-meetings tool. It is not for competitive gaming, and saying so up front is cheaper than disappointing someone who expected otherwise.

What is actually verified

v0.1.0 is the first end-to-end feature set, and the status line on it is "working, tested scaffold" rather than "polished release" — deliberately.

Verified without hardware: the desktop test suite passes under cargo test/clippy/fmt, and every platform-independent Android module (AudioPacket, ControlMessage, Crypto, JitterBuffer) is unit-tested on a plain JVM. That includes a cross-implementation known-answer test where Android's Cipher.getInstance("ChaCha20-Poly1305") decrypts a ciphertext produced by the Rust side and recovers the exact plaintext — which is the only way to know two independent crypto implementations actually agree.

The Linux mute-control path is exercised against a live pipewire-pulse server, asserting the real effect through wpctl rather than just that the call returned Ok.

What I learned

Real-time audio punishes averages. Every bug in this project that mattered was a tail — the 341ms stall in an otherwise 10ms cadence, the one fragmented packet in a pair, the single stall that permanently displaced the play position. The mean throughput was fine throughout. Nothing you learn from an average tells you the stream is unlistenable.

The other thing: latency is a budget with an owner for every line. Once the Bluetooth line item was written down at 100–200ms, most of the tempting micro-optimisations upstream stopped being worth doing, and the two that genuinely mattered — pacing and backlog-shedding — became obvious. Writing the budget down honestly was worth more than any individual fix in it.