Notes. Traced in a patched go2rtc fork feeding a video pipeline, against a source fed by an external exec/pipe process that could hang on respawn. The mutex cause is measured (goroutine dump + regression test); the application-side trigger sequence is a plausible mechanism, labeled below. No warranty; corrections welcome.

At 10,000 feet

Setting: go2rtc (a popular streaming relay) wraps each source in a “producer” that consumers attach to and a reconnect path redials on drop; some sources are external processes read over a pipe (exec/pipe producers).

Problem: the reconnect path holds the producer’s mutex across the redial. If the redial blocks — a respawned process that hangs before its first byte, with no timeout — the stall watchdog, every consumer’s attach, and stop() all block behind it. The producer is stranded permanently.

What this note establishes (measured): one hung redial under the held mutex froze the whole producer: 0 frames for ~194 s, no recovery pre-fix; terminal states lasted many hours in the field. A patched build (redial off the mutex, bounded pipe probe, process-group kills) self-recovered in ~6 s.

Takeaway: never hold a producer-wide lock across an unbounded dial. Dial outside the lock, time out the first-byte probe (pipe sources included), kill the process group after a grace period, and force-reset the reconnect loop after N consecutive instant failures.

The finding: the mutex is held across a redial that can never return

reconnect() acquires the producer mutex p.mu, then (re)creates the producer — for an exec/pipe source, spawning the child and blocking on its output. A child that hangs before its first byte holds the mutex forever, freezing the stall watchdog (why logs go silent mid-wedge), every consumer Dial, and stop(); a goroutine dump showed the reconnect goroutine parked in the blocking open with others queued behind it.

Application-side trigger (plausible mechanism, not fully proven): the last consumer drops during a stall, resetting the producer; a fresh consumer then dials while reconnect is mid-redial against a dead source. Reconstructed from code and the dump, not captured live.

The measured wedge, by the numbers

On a repro harness, frames emitted after the source stops responding:

build first-byte probe bound frames recovery
pre-fix 30 s default (but redial under mutex) 0 for ~194 s never (terminal)
partial fix (redial off mutex) 30 s default 0 for ~136 s ~29 s
partial fix, probe disabled none (starttimeout=3600) 0 for the whole run never (terminal)
full fix producer-level backstop 0 until first abandon ~63 s, bounded abandons

Replayed in throwaway containers, the old binary wedged terminally (no logs, hung process tree, consumers dead even after the source recovered); the patched binary recovered in ~6 s.

The fixes

  1. Redial outside the lockreconnect() releases p.mu first; a stale dial landing late is stopped and discarded. The killer fix.
  2. Bound the pipe first-byte probe — the pipe/exec flow had no timeout (only RTSP did); it now shares the 30 s start-timeout.
  3. Kill the process group after a grace period — ~5 s for clean exit, then SIGKILL pid + group, sweeping grandchildren.
  4. Force-reset after 3 strikes — three consecutive instant stale-state failures reset the producer to idle; any success or different error resets the counter.

A remaining strand on the fresh-consumer Dial() route is closed by a later, separate patch (single-flight dial, epoch counter, producer-level dial backstop).

Honesty and prior art

  • Measured: the wedge, the goroutine dump, the ~194 s pre-fix run, the ~6 s patched recovery; a mutex-deadlock regression test passes under the race detector.
  • Plausible mechanism (not fully proven): the exact consumer/reconnect interleaving reaching the terminal state.
  • Prior art (partial novelty): upstream go2rtc issues cover the exec/reconnect hang symptom class — credit there. No upstream report identifies the root cause: the mutex held across an unbounded redial.

Licensed under the site footer’s CC BY 4.0. If this saved you a debugging session, the optional thanks link in the footer is appreciated, no obligation.

Validation (2026-08-31)

Version: go2rtc HEAD c245815 (2026-07-13), app.Version = "1.9.14", 1 commit past the v1.9.14 tag (github.com/AlexxIT/go2rtc).

Source-confirmed: YES (not fixed upstream). Both Dial() (producer.go:58/62) and reconnect() (producer.go:177/187) call GetProducer(p.url) while holding p.mu. For an exec: pipe source the chain ends in exec.handlePipe (exec/exec.go:155)magic.OpenPeek(4), blocking on the pipe read with no timeout (only handleRTSP applies the 30 s starttimeout). stop() (producer.go:246) also needs the mutex. Symptom class matches upstream issues #1204/#1243/#1084.

Runtime repro: REPRODUCED. HEAD built in Docker (golang:1.24-alpine); config test: exec:sleep 3600 (opens stdout, never writes). All concurrent /api/stream.mp4?src=test consumers hung; a SIGQUIT goroutine dump showed the lock holder parked in Dial → GetProducer → exec.handlePipe → magic.Open → Peek/Read, with waiters queued in sync.Mutex.Lock at producer.go:58.

Verdict: validated — source-confirmed and runtime-reproduced on go2rtc 1.9.14 (HEAD, current); unfixed upstream; the four fixes address the reproduced failure.