Reverb ASR+Diarization: Open-Source Speech Recognition for Long-Form Audio
WHY IT MATTERS
A new open-source ASR and diarization project called Reverb targets long-form audio transcription with speaker separation. It was submitted to Show HN.
What Happened
A new open-source project named Reverb was submitted to Hacker News under Show HN, offering combined automatic speech recognition (ASR) and speaker diarization aimed at long-form audio. The project targets meeting recordings, podcasts, and call analytics pipelines where transcription and speaker separation must happen in a single pass over extended inputs. Reverb is distributed as open-source code and positioned against existing tooling that treats ASR and diarization as separate, loosely coupled stages.
Why It Matters
Long-form, diarized transcription has remained a weak point in open tooling. Whisper-class models handle transcription well but do not natively attribute speech to speakers; standalone diarization libraries such as pyannote require orchestration, alignment logic, and careful handling of chunk boundaries where speaker identity shifts mid-segment. The result is that teams building meeting or call analytics either assemble fragile internal pipelines or pay per-minute commercial APIs. An open, integrated ASR+diarization stack lowers the integration burden and removes a recurring cost line for high-volume audio processing. It also makes on-premise and air-gapped deployment viable for regulated workloads where audio cannot leave controlled infrastructure.
Technical Details
Reverb couples an ASR component with speaker diarization in a single pipeline designed for long-form inputs rather than short clips. The typical integration path mirrors existing tooling: audio ingestion, resampling, inference, and post-processing to emit timestamped, speaker-labeled segments. Realistic deployment expectations should account for the standard constraints of open ASR: GPU memory scales with model size and batch length, and diarization accuracy degrades in overlapping speech, heavy crosstalk, and low-SNR conditions. Word error rate and diarization error rate are the two metrics that determine downstream usability, and both should be benchmarked against the specific domain — call center audio differs substantially from studio podcasts. Teams should verify segmentation stability across long recordings, since speaker-turn boundaries are the most common failure point in multi-hour inputs.
Operational Impact
For builders, the practical change is consolidation: a single pipeline replaces the glue code that currently stitches Whisper-family transcription to a separate diarization model, reducing the surface area for bugs at segment boundaries. Cost structure shifts from per-minute API billing to fixed GPU inference, which changes unit economics for high-volume batch processing — archival transcription, compliance review, and bulk call analysis become throughput-bound rather than budget-bound. Latency-sensitive applications such as live meeting assistants still favor hosted or streaming solutions, but offline and batch workflows gain a credible self-hosted option. Operators inheriting existing pipelines should expect moderate migration effort to revalidate speaker attribution accuracy on their own audio, since diarization performance is domain-sensitive and does not transfer cleanly from benchmark corpora.
SOURCE
HackerNews
SHARE
MORE FROM STUFFINSIDER
VoiceStudio: Open-Source Local ElevenLabs Alternative With 646 Languages
Oct 10OPEN SOURCEGoogle AI Edge Releases ml-drift for GPU-Accelerated ML Inference
Oct 9OPEN SOURCEopenGym Self-Hosted Workout Tracker Gains 1,494 GitHub Stars
Oct 7OPEN SOURCEDeepGEMM: DeepSeek's Efficient GPU BLAS Kernel Library Gains 363 Stars
Oct 6