One C/C++ Library, Sixteen Speech-to-Text Model Families
You've probably been here: you find a speech-to-text model you want to use, and it comes with its own Python runtime, its own dependency tree, and its own opinions about how inference should work. Now multiply that by the number of model families you need to support. transcribe.cpp is a C/C++ inference library that tries to collapse that mess into a single runtime, running sixteen STT model families on top of ggml.
What It Does
transcribe.cpp is a C/C++ speech-to-text inference library. It runs GGUF-format models on the ggml runtime, which means you get Metal, Vulkan, and CUDA backends for GPU inference along with a tinyBLAS-accelerated CPU path. The library covers 16 model families and 60+ variants, supporting both streaming and batch transcription.
The model catalog is broad and spans very different architectures. There's Canary (including the 180m flash variant and the 1b v2), Canary-Qwen 2.5B, Cohere Transcribe, Fun-ASR-Nano, GigaAM-v3 in both CTC and RNN-T flavors, Granite Speech 4, 4.1, and 5.0 TurboCTC, MedASR, Moonshine, Moonshine Streaming, MOSS-Transcribe-Diarize, Multitalker Parakeet Streaming, two Nemotron streaming families, a large Parakeet lineup, Qwen3-ASR, SenseVoice Small, and Voxtral. Capabilities vary by family: some support translation, some produce token or word timestamps, some handle diarization, and several are built for streaming.
Why It's Cool
-
One runtime, many architectures. This is the whole point, and it's a genuinely hard thing to pull off. Parakeet, Moonshine, Voxtral, and Granite Speech don't share an inference graph. Wrapping them all behind the same ggml-based library means you don't rewrite your integration every time you swap models.
-
The backend story is practical, not aspirational. Metal, Vulkan, and CUDA cover most of the hardware people actually deploy on, and the tinyBLAS CPU path means you aren't dead in the water without a GPU. For a library like this, that's the difference between "interesting" and "usable."
-
Verification is stated up front. The README claims every transcription model published under
handy-computeron Hugging Face is numerically verified and WER-tested against its reference implementation. That's the kind of claim that matters when you're picking a model to ship. Speech-to-text has a nasty habit of producing output that looks plausible but is quietly wrong, and a numerical check against the reference is how you catch that. -
Streaming and batch, not one or the other. A lot of libraries pick a lane. Here, streaming shows up across Moonshine Streaming, Multitalker Parakeet Streaming, both Nemotron families, and parts of Parakeet. If you need real-time transcription, you're not stuck bolting a chunking scheme onto a batch-only model.
-
The capability matrix is honest. Rather than claiming every model does everything, the catalog lists capabilities per family. GigaAM-v3 gets token timestamps. Granite Speech 4/4.1 gets diarize, translate, and word timestamps. MOSS-Transcribe-Diarize gets diarization and segment timestamps. MedASR is domain-specific. You can see at a glance what you're actually getting.
-
Per-model documentation. Each family links to its own doc under
docs/models/. For a library with this many moving parts, that's the difference between a weekend of reverse-engineering and an afternoon of reading.
How to Try It
Start at the repository: github.com/handy-computer/transcribe.cpp.
- Clone the repo and check the build instructions. You'll want to confirm which backend you're targeting (Metal, Vulkan, CUDA, or the CPU path) before you build, since that affects your toolchain.
- Pick a model family from the catalog table and read its doc under
docs/models/. Capabilities differ enough between families that it's worth a few minutes to match the model to your use case. - Grab the corresponding GGUF model from the handy-computer Hugging Face org.
- Build the library and link it into your C or C++ project, or run whatever example binaries ship with the repo.
- Test against your own audio before committing to a model. The WER claims are against reference implementations, but your audio and your accents are the only benchmark that ultimately matters.
If you're evaluating multiple families, the shared runtime is the real payoff here. Swapping from Parakeet to Moonshine or Voxtral becomes a model-file change rather than an integration rewrite.
Final Thoughts
transcribe.cpp is aimed at people embedding speech-to-text into native applications, edge deployments, or anywhere a Python runtime is unwelcome. The breadth is the headline, but the verification work and the per-family docs are what make it something you'd actually ship on rather than just experiment with. If you've been juggling separate inference stacks for every STT model you support, this is worth a serious look.
Follow @githubprojects for more developer tools and open source projects.