Parakeet TDT v3 + Streaming Sortformer β ONNX bundle for Vernacula
Combined ONNX shipping bundle used by Vernacula as its default ASR + diarization + VAD stack. Three upstream models are co-located here so a single download brings up the full pipeline.
- Conversion scripts:
scripts/nemo_export/ - Vernacula: github.com/christopherthompson81/vernacula
- Upstream models: Parakeet TDT v3, Streaming Sortformer, Silero VAD
Highlights
- DFT-basis mel frontend replaces
torch.stft. ORT's STFT op diverged from NeMo (cosine β 0.23 on first inspection); the replacement uses precomputed cos/sin basis matrices as Conv1D weights, with center-padded windows and standard ops only. Restored bit-for-bit parity to the NeMo reference. - Streaming Sortformer 6β3 ONNX contract. NeMo's
concat_and_pad()isn't ONNX-traceable; the custom exporter replaces dynamic per-batch slicing with fixed-shape ops at 992 chunk frames (124 subsampled at 8Γ downsampling). Inputs:chunk, chunk_lengths, spkcache, spkcache_lengths, fifo, fifo_lengths. Outputs:spkcache_fifo_chunk_preds, chunk_pre_encode_embs, chunk_pre_encode_lengths. - CoreML-compilable Sortformer variant for Apple Silicon. The stock diarization graph cannot be compiled by ONNX Runtime's CoreML EP at all (
Failed to create MLModel β¦ error code: -14), so the Neural Engine was unreachable on macOS.diar_streaming_sortformer_4spk-v2.1.coreml.onnxis a fully static re-export plus four value-preserving graph rewrites that compile as a single CoreML partition β 51.5 ms/chunk vs 171.8 ms on CPU (M5, ORT 1.29.0). It is a steady-state-only graph with a different input contract; read Sortformer CoreML variant before using it. - Dynamic-batch encoder + dynamic-batch joint decoder for Parakeet TDT (preprocessor still batch-1 post-export). INT8 variants of encoder, decoder-joint, and Sortformer ship for CPU-only inference.
- Chunk-by-chunk parity diagnostic (
compare_sortformer_chunk_outputs.py) compares NeMo vs ONNX state evolution across streaming chunks to localise drift to either model output or carry-state. - Preprocessor export sweep (
tune_nemo128_export.py) scores wrapper / custom / DFT modes against a legacy reference with feature-level and encoder-output deltas β the tooling that picked the DFT path in the first bullet.
Contents
| File | Source | Purpose |
|---|---|---|
encoder-model.onnx (+ .data) |
nvidia/parakeet-tdt-0.6b-v3 |
Parakeet TDT FastConformer encoder |
encoder-model.int8.onnx |
(quantized from above) | INT8-quantized encoder for CPU |
decoder_joint-model.onnx |
Parakeet TDT v3 | Joint decoder + prediction network |
decoder_joint-model.int8.onnx |
(quantized) | INT8-quantized decoder for CPU |
vocab.txt |
Parakeet TDT v3 | Subword vocabulary |
nemo128.onnx |
NeMo preprocessor | 80-mel log-FBANK frontend (128-dim hop config) |
diar_streaming_sortformer_4spk-v2.1.onnx |
nvidia/diar_streaming_sortformer_4spk-v2.1 |
Streaming 4-speaker diarization |
diar_streaming_sortformer_4spk-v2.1_int8.onnx |
(quantized) | INT8-quantized diarization |
diar_streaming_sortformer_4spk-v2.1.coreml.onnx |
Sortformer | Static steady-state diarization graph for the CoreML EP β different contract, see below |
sortformer/diar_streaming_sortformer_4spk-v2.1.onnx |
Sortformer | Same graph in subdir layout for legacy clients |
silero_vad.onnx |
snakers4/silero-vad | Voice activity detection |
config.json, manifest.json |
Vernacula | Runtime config + per-file MD5 hashes |
Sortformer CoreML variant
diar_streaming_sortformer_4spk-v2.1.coreml.onnx exists because ONNX Runtime's CoreML
execution provider is the only route to the Apple Neural Engine, and the stock graph
cannot be compiled by it at all β it slices by tensor values, which makes every
downstream shape data-dependent, and CoreML's MIL runtime rejects unbounded dimensions
outright. The variant fixes the shapes at export time and applies four rewrites
(constant-fold at ORT_ENABLE_BASIC, delete 51 all-False Where masks, rewrite 34
constant zero Pads as Concats, pre-transpose 180 Gemm weights) to reach a single
CoreML partition instead of 71.
Measured on an M5 (Mac17,3, macOS 26.6.2), ORT 1.29.0 β the version Vernacula builds against β chunk=992 / spkcache=188 / fifo=124:
| model | EP | inference | cold load | warm load |
|---|---|---|---|---|
| stock dynamic | CPU | 171.8 ms | 0.50 s | β |
| stock dynamic | CoreML | fails to compile | β | β |
| CoreML variant | CPU | 154.0 ms | β | β |
| CoreML variant | CoreML | 51.5 ms | 2.2 s | 0.16 s |
3.34Γ vs the stock graph on CPU. The CoreML EP takes the whole graph as a single
partition (1751 of 1751 nodes -- the file holds 1788, and ORT's BASIC fold trims it to
1751 at load), measured at ORT_ENABLE_BASIC (see note 3). Outputs match the stock graph to 4.470E-07 (preds,
rms 7.616E-08) running CoreML against stock-on-CPU, and to 3.576E-07 CPU-to-CPU. The
static shapes alone also make it ~10% faster on plain CPU, so it is not purely a macOS
artifact.
The stock graph's CoreML failure reproduces on 1.29.0 exactly as described above β
Input: _ConstantOfShape_1_output_0 has unbounded dimension which is not supported.
An earlier round on ORT 1.24.4 measured 52.3 ms for the variant on CoreML, alongside
196.3 ms (CPU) and 113.0 ms (WebGPU) for the stock graph and 101.7 ms for the variant on
WebGPU. Those CPU/WebGPU baselines disagree with other figures recorded for the same
machine and ORT (163.5 ms / 94.0 ms) and have not been reconciled; treat the 1.29.0 table
above as authoritative and the 1.24.4 baselines as indicative only. WebGPU was not
re-measured on 1.29.0 β the provider is absent from the onnxruntime PyPI wheel and ships
only in the packaged app.
This is not a drop-in replacement. Four things differ:
| stock | CoreML variant | |
|---|---|---|
| inputs | 6 (chunk, spkcache, fifo + 3 *_lengths) |
3 (chunk, spkcache, fifo) |
| shapes | dynamic | fixed [1,992,128] / [1,188,512] / [1,124,512] |
spkcache_fifo_chunk_preds |
[batch, time_out, 4] |
[1, 436, 4] |
| graph optimization level | any | ORT_ENABLE_BASIC or lower |
All three lengths are baked in, so this graph is steady-state only. Pass full-size buffers whose real content fills them:
spkcachegenuinely 188 frames,fifogenuinely 124,chunkgenuinely 992.β Warm-up is NOT safe to send here, contrary to what an earlier version of this note said. A streaming runtime builds the cache and FIFO up from empty, and the baked lengths claim all 188 / 124 frames are real, so zero-padding them makes the graph attend to padding as audio β the same failure as note 2, in a different input. Vernacula's own caller routes every chunk to the stock graph until both buffers are genuinely full.
Never feed it the final, short chunk of a recording. With
chunk_lengthsbaked to 992 the graph attends to that chunk's zero-padded tail as real audio, which moves its speaker probabilities by up to 0.54 (rms 0.24) on a 0..1 scale β enough to flip speaker assignments. Route that one chunk per recording to the stock dynamic graph, which is in this same repo.Create the session at
ORT_ENABLE_BASICor lower. AtORT_ENABLE_EXTENDEDand above, ORT fails to load the file outright βAddInitializedOrtValue Attempt to replace the existing tensor, fromMatMulAddFusionre-running over an already-optimized graph. Confirmed on both 1.26.0 and 1.29.0.β This is masked when the CoreML EP is registered. CoreML claims the whole graph before the CPU fusions run, so on macOS all four levels appear to load β but the moment the graph falls back to the CPU EP (CoreML absent from the build, a non-Apple machine, or
MLComputeUnitsdeclining it) the same session options throw. Pin BASIC regardless of platform; it costs nothing, since the file is already an optimized graph.level CPU EP CoreML EP registered ORT_DISABLE_ALLloads loads ORT_ENABLE_BASICloads loads ORT_ENABLE_EXTENDEDfails loads ORT_ENABLE_ALLfails loads A separate 1.24.4 measurement found EXTENDED also shattering CoreML partitioning (69 β 191) on the earlier four-input graph.
CoreML partitioning is ORT-version dependent β re-validate on any ORT upgrade. Validated on 1.24.4 and 1.29.0, which both reach a single partition. An earlier note here claimed 1.29.0 split the graph into 194 partitions and diverged at ~1e-2; that was measured against a different build of the graph and does not reproduce β 1.29.0 gives 1 partition and
4.470E-07. Set the CoreML EP'sModelCacheDirectoryor every load pays the 2.2 s compile instead of 0.16 s.
fp16 was measured faster still (22.3 ms) and is not shipped, though no longer for the
original reason. chunk_pre_encode_embs does diverge to 9.8e-02, and those embeddings feed
back into the speaker cache and FIFO β but the end-to-end check that was missing has since
been run, and fidelity DER against NeMo is 0.000% on real speech (3 frames in 3378 flip
their binarized speaker set, all absorbed by the median filter). What holds fp16 back now is
throughput: it is 1.66Γ faster on CUDA but 25% slower on CPU, since there are no native
fp16 CPU kernels, so it can only ever be an execution-provider-gated variant rather than a
replacement.
Reproduce it with:
python scripts/nemo_export/export_sortformer_nemo_to_onnx.py \
--nemo <path>/diar_streaming_sortformer_4spk-v2.1.nemo \
--output sortformer_coreml.onnx --opset 17 \
--coreml-static-batch1 --coreml-const-lengths --coreml-const-chunk-length \
--chunk-frames 992 --fixed-spkcache-frames 188 --fixed-fifo-frames 124 --overwrite
python scripts/nemo_export/coreml_optimize_sortformer.py \
--input sortformer_coreml.onnx \
--output diar_streaming_sortformer_4spk-v2.1.coreml.onnx \
--verify --reference diar_streaming_sortformer_4spk-v2.1.onnx
The techniques generalize; see
docs/coreml_onnx_playbook.md.
A note on the Parakeet encoder and CoreML
Static-shape CoreML re-exports of the Parakeet encoder were built, published here, and
then withdrawn. They worked β single CoreML partition, exact to the stock graph, and a real
1.56Γ on inference β but a bucket session costs ~2.9 s to open and a static-shape design
needs several of them, so the Neural Engine's win went straight back into loading. On a
10-minute recording the encoder took 19.5 s through the buckets against 10.8 s on the CPU
EP and 7.1 s on WebGPU, which runs the stock encoder-model.onnx unmodified.
They also cost ~4.4 GB of compiled CoreML cache each, and real speech segments fill a bucket only about half way.
So for ASR on Apple Silicon, use encoder-model.onnx with the WebGPU execution
provider: dynamic shapes, no re-export, no bucketing. The diarization variant above is
different β its shapes really are fixed, and it is a clear win on the Neural Engine.
Method and measurements:
docs/coreml_onnx_playbook.md
(Technique 6) and
the investigation log.
Export provenance
Exported via scripts/nemo_export/
in the Vernacula repo, which contains:
export_parakeet_nemo_to_onnx.pyβ Parakeet.nemoβ split ONNX with TDT decoder state wired explicitlyexport_sortformer_nemo_to_onnx.pyβ Streaming Sortformer.nemoβ six-input / three-output ONNX contractexport_silero_vad_to_onnx.pyβ Silero VAD β ONNXexport_parakeet_coreml_encoder.pyβ Parakeet encoder β static-shape CoreML buckets with the padding mask hoisted to an input (kept for reference; the artifacts it makes are not shipped, see above)
The Parakeet export traces the RNNT/TDT decoder loop into a separate joint graph so each step is a fixed-shape ORT call. Sortformer is exported as a streaming graph that takes incoming frames + carry state and returns diarization logits chunk-by-chunk.
License
This bundle aggregates three upstream models under three different licenses. Each component retains its upstream license; redistribution here does not change those terms.
| Component | Upstream license |
|---|---|
| Parakeet TDT v3 weights | CC-BY-4.0 |
| Streaming Sortformer weights | NVIDIA Open Model License |
| Silero VAD weights | MIT |
NeMo mel-frontend code (nemo128.onnx) |
Apache-2.0 |
If you redistribute this bundle, propagate all four licenses with it.
Using these files
The cleanest path is via Vernacula, which downloads, caches, and validates
this package against manifest.json automatically. Outside Vernacula, pull
the package with huggingface_hub and load each .onnx with onnxruntime
directly β input / output tensor contracts are documented in
scripts/nemo_export/README.md.
from huggingface_hub import snapshot_download
path = snapshot_download(repo_id="christopherthompson81/sortformer_parakeet_onnx")
Limitations
These graphs preserve the numerical behavior of the upstream PyTorch checkpoints distributed via NVIDIA NeMo. Accuracy, language coverage, and known failure modes inherit from the upstream model cards (Parakeet, Sortformer) β see those for the authoritative discussion. INT8 variants trade a small amount of WER for ~2Γ CPU throughput; use the float32 variants where accuracy is the priority.
Citation
For the underlying models, please cite the upstream authors. See:
Acknowledgments
- Original Parakeet TDT v3 and Streaming Sortformer: NVIDIA NeMo team
- Original Silero VAD: Silero Team (snakers4)
- ONNX repackaging: Chris Thompson for Vernacula
Issues with the ONNX export specifically: open an issue on the Vernacula repo. Issues with the underlying models: see the upstream model cards.
See also
- Vernacula on GitHub β the speech pipeline app this package is built for
- Conversion scripts (
scripts/nemo_export/) β the export pipelines that produced these files nvidia/parakeet-tdt-0.6b-v3β upstream Parakeet model cardnvidia/diar_streaming_sortformer_4spk-v2.1β upstream Sortformer model card- Silero VAD on GitHub β upstream VAD source
- NVIDIA NeMo on GitHub β toolkit used to train and export the NVIDIA models
- Other Vernacula model packages
- Downloads last month
- 269
Model tree for christopherthompson81/sortformer_parakeet_onnx
Base model
nvidia/diar_streaming_sortformer_4spk-v2.1