Parakeet TDT v3 + Streaming Sortformer β€” ONNX bundle for Vernacula

Combined ONNX shipping bundle used by Vernacula as its default ASR + diarization + VAD stack. Three upstream models are co-located here so a single download brings up the full pipeline.

Highlights

  • DFT-basis mel frontend replaces torch.stft. ORT's STFT op diverged from NeMo (cosine β‰ˆ 0.23 on first inspection); the replacement uses precomputed cos/sin basis matrices as Conv1D weights, with center-padded windows and standard ops only. Restored bit-for-bit parity to the NeMo reference.
  • Streaming Sortformer 6β†’3 ONNX contract. NeMo's concat_and_pad() isn't ONNX-traceable; the custom exporter replaces dynamic per-batch slicing with fixed-shape ops at 992 chunk frames (124 subsampled at 8Γ— downsampling). Inputs: chunk, chunk_lengths, spkcache, spkcache_lengths, fifo, fifo_lengths. Outputs: spkcache_fifo_chunk_preds, chunk_pre_encode_embs, chunk_pre_encode_lengths.
  • CoreML-compilable Sortformer variant for Apple Silicon. The stock diarization graph cannot be compiled by ONNX Runtime's CoreML EP at all (Failed to create MLModel … error code: -14), so the Neural Engine was unreachable on macOS. diar_streaming_sortformer_4spk-v2.1.coreml.onnx is a fully static re-export plus four value-preserving graph rewrites that compile as a single CoreML partition β€” 51.5 ms/chunk vs 171.8 ms on CPU (M5, ORT 1.29.0). It is a steady-state-only graph with a different input contract; read Sortformer CoreML variant before using it.
  • Dynamic-batch encoder + dynamic-batch joint decoder for Parakeet TDT (preprocessor still batch-1 post-export). INT8 variants of encoder, decoder-joint, and Sortformer ship for CPU-only inference.
  • Chunk-by-chunk parity diagnostic (compare_sortformer_chunk_outputs.py) compares NeMo vs ONNX state evolution across streaming chunks to localise drift to either model output or carry-state.
  • Preprocessor export sweep (tune_nemo128_export.py) scores wrapper / custom / DFT modes against a legacy reference with feature-level and encoder-output deltas β€” the tooling that picked the DFT path in the first bullet.

Contents

File Source Purpose
encoder-model.onnx (+ .data) nvidia/parakeet-tdt-0.6b-v3 Parakeet TDT FastConformer encoder
encoder-model.int8.onnx (quantized from above) INT8-quantized encoder for CPU
decoder_joint-model.onnx Parakeet TDT v3 Joint decoder + prediction network
decoder_joint-model.int8.onnx (quantized) INT8-quantized decoder for CPU
vocab.txt Parakeet TDT v3 Subword vocabulary
nemo128.onnx NeMo preprocessor 80-mel log-FBANK frontend (128-dim hop config)
diar_streaming_sortformer_4spk-v2.1.onnx nvidia/diar_streaming_sortformer_4spk-v2.1 Streaming 4-speaker diarization
diar_streaming_sortformer_4spk-v2.1_int8.onnx (quantized) INT8-quantized diarization
diar_streaming_sortformer_4spk-v2.1.coreml.onnx Sortformer Static steady-state diarization graph for the CoreML EP β€” different contract, see below
sortformer/diar_streaming_sortformer_4spk-v2.1.onnx Sortformer Same graph in subdir layout for legacy clients
silero_vad.onnx snakers4/silero-vad Voice activity detection
config.json, manifest.json Vernacula Runtime config + per-file MD5 hashes

Sortformer CoreML variant

diar_streaming_sortformer_4spk-v2.1.coreml.onnx exists because ONNX Runtime's CoreML execution provider is the only route to the Apple Neural Engine, and the stock graph cannot be compiled by it at all β€” it slices by tensor values, which makes every downstream shape data-dependent, and CoreML's MIL runtime rejects unbounded dimensions outright. The variant fixes the shapes at export time and applies four rewrites (constant-fold at ORT_ENABLE_BASIC, delete 51 all-False Where masks, rewrite 34 constant zero Pads as Concats, pre-transpose 180 Gemm weights) to reach a single CoreML partition instead of 71.

Measured on an M5 (Mac17,3, macOS 26.6.2), ORT 1.29.0 β€” the version Vernacula builds against β€” chunk=992 / spkcache=188 / fifo=124:

model EP inference cold load warm load
stock dynamic CPU 171.8 ms 0.50 s β€”
stock dynamic CoreML fails to compile β€” β€”
CoreML variant CPU 154.0 ms β€” β€”
CoreML variant CoreML 51.5 ms 2.2 s 0.16 s

3.34Γ— vs the stock graph on CPU. The CoreML EP takes the whole graph as a single partition (1751 of 1751 nodes -- the file holds 1788, and ORT's BASIC fold trims it to 1751 at load), measured at ORT_ENABLE_BASIC (see note 3). Outputs match the stock graph to 4.470E-07 (preds, rms 7.616E-08) running CoreML against stock-on-CPU, and to 3.576E-07 CPU-to-CPU. The static shapes alone also make it ~10% faster on plain CPU, so it is not purely a macOS artifact.

The stock graph's CoreML failure reproduces on 1.29.0 exactly as described above β€” Input: _ConstantOfShape_1_output_0 has unbounded dimension which is not supported.

An earlier round on ORT 1.24.4 measured 52.3 ms for the variant on CoreML, alongside 196.3 ms (CPU) and 113.0 ms (WebGPU) for the stock graph and 101.7 ms for the variant on WebGPU. Those CPU/WebGPU baselines disagree with other figures recorded for the same machine and ORT (163.5 ms / 94.0 ms) and have not been reconciled; treat the 1.29.0 table above as authoritative and the 1.24.4 baselines as indicative only. WebGPU was not re-measured on 1.29.0 β€” the provider is absent from the onnxruntime PyPI wheel and ships only in the packaged app.

This is not a drop-in replacement. Four things differ:

stock CoreML variant
inputs 6 (chunk, spkcache, fifo + 3 *_lengths) 3 (chunk, spkcache, fifo)
shapes dynamic fixed [1,992,128] / [1,188,512] / [1,124,512]
spkcache_fifo_chunk_preds [batch, time_out, 4] [1, 436, 4]
graph optimization level any ORT_ENABLE_BASIC or lower
  1. All three lengths are baked in, so this graph is steady-state only. Pass full-size buffers whose real content fills them: spkcache genuinely 188 frames, fifo genuinely 124, chunk genuinely 992.

    ⚠ Warm-up is NOT safe to send here, contrary to what an earlier version of this note said. A streaming runtime builds the cache and FIFO up from empty, and the baked lengths claim all 188 / 124 frames are real, so zero-padding them makes the graph attend to padding as audio β€” the same failure as note 2, in a different input. Vernacula's own caller routes every chunk to the stock graph until both buffers are genuinely full.

  2. Never feed it the final, short chunk of a recording. With chunk_lengths baked to 992 the graph attends to that chunk's zero-padded tail as real audio, which moves its speaker probabilities by up to 0.54 (rms 0.24) on a 0..1 scale β€” enough to flip speaker assignments. Route that one chunk per recording to the stock dynamic graph, which is in this same repo.

  3. Create the session at ORT_ENABLE_BASIC or lower. At ORT_ENABLE_EXTENDED and above, ORT fails to load the file outright β€” AddInitializedOrtValue Attempt to replace the existing tensor, from MatMulAddFusion re-running over an already-optimized graph. Confirmed on both 1.26.0 and 1.29.0.

    ⚠ This is masked when the CoreML EP is registered. CoreML claims the whole graph before the CPU fusions run, so on macOS all four levels appear to load β€” but the moment the graph falls back to the CPU EP (CoreML absent from the build, a non-Apple machine, or MLComputeUnits declining it) the same session options throw. Pin BASIC regardless of platform; it costs nothing, since the file is already an optimized graph.

    level CPU EP CoreML EP registered
    ORT_DISABLE_ALL loads loads
    ORT_ENABLE_BASIC loads loads
    ORT_ENABLE_EXTENDED fails loads
    ORT_ENABLE_ALL fails loads

    A separate 1.24.4 measurement found EXTENDED also shattering CoreML partitioning (69 β†’ 191) on the earlier four-input graph.

  4. CoreML partitioning is ORT-version dependent β€” re-validate on any ORT upgrade. Validated on 1.24.4 and 1.29.0, which both reach a single partition. An earlier note here claimed 1.29.0 split the graph into 194 partitions and diverged at ~1e-2; that was measured against a different build of the graph and does not reproduce β€” 1.29.0 gives 1 partition and 4.470E-07. Set the CoreML EP's ModelCacheDirectory or every load pays the 2.2 s compile instead of 0.16 s.

fp16 was measured faster still (22.3 ms) and is not shipped, though no longer for the original reason. chunk_pre_encode_embs does diverge to 9.8e-02, and those embeddings feed back into the speaker cache and FIFO β€” but the end-to-end check that was missing has since been run, and fidelity DER against NeMo is 0.000% on real speech (3 frames in 3378 flip their binarized speaker set, all absorbed by the median filter). What holds fp16 back now is throughput: it is 1.66Γ— faster on CUDA but 25% slower on CPU, since there are no native fp16 CPU kernels, so it can only ever be an execution-provider-gated variant rather than a replacement.

Reproduce it with:

python scripts/nemo_export/export_sortformer_nemo_to_onnx.py \
  --nemo <path>/diar_streaming_sortformer_4spk-v2.1.nemo \
  --output sortformer_coreml.onnx --opset 17 \
  --coreml-static-batch1 --coreml-const-lengths --coreml-const-chunk-length \
  --chunk-frames 992 --fixed-spkcache-frames 188 --fixed-fifo-frames 124 --overwrite

python scripts/nemo_export/coreml_optimize_sortformer.py \
  --input sortformer_coreml.onnx \
  --output diar_streaming_sortformer_4spk-v2.1.coreml.onnx \
  --verify --reference diar_streaming_sortformer_4spk-v2.1.onnx

The techniques generalize; see docs/coreml_onnx_playbook.md.

A note on the Parakeet encoder and CoreML

Static-shape CoreML re-exports of the Parakeet encoder were built, published here, and then withdrawn. They worked β€” single CoreML partition, exact to the stock graph, and a real 1.56Γ— on inference β€” but a bucket session costs ~2.9 s to open and a static-shape design needs several of them, so the Neural Engine's win went straight back into loading. On a 10-minute recording the encoder took 19.5 s through the buckets against 10.8 s on the CPU EP and 7.1 s on WebGPU, which runs the stock encoder-model.onnx unmodified.

They also cost ~4.4 GB of compiled CoreML cache each, and real speech segments fill a bucket only about half way.

So for ASR on Apple Silicon, use encoder-model.onnx with the WebGPU execution provider: dynamic shapes, no re-export, no bucketing. The diarization variant above is different β€” its shapes really are fixed, and it is a clear win on the Neural Engine.

Method and measurements: docs/coreml_onnx_playbook.md (Technique 6) and the investigation log.

Export provenance

Exported via scripts/nemo_export/ in the Vernacula repo, which contains:

  • export_parakeet_nemo_to_onnx.py β€” Parakeet .nemo β†’ split ONNX with TDT decoder state wired explicitly
  • export_sortformer_nemo_to_onnx.py β€” Streaming Sortformer .nemo β†’ six-input / three-output ONNX contract
  • export_silero_vad_to_onnx.py β€” Silero VAD β†’ ONNX
  • export_parakeet_coreml_encoder.py β€” Parakeet encoder β†’ static-shape CoreML buckets with the padding mask hoisted to an input (kept for reference; the artifacts it makes are not shipped, see above)

The Parakeet export traces the RNNT/TDT decoder loop into a separate joint graph so each step is a fixed-shape ORT call. Sortformer is exported as a streaming graph that takes incoming frames + carry state and returns diarization logits chunk-by-chunk.

License

This bundle aggregates three upstream models under three different licenses. Each component retains its upstream license; redistribution here does not change those terms.

Component Upstream license
Parakeet TDT v3 weights CC-BY-4.0
Streaming Sortformer weights NVIDIA Open Model License
Silero VAD weights MIT
NeMo mel-frontend code (nemo128.onnx) Apache-2.0

If you redistribute this bundle, propagate all four licenses with it.

Using these files

The cleanest path is via Vernacula, which downloads, caches, and validates this package against manifest.json automatically. Outside Vernacula, pull the package with huggingface_hub and load each .onnx with onnxruntime directly β€” input / output tensor contracts are documented in scripts/nemo_export/README.md.

from huggingface_hub import snapshot_download
path = snapshot_download(repo_id="christopherthompson81/sortformer_parakeet_onnx")

Limitations

These graphs preserve the numerical behavior of the upstream PyTorch checkpoints distributed via NVIDIA NeMo. Accuracy, language coverage, and known failure modes inherit from the upstream model cards (Parakeet, Sortformer) β€” see those for the authoritative discussion. INT8 variants trade a small amount of WER for ~2Γ— CPU throughput; use the float32 variants where accuracy is the priority.

Citation

For the underlying models, please cite the upstream authors. See:

Acknowledgments

  • Original Parakeet TDT v3 and Streaming Sortformer: NVIDIA NeMo team
  • Original Silero VAD: Silero Team (snakers4)
  • ONNX repackaging: Chris Thompson for Vernacula

Issues with the ONNX export specifically: open an issue on the Vernacula repo. Issues with the underlying models: see the upstream model cards.

See also

Downloads last month
269
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for christopherthompson81/sortformer_parakeet_onnx

Quantized
(14)
this model