Voicing-ASR
Overview
Introduction
Voicing-ASR is a family of automatic speech recognition (ASR) models designed for accurate, fast, and robust speech recognition across 52 languages and dialects. The models support both language identification and automatic speech recognition, leveraging large-scale speech training data and the advanced audio understanding capabilities of the Voicing-Omni foundation model.
This version delivers state-of-the-art performance among open-source ASR models and demonstrates competitive performance against leading proprietary commercial speech recognition APIs.
Key Features
🌍 Multilingual Support Supports language identification and speech recognition across 52 languages and dialects, making it suitable for multilingual and global applications.
⚡ High Quality and Low Latency Provides accurate and robust transcription across a wide range of acoustic conditions, including noisy environments, diverse speakers, and challenging speech patterns.
🎯 Strong Recognition Performance Achieves competitive results across both open-source and internal evaluation benchmarks, with model delivering state-of-the-art performance among open-source ASR systems.
🚀 Production-Ready Inference Toolkit Alongside the model architectures and weights, Voicing-ASR provides a comprehensive inference framework designed for real-world deployment.
The toolkit supports:
- vLLM-based batch inference
- Asynchronous inference and serving
- Streaming ASR
- Timestamp prediction
- High-throughput inference
- Scalable production serving
Voicing-ASR is designed to provide a complete solution for researchers and developers looking to integrate high-quality multilingual speech recognition into production applications.
- Downloads last month
- 62