FunASR is a fundamental, end-to-end speech recognition toolkit designed for industrial applications. It boasts exceptional speed, up to 170x realtime, significantly outperforming Whisper, and supports over 50 languages. Beyond basic transcription, FunASR integrates advanced functionalities such as built-in speaker diarization, emotion detection, and real-time streaming capabilities. It also offers an OpenAI-compatible API for easy deployment and integration with AI agents, making it a versatile and cost-effective solution for a wide range of audio processing needs.
Key Features
01Support for 50+ Languages
0216,248 GitHub stars
03170x Realtime Speech Recognition (GPU)
04Emotion Detection and Audio Events
05OpenAI-Compatible API and Agent Integration
06Built-in Speaker Diarization