Sintra AI
Home
Live Feed
Automation Hub
Prompt Library256
AI News554
Weekly Digest
Topic Hubs
AI History
AI Labs
Research
Learning Paths
Guides
Resources
Concepts
Videos
AI Tools74
Models
Claude
Google AI
Cost Calc
All tools/🎙️ Audio & Voice

Mega-ASR

Tsinghua University (xzf-thu)

Free

Foundation ASR built for real-world audio where Whisper breaks down

About

Mega-ASR is an open-source speech recognition model purpose-built for real-world audio — noise, far-field speech, echo, reverberation, recording artifacts, electronic distortion, and transmission dropout — covering 54 compound acoustic scenarios in a single model. Built on a Whisper backbone with LoRA routing, it achieves up to 30% WER reduction over SOTA in challenging conditions. Released May 19, 2026 with Apache 2.0 weights on Hugging Face. **Step-by-step:** 1) `git clone https://github.com/xzf-thu/Mega-ASR`. 2) `pip install -r requirements.txt`. 3) Download weights from HuggingFace (linked in repo). 4) Run the provided inference script with your audio file. 5) Alternatively, install via Pinokio for a one-click desktop app. **Best for:** call-center audio, field recordings, podcasts with bad mics, surveillance, any pipeline where Whisper degrades.

★
Key differentiator

Up to 30% WER reduction vs Whisper on noisy, echoing, or far-field recordings.

Free (Apache 2.0, self-hosted) · Under 5 GB on consumer GPU

transcriptionnoisy audioopen-sourceWhisperreal-world
Open Mega-ASR

More Audio & Voice tools

ElevenLabs

The leading AI voice platform for text-to-speech and voice cloning

Freemium

Suno

AI music generator that creates complete songs from a text prompt

Freemium

Udio

AI music creation with granular style and production control

Freemium