Speech and audio understanding
Transcribe, clean and route spoken media across ASR and audio-processing systems.
What we currently think
The best speech workflow depends on language, noise, timestamps, speaker structure and what the transcript must support next.
Start with your question
- Which transcription path fits this recording?
- What audio defect should be repaired first?
- What structure must survive into the next workflow?
Start here
Qwen3-ASR Speech Recognition Workflows (Overview)
How we are piloting Alibaba Cloud's Qwen3-ASR model for multilingual speech recognition while keeping Whisper in the stack for timestamps, offline capture, and open-source tooling.
GPT-4o Transcribe Speech-to-Text Workflows (Overview)
How we are trialling OpenAI's gpt-4o-transcribe STT model alongside Whisper to accelerate edits, call recaps, and UGC compliance checks across creative production.
Diffusion Speech Denoising in 2025 -- StoRM, SGMSE+, UNIVERSE++, Schrodinger Bridges, and Streaming Variants
A field guide to 2025's diffusion-first speech denoisers: StoRM's predictive guidance loop, SGMSE/SGMSE+ refinements, UNIVERSE++'s universal training, few-step Schrodinger-Bridge offshoots, causal streaming diffusion, and MossFormer2's transformer/FSMN hybrid for classic separation-first pipelines.
Continue exploring
NVIDIA NeMo Speech Collection First Technical Read and Production Reality Check
A first engineering read of NVIDIA NeMo for speech workflows as of February 2026: what is released now, where it fits in an AI video pipeline, where it does not, and what to validate before 24GB production adoption.
SpeechBrain Conversational AI Toolkit Workflows (Overview)
How the open-source SpeechBrain ecosystem helps us prototype multilingual speech models, while Whisper stays our production default for timestamped subtitles and on-set capture.