Local transcription & batch tasks
Drop audio or video into the window and start. ASRbox preflights everything first: audio track presence, format support, duration and chunking — problems are explained before any work begins. Multiple files can be queued at once; tasks can be paused, resumed and retried.
- 15 local models: Whisper, Faster Whisper (word-level timestamps), MLX Whisper, SenseVoice, Qwen3-ASR, MOSS-Transcribe-Diarize (speaker diarization, 50+ languages)
- Model sizes from 145 MB to 6.2 GB — pick your accuracy/speed trade-off
- Resumable downloads: pause, resume and retry model downloads