infosave/whisper-large-v3-turbo-cmf

🤗 Hugging Face sourceautomatic-speech-recognitionmit3.1 GBother✓ 4 checksumsupdated today
Needs seeder →

Whisper large-v3 and large-v3-turbo for Cortiq (CMF)

Ready-to-run CMF v2 Whisper checkpoints for Cortiq 0.8.4. These are converted model files, not Transformers checkpoints.

🎙️ Try it in your browser: record from a microphone or upload audio in the Cortiq Whisper Space.

Install

Requires Rust 1.88 or newer.

cargo install cortiq-cli --version 0.8.4 --locked

Prebuilt binaries: Cortiq 0.8.4 release.

Choose a model

File Size CPU RTX PRO 4000 Vulkan Notes
whisper-large-v3-turbo-q8.cmf 826 MB 21.7 s 11.1 s Fastest measured Turbo variant; Q8 reference.
whisper-large-v3-turbo-q4tp-mixed.cmf 602 MB 25.8 s 11.9 s Mixed Q8 attention/output and Q4TP feed-forward.
whisper-large-v3-turbo-q4t-uniform.cmf 477 MB 24.3 s 12.8 s Smallest file.
whisper-large-v3-q4tp-mixed.cmf 1.16 GB 37.7 s 27.5 s Original large-v3, not Turbo.

Timing notes

Times are single-run CLI wall times on a 10.435-second English LibriSpeech WAV, using Cortiq 0.8.4 on a Xeon E5-2690 v4 (CPU, CMF_THREADS=27) and an NVIDIA RTX PRO 4000 Blackwell (Vulkan). They include model loading and startup; the Vulkan pipeline cache was warm. These sample timings are not a corpus benchmark or a word-error-rate (WER) result.

Whisper now uses Cortiq's shared worker pool and honors CMF_THREADS. On this host, changing from the 0.8.3 Whisper cap (7 workers) to 27 workers changed Turbo Q4T from 31.02 to 24.30 s on CPU and 16.08 to 12.80 s on Vulkan; Turbo Q4TP mixed changed from 33.72 to 25.82 s on CPU and 13.27 to 11.93 s on Vulkan. Results are hardware- and audio-dependent.

Transcribe

Download a .cmf file from the table, then run:

cortiq verify whisper-large-v3-turbo-q4tp-mixed.cmf
cortiq transcribe whisper-large-v3-turbo-q4tp-mixed.cmf speech.wav --language en
cortiq transcribe whisper-large-v3-turbo-q4tp-mixed.cmf interview.wav --language ru

Input is WAV audio. Cortiq supports CPU; Apple Silicon uses Metal, and Vulkan/wgpu is available on supported systems (CMF_GPU=wgpu). For English translation, add --task translate.

The mixed q4tp variant keeps attention and the shared decoder token/output matrix at Q8_2f, and uses Q4TP for feed-forward weights. Recognition quality varies by language and audio; validate WER on your target data before production use.

Sources: OpenAI Whisper large-v3-turbo and large-v3. This is a format conversion, not a fine-tune. See LICENSE.