Whisper large-v3 and large-v3-turbo for Cortiq (CMF)
Ready-to-run CMF v2 Whisper checkpoints for Cortiq 0.8.4. These are converted model files, not Transformers checkpoints.
🎙️ Try it in your browser: record from a microphone or upload audio in the Cortiq Whisper Space.
Install
Requires Rust 1.88 or newer.
cargo install cortiq-cli --version 0.8.4 --locked
Prebuilt binaries: Cortiq 0.8.4 release.
Choose a model
| File | Size | CPU | RTX PRO 4000 Vulkan | Notes |
|---|---|---|---|---|
| whisper-large-v3-turbo-q8.cmf | 826 MB | 21.7 s | 11.1 s | Fastest measured Turbo variant; Q8 reference. |
| whisper-large-v3-turbo-q4tp-mixed.cmf | 602 MB | 25.8 s | 11.9 s | Mixed Q8 attention/output and Q4TP feed-forward. |
| whisper-large-v3-turbo-q4t-uniform.cmf | 477 MB | 24.3 s | 12.8 s | Smallest file. |
| whisper-large-v3-q4tp-mixed.cmf | 1.16 GB | 37.7 s | 27.5 s | Original large-v3, not Turbo. |
Timing notes
Times are single-run CLI wall times on a 10.435-second English LibriSpeech WAV, using Cortiq 0.8.4 on a Xeon E5-2690 v4 (CPU, CMF_THREADS=27) and an NVIDIA RTX PRO 4000 Blackwell (Vulkan). They include model loading and startup; the Vulkan pipeline cache was warm. These sample timings are not a corpus benchmark or a word-error-rate (WER) result.
Whisper now uses Cortiq's shared worker pool and honors CMF_THREADS. On this host, changing from the 0.8.3 Whisper cap (7 workers) to 27 workers changed Turbo Q4T from 31.02 to 24.30 s on CPU and 16.08 to 12.80 s on Vulkan; Turbo Q4TP mixed changed from 33.72 to 25.82 s on CPU and 13.27 to 11.93 s on Vulkan. Results are hardware- and audio-dependent.
Transcribe
Download a .cmf file from the table, then run:
cortiq verify whisper-large-v3-turbo-q4tp-mixed.cmf
cortiq transcribe whisper-large-v3-turbo-q4tp-mixed.cmf speech.wav --language en
cortiq transcribe whisper-large-v3-turbo-q4tp-mixed.cmf interview.wav --language ru
Input is WAV audio. Cortiq supports CPU; Apple Silicon uses Metal, and Vulkan/wgpu is available on supported systems (CMF_GPU=wgpu). For English translation, add --task translate.
The mixed q4tp variant keeps attention and the shared decoder token/output matrix at Q8_2f, and uses Q4TP for feed-forward weights. Recognition quality varies by language and audio; validate WER on your target data before production use.
Sources: OpenAI Whisper large-v3-turbo and large-v3. This is a format conversion, not a fine-tune. See LICENSE.