LocalVQE GGUF
GGUF weights of LocalVQE v1.3 (LocalAI) for s2s.cpp, a C++17/GGML voice assistant. The model removes from the microphone what the loudspeaker played, and cleans noise and reverberation in the same pass: the assistant hears its user over laptop speakers, no headphones, on every browser.
Files
| file | size |
|---|---|
| localvqe-v1.3-F32.gguf | 19 MB |
DeepVQE derivative, 4.8M parameters, 16 kHz, hops of 256 samples. It estimates the echo delay itself over about one second, so the reference only has to be roughly in step with the microphone.
Not the upstream GGUF
LocalAI-io/LocalVQE publishes its own GGUF files for its own engine. This one is a separate port with its own layout, read by src/localvqe.cpp of s2s.cpp: the streams of every connection share one batch on the device, their layer histories kept in VRAM, one slot per stream. The two files are not interchangeable.
convert.py builds it from the PyTorch checkpoint localvqe-v1.3-4.8M.pt: the alignment temperature is folded into its last convolution, the S4D bottleneck is stored as its real and imaginary parts, every dimension is an lv.* metadata key.
Quick start
git clone https://github.com/ServeurpersoCom/s2s.cpp.git
cd s2s.cpp
git submodule update --init
./buildcuda.sh
./models.sh # this file and the rest of the pipeline -> models/
./server.sh # then open http://localhost:8088, echo cancellation: server
Parity
test-localvqe streams a synthetic call through the port and the upstream PyTorch model:
| backend | output cossim | echo removed, far end alone |
|---|---|---|
| CPU | 1.000000000 | 56 dB |
| CUDA | 0.999999737 | 56 dB |
| Vulkan | 0.999998438 | 55 dB |
Four streams sharing the batch give the same output as one.
License
Upstream model: LocalVQE by LocalAI (Richard Palethorpe), Apache 2.0, derived from DeepVQE (Indenbom et al., Interspeech 2023)
GGUF tooling: s2s.cpp, MIT