Serveurperso/LocalVQE-GGUF

🤗 Hugging Face sourceaudio-to-audioapache-2.019 MBGGUF✓ 1 checksumupdated today
Needs seeder →

LocalVQE GGUF

GGUF weights of LocalVQE v1.3 (LocalAI) for s2s.cpp, a C++17/GGML voice assistant. The model removes from the microphone what the loudspeaker played, and cleans noise and reverberation in the same pass: the assistant hears its user over laptop speakers, no headphones, on every browser.

Files

file size
localvqe-v1.3-F32.gguf 19 MB

DeepVQE derivative, 4.8M parameters, 16 kHz, hops of 256 samples. It estimates the echo delay itself over about one second, so the reference only has to be roughly in step with the microphone.

Not the upstream GGUF

LocalAI-io/LocalVQE publishes its own GGUF files for its own engine. This one is a separate port with its own layout, read by src/localvqe.cpp of s2s.cpp: the streams of every connection share one batch on the device, their layer histories kept in VRAM, one slot per stream. The two files are not interchangeable.

convert.py builds it from the PyTorch checkpoint localvqe-v1.3-4.8M.pt: the alignment temperature is folded into its last convolution, the S4D bottleneck is stored as its real and imaginary parts, every dimension is an lv.* metadata key.

Quick start

git clone https://github.com/ServeurpersoCom/s2s.cpp.git
cd s2s.cpp
git submodule update --init
./buildcuda.sh
./models.sh      # this file and the rest of the pipeline -> models/
./server.sh      # then open http://localhost:8088, echo cancellation: server

Parity

test-localvqe streams a synthetic call through the port and the upstream PyTorch model:

backend output cossim echo removed, far end alone
CPU 1.000000000 56 dB
CUDA 0.999999737 56 dB
Vulkan 0.999998438 55 dB

Four streams sharing the batch give the same output as one.

License

Upstream model: LocalVQE by LocalAI (Richard Palethorpe), Apache 2.0, derived from DeepVQE (Indenbom et al., Interspeech 2023)

GGUF tooling: s2s.cpp, MIT