drbaph/NovaSR

认证创作者 drbaph 已认证
🤗 Hugging Face 来源audio-to-audioapache-2.01 MBother✓ 5 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo drbaph/NovaSR ./model-folder
需要做种者 →

Custom Node for ComfyUI

https://github.com/Saganaki22/ComfyUI-NovaSR

NovaSR: Pushing the Limits of Extreme Efficiency in Audio Super-Resolution

This is the model for NovaSR, a tiny 50kb audio upsampling model that upscales muffled 16khz audio into clear and crisp 48khz audio at speeds from 100-3500x realtime.

Audio Samples

Before Processing (16kHz): Your browser does not support the audio element.

After Processing (48kHz): Your browser does not support the audio element.

Details

  • Model Size: 52kb for pytorch version
  • Input Rate: 16kHz
  • Output Rate: 48kHz
  • Inference Speed: 300-3500x realtime depending on gpu
  • Mono

Comparisons

Comparisons were done on A100 gpu. Higher realtime means faster processing speeds. Comparison on CPU are coming soon.

Model Speed (Real-Time) Model Size
NovaSR 3600x realtime ~52 KB
FlowHigh 20x realtime ~450 MB
FlashSR 14x realtime ~1000 MB
AudioSR 0.6x realtime ~6000 MB

Usage

Please check out the github repo for usage: https://github.com/Saganaki22/ComfyUI-NovaSR

Original Repo: https://github.com/ysharma3501/NovaSR

If you find the model/code helpful, stars or likes would be appreciated.

Thank you.