ethanfel/FoleyTune-LoRAs

🤗 Hugging Face 来源apache-2.0913 MBother✓ 4 个校验和今天更新
已有模型文件?提交模型种子

如果你有完整的模型文件并有权分享,请把示例文件夹路径替换为你的文件路径,再运行这条命令。它会校验文件、制作种子,并将磁力链接和校验和提交给 Pirate Face。请让种子客户端持续做种,方便其他人从节点下载。Pirate Face 不接收模型文件。你可以从账户页面获取社区密钥。也可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo ethanfel/FoleyTune-LoRAs ./model-folder
需要做种者 →

FoleyTune LoRAs

LoRA adapters trained with ComfyUI-HunyuanVideo-FoleyTune for video-to-audio generation on the HunyuanVideo audio transformer.

Adult content. Some adapters in this repository are trained on adult audio datasets. Browse subdirectories for details.

⚠️ Render CFG depends on the folder. Some adapters must be rendered at CFG = 1.0 (they saturate at higher guidance); others work at any normal CFG. Check the folder / its README before rendering.

Categories

Folder Description Adapters
type-a/ Adult audio LoRAs 5 blowjob adapters
CFG_1.0/ Render @ CFG 1.0 ONLY (saturate at higher CFG) 1 sex/moaning + skin-slap merge
standard_cfg/ Standard-CFG adapters (any normal guidance) 1 blowjob

The CFG_1.0/ and standard_cfg/ adapters use the newer all_blocks_sync_io target + Prodigy+ ScheduleFree (c=20) recipe (see each folder's README), and render best with TF32 off, raw output (no bandwidth extension). The type-a/ config below is the older sweep.

Base Training Config

All adapters share this foundation unless noted in their category README:

  • Base model: HunyuanVideo audio transformer
  • LoRA target: all_attn_mlp (attention + MLP layers)
  • Rank: 64, Alpha: 64
  • Optimizer: Prodigy (adaptive learning rate)
  • Schedule: Cosine with curriculum switching
  • Precision: bf16
  • Visual dropout: 0.5

Metrics

  • PBC (Per-Band Correlation): frequency-band tracking accuracy. Higher is better.
  • SC (Spectral Convergence): overall spectral fidelity. Lower is better.
  • MCD (Mel-Cepstral Distortion): timbral accuracy. Lower is better.
  • TV (Temporal Variance): dynamic range over time. Higher means richer variation.

Links

Usage

Load in ComfyUI with the FoleyTune LoRA Loader node (supports both .safetensors and .pt), or manually:

import torch
from safetensors.torch import load_file

state_dict = load_file("type-a/typeA_blowjob_rank1_sigma07_15k_PBC0661_SC1058.safetensors")
# Apply to HunyuanVideo audio transformer with load_lora()