rzgar/Bernini-R-S2V

🤗 On Hugging Faceimage-to-imageapache-2.0143 GBother✓ Checksum-verifiedupdated 0d ago
Magnet

Bernini-R-S2V

ComfyUI Bernini-R-S2V custom node v2 (update)

Bernini S2V Conditioning v2 - Bernini in-context video/image conditioning with masked lip-sync for one or two speakers.

Your browser does not support the video tag.

Unzip ComfyUI-WanBerniniS2V_v2.zip into ComfyUI/custom_nodes/, then restart ComfyUI.

or save the Python files in: ComfyUI/custom_nodes/ComfyUI-WanBerniniS2V_v2/

Disable or remove the older ComfyUI-WanBerniniS2V folder if you only want v2.

One speaker

  • audio_1 + mask_1
  • Leave audio_2 / mask_2 unwired

Two speakers (dialog)

  • audio_1 + mask_1 - first speaker
  • audio_2 + mask_2 - second speaker
  • speaker_2_start_frame = -1 - second audio starts when the first clip ends

Masks

Paint on the output frame where each speaker's face appears. Masks control lip-sync placement only, they are not tied to reference_image_N slots.

________

Your browser does not support the video tag.

Speech-driven video on Bernini-R , T2V, I2V, and V2V with lip-sync.

This model adds single-speaker audio support to Bernini-R, so you can drive video with speech in text-to-video, image-to-video, and video-to-video setups.

It is not state-of-the-art audio-to-video, but it removes the need for post-processing or extra models just to add speech to Wan videos.

For basic talking-head work, or longer videos built from short sequences, it is a handy all-in-one option on top of Bernini's motion and editing strengths.

| File | Role |

|---|---|

| 'wan2.2_bernini_r_high_noise_fp16_s2v.safetensors'| High noise |

| 'wan2.2_bernini_r_low_noise_fp16_s2v.safetensors'| Low noise |

| 'wan2.2_bernini_r_high_noise_fp8_scaled_s2v.safetensors' | High noise |

| 'wan2.2_bernini_r_low_noise_fp8_scaled_s2v.safetensors' | Low noise |

| 'wan2.2_bernini_r_high_noise_int8_convrot_s2v.safetensors' | High noise |

| 'wan2.2_bernini_r_low_noise_int8_convrot_s2v.safetensors' | Low noise |

ComfyUI detects these as WAN22_S2V / 'WanModel_S2V' (audio keys trigger S2V model type).

Usage

1. Download the checkpoints.

2. Place in:

   ComfyUI/models/diffusion_models/

3. Add wav2vec2 to:

   ComfyUI/models/audio_encoders/

4. Install ComfyUI-WanBerniniS2V custom node:

   (Create a folder in 'ComfyUI/custom_nodes/'  named 'ComfyUI-WanBerniniS2V', then save the Python files in that folder.)

5. Restart ComfyUI.

6. Search for the Bernini S2V Conditioning node or use the example workfllow

Audio tips

| Setting | Recommendation |

|---|---|

| Channels | Mono, wav2vec2 downmixes stereo internally |

| Sample rate | 44.1 khz or 48 khz (resampled to 16 kHz) |

| Content | Clear speech, less background music = better sync |

| Length | in my testing, max 9 to 15 seconds |