Wan-AI/Wan2.2-S2V-14B · speech-to-video (Korean talking-head)
| Condition · Value | |
|---|---|
| GPU | RTX 3090 24GB |
| Power cap (W) | 300 |
| Resolution | 832x480 |
| Frames | 49 |
| Steps | 30 |
| Seed | 42 |
| Precision profile | fp8 offload, persist=13e9 resident |
| Audio encoder | wav2vec2-korean-swap (Korean) |
| Metric · Value | |
|---|---|
| Wall time (s) | 884.9 |
| Per-step time (s) | 28.9 |
| Peak VRAM (GB) | 17.97 |
| Peak host RAM (GB) | 46.5 |
| Load time (s) | 11.1 |
| Repeat count | 1 |
| Render stability | STABLE (sharp_ratio 1.004, motion 1.906) |
Wan2.2-S2V-14B renders a Korean audio-driven talking-head (832x480/49f/30 steps) on a single 24GB RTX 3090 in 884.9s at GPU1=300W: peak VRAM 17.97GB (accurate torch.cuda.max_memory_allocated), 46.5GB host RAM, quality STABLE. Korean wav2vec2-korean-swap encoder. First model with an accurate in-process VRAM number (nvidia-smi VRAM is unreliable on WSL2 GPU-PV). Fix applied: librosa.load->soundfile (numba/NumPy-2.2 incompat). Video shown in veizik dashboard gallery + [internal] drive.
Run the same fixed configuration on your card and open a result PR — see Run it yourself. Raw run manifests are attached to the internal evidence record and are released alongside external reproduction tooling.