GeneFace++ 默认可能用 512 或更高:
# 配置示例
img_size: 256推理时也建议:
batch_size = 1尤其 video rendering 时。
开启自动混合精度:
with torch.cuda.amp.autocast():
...或模型直接转 half:
model.half()torch.inference_mode()比 no_grad 更省:
with torch.inference_mode():
output = model(input)不要一次性渲染整段视频:
batch_size: 1 or 2img_size: 256如果代码支持:
model.use_checkpointing = True用时间换显存。
scaler = torch.cuda.amp.GradScaler()如果有多卡:
torchrun --nproc_per_node=2 train.pytorch.cuda.empty_cache()GeneFace++ 可换:
| 设置 | 显存 |
|---|---|
| 512 + bs=4 | 24G+ |
| 256 + bs=1 + FP16 | 8–12G |
| 256 + bs=1 + AMP + ckpt | 6–8G |
256 分辨率 + batch=1 + FP16 + inference_mode + 分段渲染
如果你愿意,可以:
我可以直接帮你改配置。