fix(1978): MuseTalk 封装强制替换为 TTS 驱动音轨 + 音频长于视频时循环画面 (#1997)
CI/CD Pipeline / Check if frontend-only change (push) Has been skipped
CI/CD Pipeline / Check push changed paths (pull_request) Has been skipped
CI/CD Pipeline / PR Build API Image (push) Has been skipped
CI/CD Pipeline / Dedup Check - skip PR tests when covered by push pipeline (push) Successful in 1s
CI/CD Pipeline / Dedup Check - skip PR tests when covered by push pipeline (pull_request) Successful in 4s
CI/CD Pipeline / Check if frontend-only change (pull_request) Successful in 3s
CI/CD Pipeline / PR Build Worker Image (push) Has been skipped
CI/CD Pipeline / PR Build Web Image (push) Has been skipped
CI/CD Pipeline / Frontend Lint (push) Has been skipped
CI/CD Pipeline / Check push changed paths (push) Successful in 8s
CI/CD Pipeline / Validate - Style (pull_request) Has been skipped
CI/CD Pipeline / Validate - Security (pull_request) Has been skipped
CI/CD Pipeline / Validate - Python (mypy + alembic) (pull_request) Has been skipped
CI/CD Pipeline / Unit Tests (pull_request) Has been skipped
CI/CD Pipeline / Integration Tests (pull_request) Has been skipped
CI/CD Pipeline / Frontend Lint (pull_request) Has been skipped
CI/CD Pipeline / Frontend Unit Tests (pull_request) Has been skipped
CI/CD Pipeline / Build Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / PR Build Web Image (pull_request) Successful in 43s
CI/CD Pipeline / PR Build API Image (pull_request) Successful in 53s
CI/CD Pipeline / PR Build Worker Image (pull_request) Successful in 59s
CI/CD Pipeline / Build Production API Image (pull_request) Has been skipped
CI/CD Pipeline / Build Production Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging API Image (push) Successful in 49s
CI/CD Pipeline / Build Production Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Web Image (push) Successful in 22s
CI/CD Pipeline / CI Gate (pull_request) Successful in 2s
CI/CD Pipeline / Deploy Staging (Watchtower auto-deploy) (pull_request) Has been skipped
CI/CD Pipeline / Deploy Production (pull_request) Has been skipped
CI/CD Pipeline / Staging E2E Tests (pull_request) Has been skipped
CI/CD Pipeline / Staging API Integration Tests (pull_request) Has been skipped
CI/CD Pipeline / ACR Image Cleanup (pull_request) Has been skipped
CI/CD Pipeline / Production Browser E2E (pull_request) Has been skipped
CI/CD Pipeline / Canary Release to Production (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Worker Image (push) Successful in 32s
CI/CD Pipeline / Retag skipped Staging API Image (push) Has been skipped
CI/CD Pipeline / Retag skipped Staging Web Image (push) Has been skipped
CI/CD Pipeline / Retag skipped Staging Worker Image (push) Has been skipped
Preview Deploy / Deploy Preview Environment (pull_request) Successful in 1m33s
CI/CD Pipeline / Deploy Staging (Watchtower auto-deploy) (push) Successful in 46s
PR Automation / Auto Approve on CI Green (pull_request) Successful in 3m8s
PR Automation / Auto Merge on CI Green + Approved (pull_request) Has been skipped
CI/CD Pipeline / Integration Tests (push) Successful in 3m36s
CI/CD Pipeline / Validate - Python (mypy + alembic) (push) Successful in 3m43s
CI/CD Pipeline / Frontend Unit Tests (push) Successful in 4m15s
CI/CD Pipeline / Validate - Style (push) Successful in 4m22s
CI/CD Pipeline / ACR Image Cleanup (push) Successful in 2m22s
CI/CD Pipeline / Staging API Integration Tests (push) Successful in 3m48s
CI/CD Pipeline / Staging E2E Tests (push) Failing after 3m55s
AI Code Review / AI Code Review (pull_request) Successful in 6m38s
CI/CD Pipeline / Validate - Security (push) Successful in 7m26s
CI/CD Pipeline / Unit Tests (push) Successful in 1h52m21s
CI/CD Pipeline / Build Production API Image (push) Has been skipped
CI/CD Pipeline / Build Production Web Image (push) Has been skipped
CI/CD Pipeline / Build Production Worker Image (push) Has been skipped
CI/CD Pipeline / CI Gate (push) Has been skipped
CI/CD Pipeline / Canary Release to Production (push) Has been skipped
CI/CD Pipeline / Deploy Production (push) Has been skipped
CI/CD Pipeline / Production Browser E2E (push) Has been skipped

Co-authored-by: backend-dev <dev@xiaoxiajianji.com>
Co-committed-by: backend-dev <dev@xiaoxiajianji.com>
This commit was merged in pull request #1997.
This commit is contained in:
2026-09-20 02:27:44 +08:00
committed by auto-approve-bot
parent d959dd874f
commit b0b81a5d60
3 changed files with 592 additions and 12 deletions
+33 -2
View File
@@ -48,8 +48,33 @@ vim .env
| `MUSE_AUDIO_MAX_MB` | 音频上传大小限制 MB | `20` |
| `MUSE_DEFAULT_FPS` | 视频 fps 兜底值 | `25.0` |
| `MUSE_TEMP_DIR` | 临时文件目录 | `/tmp/musetalk_$$` |
| `MUSE_VIDEO_ENCODER` | 循环视频时的编码器:`auto`(优先 h264_nvenc,失败回退 libx264)/`h264_nvenc`/`libx264` | `auto` |
| `MUSE_ENABLE_VIDEO_LOOP` | 驱动音频比视频长时循环视频补齐画面,`0` 关闭 | `1` |
### 2.2 启动服务
### 2.2 更新部署(音轨修复,必做)
> ⚠️ 2026-09-20 修复严重 bug:旧版封装保留了源视频音轨,结果口型配的是原声而不是 TTS 驱动音频。RTX2060 机器必须重新拉取 `musetalk_server.py` 并重启:
```bash
# 在 RTX2060 上备份旧文件并拉取新版本(按实际部署路径调整)
cp ~/projects/MuseTalk/musetalk_server.py ~/projects/MuseTalk/musetalk_server.py.bak
# 从仓库 raw 地址下载最新版(替换为你的仓库地址/分支)
wget -O ~/projects/MuseTalk/musetalk_server.py \
"https://git.xiaoxiajianji.com/xiaoxia/xiaoxia-saas/raw/branch/develop/deploy/gpu_worker/musetalk_server.py"
# 重启服务
sudo systemctl restart musetalk-server
sudo systemctl status musetalk-server
curl http://127.0.0.1:7861/health
```
修复后封装逻辑:
- 最终 mux 强制 `-map 0:v -map 1:a`:视频流只取 MuseTalk 无声画面,音轨只取 TTS 驱动音频,杜绝 ffmpeg 默认行为带入源视频音轨
- 驱动音频不长于视频时:`-c:v copy -c:a aac -shortest`,无损秒封装
- 驱动音频长于视频时(如 TTS 15s vs 视频 9s):`-stream_loop -1` 循环画面,RTX2060 走 `h264_nvenc` 硬件重编码(NVENC 失败自动回退 libx264),`-t` 精确卡到音频时长
### 2.3 启动服务
```bash
# 前台运行(调试用)
@@ -60,7 +85,7 @@ sudo systemctl start musetalk-server
sudo systemctl enable musetalk-server
```
### 2.3 验证健康检查
### 2.4 验证健康检查
```bash
curl http://127.0.0.1:7861/health
@@ -184,3 +209,9 @@ MuseTalk 健康检查通过: {...}
新增:
- `/cancel` 端点:终止当前推理任务,清理临时文件
- `/health` 端点:返回 GPU 显存信息和当前任务状态
2026-09-20 追加修复(音轨正确性,上线阻断级):
9. **音轨未替换(严重)**:旧最终封装让 ffmpeg 默认选流,结果保留了源视频自带音轨(与画面相关系数 0.9998,与 TTS 无关)。改为 `_mux_video_with_audio()` 统一封装,强制 `-map 0:v:0 -map 1:a:0`,画面取 MuseTalk 无声产物、音轨只取驱动音频
10. **音视频时长不对齐**:TTS 长于原视频时 `-shortest` 会截短语音。改为探测双方时长,音频更长时 `-stream_loop -1` 循环画面 + `h264_nvenc` 硬件重编码(`MUSE_VIDEO_ENCODER=auto`,失败回退 libx264)+ `-t <音频时长>`;不循环时 `-c:v copy` 秒封装
- 开关 `MUSE_ENABLE_VIDEO_LOOP=0` 可关闭循环;请求也支持 form 参数 `enable_video_loop` 单任务覆盖
+185 -10
View File
@@ -11,6 +11,8 @@
MUSE_AUDIO_MAX_MB 音频上传大小限制 MB,默认 20
MUSE_DEFAULT_FPS 视频 fps 兜底值,默认 25.0
MUSE_TEMP_DIR 临时文件目录,默认 /tmp/musetalk_$$
MUSE_VIDEO_ENCODER 循环视频时的编码器:auto(默认,优先 h264_nvenc 兜底 libx264)/h264_nvenc/libx264
MUSE_ENABLE_VIDEO_LOOP 驱动音频比视频长时是否循环视频补齐,默认 1(开启)
接口:
GET /health 健康检查 + GPU 显存信息
@@ -57,6 +59,12 @@ class Config:
audio_max_mb: int = int(_env("MUSE_AUDIO_MAX_MB", "20"))
default_fps: float = float(_env("MUSE_DEFAULT_FPS", "25.0"))
temp_dir: str = _env("MUSE_TEMP_DIR", f"/tmp/musetalk_{os.getpid()}")
# 循环视频时编码器:auto 优先 h264_nvenc(RTX2060 支持),失败兜底 libx264
video_encoder: str = _env("MUSE_VIDEO_ENCODER", "auto") or "auto"
# 驱动音频比视频长时循环视频补齐画面
enable_video_loop: bool = _env("MUSE_ENABLE_VIDEO_LOOP", "1") not in ("0", "false", "False", "")
# 判定音视频时长差异的容差(秒),避免 ffprobe 微小误差触发无谓的循环/重编码
duration_epsilon: float = 0.25
# ── 全局状态 ──────────────────────────────────────────────────────────
@@ -159,6 +167,152 @@ def _get_video_fps(video_path: Path) -> float:
return Config.default_fps
def _get_media_duration(path: Path) -> float:
"""用 ffprobe 读媒体时长(秒),失败返回 0.0."""
try:
out = subprocess.check_output(
[
"ffprobe",
"-v",
"error",
"-show_entries",
"format=duration",
"-of",
"default=noprint_wrappers=1:nokey=1",
str(path),
],
stderr=subprocess.DEVNULL,
timeout=10,
)
duration = float(out.decode().strip())
return duration if duration > 0 else 0.0
except Exception as exc:
logger.warning("ffprobe 读时长失败 %s: %s", path, exc)
return 0.0
def _pick_video_encoder() -> str:
"""选择视频编码器:配置指定则用指定值;auto 时探测 NVENC 是否可用,不可用回退 libx264."""
configured = Config.video_encoder.strip()
if configured in ("h264_nvenc", "libx264"):
return configured
# auto:探测本机 ffmpeg 是否编译了 h264_nvenc
try:
result = subprocess.run(
["ffmpeg", "-hide_banner", "-encoders"],
stdout=subprocess.PIPE,
stderr=subprocess.DEVNULL,
timeout=10,
check=False,
)
if b"h264_nvenc" in result.stdout:
return "h264_nvenc"
except Exception as exc:
logger.warning("探测 ffmpeg 编码器失败,回退 libx264: %s", exc)
return "libx264"
def _mux_video_with_audio(
video_path: Path,
audio_path: Path,
output_path: Path,
enable_video_loop: Optional[bool] = None,
timeout: float = 300,
) -> None:
"""把无声画面视频与驱动音频封装为最终结果.
关键正确性要求:必须用 -map 0:v -map 1:a 显式指定取第一个输入(推理画面)的
视频流和第二个输入(驱动音频 TTS)的音频流,禁止 ffmpeg 默认流选择行为
(否则会把源视频自带音轨带进结果,口型与声音错位)。
时长对齐:驱动音频比视频长时(TTS 15s vs 原视频 9s 很常见),用
-stream_loop -1 循环视频画面到音频长度(NVENC 硬件重编码),-t 卡到音频时长;
音频不超过视频时直接 -c:v copy 无损快封装,-shortest 以较短流为准。
"""
video_duration = _get_media_duration(video_path)
audio_duration = _get_media_duration(audio_path)
loop_enabled = Config.enable_video_loop if enable_video_loop is None else enable_video_loop
need_loop = bool(
loop_enabled
and audio_duration > 0
and video_duration > 0
and audio_duration > video_duration + Config.duration_epsilon
)
if need_loop:
encoder = _pick_video_encoder()
# preset 随编码器选择:h264_nvenc 用 p1-p7,libx264 用词形 preset
preset = "p4" if encoder == "h264_nvenc" else "veryfast"
logger.info(
"音频(%.2fs)长于视频(%.2fs),循环视频并以 %s(%s) 重编码至音频长度",
audio_duration,
video_duration,
encoder,
preset,
)
def build_cmd(enc: str, pre: str) -> list:
return [
"ffmpeg",
"-y",
"-stream_loop",
"-1",
"-i",
str(video_path),
"-i",
str(audio_path),
"-map",
"0:v:0",
"-map",
"1:a:0",
"-c:v",
enc,
"-preset",
pre,
"-c:a",
"aac",
"-b:a",
"128k",
"-t",
f"{audio_duration:.3f}",
str(output_path),
]
try:
_run_ffmpeg(build_cmd(encoder, preset), timeout=timeout)
except RuntimeError:
# NVENC 可能因驱动/占用失败,兜底 libx264 重试一次
if encoder == "h264_nvenc":
logger.warning("h264_nvenc 封装失败,回退 libx264 重试")
_run_ffmpeg(build_cmd("libx264", "veryfast"), timeout=timeout)
else:
raise
else:
# 视频不短于音频:直接复制视频流,只把音频替换为驱动音频并转 AAC
cmd = [
"ffmpeg",
"-y",
"-i",
str(video_path),
"-i",
str(audio_path),
"-map",
"0:v:0",
"-map",
"1:a:0",
"-c:v",
"copy",
"-c:a",
"aac",
"-b:a",
"128k",
"-shortest",
str(output_path),
]
_run_ffmpeg(cmd, timeout=timeout)
def _check_file_size(file, max_mb: int, label: str) -> Optional[str]:
"""检查文件大小,超限返回错误信息,否则返回 None."""
file.seek(0, 2)
@@ -190,11 +344,18 @@ def _run_ffmpeg(cmd: list, timeout: float = 120) -> subprocess.CompletedProcess:
raise RuntimeError(f"ffmpeg 超时(>{timeout}s)") from exc
def _run_inference(video_path: Path, audio_path: Path, output_path: Path) -> None:
def _run_inference(
video_path: Path,
audio_path: Path,
output_path: Path,
enable_video_loop: Optional[bool] = None,
) -> None:
"""执行 MuseTalk 推理(可被子线程和测试独立调用).
实际部署时替换为 MuseTalk 真实推理逻辑。
此处为示例实现:提取帧 → 合并音视频。
此处为示例实现:提取帧 → 生成无声画面 → 用驱动音频封装。
enable_video_loop: 驱动音频长于视频时是否循环视频;None 走全局配置。
"""
fps = _get_video_fps(video_path)
logger.info("视频 fps: %.2f", fps)
@@ -218,27 +379,32 @@ def _run_inference(video_path: Path, audio_path: Path, output_path: Path) -> Non
if not frame_files:
raise RuntimeError("未从视频中提取到帧")
# TODO: 替换为 MuseTalk 实际推理逻辑
# TODO: 替换为 MuseTalk 实际推理逻辑。
# MuseTalk 真实产物是「无声画面视频」,音轨必须在封装阶段用驱动音频替换。
logger.warning("使用示例推理逻辑,未实际调用 MuseTalk 模型")
# 示例:从源视频生成无声画面(-an 丢弃原音轨),模拟 MuseTalk 推理产物。
# 真实部署时 silent_video_path 应替换为 MuseTalk 输出的无声视频路径。
silent_video_path = video_path.parent / "visual_silent.mp4"
_run_ffmpeg(
[
"ffmpeg",
"-y",
"-i",
str(video_path),
"-i",
str(audio_path),
"-an",
"-c:v",
"libx264",
"-c:a",
"aac",
"-shortest",
str(output_path),
"-preset",
"veryfast",
str(silent_video_path),
],
timeout=300,
)
# 统一封装:显式 -map 取推理画面 + 驱动音频;音频更长时循环视频。
_mux_video_with_audio(silent_video_path, audio_path, output_path, enable_video_loop=enable_video_loop)
if not output_path.exists() or output_path.stat().st_size < 1024:
raise RuntimeError("推理产物不存在或过小")
@@ -286,6 +452,13 @@ def inference():
audio_file = request.files["audio"]
task_id = request.form.get("task_id", f"task_{int(time.time())}")
# 可选:本次任务是否在音频长于视频时循环视频(缺省走全局配置)
loop_param = request.form.get("enable_video_loop")
if loop_param is not None:
task_enable_loop = loop_param.strip() not in ("0", "false", "False", "")
else:
task_enable_loop = None
# 文件大小检查
err = _check_file_size(video_file, Config.video_max_mb, "视频")
if err:
@@ -319,7 +492,7 @@ def inference():
def inference_thread():
try:
_run_inference(video_path, audio_path, output_path)
_run_inference(video_path, audio_path, output_path, enable_video_loop=task_enable_loop)
except Exception as exc:
result_container["error"] = str(exc)
@@ -413,6 +586,8 @@ def main():
logger.info(" 视频大小限制: %dMB", Config.video_max_mb)
logger.info(" 音频大小限制: %dMB", Config.audio_max_mb)
logger.info(" 默认 fps: %.1f", Config.default_fps)
logger.info(" 视频编码器: %s", Config.video_encoder)
logger.info(" 音频长于视频时循环视频: %s", Config.enable_video_loop)
logger.info("=" * 60)
# 检查 GPU