fix(gpu): #1970 MuseTalk worker 推理期心跳/超时 900/重试收敛/短视频前置失败
CI/CD Pipeline / Check push changed paths (pull_request) Has been skipped
CI/CD Pipeline / Dedup Check - skip PR tests when covered by push pipeline (pull_request) Successful in 0s
CI/CD Pipeline / Check if frontend-only change (pull_request) Successful in 1s
CI/CD Pipeline / Frontend Lint (pull_request) Has been skipped
CI/CD Pipeline / Frontend Unit Tests (pull_request) Has been skipped
CI/CD Pipeline / PR Build Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Deploy Staging (Watchtower auto-deploy) (pull_request) Has been skipped
CI/CD Pipeline / Staging E2E Tests (pull_request) Has been skipped
CI/CD Pipeline / Staging API Integration Tests (pull_request) Has been skipped
CI/CD Pipeline / ACR Image Cleanup (pull_request) Has been skipped
Preview Deploy / Deploy Preview Environment (pull_request) Successful in 1m23s
CI/CD Pipeline / Integration Tests (pull_request) Successful in 2m23s
PR Automation / Auto Approve on CI Green (pull_request) Successful in 3m3s
CI/CD Pipeline / PR Build Worker Image (pull_request) Successful in 3m0s
CI/CD Pipeline / PR Build API Image (pull_request) Successful in 3m1s
CI/CD Pipeline / Validate - Style (pull_request) Successful in 3m5s
CI/CD Pipeline / Validate - Python (mypy + alembic) (pull_request) Successful in 4m19s
AI Code Review / AI Code Review (pull_request) Successful in 6m27s
CI/CD Pipeline / Validate - Security (pull_request) Successful in 6m44s
CI/CD Pipeline / Unit Tests (pull_request) Failing after 8m42s
CI/CD Pipeline / Build Production API Image (pull_request) Has been skipped
CI/CD Pipeline / Build Production Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Production Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Deploy Production (pull_request) Has been skipped
CI/CD Pipeline / Canary Release to Production (pull_request) Has been skipped
CI/CD Pipeline / CI Gate (pull_request) Failing after 1s
CI/CD Pipeline / Production Browser E2E (pull_request) Has been skipped
PR Automation / Auto Merge on CI Green + Approved (pull_request) Successful in 5m54s

- GPU_TASK_TIMEOUT_SECONDS 默认 300→900(base.py + env 模板),worker
  REQUEST_TIMEOUT 默认同步 300→900,RTX2060 6G 处理 720p 长视频不再超时
- worker 新增 TaskHeartbeat daemon 线程:任务处理期间每 30s POST
  /gpu/register(task_id=...) 续任务心跳,服务端只在任务心跳真正停滞
  超过 900s(崩溃/断网)或 worker 明确上报 failed 时才回退 pending,
  长推理阻塞主循环不再导致误回退
- register schema/service 支持 task_id:_touch_task_heartbeat 只刷新
  属于该 worker 且仍 processing 的任务,已完成/已被回收重派的过期心跳忽略
- worker 本地 TASK_MAX_RETRY 2→1,且仅对瞬时错误(连接失败/超时/5xx)重试;
  4xx、结果过小等确定性失败不本地重试,服务端 MAX_ATTEMPTS=3 不变,
  消除 3×3=9 次推理放大
- <3s 输入视频(MuseTalk division by zero)下载后 ffprobe 前置校验,
  直接上报 failed"视频过短",不调用推理;ffprobe 不可用时不拦截
- 新增 11 个单测(worker 独立脚本按路径加载),全量 15839 passed
This commit is contained in:
xiaoxia
2026-09-19 14:29:22 +08:00
parent a1f25a4426
commit a8f1069cd2
12 changed files with 514 additions and 59 deletions
+1
View File
@@ -99,6 +99,7 @@ def register_worker(
gpu_name=body.gpu_name,
free_vram_mb=body.free_vram_mb,
capabilities=body.capabilities,
task_id=body.task_id,
)
return GpuWorkerRegisterResponse(ok=True, server_time=datetime.now(UTC), message="ok")
+8
View File
@@ -22,6 +22,14 @@ class GpuWorkerRegisterRequest(BaseModel):
gpu_name: str = Field("", max_length=200, description="GPU 型号,如 'NVIDIA GeForce RTX 2060'")
free_vram_mb: int = Field(0, ge=0, description="当前空闲显存(MB)")
capabilities: str = Field("musetalk", max_length=500, description="能力列表,逗号分隔,如 'musetalk'")
task_id: Optional[str] = Field(
None,
max_length=64,
description=(
"当前正在处理的任务 ID。Worker 推理期间定期心跳时携带,"
"服务端同步刷新该任务 last_heartbeat_at,防止长推理被误判超时;空闲时不传"
),
)
class GpuWorkerRegisterResponse(BaseModel):
+43 -3
View File
@@ -49,7 +49,15 @@ class GpuLipsyncService:
gpu_name: str = "",
free_vram_mb: int = 0,
capabilities: str = "musetalk",
task_id: Optional[str] = None,
) -> GpuWorkerModel:
"""Worker 注册/心跳。
task_id 非空时(Worker 推理期间的任务级心跳),同步把对应 processing
任务的 last_heartbeat_at 续到当前时间,使长推理不会被
``_recover_timed_out_tasks`` 误回退。任务已结束 / 不属于该 worker
(如已被超时回收重新派发)时忽略,不报错。
"""
now = datetime.now(UTC)
worker = self.db.query(GpuWorkerModel).filter(GpuWorkerModel.worker_id == worker_id).one_or_none()
if worker is None:
@@ -69,6 +77,8 @@ class GpuLipsyncService:
worker.free_vram_mb = free_vram_mb
worker.capabilities = capabilities or worker.capabilities
worker.last_heartbeat_at = now
if task_id:
self._touch_task_heartbeat(task_id, worker_id, now)
self.db.commit()
return worker
@@ -78,8 +88,10 @@ class GpuLipsyncService:
"""原子地认领一条最早的 pending 任务,返回给 worker;无任务返回 None.
同时会:
- 把 processing 状态且超时(超过 gpu_task_timeout_seconds 无心跳)的任务
回退为 pending(attempt++,超过 MAX_ATTEMPTS 置 failed),让其它 worker 认领。
- 把 processing 状态且真正超时(任务心跳停滞超过
gpu_task_timeout_seconds;Worker 推理期会通过 register(task_id=...)
续心跳,长推理不会误判)的任务回退为 pending(attempt++,超过
MAX_ATTEMPTS 置 failed),让其它 worker 认领。
- 刷新 worker 心跳。
"""
now = datetime.now(UTC)
@@ -238,6 +250,28 @@ class GpuLipsyncService:
def _result_key(self, task_id: str) -> str:
return f"{self.RESULT_PREFIX}{task_id}.mp4"
def _touch_task_heartbeat(self, task_id: str, worker_id: str, now: datetime) -> None:
"""Worker 推理期间的任务级心跳:只刷新属于该 worker 且仍在 processing 的任务。
任务不存在 / 已被超时回收重新派发 / 已完成 → 静默忽略(此时旧 worker 的
结果上报会被结果接口按最终态处理)。
"""
task = self.db.get(GpuLipsyncTaskModel, task_id)
if task is None:
return
if task.status != "processing" or task.worker_id != worker_id:
logger.info(
"忽略过期任务心跳 task=%s worker=%s(status=%s owner=%s)",
task_id,
worker_id,
task.status,
task.worker_id,
)
return
task.last_heartbeat_at = now
task.updated_at = now
self.db.flush()
def _touch_worker(self, worker_id: str, now: datetime) -> None:
if not worker_id:
return
@@ -260,7 +294,13 @@ class GpuLipsyncService:
self.db.flush()
def _recover_timed_out_tasks(self, now: datetime) -> None:
"""扫描 processing 状态且超时(无心跳)的任务,回退 pending 或失败."""
"""扫描 processing 状态且真正超时的任务,回退 pending 或失败。
判定只看任务自身 last_heartbeat_at:claim 时写入,Worker 推理期间通过
/gpu/register(task_id=...) 每 30s 续期。因此仅在 Worker 崩溃/断网
(任务心跳停滞超过 gpu_task_timeout_seconds)时才回收,
不会因 Worker 主循环忙于推理而误回退。
"""
timeout = self.settings.gpu_task_timeout_seconds
cutoff = now - timedelta(seconds=timeout)
stuck_tasks = (