fix(voice-clone): 处理中卡死兜底 — worker重启/Celery消息丢失后processing记录永久卡住问题

根因:
- 声音克隆走 API提交CosyVoice→Celery worker轮询5分钟→mark ready/failed
- 当worker容器重启、进程OOM或Celeryprefetch消息丢失时,已进入processing状态的voice_clone_profile没有兜底恢复机制
- 对比ingest链路有cleanup_stale_ingest_jobs beat任务+worker_ready启动恢复,voice_clone链路缺失
- 结果: 用户看到"克隆处理中,请稍候..."永久转圈,刷新也不变

修复:
1. SQLAlchemyVoiceCloneProfileRepository新增cleanup_stale_processing方法:
   扫描updated_at超过10分钟的processing记录(正常克隆<5分钟),标记failed并提示用户重试
2. _startup.py新增recover_stale_voice_clones_on_startup+worker_ready信号:
   worker启动时一次性扫描恢复,用户重启/部署后立即解锁卡死任务
3. cleanup.py新增scheduled_cleanup_stale_voice_clones beat任务:
   每5分钟巡检兜底,防止运行期OOM/消息丢失导致的新卡死
4. celery_app.py beat_schedule注册新任务
5. voice_clone.py任务入口加received日志,便于日志排查
This commit is contained in:
CI Bot
2026-09-26 11:34:43 +08:00
committed by saas-backend
parent 534e4fcc36
commit a33532113b
5 changed files with 130 additions and 0 deletions
@@ -136,6 +136,39 @@ class SQLAlchemyVoiceCloneProfileRepository:
)
return {voice_id: profile_id for voice_id, profile_id in rows}
def cleanup_stale_processing(self, timeout_minutes: int = 10) -> int:
"""清理超时卡在 processing 的克隆档案。
worker 重启、Celery 任务丢失或 OOM 被杀时,processing 档案会永久卡住。
updated_at < NOW() - timeout_minutes 的 processing 记录,标记为 failed
并附带明确错误信息,用户可在前端点击「重试」。
Args:
timeout_minutes: 超时分钟数,默认 10 分钟(正常克隆 < 5 分钟)
Returns:
清理的记录数
"""
from datetime import datetime, timedelta, UTC
cutoff = datetime.now(UTC) - timedelta(minutes=timeout_minutes)
models = (
self.session.query(VoiceCloneProfileModel)
.filter(
VoiceCloneProfileModel.status == "processing",
VoiceCloneProfileModel.updated_at < cutoff,
)
.all()
)
count = 0
for model in models:
model.status = "failed"
model.error_message = f"克隆任务执行超时(超过 {timeout_minutes} 分钟未更新,可能因服务重启中断),请重试"
count += 1
if count > 0:
self.session.commit()
return count
@staticmethod
def _model_to_entity(model: VoiceCloneProfileModel) -> VoiceCloneProfile:
return VoiceCloneProfile(