fix(worker+api): resolve batch title race + auto-queue pending tasks (#2098)
CI/CD Pipeline / Check push changed paths (pull_request) Has been skipped
CI/CD Pipeline / Dedup Check - skip PR tests when covered by push pipeline (pull_request) Successful in 1s
CI/CD Pipeline / Check if frontend-only change (pull_request) Successful in 2s
CI/CD Pipeline / Frontend Lint (pull_request) Has been skipped
CI/CD Pipeline / Frontend Unit Tests (pull_request) Has been skipped
CI/CD Pipeline / PR Build Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Deploy Staging (Watchtower auto-deploy) (pull_request) Has been skipped
CI/CD Pipeline / Staging E2E Tests (pull_request) Has been skipped
CI/CD Pipeline / Staging API Integration Tests (pull_request) Has been skipped
CI/CD Pipeline / ACR Image Cleanup (pull_request) Has been skipped
CI/CD Pipeline / PR Build API Image (pull_request) Successful in 1m3s
Preview Deploy / Deploy Preview Environment (pull_request) Successful in 2m23s
PR Automation / Auto Approve on CI Green (pull_request) Successful in 3m3s
CI/CD Pipeline / PR Build Worker Image (pull_request) Successful in 4m8s
AI Code Review / AI Code Review (pull_request) Successful in 7m3s
CI/CD Pipeline / Integration Tests (pull_request) Successful in 8m12s
CI/CD Pipeline / Validate - Style (pull_request) Successful in 8m39s
CI/CD Pipeline / Validate - Python (mypy + alembic) (pull_request) Successful in 9m47s
PR Automation / Auto Merge on CI Green + Approved (pull_request) Successful in 10m51s
CI/CD Pipeline / Unit Tests (pull_request) Failing after 14m41s
CI/CD Pipeline / Validate - Security (pull_request) Successful in 24m23s
CI/CD Pipeline / Build Production API Image (pull_request) Has been skipped
CI/CD Pipeline / Build Production Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Production Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Deploy Production (pull_request) Has been skipped
CI/CD Pipeline / Canary Release to Production (pull_request) Has been skipped
CI/CD Pipeline / CI Gate (pull_request) Failing after 2s
CI/CD Pipeline / Production Browser E2E (pull_request) Has been skipped

Two batch-rendering bugs found from staging logs (6-video batch 17:30-17:43):

Bug A (P1) — title_config race when same plan renders concurrently:
  _sync_task_config_to_plan wrote each task's title_config/bgm/resolution
  to the shared plan.config row, then RenderAdapter.render_plan read back
  from plan.config during rendering. If two tasks sharing the same plan
  (e.g. batch retry, preview+render, retry+new) ran concurrently, Task B's
  write could overwrite Task A's title before Task A's ffmpeg read it,
  producing videos with the wrong title text/style.

  Fix: stop mutating plan.config in the worker render path. Introduce
  task_config_override threading through render_plan / _do_render /
  UnifiedRenderService; URS reads config via _effective_config() which
  does a deep-copy plan.config merged with per-task override (title/bgm/
  export). _build_task_config_override builds the override dict (with
  field-name normalization: font_size→size, font_color→color) from
  task_info, and _download_voice_for_task handles voiceover download
  independently. plan.config is left as-authored by the API writeback;
  concurrent tasks no longer race on it.

  Files:
  - apps/worker/video_processing/unified_render_service.py: add
    override_config param + _cfg_section/_effective_config helpers;
    replace all render-time self.plan.config reads (title/subtitle/bgm/
    export/TTS) with _effective_config().
  - apps/worker/video_processing/render_adapter.py: render_plan/_do_render/
    _prepare_bgm accept task_config_override, merge into bgm/export reads,
    forward to URS.
  - apps/worker/worker_app/tasks/generation.py: replace
    _sync_task_config_to_plan (which wrote plan.config) with
    _build_task_config_override + _download_voice_for_task; pass
    task_config_override to render_plan.

Bug B (P0) — USER_PENDING_LIMIT=3 returns HTTP 429, blocks batch submit:
  Pre-check in create_generation_task rejected the whole batch with 429
  USER_QUEUE_FULL once user had ≥3 pending tasks; UI showed 'wait ~4 min'
  and prevented any more submissions. Users expect to submit a batch of
  6 and have them queue naturally (worker concurrency=2).

  Fix:
  - USER_PENDING_LIMIT 3→20 (soft cap for abuse protection, supports
    typical batch sizes of 6-10 with headroom).
  - Remove user-level 429 rejection from pre-check, retry endpoint,
    single-task confirm endpoint, and safe_enqueue (user path now logs
    a warning and continues to enqueue). Global GLOBAL_PENDING_LIMIT=20
    is retained as a hard 503 system-busy guard.
  - safe_enqueue post-enqueue check: user-over only logs, does not
    fail the task or raise.
  - Batch loop UserPendingLimitExceeded except branch now treats it as
    a successful enqueue (should not trigger in practice).

  Files:
  - apps/api/app/core/task_enqueue.py: limit 3→20; user checks log-only.
  - apps/api/app/api/routes/generation_tasks.py: remove user pre-check
    429; soften retry/single/batch except branches.

Stacked on PR #2095/#2097 which already fixed title baseline/width-scaling
and default-constant alignment with CPU video_filter_builder.

Tests:
- tests/unit/test_gpu_direct_pipeline.py: 30 passed
- tests/unit/test_render_adapter_pure.py: 18 passed
- ruff check/format clean; py_compile clean
This commit is contained in:
saas-backend-agent
2026-09-29 18:31:05 +08:00
parent dfc5e5a5b6
commit 96bcec5fdd
5 changed files with 197 additions and 116 deletions
@@ -158,6 +158,7 @@ class UnifiedRenderService:
bgm_path: str | None = None, # BGM 本地文件路径
voiceover_audio_path: str | None = None, # 配音素材库音频本地路径
clip_has_text: list[bool] | None = None, # 源视频片段是否有文字(来自 atom_clip.ai_tags.has_text)
override_config: dict | None = None, # Bug A: task 级 config 覆盖(title/bgm/export/subtitle),防并发竞态
):
self.plan = plan
self.clips = clips
@@ -170,6 +171,9 @@ class UnifiedRenderService:
self.asr_service = asr_service
self.bgm_path = bgm_path
self.voiceover_audio_path = voiceover_audio_path
# Bug A: task 级 config override(深拷贝),优先级高于 plan.config;
# 避免同 plan 多任务并发渲染时 _sync_task_config_to_plan 写 plan.config["title"] 互相覆盖。
self._override_config = dict(override_config) if isinstance(override_config, dict) else {}
# #1970:片段级文字检测(顺序与非 audio 的源视频片段一致);None 表示无可靠检测,保守不翻转
self._clip_has_text = clip_has_text
self._transition_engine = TransitionEngine(default_duration=transition_duration)
@@ -180,6 +184,28 @@ class UnifiedRenderService:
self._micro_plan_cache: Any = None
self._micro_plan_loaded = False
def _cfg_section(self, section: str) -> dict:
"""读取单个配置段:override_config 优先于 plan.config(Bug A 防并发竞态)。"""
base = dict((self.plan.config or {}).get(section, {}) or {})
override = self._override_config.get(section)
if isinstance(override, dict) and override:
base.update(override) # 浅合并,保留 base 中未被覆盖字段
return base
def _effective_config(self) -> dict:
"""读取完整 config:override_config 顶层段覆盖 plan.config(Bug A 防并发竞态)。"""
import copy
full = copy.deepcopy(self.plan.config or {})
for k, v in self._override_config.items():
if isinstance(v, dict):
sec = dict(full.get(k, {}) or {})
sec.update(v)
full[k] = sec
else:
full[k] = v
return full
# ── #1970 PR2 智能降重:片段级微变换 ───────────────────────────────────
def _dedup_enabled(self) -> bool:
"""读取 plan.config.dedup_enabled,缺省视为 True(向后兼容)。"""
@@ -431,7 +457,7 @@ class UnifiedRenderService:
has_audio = pass_through_has_audio
# 直通模式下也支持 BGM 混音:提取音频 → 混 BGM → 合并回视频
if self.bgm_path and pass_through_has_audio:
config = self.plan.config or {}
config = self._effective_config()
bgm_config = config.get("bgm", {}) or {}
if bgm_config.get("enabled", False):
ctx = RenderContext(work_dir=self.work_dir, plan_id=self.plan.id)
@@ -471,7 +497,7 @@ class UnifiedRenderService:
"[unified-render] pass-through BGM mix failed, skipping: plan_id=%s", self.plan.id
)
else:
config = self.plan.config or {}
config = self._effective_config()
bgm_config = config.get("bgm", {}) or {}
if not isinstance(bgm_config, dict):
bgm_config = {}
@@ -696,7 +722,7 @@ class UnifiedRenderService:
Returns:
ASS 文件路径,没有字幕时返回 None
"""
config = self.plan.config or {}
config = self._effective_config()
# #1901 统一读 "title",兼容老数据 "title_config"
title_cfg = config.get("title", {}) or {}
if not isinstance(title_cfg, dict) or not (title_cfg.get("text") or "").strip():
@@ -881,7 +907,7 @@ class UnifiedRenderService:
Returns:
是否成功添加了配音音轨
"""
config = self.plan.config or {}
config = self._effective_config()
tts_cfg = config.get("tts", {}) or {}
if not isinstance(tts_cfg, dict):
tts_cfg = {}
@@ -2275,7 +2301,7 @@ class UnifiedRenderService:
try:
from video_processing import gpu_direct_pipeline as gdp
cfg = self.plan.config or {}
cfg = self._effective_config()
video_layer = next(_lyr for _lyr in layers if _lyr.role not in ("audio",))
video_clips = [c for c in video_layer.clips if c.clip_type != "audio"]