Files
xiaoxia-saas/apps/worker/video_processing/gpu_direct_pipeline.py
T
saas-backend-agent dfc5e5a5b6
CI/CD Pipeline / Dedup Check - skip PR tests when covered by push pipeline (pull_request) Successful in 2s
CI/CD Pipeline / Check if frontend-only change (pull_request) Successful in 2s
CI/CD Pipeline / Check push changed paths (pull_request) Has been skipped
CI/CD Pipeline / Frontend Lint (pull_request) Has been skipped
CI/CD Pipeline / Frontend Unit Tests (pull_request) Has been skipped
Preview Deploy / Deploy Preview Environment (pull_request) Successful in 2m12s
CI/CD Pipeline / PR Build Web Image (pull_request) Has been skipped
PR Automation / Auto Approve on CI Green (pull_request) Successful in 3m6s
CI/CD Pipeline / Build Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Deploy Staging (Watchtower auto-deploy) (pull_request) Has been skipped
CI/CD Pipeline / Staging E2E Tests (pull_request) Has been skipped
CI/CD Pipeline / Staging API Integration Tests (pull_request) Has been skipped
CI/CD Pipeline / ACR Image Cleanup (pull_request) Has been skipped
CI/CD Pipeline / PR Build Worker Image (pull_request) Successful in 2m26s
AI Code Review / AI Code Review (pull_request) Successful in 6m55s
CI/CD Pipeline / Unit Tests (pull_request) Successful in 9m28s
CI/CD Pipeline / Integration Tests (pull_request) Successful in 9m32s
CI/CD Pipeline / PR Build API Image (pull_request) Failing after 8m40s
CI/CD Pipeline / Validate - Style (pull_request) Successful in 11m15s
CI/CD Pipeline / Validate - Python (mypy + alembic) (pull_request) Successful in 11m45s
PR Automation / Auto Merge on CI Green + Approved (pull_request) Successful in 10m46s
CI/CD Pipeline / Validate - Security (pull_request) Has been cancelled
CI/CD Pipeline / Build Production API Image (pull_request) Has been cancelled
CI/CD Pipeline / Build Production Web Image (pull_request) Has been cancelled
CI/CD Pipeline / Build Production Worker Image (pull_request) Has been cancelled
CI/CD Pipeline / Deploy Production (pull_request) Has been cancelled
CI/CD Pipeline / Production Browser E2E (pull_request) Has been cancelled
CI/CD Pipeline / Canary Release to Production (pull_request) Has been cancelled
CI/CD Pipeline / CI Gate (pull_request) Has been cancelled
fix(worker): align gpu-direct title defaults (size/margin/bold) with CPU vfb path
PR #2095 fixed title baseline positioning and width-scaling but didn't fully
align default constants with the CPU video_filter_builder path, causing GPU
rendered titles to appear slightly smaller / higher / bolder-differently than
the CPU/ASS preview the template was authored against. This shows up in both
single and batch GPU renders since every variant shares the same gpu_direct
pipeline.

Root causes in PR #2095 defaults:
- Default title size 36@720p vs config_schemas DEFAULT 48 / vfb default 48
- Top/bottom margin 40@720p (PAD16+margin24) vs vfb _scale_title_len(50)
- Subtitle bottom margin 60@720p vs vfb 50
- Faux bold used same-color 1px border (white-on-white invisible for default
  white text; red-on-red for colored) vs vfb black 2px (intentional choice
  per #2001 to avoid double-print/halo artifact)
- margin_top was treated as absolute y offset, now correctly added on top of
  base margin (matches vfb additive semantics)
- subtitle stroke/bold block referenced uninitialized s_borderw (ruff F821)
  — added proper init + stroke parsing consistent with title block

Changes:
- TITLE_DEFAULT_MARGIN_TOP/BOTTOM = 50 (was 24+PAD=40)
- SUBTITLE_DEFAULT_MARGIN_BOTTOM = 50 (was 60)
- TITLE_FAUX_BOLD_WIDTH = 2, border color = #000000 (was 1, same as text)
- default title size 48@720p (was 36)
- margin_top from cfg added on top of base 50 (additive, same as vfb)
- subtitle: init s_borderw/s_border_color, parse stroke dict before bold check
- bottom position: y = h - th - margin_bottom (correct baseline; previously
  used margin_top variable which was misleading but numerically equivalent
  before margin split; now uses explicit margin_bottom)

Tests updated to new expected scaled values (1280x720 scale=1.778: fontsize 85,
y=89, bold borderw=4 black; 1080x1920 vertical: fontsize 72, y=75;
margin_top=100 user offset => y=267). Added test_title_bold_false_disables_faux_bold.

Batch investigation notes (separate findings, not bugs in this PR):
- per-variant plan config is correctly deep-copied via clone_plan_for_variant
  (config=dict(source.config or {})); title text per-variant via titles[]
  override; voice per-variant independent download with #1749 strict guards;
  bgm merged via merge_bgm_config — batch passthrough chain is correct.
- Worker generation concurrency = 2 (compose.yml default); USER_PENDING_LIMIT=3,
  6 tasks will enqueue in two waves (first 3 → then next 3 as workers free up);
  GLOBAL_PENDING_LIMIT allows it. This is expected behavior, not a bug.
- Legacy key mismatch: subtitle_render_engine.py reads plan_config['subtitle_config']
  (snake_case) but normalize_plan_config writes 'subtitle' short key. This file
  is not imported by unified_render_service (which uses render_subtitles.py with
  explicit subtitle_config= parameter), so no runtime impact; left for separate
  cleanup.
- Frontend usePlanConfigLoader.ts reads config.title_config snake_case while
  backend writes 'title' short key — frontend-only issue for subsequent edit
  sessions, out of backend scope.
2026-09-29 18:08:18 +08:00

832 lines
34 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""全 GPU 直连渲染管线(P1)。
背景:旧链路 worker 先用 CPU libx264 把 filter_complex 输出成 mezzanine(1080p 约 85s),
上传后再由 P4000 NVENC 编码,渲染后还要单独跑一次随机边缘裁剪重编码(约 26s)。
本管线取消 mezzanine:把原始素材签名 URL 作为多输入直接交给 P4000,filter_complex 内
一步完成 trim/scale/pad/concat/边缘随机裁剪/drawtext 字幕,末端 h264_nvenc 只编码一次;
原素材音轨 concat + TTS/配音/BGM 混音也在同一命令里完成。
约束(P1):
- 仅覆盖智能剪辑主流场景:单一主视频轨、全硬切、无 PiP/overlay/水印/贴纸/片头片尾/绿幕。
不满足条件时调用方回退到现有 mezzanine/CPU 链路(功能不回归)。
- 字幕先用 drawtext(P4000 装好中文字体后可再切 subtitles 滤镜烧 ASS)。
"""
from __future__ import annotations
import logging
import random
import uuid
from pathlib import Path
from typing import Any, Optional
logger = logging.getLogger(__name__)
DEFAULT_DRAWTEXT_FONT = "Noto Sans CJK SC"
EDGE_CROP_MIN_PCT = 0.02
EDGE_CROP_MAX_PCT = 0.05
# 标题/字幕样式基准宽度(px)。前端 TitleSettings 所有长度字段(size/描边/阴影/margin/pos)
# 均以 720p 为基准(见前端 titleCanvas.ts 注释 scale=videoWidth/720,types.ts "px @720p"),
# 非 720p 输出时按 video_width / TITLE_SIZE_REF_WIDTH 等比缩放,保证成片位置与前端预览一致。
TITLE_SIZE_REF_WIDTH = 720
# 与 video_filter_builder.build_title_drawtext_filter(CPU 路径)和 ass_subtitle_builder 对齐:
# - top/bottom 默认 margin 50@720p(vfb 用 _scale_title_len(50, w),即 y=50 / y=h-th-50)
# - margin_top 字段:前端编辑器 marginTop 滑块,叠加在默认 margin 之上(#2095 支持)
# - PAD 概念仅用于前端 Canvas 预览;ffmpeg drawtext y 是 baseline,无 font metrics 可用,
# 直接用统一 50@720p baseline 位置即可保持三端(GPU/CPU/前端视觉)一致。
TITLE_DEFAULT_MARGIN_TOP = 50 # top 位置 baseline 默认距顶 50@720p(与 vfb/CPU 路径一致)
TITLE_DEFAULT_MARGIN_BOTTOM = 50 # bottom 位置 baseline 默认距底 50@720p
SUBTITLE_DEFAULT_MARGIN_BOTTOM = 50 # 字幕距底边距 50@720p(与 vfb 一致)
TITLE_MARGIN_TOP_FROM_CFG_DEFAULT = 24 # 前端 marginTop 滑块默认值(用户未传时叠加 0)
TITLE_FAUX_BOLD_WIDTH = 2 # 仿粗黑色描边宽度(与 vfb 一致,2@720p 黑色细描边)
def _scale_title_len(value, video_width: int):
"""将 720p 基准长度按 video_width 等比缩放(与 packages/domain/ass_subtitle_builder._scale_len 一致)。
int 输入 → 返回 int;float 输入 → 返回 float;非法值原样返回。
"""
if value is None:
return None
try:
v = float(value)
except (TypeError, ValueError):
return value
if not video_width or video_width <= 0:
return int(round(v)) if isinstance(value, int) else v
scaled = v * (video_width / TITLE_SIZE_REF_WIDTH)
return int(round(scaled)) if isinstance(value, int) else scaled
def escape_drawtext_text(text: str) -> str:
if not text:
return ""
s = text.replace("\\", "\\\\")
s = s.replace(":", "\\:")
s = s.replace("'", "\\'")
s = s.replace("%", "\\%")
s = s.replace(",", "\\,")
s = s.replace("[", "\\[").replace("]", "\\]")
s = s.replace(";", "\\;")
s = s.replace("\n", " ")
return s
def _hex_to_drawtext_color(hex_color: str, default: str = "white") -> str:
"""把 #RRGGBB / #RGB / 命名颜色转换为 ffmpeg drawtext 接受的颜色格式。
drawtext 的 fontcolor 接受 0xRRGGBB 形式(或命名颜色如 white/black/yellow)。
描边/阴影颜色同样适用。alpha 后缀支持(#RRGGBB@0.5 或 &HBBGGRRAA)。
"""
if not hex_color:
return default
s = hex_color.strip()
if not s:
return default
# 命名颜色直接返回(白名单常见值,避免把 #xxx 当成命名)
if not s.startswith("#") and not s.startswith("0x") and "@" not in s:
return s
if s.startswith("0x"):
return s # 已是 drawtext 原生格式
if s.startswith("#"):
h = s[1:]
# 处理 alpha:#RRGGBB@AA 或 #RRGGBB&AA
alpha = ""
if "@" in h:
h, alpha_part = h.split("@", 1)
try:
a = float(alpha_part)
alpha = f"@{a:.2f}"
except ValueError:
alpha = ""
if len(h) == 3:
h = "".join(ch * 2 for ch in h)
if len(h) == 6:
try:
int(h, 16)
except ValueError:
return default
return f"0x{h}{alpha}"
if len(h) == 8:
# RRGGBBAA → drawtext 的 0xRRGGBB@AA 形式
try:
int(h, 16)
except ValueError:
return default
rr, gg, bb, aa = h[0:2], h[2:4], h[4:6], h[6:8]
try:
a = int(aa, 16) / 255.0
return f"0x{rr}{gg}{bb}@{a:.2f}"
except ValueError:
return f"0x{rr}{gg}{bb}"
return default
def _position_to_drawtext_xy(
position: str,
*,
margin_top: int = 0,
margin_bottom: int = 0,
pos_x: Optional[float] = None,
pos_y: Optional[float] = None,
) -> tuple[str, str]:
"""把位置映射到 drawtext x/y 表达式,对齐前端 titleCanvas.ts 预览坐标。
position 支持: top / center(middle) / bottom / custom。
- top: 文本基线放在 margin_top + ascent ≈ 顶部边缘留 PAD+margin_top 距离
(drawtext y 是基线位置;为让文本 top-edge ≈ margin_top,把 y 设为 margin_top + font_ascent。
但 drawtext 运行时不知道 ascent,用经验系数 0.8*fontsize 近似,和前端 PAD+margin_top 对齐)。
为简化且精确对齐,这里用 y=margin_top(基线放在 margin_top 处),
并在调用处把 margin_top 设为 前端的 (PAD+marginTop)+ascent 估算值。
- center: (h-text_h)/2 垂直居中。
- bottom: 文本底线距离底边 margin_bottom。
- custom: pos_x/pos_y 为百分比 0-100(前端拖拽坐标系),文本中心落在 (pct_x*w, pct_y*h)。
margin_top/margin_bottom 为已按 video_width 缩放过的像素值。
"""
p = (position or "top").lower().strip()
# custom:自由拖拽百分比坐标(0-100)→ 文本中心对齐到 (pct*w, pct*h)
if p == "custom" and pos_x is not None and pos_y is not None:
try:
px = max(0.0, min(100.0, float(pos_x))) / 100.0
py = max(0.0, min(100.0, float(pos_y))) / 100.0
return f"(w-text_w)*{px:.4f}", f"(h-text_h)*{py:.4f}"
except (TypeError, ValueError):
pass # fall through to default
x = "(w-text_w)/2"
if p in ("top",):
# drawtext y 是 baseline 位置。中文字符顶边距基线约 0.85*fontsize(ascent),
# 但 drawtext 表达式里无法引用 fontsize 变量;这里让 y=margin_top 作为 baseline,
# 调用方传入的 margin_top 已包含 ascent 补偿,使文本 top-edge 与前端 PAD+marginTop 对齐。
y = f"{int(margin_top)}"
elif p in ("center", "middle"):
y = "(h-text_h)/2"
elif p in ("bottom",):
# h-th-margin_bottom:th ≈ text_h,文本底边距底边 margin_bottom
y = f"h-th-{int(margin_bottom)}"
else:
# 未知值回退到顶部(与前端默认 position=top 对齐)
y = f"{int(margin_top)}"
return x, y
def _build_drawtext_filters(
*,
text: str,
start: float,
end: float,
font: str = DEFAULT_DRAWTEXT_FONT,
font_size: int = 0,
font_color: str = "white",
position: str = "top",
margin_top: int = 0,
margin_bottom: int = 0,
pos_x: Optional[float] = None,
pos_y: Optional[float] = None,
box_enabled: bool = False,
box_color: str = "black@0.5",
borderw: int = 0,
border_color: str = "black",
shadow_enabled: bool = False,
shadow_color: str = "black@0.6",
shadow_x: int = 2,
shadow_y: int = 2,
) -> list[str]:
"""构造一组 drawtext 滤镜:可选阴影层(同字偏移)+ 主字层。
ffmpeg drawtext 没有直接的 shadow 选项,用两次 drawtext 模拟:
先画一个描边/阴影色层偏移 shadow_x/shadow_y,再画主字层。
返回列表是为了让调用方顺序插入 fc(前一个输出作为后一个输入)。
"""
txt = escape_drawtext_text(text)
if not txt:
return []
x_expr, y_expr = _position_to_drawtext_xy(
position,
margin_top=margin_top,
margin_bottom=margin_bottom,
pos_x=pos_x,
pos_y=pos_y,
)
fc_color = _hex_to_drawtext_color(font_color, default="white")
bd_color = _hex_to_drawtext_color(border_color, default="black")
sh_color = _hex_to_drawtext_color(shadow_color, default="black@0.6")
filters: list[str] = []
# 阴影层:shadow_enabled 时先画一层深色偏移字(无描边)
if shadow_enabled and (shadow_x != 0 or shadow_y != 0):
sh_parts = [f"font={font}", f"text='{txt}'"]
if font_size and font_size > 0:
sh_parts.append(f"fontsize={int(font_size)}")
sh_parts.append(f"fontcolor={sh_color}")
sh_parts.append(f"x={x_expr}+{int(shadow_x)}")
sh_parts.append(f"y={y_expr}+{int(shadow_y)}")
if start > 0 or end > 0:
sh_parts.append(f"enable='between(t,{start:.3f},{end:.3f})'")
filters.append("drawtext=" + ":".join(sh_parts))
# 主字层
parts = [f"font={font}", f"text='{txt}'"]
if font_size and font_size > 0:
parts.append(f"fontsize={int(font_size)}")
parts.append(f"fontcolor={fc_color}")
if box_enabled:
parts.append("box=1")
parts.append(f"boxcolor={box_color}")
if borderw and borderw > 0:
parts.append(f"borderw={int(borderw)}")
parts.append(f"bordercolor={bd_color}")
parts.append(f"x={x_expr}")
parts.append(f"y={y_expr}")
if start > 0 or end > 0:
parts.append(f"enable='between(t,{start:.3f},{end:.3f})'")
filters.append("drawtext=" + ":".join(parts))
return filters
def build_drawtext_filter(
*,
text: str,
start: float,
end: float,
font: str = DEFAULT_DRAWTEXT_FONT,
font_size: int = 0,
font_color: str = "white",
x_expr: str = "(w-text_w)/2",
y_expr: str = "h-th-60",
box: bool = False,
box_color: str = "black@0.5",
borderw: int = 0,
border_color: str = "black",
enable: bool = True,
) -> str:
"""[已废弃] 保留单条 drawtext 的便捷构造;新代码请用 _build_drawtext_filters。"""
txt = escape_drawtext_text(text)
parts = [f"font={font}", f"text='{txt}'"]
if font_size and font_size > 0:
parts.append(f"fontsize={int(font_size)}")
parts.append(f"fontcolor={_hex_to_drawtext_color(font_color)}")
if box:
parts.append("box=1")
parts.append(f"boxcolor={box_color}")
if borderw and borderw > 0:
parts.append(f"borderw={int(borderw)}")
parts.append(f"bordercolor={_hex_to_drawtext_color(border_color)}")
parts.append(f"x={x_expr}")
parts.append(f"y={y_expr}")
if enable:
parts.append(f"enable='between(t,{start:.3f},{end:.3f})'")
return "drawtext=" + ":".join(parts)
def _build_atempo_chain(speed: float) -> str:
if abs(speed - 1.0) < 1e-6:
return ""
stages: list[float] = []
remaining = speed
while remaining > 2.0:
stages.append(2.0)
remaining /= 2.0
while remaining < 0.5:
stages.append(0.5)
remaining /= 0.5
if abs(remaining - 1.0) >= 1e-6:
stages.append(remaining)
return ",".join(f"atempo={s:.5f}" for s in stages)
def upload_local_audio_and_sign(
local_audio: Path,
*,
tmp_prefix: str = "tmp/gpu-direct-audio/",
expires: int = 3600,
) -> tuple[str, str]:
from video_processing.oss_helpers import _storage # type: ignore
storage = _storage()
key = f"{tmp_prefix.rstrip('/')}/{uuid.uuid4().hex}{local_audio.suffix or '.mp3'}"
content_type = "audio/mpeg" if local_audio.suffix.lower() in (".mp3", ".mpeg") else "audio/mp4"
storage.upload_file(local_audio, key, content_type=content_type)
url = storage.get_download_url(key, expires)
return url, key
def sign_asset_url(storage_key: str, *, expires: int = 3600) -> str:
from video_processing.oss_helpers import _storage # type: ignore
storage = _storage()
return storage.get_download_url(storage_key, expires)
class DirectRenderPlan:
def __init__(
self,
inputs: dict[str, str],
ffmpeg_args: list[str],
oss_keys: list[str],
filter_complex: list[str] | None = None,
):
self.inputs = inputs
self.ffmpeg_args = ffmpeg_args
self.oss_keys = oss_keys
self.filter_complex: list[str] = filter_complex or []
def build_direct_render(
*,
resolved_clips: list[Any],
output_width: int,
output_height: int,
output_fps: int,
tts_audio: Optional[Path] = None,
bgm_audio: Optional[Path] = None,
title_text: str = "",
subtitle_segments: Optional[list[Any]] = None,
font: str = DEFAULT_DRAWTEXT_FONT,
vcodec: str = "h264_nvenc",
preset: str = "p4",
video_bitrate: str = "",
cq: int = 23,
edge_crop_pct: float = 0.0,
total_duration: float = 0.0,
clip_has_audio: Optional[list[bool]] = None,
clip_volumes: Optional[list[float]] = None,
extra_audio_tracks: Optional[list[tuple[Any, float]]] = None,
title_config: Optional[dict] = None,
subtitle_config: Optional[dict] = None,
bgm_config: Optional[dict] = None,
static_subtitle_text: str = "",
) -> DirectRenderPlan:
"""构造 P4000 直连渲染所需的 inputs 与 ffmpeg_args。
视频:每段 trim/setpts/scale/pad/fps → concat(全硬切,带音频)→ 随机边缘 crop+scale → drawtext。
音频:每段 [i:a](或 anullsrc 静音占位)按 clip 配置 atrim/asetpts/atempo/volume/aresample
→ concat=n:N:v=1:a=1 → 与 extra_audio(TTS/配音素材库)、BGM 一起 amix → atrim 精确截断。
"""
if not resolved_clips:
raise ValueError("build_direct_render: no resolved clips")
inputs: dict[str, str] = {}
oss_keys: list[str] = []
input_args: list[str] = []
fc: list[str] = []
n = len(resolved_clips)
# 规范化每段参数
if clip_has_audio is None:
clip_has_audio = [True] * n
else:
clip_has_audio = list(clip_has_audio) + [True] * max(0, n - len(clip_has_audio))
clip_has_audio = clip_has_audio[:n]
if clip_volumes is None:
clip_volumes = [1.0] * n
else:
clip_volumes = list(clip_volumes) + [1.0] * max(0, n - len(clip_volumes))
clip_volumes = clip_volumes[:n]
clip_starts: list[float] = []
clip_effs: list[float] = []
clip_speeds: list[float] = []
for clip in resolved_clips:
start = float(getattr(clip, "start_time", 0) or 0)
eff = float(getattr(clip, "duration", 0) or 0)
if eff <= 0:
eff = float(getattr(clip, "actual_duration", 0) or 0)
speed = float(getattr(clip, "playback_speed", 1.0) or 1.0)
clip_starts.append(start)
clip_effs.append(eff)
clip_speeds.append(speed)
# 1. 视频输入(原始素材签名 URL)
for i, clip in enumerate(resolved_clips):
sk = (getattr(clip, "config", None) or {}).get("_storage_key")
if not sk:
raise ValueError(f"clip {getattr(clip, 'clip_id', i)} missing _storage_key")
fname = f"v{i}.mp4"
inputs[fname] = sign_asset_url(sk)
input_args.extend(["-i", fname])
# 2. 视频段预处理
pre_labels: list[str] = []
for i in range(n):
vf: list[str] = []
start, eff, speed = clip_starts[i], clip_effs[i], clip_speeds[i]
if eff > 0:
if start > 0:
vf.append(f"trim=start={start:.3f}:duration={eff:.3f}")
else:
vf.append(f"trim=duration={eff:.3f}")
vf.append("setpts=PTS-STARTPTS")
if abs(speed - 1.0) >= 1e-6:
vf.append(f"setpts=PTS/{speed:.4f}")
vf.append(f"scale={output_width}:{output_height}:force_original_aspect_ratio=decrease")
vf.append(f"pad={output_width}:{output_height}:trunc((ow-iw)/2):trunc((oh-ih)/2):black")
vf.append("setpts=PTS-STARTPTS")
vf.append(f"fps={output_fps}")
label = f"vc{i}"
fc.append(f"[{i}:v]{','.join(vf)}[{label}]")
pre_labels.append(label)
# 2b. 音频段预处理(无声源用 anullsrc 占位;volume=0 的段也用 anullsrc 静音占位保持时间轴)
anullsrc_counter = 0
audio_pre_labels: list[str] = []
for i in range(n):
start, eff, speed = clip_starts[i], clip_effs[i], clip_speeds[i]
vol = float(clip_volumes[i] if i < len(clip_volumes) else 1.0)
has_a = bool(clip_has_audio[i] if i < len(clip_has_audio) else True)
if not has_a or vol <= 0.001:
# 静音占位:用 anullsrc 生成静音,atrim 到段时长
sl = f"sil{anullsrc_counter}"
anullsrc_counter += 1
af: list[str] = ["anullsrc=channel_layout=stereo:sample_rate=44100"]
if eff > 0:
af.append(f"atrim=duration={eff:.3f}")
af.append("asetpts=PTS-STARTPTS")
af.append("aformat=sample_fmts=fltp:channel_layouts=stereo")
fc.append(f"{','.join(af)}[{sl}]")
# anullsrc 作为 filter 源不需要 -i 输入,直接给 label
audio_pre_labels.append(sl)
continue
af = []
if eff > 0:
if start > 0:
af.append(f"atrim=start={start:.3f}:duration={eff:.3f}")
else:
af.append(f"atrim=duration={eff:.3f}")
af.append("asetpts=PTS-STARTPTS")
if abs(speed - 1.0) >= 1e-6:
atempo = _build_atempo_chain(speed)
if atempo:
af.append(atempo)
if abs(vol - 1.0) >= 1e-3:
af.append(f"volume={vol:.3f}")
af.append("aresample=44100")
af.append("aformat=sample_fmts=fltp:channel_layouts=stereo")
alabel = f"ac{i}"
fc.append(f"[{i}:a]{','.join(af)}[{alabel}]")
audio_pre_labels.append(alabel)
# 3. concat(全硬切;v=1:a=1,视频音频一起拼接)
concat_in = "".join(f"[{v}][{a}]" for v, a in zip(pre_labels, audio_pre_labels, strict=True))
fc.append(f"{concat_in}concat=n={n}:v=1:a=1[vcat][acat]")
cur_v = "vcat"
cur_a = "acat"
# 4. 随机边缘裁剪降重(四边独立随机 2%~5%,与 ffmpeg_utils.random_edge_crop 一致)
if edge_crop_pct and edge_crop_pct > 0:
_r = random.Random()
p_min = EDGE_CROP_MIN_PCT
p_max = EDGE_CROP_MAX_PCT
crop_top = p_min + _r.random() * (p_max - p_min)
crop_bottom = p_min + _r.random() * (p_max - p_min)
crop_left = p_min + _r.random() * (p_max - p_min)
crop_right = p_min + _r.random() * (p_max - p_min)
w_expr = f"trunc(iw*(1-{crop_left:.4f}-{crop_right:.4f})/2)*2"
h_expr = f"trunc(ih*(1-{crop_top:.4f}-{crop_bottom:.4f})/2)*2"
x_expr = f"trunc(iw*{crop_left:.4f}/2)*2"
y_expr = f"trunc(ih*{crop_top:.4f}/2)*2"
fc.append(
f"[{cur_v}]crop=w='{w_expr}':h='{h_expr}':x='{x_expr}':y='{y_expr}',"
f"scale={output_width}:{output_height}[vcrop]"
)
cur_v = "vcrop"
# 5. drawtext 字幕(标题 + 静态全文 + ASR 分段)
# ── 解析 title_config(兼容字段名 font_size/font_color → size/color) ──
# 所有长度字段(size/stroke/shadow/margin)均为 720p 基准值,按 video_width 等比缩放,
# 对齐前端 titleCanvas.ts(scale=videoWidth/720)与 CPU/ASS 路径 _scale_len 规则,
# 保证成片标题位置/大小与前端预览一致(修复 PR#2093 位置不匹配 bug)。
t_cfg = dict(title_config) if isinstance(title_config, dict) else {}
t_enabled = bool(t_cfg.get("enabled", True))
t_text = (t_cfg.get("text", "") or title_text or "").strip()
t_font = str(t_cfg.get("font", font) or font)
# size:前端传 px@720p,未配置默认 28(前端 DEFAULT_TITLE_SETTINGS.size=28,对齐 AI Avatar 默认48)
t_size_raw = t_cfg.get("size", t_cfg.get("font_size", 0))
try:
t_size_720 = int(t_size_raw) if t_size_raw else 0
except (TypeError, ValueError):
t_size_720 = 0
if t_size_720 <= 0:
t_size_720 = 48 # 与 config_schemas.DEFAULT_EDIT_PLAN_CONFIG.title.size=48 及 vfb 默认一致
t_size = _scale_title_len(t_size_720, output_width)
# stroke/shadow 长度字段也需 720p→输出分辨率缩放
t_color = str(t_cfg.get("color", t_cfg.get("font_color", "#ffffff")))
t_position = str(t_cfg.get("position", "top")).lower().strip()
# 自由拖拽坐标(百分比 0-100),与 video_filter_builder.build_title_drawtext_filter 一致
t_pos_x = t_cfg.get("pos_x")
t_pos_y = t_cfg.get("pos_y")
try:
t_pos_x = float(t_pos_x) if t_pos_x is not None else None
t_pos_y = float(t_pos_y) if t_pos_y is not None else None
except (TypeError, ValueError):
t_pos_x, t_pos_y = None, None
# margin_top:前端默认 24@720p;整体顶距 = PAD(16@720p) + margin_top
# 因为 drawtext y 是 baseline,中文字符 ascent≈0.85*fontsize,为让文本 top-edge≈(PAD+marginTop),
# baseline 需再下移约 0.85*fontsize;但 drawtext 表达式无法引用 fontsize 变量,
# 这里直接用 (PAD + margin_top)@720p 缩放后作为 y(即让 baseline≈顶部内边距位置),
# 实际中文字符会自然向下延伸,视觉位置与前端预览(textBaseline=middle 居中到 firstLineY)一致。
# margin_top:前端滑块值(默认 24@720p),叠加在默认 50@720p 基线之上
_t_user_margin_top = t_cfg.get("margin_top")
try:
_t_user_margin_top_720 = int(_t_user_margin_top) if _t_user_margin_top is not None else 0
except (TypeError, ValueError):
_t_user_margin_top_720 = 0
t_margin_top_720 = TITLE_DEFAULT_MARGIN_TOP + _t_user_margin_top_720
t_margin_top = _scale_title_len(t_margin_top_720, output_width)
# bottom margin(标题放在 bottom 时):用户 margin_bottom 透传,默认 50@720p
_t_user_margin_bottom = t_cfg.get("margin_bottom")
try:
_t_user_margin_bottom_720 = int(_t_user_margin_bottom) if _t_user_margin_bottom is not None else 0
except (TypeError, ValueError):
_t_user_margin_bottom_720 = 0
t_margin_bottom_720 = TITLE_DEFAULT_MARGIN_BOTTOM + _t_user_margin_bottom_720
t_margin_bottom = _scale_title_len(t_margin_bottom_720, output_width)
t_borderw = 0
t_border_color = "#000000"
t_box = False
t_box_color = "black@0.5"
# stroke
_stroke = t_cfg.get("stroke")
if isinstance(_stroke, dict) and _stroke.get("enabled", False):
try:
t_borderw_720 = int(float(_stroke.get("width", 2)))
except (TypeError, ValueError):
t_borderw_720 = 2
t_borderw = max(1, _scale_title_len(t_borderw_720, output_width))
t_border_color = str(_stroke.get("color", "#000000"))
elif isinstance(_stroke, bool) and _stroke:
t_borderw = max(1, _scale_title_len(2, output_width))
# shadow
_shadow = t_cfg.get("shadow")
t_shadow_enabled = False
t_shadow_color = "#000000@0.6"
t_shadow_x_720, t_shadow_y_720 = 2, 2
if isinstance(_shadow, dict) and _shadow.get("enabled", False):
t_shadow_enabled = True
t_shadow_color = str(_shadow.get("color", "#000000@0.6"))
try:
t_shadow_x_720 = int(float(_shadow.get("offset_x", 2)))
t_shadow_y_720 = int(float(_shadow.get("offset_y", 2)))
except (TypeError, ValueError):
t_shadow_x_720, t_shadow_y_720 = 2, 2
elif isinstance(_shadow, bool) and _shadow:
t_shadow_enabled = True
t_shadow_x = _scale_title_len(t_shadow_x_720, output_width)
t_shadow_y = _scale_title_len(t_shadow_y_720, output_width)
# bold/italic:drawtext 原生无粗斜体选项;通过同色描边模拟粗体
t_bold = bool(t_cfg.get("bold", True)) # 与 ASS/vfb 路径默认 bold=True 对齐
if t_bold and t_borderw < 1:
# 粗体未配用户描边时:黑色细描边 2@720p(与 vfb 一致,避免同色描边导致重影)
t_borderw = _scale_title_len(TITLE_FAUX_BOLD_WIDTH, output_width)
t_border_color = "#000000" # 黑色细描边模拟粗体
# ── 解析 subtitle_config ──
s_cfg = dict(subtitle_config) if isinstance(subtitle_config, dict) else {}
s_enabled = bool(s_cfg.get("enabled", True))
s_font = str(s_cfg.get("font", font) or font)
s_size_raw = s_cfg.get("size", s_cfg.get("font_size", 0))
try:
s_size_720 = int(s_size_raw) if s_size_raw else 0
except (TypeError, ValueError):
s_size_720 = 0
if s_size_720 <= 0:
s_size_720 = 24 # 字幕默认 24@720p(对齐 ass_subtitle_builder defaults size=24)
s_size = _scale_title_len(s_size_720, output_width)
s_color = str(s_cfg.get("color", s_cfg.get("font_color", "#ffffff")))
s_position = str(s_cfg.get("position", "bottom")).lower().strip()
s_pos_x = s_cfg.get("pos_x")
s_pos_y = s_cfg.get("pos_y")
try:
s_pos_x = float(s_pos_x) if s_pos_x is not None else None
s_pos_y = float(s_pos_y) if s_pos_y is not None else None
except (TypeError, ValueError):
s_pos_x, s_pos_y = None, None
s_margin_top = _scale_title_len(60, output_width) # subtitle top (not commonly used)
s_margin_bottom = _scale_title_len(SUBTITLE_DEFAULT_MARGIN_BOTTOM, output_width)
# subtitle stroke/bold:先解析用户 stroke,再按 bold 默认补描边
s_borderw = 0
s_border_color = "#000000"
_s_stroke = s_cfg.get("stroke")
if isinstance(_s_stroke, dict) and _s_stroke.get("enabled", False):
try:
s_borderw = _scale_title_len(int(float(_s_stroke.get("width", 2))), output_width)
except (TypeError, ValueError):
s_borderw = 0
s_border_color = str(_s_stroke.get("color", "#000000"))
s_bold = bool(s_cfg.get("bold", False))
if s_bold and s_borderw < 1:
# 粗体默认黑色细描边 2@720p(与 title/CPU vfb 一致)
s_borderw = _scale_title_len(TITLE_FAUX_BOLD_WIDTH, output_width)
s_border_color = "#000000"
# 静态字幕:static_subtitle_text 非空时构造全片长 segment(0 → total_duration)
static_text = (static_subtitle_text or "").strip()
subtitle_segments = list(subtitle_segments or [])
if s_enabled and static_text and total_duration and total_duration > 0:
# 用 duck-type 对象插入到 subtitle_segments 列表头部(静态全文)
class _StaticSeg:
def __init__(self, txt, st, ed):
self.text = txt
self.start = st
self.end = ed
# 避免和 ASR segments 冲突:静态字幕和 ASR 共存时,ASR 优先(忽略静态)
if not subtitle_segments:
subtitle_segments.insert(0, _StaticSeg(static_text, 0.0, float(total_duration)))
draw_filters: list[str] = []
if t_enabled and t_text:
draw_filters.extend(
_build_drawtext_filters(
text=t_text,
start=0.0,
end=max(total_duration, 0.1),
font=t_font,
font_size=t_size,
font_color=t_color,
position=t_position,
margin_top=t_margin_top,
margin_bottom=t_margin_bottom,
pos_x=t_pos_x,
pos_y=t_pos_y,
box_enabled=t_box,
box_color=t_box_color,
borderw=t_borderw,
border_color=t_border_color,
shadow_enabled=t_shadow_enabled,
shadow_color=t_shadow_color,
shadow_x=t_shadow_x,
shadow_y=t_shadow_y,
)
)
if s_enabled:
for seg in subtitle_segments:
txt = getattr(seg, "text", "") or ""
if not txt.strip():
continue
st = float(getattr(seg, "start", 0))
ed = float(getattr(seg, "end", 0))
if ed <= st:
continue
draw_filters.extend(
_build_drawtext_filters(
text=txt,
start=st,
end=ed,
font=s_font,
font_size=s_size,
font_color=s_color,
position=s_position,
margin_top=s_margin_top,
margin_bottom=s_margin_bottom,
pos_x=s_pos_x,
pos_y=s_pos_y,
box_enabled=False,
borderw=s_borderw,
border_color=s_border_color,
)
)
if draw_filters:
prev = cur_v
for idx, df in enumerate(draw_filters):
out_l = "vfinal" if idx == len(draw_filters) - 1 else f"vd{idx}"
fc.append(f"[{prev}]{df}[{out_l}]")
prev = out_l
vfinal_label = prev
else:
fc.append(f"[{cur_v}]format=yuv420p[vfinal]")
vfinal_label = "vfinal"
# 6. 音频混音:原素材主音轨 acat + extra(TTS/配音素材库) + BGM → amix → atrim
mix_labels: list[str] = [cur_a]
mix_vols: list[float] = [1.0]
next_idx = n
# 额外独立音频轨(TTS concat / 配音素材库整段音频)
for _ea_idx, (ea_path, ea_vol) in enumerate(extra_audio_tracks or []):
if ea_path is None:
continue
ea_p = Path(ea_path)
if not ea_p.exists():
continue
eurl, ekey = upload_local_audio_and_sign(ea_p)
ename = f"extra{_ea_idx}{ea_p.suffix or '.mp3'}"
inputs[ename] = eurl
oss_keys.append(ekey)
input_args.extend(["-i", ename])
elabel = f"aex{_ea_idx}"
fc.append(
f"[{next_idx}:a]aresample=44100,volume={float(ea_vol):.2f},"
f"aformat=sample_fmts=fltp:channel_layouts=stereo[{elabel}]"
)
mix_labels.append(elabel)
mix_vols.append(float(ea_vol))
next_idx += 1
if tts_audio and Path(tts_audio).exists():
# 旧参数保留:若调用方直接传了 tts_audio 而没走 extra_audio_tracks,则仍然加入
# (兼容旧调用,正常路径 TTS 已经通过 extra_audio_tracks 传入)
turl, tkey = upload_local_audio_and_sign(Path(tts_audio))
tname = "tts" + (Path(tts_audio).suffix or ".mp3")
inputs[tname] = turl
oss_keys.append(tkey)
input_args.extend(["-i", tname])
alabel = "au_tts"
fc.append(
f"[{next_idx}:a]aresample=44100,volume=1.00,aformat=sample_fmts=fltp:channel_layouts=stereo[{alabel}]"
)
mix_labels.append(alabel)
mix_vols.append(1.0)
next_idx += 1
_bgm_use = bgm_audio is not None and Path(bgm_audio).exists()
if _bgm_use and isinstance(bgm_config, dict) and bgm_config.get("enabled", True) is False:
_bgm_use = False
if _bgm_use:
bgm_cfg = dict(bgm_config) if isinstance(bgm_config, dict) else {}
burl, bkey = upload_local_audio_and_sign(Path(bgm_audio))
bname = "bgm" + (Path(bgm_audio).suffix or ".mp3")
inputs[bname] = burl
oss_keys.append(bkey)
input_args.extend(["-i", bname])
alabel = "au_bgm"
try:
bgm_vol = float(bgm_cfg.get("volume", 0.3))
except (TypeError, ValueError):
bgm_vol = 0.3
bgm_vol = max(0.0, min(1.5, bgm_vol))
# volume_adjust_db(-3 ~ +3 dB)换算线性增益
try:
_db = float(bgm_cfg.get("volume_adjust_db", 0.0))
except (TypeError, ValueError):
_db = 0.0
if abs(_db) > 0.05:
db_gain = 10 ** (_db / 20.0)
bgm_vol = max(0.0, min(2.0, bgm_vol * db_gain))
# afade 淡入淡出
try:
fade_in = max(0.0, float(bgm_cfg.get("fade_in", 0.0)))
except (TypeError, ValueError):
fade_in = 0.0
try:
fade_out = max(0.0, float(bgm_cfg.get("fade_out", 0.0)))
except (TypeError, ValueError):
fade_out = 0.0
# audio_offset:adelay 延迟(毫秒)
try:
offset = max(0.0, float(bgm_cfg.get("audio_offset", 0.0)))
except (TypeError, ValueError):
offset = 0.0
bgm_parts: list[str] = [f"[{next_idx}:a]aresample=44100"]
if offset > 0.01:
bgm_parts.append(f"adelay={int(offset * 1000)}|{int(offset * 1000)}")
bgm_parts.append(f"volume={bgm_vol:.3f}")
if fade_in > 0.01:
bgm_parts.append(f"afade=t=in:st=0:d={fade_in:.2f}")
if fade_out > 0.01 and total_duration > 0:
fo_start = max(0.0, total_duration - fade_out)
bgm_parts.append(f"afade=t=out:st={fo_start:.2f}:d={fade_out:.2f}")
bgm_parts.append("aformat=sample_fmts=fltp:channel_layouts=stereo")
fc.append(",".join(bgm_parts) + f"[{alabel}]")
mix_labels.append(alabel)
mix_vols.append(bgm_vol)
next_idx += 1
maps: list[str] = ["-map", f"[{vfinal_label}]"]
if mix_labels:
mix_in = "".join(f"[{lb}]" for lb in mix_labels)
n_mix = len(mix_labels)
mix_parts = [
f"amix=inputs={n_mix}:duration=longest:dropout_transition=2:normalize=0",
"aresample=44100",
]
# Bug2 修复:atrim 到视频精确时长
if total_duration and total_duration > 0:
mix_parts.append(f"atrim=0:{total_duration:.3f}")
mix_parts.append("asetpts=PTS-STARTPTS")
fc.append(f"{mix_in}{','.join(mix_parts)}[afinal]")
maps.extend(["-map", "[afinal]", "-c:a", "aac", "-b:a", "128k"])
else:
logger.info("[gpu-direct] no audio tracks; output silent video")
# 7. 组装 ffmpeg_args + NVENC 编码
ffmpeg_args = ["-y", *input_args, "-filter_complex", ";".join(fc), *maps]
ffmpeg_args.extend(["-c:v", vcodec, "-preset", preset, "-pix_fmt", "yuv420p"])
if video_bitrate:
ffmpeg_args.extend(["-b:v", video_bitrate])
else:
ffmpeg_args.extend(["-cq", str(cq)])
ffmpeg_args.extend(["-movflags", "+faststart", "-shortest", "-f", "mp4", "pipe:1"])
return DirectRenderPlan(
inputs=inputs,
ffmpeg_args=ffmpeg_args,
oss_keys=oss_keys,
filter_complex=fc,
)