Files
xiaoxia-saas/apps/worker/video_processing/gpu_direct_pipeline.py
T
Coze Agent 11c554e43a
CI/CD Pipeline / Check push changed paths (pull_request) Has been skipped
CI/CD Pipeline / Dedup Check - skip PR tests when covered by push pipeline (pull_request) Successful in 1s
CI/CD Pipeline / Check if frontend-only change (pull_request) Successful in 1s
CI/CD Pipeline / Frontend Lint (pull_request) Has been skipped
CI/CD Pipeline / Frontend Unit Tests (pull_request) Has been skipped
CI/CD Pipeline / PR Build Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Deploy Staging (Watchtower auto-deploy) (pull_request) Has been skipped
CI/CD Pipeline / Staging E2E Tests (pull_request) Has been skipped
CI/CD Pipeline / Staging API Integration Tests (pull_request) Has been skipped
CI/CD Pipeline / ACR Image Cleanup (pull_request) Has been skipped
CI/CD Pipeline / PR Build API Image (pull_request) Successful in 51s
Preview Deploy / Deploy Preview Environment (pull_request) Successful in 1m49s
CI/CD Pipeline / PR Build Worker Image (pull_request) Successful in 2m32s
PR Automation / Auto Approve on CI Green (pull_request) Successful in 3m6s
CI/CD Pipeline / Integration Tests (pull_request) Successful in 3m18s
CI/CD Pipeline / Validate - Python (mypy + alembic) (pull_request) Successful in 3m38s
CI/CD Pipeline / Validate - Style (pull_request) Successful in 3m45s
AI Code Review / AI Code Review (pull_request) Successful in 6m41s
CI/CD Pipeline / Unit Tests (pull_request) Successful in 7m15s
CI/CD Pipeline / Validate - Security (pull_request) Successful in 9m14s
CI/CD Pipeline / Build Production Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Production Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Build Production API Image (pull_request) Has been skipped
CI/CD Pipeline / Deploy Production (pull_request) Has been skipped
CI/CD Pipeline / Canary Release to Production (pull_request) Has been skipped
CI/CD Pipeline / Production Browser E2E (pull_request) Has been skipped
CI/CD Pipeline / CI Gate (pull_request) Successful in 1s
PR Automation / Auto Merge on CI Green + Approved (pull_request) Successful in 6m48s
ACR Cleanup / ACR Image Cleanup (pull_request_target) Successful in 31s
Preview Cleanup / Cleanup Preview Environment (pull_request) Successful in 1m1s
fix(gpu-direct): passthrough template title/subtitle/BGM config
Problem: GPU direct render MVP hardcoded title/subtitle/BGM styles and
ignored template config fields:
- P0 static subtitle text (subtitle.text) was never rendered
- P0 title styles (font/size/color/position/stroke/shadow/bold) all hardcoded
- P1 BGM volume hardcoded at 0.35, no fade/offset/adjust-db
- P1 ASR subtitle styles (color/size/position/font) ignored
- P2 extra_audio_tracks (TTS/voiceover) volume hardcoded at 1.0

Changes:
1. build_direct_render() now accepts title_config/subtitle_config/
   bgm_config dicts + static_subtitle_text.
2. Added helpers _hex_to_drawtext_color (#RGB/#RRGGBB/#RRGGBBAA/named),
   _position_to_drawtext_xy (top/center/bottom), _build_drawtext_filters
   (shadow layer + main layer, stroke via borderw/bordercolor, bold
   simulated via same-color borderw).
3. Title parses font/size/color/position/stroke/shadow/bold with safe
   defaults; legacy title_text still used as fallback.
4. Static subtitle_text renders as full-span (0→total_duration) segment;
   when ASR segments are present they take priority.
5. BGM reads volume (default 0.3), volume_adjust_db (linear gain),
   fade_in/fade_out (afade), audio_offset (adelay ms|ms), and honors
   enabled=false even if path supplied. Volume clamped to [0,1.5].
6. _try_gpu_direct() now passes cfg['title']/['subtitle']/['bgm'] whole
   dicts through, computes static_subtitle_text for non-auto subs,
   and sources extra_audio_tracks volume from audio_tracks_config
   (voiceover tracks) instead of hardcoding 1.0.
7. DirectRenderPlan exposes filter_complex for testing.
8. Added 24 unit tests covering: default backward-compat, color hex/BGR
   conversion (short/long/alpha/invalid), position (top/center/bottom),
   stroke, shadow (two drawtext layers), font override, enabled flag,
   static subtitle full-span, ASR style passthrough, BGM
   volume/fade/offset/adjust-db/enabled=false, extra audio volume,
   bold simulation, ASR vs static priority, volume clamping.

Fixes template-style rendering not taking effect in GPU pipeline.
2026-09-29 14:44:11 +08:00

695 lines
26 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""全 GPU 直连渲染管线(P1)。
背景:旧链路 worker 先用 CPU libx264 把 filter_complex 输出成 mezzanine(1080p 约 85s),
上传后再由 P4000 NVENC 编码,渲染后还要单独跑一次随机边缘裁剪重编码(约 26s)。
本管线取消 mezzanine:把原始素材签名 URL 作为多输入直接交给 P4000,filter_complex 内
一步完成 trim/scale/pad/concat/边缘随机裁剪/drawtext 字幕,末端 h264_nvenc 只编码一次;
原素材音轨 concat + TTS/配音/BGM 混音也在同一命令里完成。
约束(P1):
- 仅覆盖智能剪辑主流场景:单一主视频轨、全硬切、无 PiP/overlay/水印/贴纸/片头片尾/绿幕。
不满足条件时调用方回退到现有 mezzanine/CPU 链路(功能不回归)。
- 字幕先用 drawtext(P4000 装好中文字体后可再切 subtitles 滤镜烧 ASS)。
"""
from __future__ import annotations
import logging
import random
import uuid
from pathlib import Path
from typing import Any, Optional
logger = logging.getLogger(__name__)
DEFAULT_DRAWTEXT_FONT = "Noto Sans CJK SC"
EDGE_CROP_MIN_PCT = 0.02
EDGE_CROP_MAX_PCT = 0.05
def escape_drawtext_text(text: str) -> str:
if not text:
return ""
s = text.replace("\\", "\\\\")
s = s.replace(":", "\\:")
s = s.replace("'", "\\'")
s = s.replace("%", "\\%")
s = s.replace(",", "\\,")
s = s.replace("[", "\\[").replace("]", "\\]")
s = s.replace(";", "\\;")
s = s.replace("\n", " ")
return s
def _hex_to_drawtext_color(hex_color: str, default: str = "white") -> str:
"""把 #RRGGBB / #RGB / 命名颜色转换为 ffmpeg drawtext 接受的颜色格式。
drawtext 的 fontcolor 接受 0xRRGGBB 形式(或命名颜色如 white/black/yellow)。
描边/阴影颜色同样适用。alpha 后缀支持(#RRGGBB@0.5 或 &HBBGGRRAA)。
"""
if not hex_color:
return default
s = hex_color.strip()
if not s:
return default
# 命名颜色直接返回(白名单常见值,避免把 #xxx 当成命名)
if not s.startswith("#") and not s.startswith("0x") and "@" not in s:
return s
if s.startswith("0x"):
return s # 已是 drawtext 原生格式
if s.startswith("#"):
h = s[1:]
# 处理 alpha:#RRGGBB@AA 或 #RRGGBB&AA
alpha = ""
if "@" in h:
h, alpha_part = h.split("@", 1)
try:
a = float(alpha_part)
alpha = f"@{a:.2f}"
except ValueError:
alpha = ""
if len(h) == 3:
h = "".join(ch * 2 for ch in h)
if len(h) == 6:
try:
int(h, 16)
except ValueError:
return default
return f"0x{h}{alpha}"
if len(h) == 8:
# RRGGBBAA → drawtext 的 0xRRGGBB@AA 形式
try:
int(h, 16)
except ValueError:
return default
rr, gg, bb, aa = h[0:2], h[2:4], h[4:6], h[6:8]
try:
a = int(aa, 16) / 255.0
return f"0x{rr}{gg}{bb}@{a:.2f}"
except ValueError:
return f"0x{rr}{gg}{bb}"
return default
def _position_to_drawtext_xy(position: str, *, margin: int = 40) -> tuple[str, str]:
"""把 top/center/bottom 位置映射到 drawtext x/y 表达式。
返回 (x_expr, y_expr)。默认居中对齐。margin 为距离视频边缘的像素。
"""
p = (position or "bottom").lower().strip()
x = "(w-text_w)/2"
if p in ("top",):
y = f"{margin}"
elif p in ("center", "middle"):
y = "(h-text_h)/2"
elif p in ("bottom",):
y = f"h-th-{margin}"
else:
# 未知值回退到底部
y = f"h-th-{margin}"
return x, y
def _build_drawtext_filters(
*,
text: str,
start: float,
end: float,
font: str = DEFAULT_DRAWTEXT_FONT,
font_size: int = 0,
font_color: str = "white",
position: str = "bottom",
margin: int = 40,
box_enabled: bool = False,
box_color: str = "black@0.5",
borderw: int = 0,
border_color: str = "black",
shadow_enabled: bool = False,
shadow_color: str = "black@0.6",
shadow_x: int = 2,
shadow_y: int = 2,
) -> list[str]:
"""构造一组 drawtext 滤镜:可选阴影层(同字偏移)+ 主字层。
ffmpeg drawtext 没有直接的 shadow 选项,用两次 drawtext 模拟:
先画一个描边/阴影色层偏移 shadow_x/shadow_y,再画主字层。
返回列表是为了让调用方顺序插入 fc(前一个输出作为后一个输入)。
"""
txt = escape_drawtext_text(text)
if not txt:
return []
x_expr, y_expr = _position_to_drawtext_xy(position, margin=margin)
fc_color = _hex_to_drawtext_color(font_color, default="white")
bd_color = _hex_to_drawtext_color(border_color, default="black")
sh_color = _hex_to_drawtext_color(shadow_color, default="black@0.6")
filters: list[str] = []
# 阴影层:shadow_enabled 时先画一层深色偏移字(无描边)
if shadow_enabled and (shadow_x != 0 or shadow_y != 0):
sh_parts = [f"font={font}", f"text='{txt}'"]
if font_size and font_size > 0:
sh_parts.append(f"fontsize={int(font_size)}")
sh_parts.append(f"fontcolor={sh_color}")
sh_parts.append(f"x={x_expr}+{int(shadow_x)}")
sh_parts.append(f"y={y_expr}+{int(shadow_y)}")
if start > 0 or end > 0:
sh_parts.append(f"enable='between(t,{start:.3f},{end:.3f})'")
filters.append("drawtext=" + ":".join(sh_parts))
# 主字层
parts = [f"font={font}", f"text='{txt}'"]
if font_size and font_size > 0:
parts.append(f"fontsize={int(font_size)}")
parts.append(f"fontcolor={fc_color}")
if box_enabled:
parts.append("box=1")
parts.append(f"boxcolor={box_color}")
if borderw and borderw > 0:
parts.append(f"borderw={int(borderw)}")
parts.append(f"bordercolor={bd_color}")
parts.append(f"x={x_expr}")
parts.append(f"y={y_expr}")
if start > 0 or end > 0:
parts.append(f"enable='between(t,{start:.3f},{end:.3f})'")
filters.append("drawtext=" + ":".join(parts))
return filters
def build_drawtext_filter(
*,
text: str,
start: float,
end: float,
font: str = DEFAULT_DRAWTEXT_FONT,
font_size: int = 0,
font_color: str = "white",
x_expr: str = "(w-text_w)/2",
y_expr: str = "h-th-60",
box: bool = False,
box_color: str = "black@0.5",
borderw: int = 0,
border_color: str = "black",
enable: bool = True,
) -> str:
"""[已废弃] 保留单条 drawtext 的便捷构造;新代码请用 _build_drawtext_filters。"""
txt = escape_drawtext_text(text)
parts = [f"font={font}", f"text='{txt}'"]
if font_size and font_size > 0:
parts.append(f"fontsize={int(font_size)}")
parts.append(f"fontcolor={_hex_to_drawtext_color(font_color)}")
if box:
parts.append("box=1")
parts.append(f"boxcolor={box_color}")
if borderw and borderw > 0:
parts.append(f"borderw={int(borderw)}")
parts.append(f"bordercolor={_hex_to_drawtext_color(border_color)}")
parts.append(f"x={x_expr}")
parts.append(f"y={y_expr}")
if enable:
parts.append(f"enable='between(t,{start:.3f},{end:.3f})'")
return "drawtext=" + ":".join(parts)
def _build_atempo_chain(speed: float) -> str:
if abs(speed - 1.0) < 1e-6:
return ""
stages: list[float] = []
remaining = speed
while remaining > 2.0:
stages.append(2.0)
remaining /= 2.0
while remaining < 0.5:
stages.append(0.5)
remaining /= 0.5
if abs(remaining - 1.0) >= 1e-6:
stages.append(remaining)
return ",".join(f"atempo={s:.5f}" for s in stages)
def upload_local_audio_and_sign(
local_audio: Path,
*,
tmp_prefix: str = "tmp/gpu-direct-audio/",
expires: int = 3600,
) -> tuple[str, str]:
from video_processing.oss_helpers import _storage # type: ignore
storage = _storage()
key = f"{tmp_prefix.rstrip('/')}/{uuid.uuid4().hex}{local_audio.suffix or '.mp3'}"
content_type = "audio/mpeg" if local_audio.suffix.lower() in (".mp3", ".mpeg") else "audio/mp4"
storage.upload_file(local_audio, key, content_type=content_type)
url = storage.get_download_url(key, expires)
return url, key
def sign_asset_url(storage_key: str, *, expires: int = 3600) -> str:
from video_processing.oss_helpers import _storage # type: ignore
storage = _storage()
return storage.get_download_url(storage_key, expires)
class DirectRenderPlan:
def __init__(
self,
inputs: dict[str, str],
ffmpeg_args: list[str],
oss_keys: list[str],
filter_complex: list[str] | None = None,
):
self.inputs = inputs
self.ffmpeg_args = ffmpeg_args
self.oss_keys = oss_keys
self.filter_complex: list[str] = filter_complex or []
def build_direct_render(
*,
resolved_clips: list[Any],
output_width: int,
output_height: int,
output_fps: int,
tts_audio: Optional[Path] = None,
bgm_audio: Optional[Path] = None,
title_text: str = "",
subtitle_segments: Optional[list[Any]] = None,
font: str = DEFAULT_DRAWTEXT_FONT,
vcodec: str = "h264_nvenc",
preset: str = "p4",
video_bitrate: str = "",
cq: int = 23,
edge_crop_pct: float = 0.0,
total_duration: float = 0.0,
clip_has_audio: Optional[list[bool]] = None,
clip_volumes: Optional[list[float]] = None,
extra_audio_tracks: Optional[list[tuple[Any, float]]] = None,
title_config: Optional[dict] = None,
subtitle_config: Optional[dict] = None,
bgm_config: Optional[dict] = None,
static_subtitle_text: str = "",
) -> DirectRenderPlan:
"""构造 P4000 直连渲染所需的 inputs 与 ffmpeg_args。
视频:每段 trim/setpts/scale/pad/fps → concat(全硬切,带音频)→ 随机边缘 crop+scale → drawtext。
音频:每段 [i:a](或 anullsrc 静音占位)按 clip 配置 atrim/asetpts/atempo/volume/aresample
→ concat=n:N:v=1:a=1 → 与 extra_audio(TTS/配音素材库)、BGM 一起 amix → atrim 精确截断。
"""
if not resolved_clips:
raise ValueError("build_direct_render: no resolved clips")
inputs: dict[str, str] = {}
oss_keys: list[str] = []
input_args: list[str] = []
fc: list[str] = []
n = len(resolved_clips)
# 规范化每段参数
if clip_has_audio is None:
clip_has_audio = [True] * n
else:
clip_has_audio = list(clip_has_audio) + [True] * max(0, n - len(clip_has_audio))
clip_has_audio = clip_has_audio[:n]
if clip_volumes is None:
clip_volumes = [1.0] * n
else:
clip_volumes = list(clip_volumes) + [1.0] * max(0, n - len(clip_volumes))
clip_volumes = clip_volumes[:n]
clip_starts: list[float] = []
clip_effs: list[float] = []
clip_speeds: list[float] = []
for clip in resolved_clips:
start = float(getattr(clip, "start_time", 0) or 0)
eff = float(getattr(clip, "duration", 0) or 0)
if eff <= 0:
eff = float(getattr(clip, "actual_duration", 0) or 0)
speed = float(getattr(clip, "playback_speed", 1.0) or 1.0)
clip_starts.append(start)
clip_effs.append(eff)
clip_speeds.append(speed)
# 1. 视频输入(原始素材签名 URL)
for i, clip in enumerate(resolved_clips):
sk = (getattr(clip, "config", None) or {}).get("_storage_key")
if not sk:
raise ValueError(f"clip {getattr(clip, 'clip_id', i)} missing _storage_key")
fname = f"v{i}.mp4"
inputs[fname] = sign_asset_url(sk)
input_args.extend(["-i", fname])
# 2. 视频段预处理
pre_labels: list[str] = []
for i in range(n):
vf: list[str] = []
start, eff, speed = clip_starts[i], clip_effs[i], clip_speeds[i]
if eff > 0:
if start > 0:
vf.append(f"trim=start={start:.3f}:duration={eff:.3f}")
else:
vf.append(f"trim=duration={eff:.3f}")
vf.append("setpts=PTS-STARTPTS")
if abs(speed - 1.0) >= 1e-6:
vf.append(f"setpts=PTS/{speed:.4f}")
vf.append(f"scale={output_width}:{output_height}:force_original_aspect_ratio=decrease")
vf.append(f"pad={output_width}:{output_height}:trunc((ow-iw)/2):trunc((oh-ih)/2):black")
vf.append("setpts=PTS-STARTPTS")
vf.append(f"fps={output_fps}")
label = f"vc{i}"
fc.append(f"[{i}:v]{','.join(vf)}[{label}]")
pre_labels.append(label)
# 2b. 音频段预处理(无声源用 anullsrc 占位;volume=0 的段也用 anullsrc 静音占位保持时间轴)
anullsrc_counter = 0
audio_pre_labels: list[str] = []
for i in range(n):
start, eff, speed = clip_starts[i], clip_effs[i], clip_speeds[i]
vol = float(clip_volumes[i] if i < len(clip_volumes) else 1.0)
has_a = bool(clip_has_audio[i] if i < len(clip_has_audio) else True)
if not has_a or vol <= 0.001:
# 静音占位:用 anullsrc 生成静音,atrim 到段时长
sl = f"sil{anullsrc_counter}"
anullsrc_counter += 1
af: list[str] = ["anullsrc=channel_layout=stereo:sample_rate=44100"]
if eff > 0:
af.append(f"atrim=duration={eff:.3f}")
af.append("asetpts=PTS-STARTPTS")
af.append("aformat=sample_fmts=fltp:channel_layouts=stereo")
fc.append(f"{','.join(af)}[{sl}]")
# anullsrc 作为 filter 源不需要 -i 输入,直接给 label
audio_pre_labels.append(sl)
continue
af = []
if eff > 0:
if start > 0:
af.append(f"atrim=start={start:.3f}:duration={eff:.3f}")
else:
af.append(f"atrim=duration={eff:.3f}")
af.append("asetpts=PTS-STARTPTS")
if abs(speed - 1.0) >= 1e-6:
atempo = _build_atempo_chain(speed)
if atempo:
af.append(atempo)
if abs(vol - 1.0) >= 1e-3:
af.append(f"volume={vol:.3f}")
af.append("aresample=44100")
af.append("aformat=sample_fmts=fltp:channel_layouts=stereo")
alabel = f"ac{i}"
fc.append(f"[{i}:a]{','.join(af)}[{alabel}]")
audio_pre_labels.append(alabel)
# 3. concat(全硬切;v=1:a=1,视频音频一起拼接)
concat_in = "".join(f"[{v}][{a}]" for v, a in zip(pre_labels, audio_pre_labels, strict=True))
fc.append(f"{concat_in}concat=n={n}:v=1:a=1[vcat][acat]")
cur_v = "vcat"
cur_a = "acat"
# 4. 随机边缘裁剪降重(四边独立随机 2%~5%,与 ffmpeg_utils.random_edge_crop 一致)
if edge_crop_pct and edge_crop_pct > 0:
_r = random.Random()
p_min = EDGE_CROP_MIN_PCT
p_max = EDGE_CROP_MAX_PCT
crop_top = p_min + _r.random() * (p_max - p_min)
crop_bottom = p_min + _r.random() * (p_max - p_min)
crop_left = p_min + _r.random() * (p_max - p_min)
crop_right = p_min + _r.random() * (p_max - p_min)
w_expr = f"trunc(iw*(1-{crop_left:.4f}-{crop_right:.4f})/2)*2"
h_expr = f"trunc(ih*(1-{crop_top:.4f}-{crop_bottom:.4f})/2)*2"
x_expr = f"trunc(iw*{crop_left:.4f}/2)*2"
y_expr = f"trunc(ih*{crop_top:.4f}/2)*2"
fc.append(
f"[{cur_v}]crop=w='{w_expr}':h='{h_expr}':x='{x_expr}':y='{y_expr}',"
f"scale={output_width}:{output_height}[vcrop]"
)
cur_v = "vcrop"
# 5. drawtext 字幕(标题 + 静态全文 + ASR 分段)
# ── 解析 title_config(兼容字段名 font_size/font_color → size/color) ──
t_cfg = dict(title_config) if isinstance(title_config, dict) else {}
t_enabled = bool(t_cfg.get("enabled", True))
t_text = (t_cfg.get("text", "") or title_text or "").strip()
t_font = str(t_cfg.get("font", font) or font)
t_size_raw = t_cfg.get("size", t_cfg.get("font_size", 0))
try:
t_size = int(t_size_raw) if t_size_raw else 0
except (TypeError, ValueError):
t_size = 0
if t_size <= 0:
t_size = max(int(output_height * 0.05), 24)
t_color = str(t_cfg.get("color", t_cfg.get("font_color", "#ffffff")))
t_position = str(t_cfg.get("position", "bottom")).lower()
t_margin = max(40, int(output_height * 0.05))
t_borderw = 0
t_border_color = "#000000"
t_box = False
t_box_color = "black@0.5"
# stroke
_stroke = t_cfg.get("stroke")
if isinstance(_stroke, dict) and _stroke.get("enabled", False):
try:
t_borderw = int(float(_stroke.get("width", 2)))
except (TypeError, ValueError):
t_borderw = 2
t_border_color = str(_stroke.get("color", "#000000"))
elif isinstance(_stroke, bool) and _stroke:
t_borderw = 2
# shadow
_shadow = t_cfg.get("shadow")
t_shadow_enabled = False
t_shadow_color = "#000000@0.6"
t_shadow_x, t_shadow_y = 2, 2
if isinstance(_shadow, dict) and _shadow.get("enabled", False):
t_shadow_enabled = True
t_shadow_color = str(_shadow.get("color", "#000000@0.6"))
try:
t_shadow_x = int(float(_shadow.get("offset_x", 2)))
t_shadow_y = int(float(_shadow.get("offset_y", 2)))
except (TypeError, ValueError):
t_shadow_x, t_shadow_y = 2, 2
elif isinstance(_shadow, bool) and _shadow:
t_shadow_enabled = True
# bold/italic:drawtext 原生无粗斜体选项;通过加大 borderw 模拟粗体
t_bold = bool(t_cfg.get("bold", False))
if t_bold and t_borderw < 1:
t_borderw = 1
t_border_color = t_color # 用文字色描边模拟加粗
# ── 解析 subtitle_config ──
s_cfg = dict(subtitle_config) if isinstance(subtitle_config, dict) else {}
s_enabled = bool(s_cfg.get("enabled", True))
s_font = str(s_cfg.get("font", font) or font)
s_size_raw = s_cfg.get("size", s_cfg.get("font_size", 0))
try:
s_size = int(s_size_raw) if s_size_raw else 0
except (TypeError, ValueError):
s_size = 0
if s_size <= 0:
s_size = max(int(output_height * 0.04), 20)
s_color = str(s_cfg.get("color", s_cfg.get("font_color", "#ffffff")))
s_position = str(s_cfg.get("position", "bottom")).lower()
s_margin = max(60, int(output_height * 0.06))
s_borderw = 2 # 字幕默认描边保证可读性
s_border_color = "#000000"
# 静态字幕:static_subtitle_text 非空时构造全片长 segment(0 → total_duration)
static_text = (static_subtitle_text or "").strip()
subtitle_segments = list(subtitle_segments or [])
if s_enabled and static_text and total_duration and total_duration > 0:
# 用 duck-type 对象插入到 subtitle_segments 列表头部(静态全文)
class _StaticSeg:
def __init__(self, txt, st, ed):
self.text = txt
self.start = st
self.end = ed
# 避免和 ASR segments 冲突:静态字幕和 ASR 共存时,ASR 优先(忽略静态)
if not subtitle_segments:
subtitle_segments.insert(0, _StaticSeg(static_text, 0.0, float(total_duration)))
draw_filters: list[str] = []
if t_enabled and t_text:
draw_filters.extend(
_build_drawtext_filters(
text=t_text,
start=0.0,
end=max(total_duration, 0.1),
font=t_font,
font_size=t_size,
font_color=t_color,
position=t_position,
margin=t_margin,
box_enabled=t_box,
box_color=t_box_color,
borderw=t_borderw,
border_color=t_border_color,
shadow_enabled=t_shadow_enabled,
shadow_color=t_shadow_color,
shadow_x=t_shadow_x,
shadow_y=t_shadow_y,
)
)
if s_enabled:
for seg in subtitle_segments:
txt = getattr(seg, "text", "") or ""
if not txt.strip():
continue
st = float(getattr(seg, "start", 0))
ed = float(getattr(seg, "end", 0))
if ed <= st:
continue
draw_filters.extend(
_build_drawtext_filters(
text=txt,
start=st,
end=ed,
font=s_font,
font_size=s_size,
font_color=s_color,
position=s_position,
margin=s_margin,
box_enabled=False,
borderw=s_borderw,
border_color=s_border_color,
)
)
if draw_filters:
prev = cur_v
for idx, df in enumerate(draw_filters):
out_l = "vfinal" if idx == len(draw_filters) - 1 else f"vd{idx}"
fc.append(f"[{prev}]{df}[{out_l}]")
prev = out_l
vfinal_label = prev
else:
fc.append(f"[{cur_v}]format=yuv420p[vfinal]")
vfinal_label = "vfinal"
# 6. 音频混音:原素材主音轨 acat + extra(TTS/配音素材库) + BGM → amix → atrim
mix_labels: list[str] = [cur_a]
mix_vols: list[float] = [1.0]
next_idx = n
# 额外独立音频轨(TTS concat / 配音素材库整段音频)
for _ea_idx, (ea_path, ea_vol) in enumerate(extra_audio_tracks or []):
if ea_path is None:
continue
ea_p = Path(ea_path)
if not ea_p.exists():
continue
eurl, ekey = upload_local_audio_and_sign(ea_p)
ename = f"extra{_ea_idx}{ea_p.suffix or '.mp3'}"
inputs[ename] = eurl
oss_keys.append(ekey)
input_args.extend(["-i", ename])
elabel = f"aex{_ea_idx}"
fc.append(
f"[{next_idx}:a]aresample=44100,volume={float(ea_vol):.2f},"
f"aformat=sample_fmts=fltp:channel_layouts=stereo[{elabel}]"
)
mix_labels.append(elabel)
mix_vols.append(float(ea_vol))
next_idx += 1
if tts_audio and Path(tts_audio).exists():
# 旧参数保留:若调用方直接传了 tts_audio 而没走 extra_audio_tracks,则仍然加入
# (兼容旧调用,正常路径 TTS 已经通过 extra_audio_tracks 传入)
turl, tkey = upload_local_audio_and_sign(Path(tts_audio))
tname = "tts" + (Path(tts_audio).suffix or ".mp3")
inputs[tname] = turl
oss_keys.append(tkey)
input_args.extend(["-i", tname])
alabel = "au_tts"
fc.append(
f"[{next_idx}:a]aresample=44100,volume=1.00,aformat=sample_fmts=fltp:channel_layouts=stereo[{alabel}]"
)
mix_labels.append(alabel)
mix_vols.append(1.0)
next_idx += 1
_bgm_use = bgm_audio is not None and Path(bgm_audio).exists()
if _bgm_use and isinstance(bgm_config, dict) and bgm_config.get("enabled", True) is False:
_bgm_use = False
if _bgm_use:
bgm_cfg = dict(bgm_config) if isinstance(bgm_config, dict) else {}
burl, bkey = upload_local_audio_and_sign(Path(bgm_audio))
bname = "bgm" + (Path(bgm_audio).suffix or ".mp3")
inputs[bname] = burl
oss_keys.append(bkey)
input_args.extend(["-i", bname])
alabel = "au_bgm"
try:
bgm_vol = float(bgm_cfg.get("volume", 0.3))
except (TypeError, ValueError):
bgm_vol = 0.3
bgm_vol = max(0.0, min(1.5, bgm_vol))
# volume_adjust_db(-3 ~ +3 dB)换算线性增益
try:
_db = float(bgm_cfg.get("volume_adjust_db", 0.0))
except (TypeError, ValueError):
_db = 0.0
if abs(_db) > 0.05:
db_gain = 10 ** (_db / 20.0)
bgm_vol = max(0.0, min(2.0, bgm_vol * db_gain))
# afade 淡入淡出
try:
fade_in = max(0.0, float(bgm_cfg.get("fade_in", 0.0)))
except (TypeError, ValueError):
fade_in = 0.0
try:
fade_out = max(0.0, float(bgm_cfg.get("fade_out", 0.0)))
except (TypeError, ValueError):
fade_out = 0.0
# audio_offset:adelay 延迟(毫秒)
try:
offset = max(0.0, float(bgm_cfg.get("audio_offset", 0.0)))
except (TypeError, ValueError):
offset = 0.0
bgm_parts: list[str] = [f"[{next_idx}:a]aresample=44100"]
if offset > 0.01:
bgm_parts.append(f"adelay={int(offset * 1000)}|{int(offset * 1000)}")
bgm_parts.append(f"volume={bgm_vol:.3f}")
if fade_in > 0.01:
bgm_parts.append(f"afade=t=in:st=0:d={fade_in:.2f}")
if fade_out > 0.01 and total_duration > 0:
fo_start = max(0.0, total_duration - fade_out)
bgm_parts.append(f"afade=t=out:st={fo_start:.2f}:d={fade_out:.2f}")
bgm_parts.append("aformat=sample_fmts=fltp:channel_layouts=stereo")
fc.append(",".join(bgm_parts) + f"[{alabel}]")
mix_labels.append(alabel)
mix_vols.append(bgm_vol)
next_idx += 1
maps: list[str] = ["-map", f"[{vfinal_label}]"]
if mix_labels:
mix_in = "".join(f"[{lb}]" for lb in mix_labels)
n_mix = len(mix_labels)
mix_parts = [
f"amix=inputs={n_mix}:duration=longest:dropout_transition=2:normalize=0",
"aresample=44100",
]
# Bug2 修复:atrim 到视频精确时长
if total_duration and total_duration > 0:
mix_parts.append(f"atrim=0:{total_duration:.3f}")
mix_parts.append("asetpts=PTS-STARTPTS")
fc.append(f"{mix_in}{','.join(mix_parts)}[afinal]")
maps.extend(["-map", "[afinal]", "-c:a", "aac", "-b:a", "128k"])
else:
logger.info("[gpu-direct] no audio tracks; output silent video")
# 7. 组装 ffmpeg_args + NVENC 编码
ffmpeg_args = ["-y", *input_args, "-filter_complex", ";".join(fc), *maps]
ffmpeg_args.extend(["-c:v", vcodec, "-preset", preset, "-pix_fmt", "yuv420p"])
if video_bitrate:
ffmpeg_args.extend(["-b:v", video_bitrate])
else:
ffmpeg_args.extend(["-cq", str(cq)])
ffmpeg_args.extend(["-movflags", "+faststart", "-shortest", "-f", "mp4", "pipe:1"])
return DirectRenderPlan(
inputs=inputs,
ffmpeg_args=ffmpeg_args,
oss_keys=oss_keys,
filter_complex=fc,
)