79e6bd52b8
CI/CD Pipeline / Dedup Check - skip PR tests when covered by push pipeline (pull_request) Successful in 2s
CI/CD Pipeline / Check if frontend-only change (pull_request) Successful in 1s
Preview Deploy / Deploy Preview Environment (pull_request) Failing after 57s
CI/CD Pipeline / Frontend Lint (pull_request) Failing after 50s
CI/CD Pipeline / Frontend Unit Tests (pull_request) Successful in 1m41s
CI/CD Pipeline / PR Build API Image (pull_request) Successful in 2m6s
CI/CD Pipeline / PR Build Web Image (pull_request) Failing after 2m16s
CI/CD Pipeline / Validate - Style (pull_request) Has been cancelled
CI/CD Pipeline / Validate - Security (pull_request) Has been cancelled
CI/CD Pipeline / Validate - Python (mypy + alembic) (pull_request) Has been cancelled
CI/CD Pipeline / Unit Tests (pull_request) Has been cancelled
CI/CD Pipeline / Integration Tests (pull_request) Has been cancelled
CI/CD Pipeline / PR Build Worker Image (pull_request) Has been cancelled
CI/CD Pipeline / Build Production API Image (pull_request) Has been cancelled
CI/CD Pipeline / Build Production Web Image (pull_request) Has been cancelled
CI/CD Pipeline / Build Production Worker Image (pull_request) Has been cancelled
CI/CD Pipeline / Deploy Production (pull_request) Has been cancelled
CI/CD Pipeline / Production Browser E2E (pull_request) Has been cancelled
CI/CD Pipeline / Canary Release to Production (pull_request) Has been cancelled
CI/CD Pipeline / CI Gate (pull_request) Has been cancelled
AI Code Review / AI Code Review (pull_request) Has been cancelled
PR Automation / Auto Approve on CI Green (pull_request) Has been cancelled
PR Automation / Auto Merge on CI Green + Approved (pull_request) Has been cancelled
CI/CD Pipeline / Staging API Integration Tests (pull_request) Failing after 109h59m25s
CI/CD Pipeline / Retag skipped Staging API Image (pull_request) Failing after 109h59m42s
CI/CD Pipeline / Build Staging Worker Image (pull_request) Failing after 109h59m30s
CI/CD Pipeline / Build Staging Web Image (pull_request) Failing after 109h59m31s
CI/CD Pipeline / ACR Image Cleanup (pull_request) Failing after 109h59m8s
CI/CD Pipeline / Retag skipped Staging Worker Image (pull_request) Failing after 109h59m25s
CI/CD Pipeline / Retag skipped Staging Web Image (pull_request) Failing after 109h59m25s
CI/CD Pipeline / Build Staging API Image (pull_request) Failing after 109h59m31s
CI/CD Pipeline / Check push changed paths (pull_request) Failing after 109h59m38s
CI/CD Pipeline / Staging E2E Tests (pull_request) Failing after 109h59m9s
CI/CD Pipeline / Deploy Staging (Watchtower auto-deploy) (pull_request) Failing after 109h59m18s
修复5个问题:
1. 最终视频标题一致性
- 后端 drawtext 默认 font_size 从 36 改为 28(对齐前端 DEFAULT_TITLE_CONFIG.size)
- 后端 drawtext 默认 position 从 'top' 改为 'bottom'(对齐前端默认值)
- 前端 contract.ts font_size fallback 从 36 改为 28
2. 封面与视频一致性
- 渲染完成后自动用 output_cover_url 更新前端封面配置
- 前端封面标题预览 fontSize 改用 size*0.55 缩放(与对口型预览对齐)
- 封面默认阴影改为 none(与后端 drawtext 无 shadow 行为对齐)
3. B-roll 位置一致性
- ModalBRollEditor 使用 lipsyncJob.script_text(TTS 时锁定的文案)
而非 state.scriptText(用户可能已修改的文案)
4. B-roll 时长精确化
- 后端 lipsync_tts.py 新增 _compute_sentence_timings():
TTS 音频用 ffmpeg silencedetect 检测静音点,与句子边界对齐
- 新增 sentence_timings JSON 字段到 LipsyncJobModel + Alembic migration
- 前端 sentences.ts 不再做字数比例估算,改为接收后端精确时间戳
- ModalBRollEditor 接收 sentenceTimings prop
5. 标题拖拽位置修复
- PanelLipsyncPreview 拖拽标题发送百分比坐标(0-100) + position:'custom'
- 后端 build_title_drawtext_filter 的 custom 位置改用 drawtext 表达式
x=(w-text_w)*{pct_x} y=(h-text_h)*{pct_y}(百分比,与视频尺寸解耦)
文件变更:
- apps/api/app/tasks/lipsync_tts.py: 句子时间戳计算
- apps/api/app/schemas/lipsync.py: sentence_timings 字段
- packages/adapters/sqlalchemy_impl/models.py: sentence_timings 列
- packages/domain/video_filter_builder.py: 百分比坐标 + 默认值对齐
- alembic/versions/075_add_sentence_timings.py: 数据库迁移
- apps/web/src/pages/ai-avatar/types.ts: SentenceTiming + RenderJob.output_cover_url
- apps/web/src/pages/ai-avatar/utils/sentences.ts: 移除字数估算
- apps/web/src/pages/ai-avatar/utils/contract.ts: font_size 默认值
- apps/web/src/pages/ai-avatar/components/PanelLipsyncPreview.tsx: 拖拽百分比
- apps/web/src/pages/ai-avatar/components/PanelCoverAndGenerate.tsx: 标题缩放
- apps/web/src/pages/ai-avatar/components/ModalBRollEditor.tsx: 后端时间戳
- apps/web/src/pages/ai-avatar/AiAvatarPage.tsx: 封面更新 + 传参
- tests/unit/test_video_filter_builder.py: 测试更新
102 lines
3.8 KiB
Python
102 lines
3.8 KiB
Python
"""对口型 API Schema 定义 — #1796 / #1809 / #1822.
|
||
|
||
支持两种输入模式(二选一):
|
||
1. TTS 直生模式(推荐):传 voice_id + script_text(+ speed/emotion),
|
||
后端内部先调 CosyVoice 合成音频,再提交 MediaKit 对口型。
|
||
2. 直接音频模式:传 video_url + audio_url(音频已由调用方准备好)。
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
from datetime import datetime
|
||
from typing import Optional
|
||
|
||
from pydantic import BaseModel, Field, model_validator
|
||
|
||
|
||
class LipsyncJobResponse(BaseModel):
|
||
"""对口型任务响应."""
|
||
|
||
id: str
|
||
user_id: str
|
||
project_id: str
|
||
video_url: str
|
||
audio_url: str
|
||
enable_video_loop: bool
|
||
voice_id: str = ""
|
||
script_text: str = ""
|
||
speed: float = 1.0
|
||
emotion: str = ""
|
||
mediakit_task_id: str
|
||
status: str
|
||
output_video_url: str
|
||
output_duration: float
|
||
error_message: str
|
||
error_code: str
|
||
sentence_timings: Optional[list] = None
|
||
submitted_at: Optional[datetime] = None
|
||
completed_at: Optional[datetime] = None
|
||
created_at: datetime
|
||
updated_at: datetime
|
||
|
||
class Config:
|
||
from_attributes = True
|
||
|
||
|
||
class CreateLipsyncJobRequest(BaseModel):
|
||
"""创建对口型任务请求.
|
||
|
||
两种模式(二选一):
|
||
- TTS 直生:voice_id + script_text 必填(+ 可选 speed/emotion);audio_url 留空。
|
||
- 直接音频:video_url + audio_url 必填。
|
||
"""
|
||
|
||
video_url: str = Field(..., description="人物视频 URL(MP4,≤30min,单人真人)")
|
||
|
||
# 模式 2:直接音频
|
||
audio_url: str = Field("", description="驱动音频 URL(mp3/aac/wav/m4a/flac);直生模式留空")
|
||
|
||
# 模式 1:TTS 直生
|
||
voice_id: str = Field("", description="音色 ID(预置音色或克隆音色 profile UUID)")
|
||
script_text: str = Field("", description="要合成的文案(直生模式必填,最长 5000 字符)")
|
||
speed: float = Field(1.0, ge=0.5, le=2.0, description="语速(0.5-2.0),默认 1.0")
|
||
emotion: str = Field("", description="情绪(natural/excited/calm/friendly 或中文 自然/兴奋/沉稳/亲切)")
|
||
|
||
enable_video_loop: bool = Field(False, description="音频长于视频时是否循环画面")
|
||
project_id: str = Field("", description="项目 ID(可选)")
|
||
|
||
@model_validator(mode="after")
|
||
def _validate_input_mode(self) -> "CreateLipsyncJobRequest":
|
||
video = (self.video_url or "").strip()
|
||
if not video:
|
||
raise ValueError("video_url 不能为空")
|
||
if not video.startswith(("http://", "https://")):
|
||
raise ValueError("video_url 必须是 HTTP/HTTPS URL")
|
||
lower = video.lower().split("?")[0]
|
||
if not lower.endswith(".mp4"):
|
||
raise ValueError("video_url 仅支持 MP4 格式")
|
||
|
||
has_audio = bool((self.audio_url or "").strip())
|
||
has_tts = bool((self.voice_id or "").strip()) and bool((self.script_text or "").strip())
|
||
|
||
if not has_audio and not has_tts:
|
||
raise ValueError(
|
||
"必须提供驱动音频:要么传 audio_url(直接音频模式),"
|
||
"要么同时传 voice_id + script_text(TTS 直生模式)"
|
||
)
|
||
|
||
if has_tts and len(self.script_text) > 5000:
|
||
raise ValueError("script_text 最长 5000 字符")
|
||
|
||
if has_audio:
|
||
au = self.audio_url.strip()
|
||
if not au.startswith(("http://", "https://")):
|
||
raise ValueError("audio_url 必须是 HTTP/HTTPS URL")
|
||
au_lower = au.lower().split("?")[0]
|
||
allowed = (".mp3", ".aac", ".wav", ".m4a", ".flac")
|
||
if not any(au_lower.endswith(ext) for ext in allowed):
|
||
raise ValueError(f"audio_url 格式不支持,仅支持: {', '.join(allowed)}")
|
||
self.audio_url = au
|
||
|
||
return self
|