Files
xiaoxia-saas/apps/api/app/schemas/lipsync.py
T
LingYing Agent 79e6bd52b8
CI/CD Pipeline / Dedup Check - skip PR tests when covered by push pipeline (pull_request) Successful in 2s
CI/CD Pipeline / Check if frontend-only change (pull_request) Successful in 1s
Preview Deploy / Deploy Preview Environment (pull_request) Failing after 57s
CI/CD Pipeline / Frontend Lint (pull_request) Failing after 50s
CI/CD Pipeline / Frontend Unit Tests (pull_request) Successful in 1m41s
CI/CD Pipeline / PR Build API Image (pull_request) Successful in 2m6s
CI/CD Pipeline / PR Build Web Image (pull_request) Failing after 2m16s
CI/CD Pipeline / Validate - Style (pull_request) Has been cancelled
CI/CD Pipeline / Validate - Security (pull_request) Has been cancelled
CI/CD Pipeline / Validate - Python (mypy + alembic) (pull_request) Has been cancelled
CI/CD Pipeline / Unit Tests (pull_request) Has been cancelled
CI/CD Pipeline / Integration Tests (pull_request) Has been cancelled
CI/CD Pipeline / PR Build Worker Image (pull_request) Has been cancelled
CI/CD Pipeline / Build Production API Image (pull_request) Has been cancelled
CI/CD Pipeline / Build Production Web Image (pull_request) Has been cancelled
CI/CD Pipeline / Build Production Worker Image (pull_request) Has been cancelled
CI/CD Pipeline / Deploy Production (pull_request) Has been cancelled
CI/CD Pipeline / Production Browser E2E (pull_request) Has been cancelled
CI/CD Pipeline / Canary Release to Production (pull_request) Has been cancelled
CI/CD Pipeline / CI Gate (pull_request) Has been cancelled
AI Code Review / AI Code Review (pull_request) Has been cancelled
PR Automation / Auto Approve on CI Green (pull_request) Has been cancelled
PR Automation / Auto Merge on CI Green + Approved (pull_request) Has been cancelled
CI/CD Pipeline / Staging API Integration Tests (pull_request) Failing after 109h59m25s
CI/CD Pipeline / Retag skipped Staging API Image (pull_request) Failing after 109h59m42s
CI/CD Pipeline / Build Staging Worker Image (pull_request) Failing after 109h59m30s
CI/CD Pipeline / Build Staging Web Image (pull_request) Failing after 109h59m31s
CI/CD Pipeline / ACR Image Cleanup (pull_request) Failing after 109h59m8s
CI/CD Pipeline / Retag skipped Staging Worker Image (pull_request) Failing after 109h59m25s
CI/CD Pipeline / Retag skipped Staging Web Image (pull_request) Failing after 109h59m25s
CI/CD Pipeline / Build Staging API Image (pull_request) Failing after 109h59m31s
CI/CD Pipeline / Check push changed paths (pull_request) Failing after 109h59m38s
CI/CD Pipeline / Staging E2E Tests (pull_request) Failing after 109h59m9s
CI/CD Pipeline / Deploy Staging (Watchtower auto-deploy) (pull_request) Failing after 109h59m18s
fix(ai-avatar): 端到端一致性修复(标题/封面/B-roll位置/B-roll时长)
修复5个问题:

1. 最终视频标题一致性
   - 后端 drawtext 默认 font_size 从 36 改为 28(对齐前端 DEFAULT_TITLE_CONFIG.size)
   - 后端 drawtext 默认 position 从 'top' 改为 'bottom'(对齐前端默认值)
   - 前端 contract.ts font_size fallback 从 36 改为 28

2. 封面与视频一致性
   - 渲染完成后自动用 output_cover_url 更新前端封面配置
   - 前端封面标题预览 fontSize 改用 size*0.55 缩放(与对口型预览对齐)
   - 封面默认阴影改为 none(与后端 drawtext 无 shadow 行为对齐)

3. B-roll 位置一致性
   - ModalBRollEditor 使用 lipsyncJob.script_text(TTS 时锁定的文案)
     而非 state.scriptText(用户可能已修改的文案)

4. B-roll 时长精确化
   - 后端 lipsync_tts.py 新增 _compute_sentence_timings():
     TTS 音频用 ffmpeg silencedetect 检测静音点,与句子边界对齐
   - 新增 sentence_timings JSON 字段到 LipsyncJobModel + Alembic migration
   - 前端 sentences.ts 不再做字数比例估算,改为接收后端精确时间戳
   - ModalBRollEditor 接收 sentenceTimings prop

5. 标题拖拽位置修复
   - PanelLipsyncPreview 拖拽标题发送百分比坐标(0-100) + position:'custom'
   - 后端 build_title_drawtext_filter 的 custom 位置改用 drawtext 表达式
     x=(w-text_w)*{pct_x} y=(h-text_h)*{pct_y}(百分比,与视频尺寸解耦)

文件变更:
- apps/api/app/tasks/lipsync_tts.py: 句子时间戳计算
- apps/api/app/schemas/lipsync.py: sentence_timings 字段
- packages/adapters/sqlalchemy_impl/models.py: sentence_timings 列
- packages/domain/video_filter_builder.py: 百分比坐标 + 默认值对齐
- alembic/versions/075_add_sentence_timings.py: 数据库迁移
- apps/web/src/pages/ai-avatar/types.ts: SentenceTiming + RenderJob.output_cover_url
- apps/web/src/pages/ai-avatar/utils/sentences.ts: 移除字数估算
- apps/web/src/pages/ai-avatar/utils/contract.ts: font_size 默认值
- apps/web/src/pages/ai-avatar/components/PanelLipsyncPreview.tsx: 拖拽百分比
- apps/web/src/pages/ai-avatar/components/PanelCoverAndGenerate.tsx: 标题缩放
- apps/web/src/pages/ai-avatar/components/ModalBRollEditor.tsx: 后端时间戳
- apps/web/src/pages/ai-avatar/AiAvatarPage.tsx: 封面更新 + 传参
- tests/unit/test_video_filter_builder.py: 测试更新
2026-09-12 01:23:22 +08:00

102 lines
3.8 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""对口型 API Schema 定义 — #1796 / #1809 / #1822.
支持两种输入模式(二选一):
1. TTS 直生模式(推荐):传 voice_id + script_text(+ speed/emotion),
后端内部先调 CosyVoice 合成音频,再提交 MediaKit 对口型。
2. 直接音频模式:传 video_url + audio_url(音频已由调用方准备好)。
"""
from __future__ import annotations
from datetime import datetime
from typing import Optional
from pydantic import BaseModel, Field, model_validator
class LipsyncJobResponse(BaseModel):
"""对口型任务响应."""
id: str
user_id: str
project_id: str
video_url: str
audio_url: str
enable_video_loop: bool
voice_id: str = ""
script_text: str = ""
speed: float = 1.0
emotion: str = ""
mediakit_task_id: str
status: str
output_video_url: str
output_duration: float
error_message: str
error_code: str
sentence_timings: Optional[list] = None
submitted_at: Optional[datetime] = None
completed_at: Optional[datetime] = None
created_at: datetime
updated_at: datetime
class Config:
from_attributes = True
class CreateLipsyncJobRequest(BaseModel):
"""创建对口型任务请求.
两种模式(二选一):
- TTS 直生:voice_id + script_text 必填(+ 可选 speed/emotion);audio_url 留空。
- 直接音频:video_url + audio_url 必填。
"""
video_url: str = Field(..., description="人物视频 URL(MP4,≤30min,单人真人)")
# 模式 2:直接音频
audio_url: str = Field("", description="驱动音频 URL(mp3/aac/wav/m4a/flac);直生模式留空")
# 模式 1:TTS 直生
voice_id: str = Field("", description="音色 ID(预置音色或克隆音色 profile UUID)")
script_text: str = Field("", description="要合成的文案(直生模式必填,最长 5000 字符)")
speed: float = Field(1.0, ge=0.5, le=2.0, description="语速(0.5-2.0),默认 1.0")
emotion: str = Field("", description="情绪(natural/excited/calm/friendly 或中文 自然/兴奋/沉稳/亲切)")
enable_video_loop: bool = Field(False, description="音频长于视频时是否循环画面")
project_id: str = Field("", description="项目 ID(可选)")
@model_validator(mode="after")
def _validate_input_mode(self) -> "CreateLipsyncJobRequest":
video = (self.video_url or "").strip()
if not video:
raise ValueError("video_url 不能为空")
if not video.startswith(("http://", "https://")):
raise ValueError("video_url 必须是 HTTP/HTTPS URL")
lower = video.lower().split("?")[0]
if not lower.endswith(".mp4"):
raise ValueError("video_url 仅支持 MP4 格式")
has_audio = bool((self.audio_url or "").strip())
has_tts = bool((self.voice_id or "").strip()) and bool((self.script_text or "").strip())
if not has_audio and not has_tts:
raise ValueError(
"必须提供驱动音频:要么传 audio_url(直接音频模式),"
"要么同时传 voice_id + script_text(TTS 直生模式)"
)
if has_tts and len(self.script_text) > 5000:
raise ValueError("script_text 最长 5000 字符")
if has_audio:
au = self.audio_url.strip()
if not au.startswith(("http://", "https://")):
raise ValueError("audio_url 必须是 HTTP/HTTPS URL")
au_lower = au.lower().split("?")[0]
allowed = (".mp3", ".aac", ".wav", ".m4a", ".flac")
if not any(au_lower.endswith(ext) for ext in allowed):
raise ValueError(f"audio_url 格式不支持,仅支持: {', '.join(allowed)}")
self.audio_url = au
return self