Files
xiaoxia-saas/apps/api/app/schemas/lipsync.py
T
xiaoxia-agent 74200949b9
CI/CD Pipeline / Check if frontend-only change (pull_request) Successful in 2s
CI/CD Pipeline / Dedup Check - skip PR tests when covered by push pipeline (pull_request) Successful in 4s
CI/CD Pipeline / PR Build API Image (pull_request) Successful in 45s
CI/CD Pipeline / PR Build Worker Image (pull_request) Successful in 45s
Preview Deploy / Deploy Preview Environment (pull_request) Successful in 1m40s
CI/CD Pipeline / Integration Tests (pull_request) Successful in 3m12s
PR Automation / Auto Approve on CI Green (pull_request) Successful in 3m10s
CI/CD Pipeline / Validate - Python (mypy + alembic) (pull_request) Successful in 3m52s
CI/CD Pipeline / Validate - Style (pull_request) Successful in 5m51s
AI Code Review / AI Code Review (pull_request) Successful in 6m48s
CI/CD Pipeline / Unit Tests (pull_request) Successful in 10m36s
CI/CD Pipeline / Validate - Security (pull_request) Successful in 12m36s
CI/CD Pipeline / CI Gate (pull_request) Successful in 0s
CI/CD Pipeline / Production Browser E2E (pull_request) Has been skipped
PR Automation / Auto Merge on CI Green + Approved (pull_request) Successful in 11m6s
ACR Cleanup / ACR Image Cleanup (pull_request_target) Successful in 18s
Preview Cleanup / Cleanup Preview Environment (pull_request) Successful in 39s
CI/CD Pipeline / Deploy Production (pull_request) Failing after 24h47m16s
CI/CD Pipeline / Staging E2E Tests (pull_request) Failing after 24h59m6s
CI/CD Pipeline / Build Staging Web Image (pull_request) Failing after 24h59m50s
CI/CD Pipeline / Frontend Lint (pull_request) Failing after 24h59m52s
CI/CD Pipeline / PR Build Web Image (pull_request) Failing after 24h59m12s
CI/CD Pipeline / Build Production API Image (pull_request) Failing after 24h46m40s
CI/CD Pipeline / Deploy Staging (Watchtower auto-deploy) (pull_request) Failing after 24h58m27s
CI/CD Pipeline / Retag skipped Staging Worker Image (pull_request) Failing after 24h58m27s
CI/CD Pipeline / Retag skipped Staging Web Image (pull_request) Failing after 24h58m27s
CI/CD Pipeline / Retag skipped Staging API Image (pull_request) Failing after 24h58m27s
CI/CD Pipeline / Build Staging Worker Image (pull_request) Failing after 24h59m10s
CI/CD Pipeline / Build Staging API Image (pull_request) Failing after 24h59m11s
CI/CD Pipeline / Check push changed paths (pull_request) Failing after 24h59m22s
CI/CD Pipeline / ACR Image Cleanup (pull_request) Failing after 24h58m27s
CI/CD Pipeline / Canary Release to Production (pull_request) Failing after 24h46m37s
CI/CD Pipeline / Build Production Worker Image (pull_request) Failing after 24h46m40s
CI/CD Pipeline / Build Production Web Image (pull_request) Failing after 24h46m40s
CI/CD Pipeline / Staging API Integration Tests (pull_request) Failing after 24h58m27s
CI/CD Pipeline / Frontend Unit Tests (pull_request) Failing after 24h59m13s
fix(#1898): TTS 情绪映射修正为 CosyVoice v3 官方英文枚举(P1 更正)
关键修复(对照阿里云百炼官方文档 https://help.aliyun.com/zh/model-studio/cosyvoice-voice-list):
- instruction 格式严格为 "你说话的情感是<情感值>。",结尾中文句号不可省略
- 情感值必须是 7 种英文枚举之一:neutral/happy/sad/angry/surprised/fearful/disgusted
- 之前 PR#1932 错误地映射成了中文描述词(如"开心愉快"),本次修正为英文枚举

改动:
- EMOTION_MAP value 从中文描述词改为英文枚举自身;新增前端中文 7 标签(中立/开心/难过/生气/惊讶/恐惧/厌恶)→英文枚举;旧英文 4 枚举归并到最接近的标准值(natural→neutral, excited/friendly→happy, calm→neutral)
- normalize_emotion 未知值默认 neutral(不返回空,保证合成不中断),仅空串/None/纯空白返回空(调用方不传 instruction 走默认自然)
- schemas/router 描述更新,voice_clones preview 白名单加入中文 7 标签
- 单测重写:89 个用例断言 instruction 情感值为纯英文 ASCII 枚举;修正旧测试中的中文描述词断言为英文枚举;未知值断言改为默认 neutral

验证:
- 15333 passed, 28 skipped(覆盖 7 英文枚举、中文 7 标签、中文别名、旧英文、大小写、空白、未知值、instruction 格式、白名单)
- ruff/black 全过
2026-09-15 14:23:11 +08:00

139 lines
6.1 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""对口型 API Schema 定义 — #1796 / #1809 / #1822 / #1845(配音前置).
支持三种输入模式:
1. TTS 直生模式(兼容旧版前端):传 voice_id + script_text(+ speed/emotion),
后端 Celery 异步做 TTS 合成 + MediaKit 提交。
2. 直接音频模式:传 video_url + audio_url(音频已由调用方准备好)。
3. 预合成音频模式(#1845 配音前置新主路径):前端先调 POST /lipsync/tts-preview
拿到 audio_url + sentence_timings,再在 create_job 时传 audio_url + audio_duration
+ sentence_timings,后端跳过 TTS 和时间戳计算,直接 ffprobe 校验后提交 MediaKit。
"""
from __future__ import annotations
from datetime import datetime
from typing import Optional
from pydantic import BaseModel, Field, model_validator
class LipsyncJobResponse(BaseModel):
"""对口型任务响应."""
id: str
user_id: str
project_id: str
video_url: str
audio_url: str
enable_video_loop: bool
voice_id: str = ""
script_text: str = ""
speed: float = 1.0
emotion: str = ""
mediakit_task_id: str
status: str
output_video_url: str
output_duration: float
error_message: str
error_code: str
sentence_timings: Optional[list] = None
submitted_at: Optional[datetime] = None
completed_at: Optional[datetime] = None
created_at: datetime
updated_at: datetime
class Config:
from_attributes = True
class CreateLipsyncJobRequest(BaseModel):
"""创建对口型任务请求.
三种模式(三选一):
- TTS 直生(旧版/降级):voice_id + script_text 必填;audio_url 留空。
- 直接音频:video_url + audio_url 必填。
- 预合成音频(#1845 新主路径):audio_url 必填 + 可选 audio_duration/sentence_timings;
后端同步 ffprobe 校验时长、写入 timings,直接提交 MediaKit。
"""
video_url: str = Field(..., description="人物视频 URL(MP4,≤30min,单人真人)")
# 模式 2/3:直接/预合成音频
audio_url: str = Field("", description="驱动音频 URL(mp3/aac/wav/m4a/flac);直生模式留空")
audio_duration: Optional[float] = Field(None, ge=0, description="预合成音频时长(秒),可选;后端会 ffprobe 校验")
sentence_timings: Optional[list] = Field(None, description="预合成接口返回的句子时间戳,可选;若传入则直接写入 job")
# 模式 1:TTS 直生
voice_id: str = Field("", description="音色 ID(预置音色或克隆音色 profile UUID)")
script_text: str = Field("", description="要合成的文案(直生模式必填,最长 5000 字符)")
speed: float = Field(1.0, ge=0.5, le=2.0, description="语速(0.5-2.0),默认 1.0")
emotion: str = Field(
"",
description="情绪(英文枚举 neutral/happy/sad/angry/surprised/fearful/disgusted,或中文 中立/开心/难过/生气/惊讶/恐惧/厌恶;空为默认自然)",
)
enable_video_loop: bool = Field(
True, description="音频长于视频时是否循环画面(AI数字人默认开启,防止音频长于视频被截断)"
)
project_id: str = Field("", description="项目 ID(可选)")
@model_validator(mode="after")
def _validate_input_mode(self) -> "CreateLipsyncJobRequest":
video = (self.video_url or "").strip()
if not video:
raise ValueError("video_url 不能为空")
if not video.startswith(("http://", "https://")):
raise ValueError("video_url 必须是 HTTP/HTTPS URL")
lower = video.lower().split("?")[0]
allowed_video_exts = (".mp4", ".mov", ".m4v", ".webm", ".avi", ".mkv", ".3gp")
if not any(lower.endswith(ext) for ext in allowed_video_exts):
raise ValueError("video_url 格式不支持,仅支持: " + ", ".join(allowed_video_exts))
has_audio = bool((self.audio_url or "").strip())
has_tts = bool((self.voice_id or "").strip()) and bool((self.script_text or "").strip())
if not has_audio and not has_tts:
raise ValueError(
"必须提供驱动音频:要么传 audio_url(直接/预合成音频模式),"
"要么同时传 voice_id + script_text(TTS 直生模式)"
)
if has_tts and len(self.script_text) > 5000:
raise ValueError("script_text 最长 5000 字符")
if has_audio:
au = self.audio_url.strip()
if not au.startswith(("http://", "https://")):
raise ValueError("audio_url 必须是 HTTP/HTTPS URL")
au_lower = au.lower().split("?")[0]
allowed = (".mp3", ".aac", ".wav", ".m4a", ".flac")
if not any(au_lower.endswith(ext) for ext in allowed):
raise ValueError(f"audio_url 格式不支持,仅支持: {', '.join(allowed)}")
self.audio_url = au
return self
# ── #1845 TTS 预合成接口 ────────────────────────────────────────────────
class AiAvatarTtsPreviewRequest(BaseModel):
"""步骤1「生成配音」预合成请求(同步 HTTP,~2-3s)."""
voice_id: str = Field(..., min_length=1, max_length=128, description="音色 ID")
script_text: str = Field(..., min_length=1, max_length=5000, description="要合成的文案")
speed: float = Field(1.0, ge=0.5, le=2.0, description="语速(0.5-2.0),默认 1.0")
emotion: str = Field(
"neutral",
max_length=32,
description="情绪(英文枚举 neutral/happy/sad/angry/surprised/fearful/disgusted,或中文 中立/开心/难过/生气/惊讶/恐惧/厌恶;默认 neutral)",
)
class AiAvatarTtsPreviewResponse(BaseModel):
"""TTS 预合成响应(临时 URL,24h 内有效,足够当前会话使用)."""
audio_url: str = Field(..., description="CosyVoice 临时音频 URL")
duration: float = Field(..., ge=0, description="音频总时长(秒),ffprobe 测得")
sentence_timings: list[dict] = Field(..., description="句子级精确时间戳")