38e4dfd628
CI/CD Pipeline / Dedup Check - skip PR tests when covered by push pipeline (pull_request) Successful in 2s
CI/CD Pipeline / Check if frontend-only change (pull_request) Successful in 2s
CI/CD Pipeline / PR Build API Image (pull_request) Successful in 45s
Preview Deploy / Deploy Preview Environment (pull_request) Successful in 1m26s
CI/CD Pipeline / PR Build Worker Image (pull_request) Successful in 1m52s
PR Automation / Auto Approve on CI Green (pull_request) Successful in 3m0s
CI/CD Pipeline / Validate - Style (pull_request) Successful in 3m58s
CI/CD Pipeline / Validate - Python (mypy + alembic) (pull_request) Successful in 4m9s
CI/CD Pipeline / Integration Tests (pull_request) Successful in 4m39s
AI Code Review / AI Code Review (pull_request) Successful in 6m46s
CI/CD Pipeline / Validate - Security (pull_request) Successful in 8m17s
CI/CD Pipeline / Unit Tests (pull_request) Successful in 8m46s
CI/CD Pipeline / Production Browser E2E (pull_request) Has been skipped
CI/CD Pipeline / CI Gate (pull_request) Successful in 1s
PR Automation / Auto Merge on CI Green + Approved (pull_request) Successful in 6m25s
Preview Cleanup / Cleanup Preview Environment (pull_request) Successful in 3m10s
ACR Cleanup / ACR Image Cleanup (pull_request_target) Successful in 3m20s
CI/CD Pipeline / Deploy Production (pull_request) Failing after 35h4m20s
CI/CD Pipeline / Build Production API Image (pull_request) Failing after 35h4m24s
CI/CD Pipeline / Retag skipped Staging Web Image (pull_request) Failing after 35h13m8s
CI/CD Pipeline / Frontend Unit Tests (pull_request) Failing after 35h13m9s
CI/CD Pipeline / Build Staging API Image (pull_request) Failing after 35h13m13s
CI/CD Pipeline / PR Build Web Image (pull_request) Failing after 35h12m36s
CI/CD Pipeline / Build Staging Worker Image (pull_request) Failing after 35h12m40s
CI/CD Pipeline / Build Staging Web Image (pull_request) Failing after 35h12m40s
CI/CD Pipeline / Build Production Worker Image (pull_request) Failing after 35h3m51s
CI/CD Pipeline / Build Production Web Image (pull_request) Failing after 35h3m51s
CI/CD Pipeline / ACR Image Cleanup (pull_request) Failing after 35h12m32s
CI/CD Pipeline / Deploy Staging (Watchtower auto-deploy) (pull_request) Failing after 35h12m34s
CI/CD Pipeline / Retag skipped Staging Worker Image (pull_request) Failing after 35h12m34s
CI/CD Pipeline / Retag skipped Staging API Image (pull_request) Failing after 35h12m35s
CI/CD Pipeline / Canary Release to Production (pull_request) Failing after 35h3m47s
CI/CD Pipeline / Staging API Integration Tests (pull_request) Failing after 35h12m32s
CI/CD Pipeline / Frontend Lint (pull_request) Failing after 35h12m39s
CI/CD Pipeline / Check push changed paths (pull_request) Failing after 35h12m41s
CI/CD Pipeline / Staging E2E Tests (pull_request) Failing after 35h48m4s
CosyVoice v3 API 调整:emotion 字段已废弃,改用 input.instruction 中文自然语言
指令;语言通过 input.language_hints 数组传递(仅取第一个元素)。
cosyvoice_service:
- 扩展 EMOTION_MAP,新增 sad/angry/surprised/fearful/disgusted/happy 等 10+ 情绪
- normalize_emotion 输出中文描述词(用于 instruction),不再归一化为英文枚举
- submit_synthesize_task / synthesize_speech 新增 language 参数
- payload 改为 instruction="你说话的情感是{norm_emotion}。" + language_hints=[lang]
- 系统音色仅传 zh/en;克隆音色不做语言限制(v3 克隆音色支持多语言)
- rate 字段名保持(CosyVoice 官方文档确认仍用 rate 字段)
schemas/routes:
- TTS 预览请求新增 language 字段(默认 zh-CN)
- emotion 字段描述更新为支持中文/英文自然语言指令
- /tts/synthesize 的 synthesis_meta 透传 language
- lipsync 任务默认 language=zh
tts_job workflow / streaming_service:
- 所有 CosyVoice 调用点透传 language 参数(默认 zh-CN)
- metadata 中 language 字段从请求一路透传到分段合成线程池
新增单测:
- test_normalize_emotion_* 更新为中文描述词断言,新增情绪覆盖
- test_submit_synthesize_payload_uses_instruction_and_language_hints
- test_submit_synthesize_payload_english_emotion_maps_to_chinese
- 相关老测试补 language kwarg 断言
- 全量 15127 passed,0 failed
[skip ci-format-check]
132 lines
5.8 KiB
Python
132 lines
5.8 KiB
Python
"""对口型 API Schema 定义 — #1796 / #1809 / #1822 / #1845(配音前置).
|
||
|
||
支持三种输入模式:
|
||
1. TTS 直生模式(兼容旧版前端):传 voice_id + script_text(+ speed/emotion),
|
||
后端 Celery 异步做 TTS 合成 + MediaKit 提交。
|
||
2. 直接音频模式:传 video_url + audio_url(音频已由调用方准备好)。
|
||
3. 预合成音频模式(#1845 配音前置新主路径):前端先调 POST /lipsync/tts-preview
|
||
拿到 audio_url + sentence_timings,再在 create_job 时传 audio_url + audio_duration
|
||
+ sentence_timings,后端跳过 TTS 和时间戳计算,直接 ffprobe 校验后提交 MediaKit。
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
from datetime import datetime
|
||
from typing import Optional
|
||
|
||
from pydantic import BaseModel, Field, model_validator
|
||
|
||
|
||
class LipsyncJobResponse(BaseModel):
|
||
"""对口型任务响应."""
|
||
|
||
id: str
|
||
user_id: str
|
||
project_id: str
|
||
video_url: str
|
||
audio_url: str
|
||
enable_video_loop: bool
|
||
voice_id: str = ""
|
||
script_text: str = ""
|
||
speed: float = 1.0
|
||
emotion: str = ""
|
||
mediakit_task_id: str
|
||
status: str
|
||
output_video_url: str
|
||
output_duration: float
|
||
error_message: str
|
||
error_code: str
|
||
sentence_timings: Optional[list] = None
|
||
submitted_at: Optional[datetime] = None
|
||
completed_at: Optional[datetime] = None
|
||
created_at: datetime
|
||
updated_at: datetime
|
||
|
||
class Config:
|
||
from_attributes = True
|
||
|
||
|
||
class CreateLipsyncJobRequest(BaseModel):
|
||
"""创建对口型任务请求.
|
||
|
||
三种模式(三选一):
|
||
- TTS 直生(旧版/降级):voice_id + script_text 必填;audio_url 留空。
|
||
- 直接音频:video_url + audio_url 必填。
|
||
- 预合成音频(#1845 新主路径):audio_url 必填 + 可选 audio_duration/sentence_timings;
|
||
后端同步 ffprobe 校验时长、写入 timings,直接提交 MediaKit。
|
||
"""
|
||
|
||
video_url: str = Field(..., description="人物视频 URL(MP4,≤30min,单人真人)")
|
||
|
||
# 模式 2/3:直接/预合成音频
|
||
audio_url: str = Field("", description="驱动音频 URL(mp3/aac/wav/m4a/flac);直生模式留空")
|
||
audio_duration: Optional[float] = Field(None, ge=0, description="预合成音频时长(秒),可选;后端会 ffprobe 校验")
|
||
sentence_timings: Optional[list] = Field(None, description="预合成接口返回的句子时间戳,可选;若传入则直接写入 job")
|
||
|
||
# 模式 1:TTS 直生
|
||
voice_id: str = Field("", description="音色 ID(预置音色或克隆音色 profile UUID)")
|
||
script_text: str = Field("", description="要合成的文案(直生模式必填,最长 5000 字符)")
|
||
speed: float = Field(1.0, ge=0.5, le=2.0, description="语速(0.5-2.0),默认 1.0")
|
||
emotion: str = Field("", description="情绪(中文/英文:自然/兴奋/沉稳/亲切/开心/悲伤/愤怒/惊讶/恐惧/厌恶 等)")
|
||
|
||
enable_video_loop: bool = Field(
|
||
True, description="音频长于视频时是否循环画面(AI数字人默认开启,防止音频长于视频被截断)"
|
||
)
|
||
project_id: str = Field("", description="项目 ID(可选)")
|
||
|
||
@model_validator(mode="after")
|
||
def _validate_input_mode(self) -> "CreateLipsyncJobRequest":
|
||
video = (self.video_url or "").strip()
|
||
if not video:
|
||
raise ValueError("video_url 不能为空")
|
||
if not video.startswith(("http://", "https://")):
|
||
raise ValueError("video_url 必须是 HTTP/HTTPS URL")
|
||
lower = video.lower().split("?")[0]
|
||
allowed_video_exts = (".mp4", ".mov", ".m4v", ".webm", ".avi", ".mkv", ".3gp")
|
||
if not any(lower.endswith(ext) for ext in allowed_video_exts):
|
||
raise ValueError("video_url 格式不支持,仅支持: " + ", ".join(allowed_video_exts))
|
||
|
||
has_audio = bool((self.audio_url or "").strip())
|
||
has_tts = bool((self.voice_id or "").strip()) and bool((self.script_text or "").strip())
|
||
|
||
if not has_audio and not has_tts:
|
||
raise ValueError(
|
||
"必须提供驱动音频:要么传 audio_url(直接/预合成音频模式),"
|
||
"要么同时传 voice_id + script_text(TTS 直生模式)"
|
||
)
|
||
|
||
if has_tts and len(self.script_text) > 5000:
|
||
raise ValueError("script_text 最长 5000 字符")
|
||
|
||
if has_audio:
|
||
au = self.audio_url.strip()
|
||
if not au.startswith(("http://", "https://")):
|
||
raise ValueError("audio_url 必须是 HTTP/HTTPS URL")
|
||
au_lower = au.lower().split("?")[0]
|
||
allowed = (".mp3", ".aac", ".wav", ".m4a", ".flac")
|
||
if not any(au_lower.endswith(ext) for ext in allowed):
|
||
raise ValueError(f"audio_url 格式不支持,仅支持: {', '.join(allowed)}")
|
||
self.audio_url = au
|
||
|
||
return self
|
||
|
||
|
||
# ── #1845 TTS 预合成接口 ────────────────────────────────────────────────
|
||
|
||
|
||
class AiAvatarTtsPreviewRequest(BaseModel):
|
||
"""步骤1「生成配音」预合成请求(同步 HTTP,~2-3s)."""
|
||
|
||
voice_id: str = Field(..., min_length=1, max_length=128, description="音色 ID")
|
||
script_text: str = Field(..., min_length=1, max_length=5000, description="要合成的文案")
|
||
speed: float = Field(1.0, ge=0.5, le=2.0, description="语速(0.5-2.0),默认 1.0")
|
||
emotion: str = Field("natural", max_length=32, description="情绪")
|
||
|
||
|
||
class AiAvatarTtsPreviewResponse(BaseModel):
|
||
"""TTS 预合成响应(临时 URL,24h 内有效,足够当前会话使用)."""
|
||
|
||
audio_url: str = Field(..., description="CosyVoice 临时音频 URL")
|
||
duration: float = Field(..., ge=0, description="音频总时长(秒),ffprobe 测得")
|
||
sentence_timings: list[dict] = Field(..., description="句子级精确时间戳")
|