fix(vlm): 图片VLM分析牛头不对马嘴 — 改用视觉模型 + prompt结构化强化 #2114
Reference in New Issue
Block a user
Delete Branch "fix/2114-vlm-vision-model-fix"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Root cause
packages/shared/ai_service.call_vision()误调用client.chat_completion(文本模型doubao-seed-1.6-250615)发送多模态content: [{type:text},{type:image_url}]。文本模型不识别图片 content-type → 返回 None →_step_image_analysis的 except 静默吞掉 → fallback 到占位{name:未识别}→ 下游 intent/copy/storyboard 完全没有图片信息,自然胡编。Fix
client.vision_completion(视觉模型doubao-1-5-vision-pro-250915,已在config.base.py配置),正确走多模态 chat/completions 协议_step_image_analysis健壮化:空 images 早返回;每张图独立 try/except;None/str/dict/异常分别打_source标记方便排查key_features/brand/category/colors/visual_style(同时向后兼容旧features字段)Files changed
Verification
Bug1 (P0): schema accepts 'full_ai' as alias for 'ai_full' (Pydantic field_validator normalizes) Bug2 (P0): concat_video_files probes each segment audio stream; Seedance gen_audio=False segments now marked has_audio=False so concat filter uses aevalsrc silence instead of failing with ffmpeg exit 234 Bug3 (P1): find_by_storage_key now queries (storage_key OR file_url) to cover historical data where the legacy file_url column held assets/<project>/<date>/... paths Bug4 (P1): duplicated-hit response no longer accesses non-existent domain Asset.file_url; new helper _get_existing_asset_url uses storage_key (fallback file_url) through storage_service.get_url() Bug5 (P1): _step_tts passes job.persona_id as voice_id (default longxiaochun_v3) and forces format='mp3' so downstream ffmpeg -map 1:a:0 works regardless of provider Bug6 (P1): call_video_generation omits ratio param in first-frame (image_url) mode; ai_client.video_generation ratio becomes Optional[str] and is omitted from payload when None, fixing 400 InvalidParameter from Seedance根因(VLM 牛头不对马嘴):packages/shared/ai_service.call_vision() 之前误用 client.chat_completion(文本模型 doubao-seed-1.6-250615)发送多模态 content list。 文本模型不认图片 content type → 返回 None → _step_image_analysis 的 except 静默吞掉 → fallback 到占位结果 {name:'未识别'} → 后续 intent/copy/storyboard 完全没图信息,自然胡编。 修复: 1) call_vision 改用 client.vision_completion(模型 doubao-1-5-vision-pro-250915), 走标准 OpenAI 多模态 chat/completions + image_url 格式 2) 超时提到 60s,温度降到 0.2,解析 ```json 代码块包裹 3) 增加详细 INFO 日志:打印 vision_model/image_url/prompt_len/原始返回前 400 字, 方便下次直接在 worker 日志排查 4) 重写 _IMAGE_ANALYSIS_PROMPT:强制结构化 JSON schema(category/name/brand/colors/ material_or_texture/key_features/visual_style/scene/target_audience_hint/ text_on_image),明确『无法判断就填无法判断,不许编造』,key_features 只能写外观可见特征、不许编功效 5) _step_image_analysis 健壮化:images 为空/None/文本返回/异常分别落 _source 标记; 每张图独立 try/except,单张失败不影响其他图 6) 下游 intent_parsing/copy_fusion/storyboard 兼容新字段 key_features/brand/ category/colors/visual_style(同时向后兼容旧 features 字段)🚀 预览环境已部署
🗑️ 预览环境已清理
PR #2114 已关闭或合并,对应的预览环境已被清理。