fix(vision): #2204 lite VLM关闭thinking模式,响应从10-12s降到<3s #2204
Reference in New Issue
Block a user
Delete Branch "fix/vision-v2-disable-thinking"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
根因定位
staging直连诊断发现:doubao-seed-2-1-lite-260915 是思考模型,单次调用产生 ~520 reasoning_tokens,耗时10-12s,导致fast路径(原设6s/8s超时)必超时,全部走到pro兜底(27-29s),3图53s/8图50s严重超标。
修复
thinking={type:disabled}+reasoning_effort=low关闭思考链。若API返回400不支持thinking参数,自动降级重试一次(去掉thinking字段)。预期
3图<10s、8图<15s目标可达。需要staging验证实际效果。
根因定位:staging直连测试发现 doubao-seed-2-1-lite-260915 是思考模型, 单次调用产生~520 reasoning_tokens,耗时10-12s,导致fast路径必超时。 修复: 1. vlm_fast_json.py 直接用 httpx 发最小 payload(不走 ai_client 包装), 显式设置 thinking={"type":"disabled"} + reasoning_effort="low" 关闭思考链, 期望响应降至 <3s;若API不支持thinking参数返回400,自动降级重试一次。 2. 超时恢复合理值:lite JSON 8s、OCR 6s、fast总超时8s(关闭thinking后预计<3s,余量充足)。 3. vlm_fallback.py 保持单次pro调用,pro走原ai_client路径(兜底场景对延迟不敏感,45s足够)。 预期:3图<10s/8图<15s目标可达。🚀 预览环境已部署
🗑️ 预览环境已清理
PR #2204 已关闭或合并,对应的预览环境已被清理。