DeepSeek 視覺模型多模態 Agent 能力逼近 Opus 4.8,4 項打成 2:2
ChainCatcher 消息,DeepSeek 公布 V4-Flash-Vision-Exp 首批 Agent 跑分。在 4 項多模態 Agent 評測中與 Opus 4.8 打成 2 勝 2 負:Agents' Last Exam 27.3 對 25.7,ZeroBench 35.0 對 34.0;ApexBench 36.5 對 39.4,Chartography 64.3 對 65.0。相比純文本版 V4-Flash-0731,視覺版在 ApexBench 從 26.2 升至 36.5,Agents' Last Exam 從 25.2 升至 27.3。官方注明純文本版會忽略多模態內容,故主要體現補上視覺輸入後的能力。加上視覺後文本 Agent 能力未明顯縮水,7 項文本評測中 6 項高於 V4-Flash-0731,如 DeepSWE 從 54.4 升至 59.3(超 Opus 4.8 的 58.0),Toolathlon 75.9 幾乎追平 Opus 4.8 的 76.2。該結果來自 DeepSeek 官方自測,非第三方獨立榜單;公開 Code Agent 文本任務使用 DeepSeek Harness 極簡模式,推理檔位為 max。