DeepSeek-V4-Flash 更新

9 条回复
145 次浏览

DeepSeek-V4-Flash 正式版 API 上线公测

Agent 能力大幅增强,基准测试远超 V4-Pro-Preview:

  • Terminal Bench 2.1: 82.7
  • NL2Repo: 54.2
  • Cybergym: 76.7
  • DeepSWE: 54.4
  • Toolathlon verified: 70.3
  • Agent Last Exam: 25.2
  • Automation Bench (Public): 25.1
  • DSBench-FullStack: 68.7
  • DSBench-Hard: 59.6

注 1:对于公开基准测试集中的 Code Agent 任务,正式版 DeepSeek-V4-Flash 使用 DeepSeek Harness 极简模式(即将发布)作为框架进行测试,并使用 max 档位,topp=0.95,temperature=1.0
注 2:DSBench-FullStack 是内部使用的全栈开发测试集,DSBench-Hard 是内部使用的 Coding Agent 难题测试集

正式版 V4-Flash 原生支持 Responses API 格式并针对性适配 Codex,具体配置方法请 参考文档

DeepSeek-V4-Flash-0731 的模型结构、尺寸和 DeepSeek-V4-Flash-preview 保持一致,仅重新进行了后训练。

注意:本次仅升级了 DeepSeek-V4-Flash 的 API 接口,DeepSeek-V4-Pro API 及 APP/WEB 端模型未做更改。
DeepSeek-V4-Pro 正式版将会尽快发布。

前排打手

后训练这么牛逼吗,flash 模型直接干翻 pro 模型,顺带把隔壁的旗舰模型也干翻了?!

马上来

FYI,flash 版只有 284B-A13B,而 pro 版有 1.6T,隔壁 glm5.2 也有 744B-A40B
不知道 1.6T 的 pro 版到底有多强think
不过 pro 的激活参数只有 49B,不知道会不会拖后腿

发表一个评论

R保持