Four tasks, total
Deepseek v4 Flash
DeepSeek省钱最省的一档,智力到一线水平。页面成型后的后续调整交给它。
- 页面成型后的后续调整
- 改文案、换图片、调间距
MODEL SELECTION
Four tasks run in sequence; the raw output is here and every step opens for you to check. No scoring, no ranking.
The short answer: use Gemini 3 Flash for new pages, then switch to Deepseek v4 Flash for follow-up edits.
COST AND TASK FIT
These are the credits each model actually spent on the same four tasks, taken from billing, not estimated. More expensive is not better — for small, well-scoped edits the cheapest tier is usually enough.
Four tasks, total
最省的一档,智力到一线水平。页面成型后的后续调整交给它。
Four tasks, total
设计能力第一梯队,Design Arena 网站生成榜排到第 21;缺点是慢,四道题跑了 34 分半。
Four tasks, total
四道题里有两道跑出全场最快(新增详情页 80 秒、SEO 96 秒);遇到会影响用户流程的取舍,它会先问一句再动手。
Four tasks, total
七个模型里只有它和 Gemini 3.7 Flash 把 robots、站点地图、给 AI 读的摘要三样都写了。代价是慢,四道题跑了 38 分钟,全场最慢。
Four tasks, total
设计稳定、有个性,做新页面的首选。其余模型的花费都以它为参照。
Four tasks, total
第三方设计榜上七个里排第二(Website 榜 #7)。信息密度和收尾最完整,代价是比基准档贵一半。
Four tasks, total
Design Arena 网站生成榜第 1,也是这里最贵的一档——同样四道题它花掉 179.6 积分,榜单方自己记录的平均生成耗时 627.9s。
暂未对外开放,放在这里只作天花板参照——它在 Design Arena 网站生成榜上排第 1,看过它的产物再回头看开放档位,才知道差距有多大、值不值那个价。
Bars use a logarithmic scale, for sensing magnitude only, in credits(6–179.6)
THIRD-PARTY DATA
All figures below are cited from public third-party leaderboards. Creght did not run these evaluations or supply any data. Follow the heading link to verify at the source.
众包盲测投票得出的 Elo 评分,专门评测「生成网站」这一类任务,与本页场景最接近。
Snapshot2026-08-18
共 161 个模型参评。数据为 2026-08-18 快照,非实时。
Bars use that leaderboard's actual range(783–1364)
由 9 项评测合成的通用能力指数,衡量推理、知识、数学与代码,不代表视觉设计能力。
Snapshot2026-08-18
通用智力指数,与「页面好不好看」不是同一回事,看设计请看左侧一列。
Bars use that leaderboard's actual range(0–100)
THE FOUR TASKS
All seven models got this same set. One task, one thing tested.
考设计。刻意不给配色、字体、排版任何指示。
Prompt, verbatim
为一家做手冲咖啡豆的独立烘焙工作室做一个单页官网首页,包含:品牌主视觉区、三款主推豆子的介绍、烘焙流程说明、门店地址与营业时间、底部联系方式。文案你来写,中文。
考它能不能把需求拆开做,而且新页面要接得上原来那套设计。
Prompt, verbatim
在现有站点上新增「豆子详情」页,为三款豆子各生成一个详情页,首页三款豆子的卡片分别链接到对应详情页;同时在导航栏加入「豆子」入口,移动端菜单也要能进。
考编码是否严谨。会不会用平台的后端和登录,还是把数据塞在页面里假装存下来了。
Prompt, verbatim
给站点加一个「预订豆子」的功能:访客选好豆子、填写数量和手机号后提交预订,提交成功后显示预订编号。预订记录要真正存下来,不能只留在页面上——刷新或者换个浏览器打开也要还在。另外加一个「我的预订」页面,只有登录后才能看到自己提交过的预订,没登录就引导去登录。
考智力。题面不给答案,看它自己想到该做什么。
Prompt, verbatim
帮我做一下 SEO 和 GEO 优化。我希望目标客户不管是在搜索引擎里搜,还是直接问 AI 助手,都能找到我们这家店。
ONE TABLE
Each cell is the credits and time that step actually took, from billing, not estimated. Click a model name to jump to its own section.
| Model | 1 从零生成首页 | 2 新增详情页 | 3 后端与登录 | 4 SEO 与 GEO | Four steps, total |
|---|---|---|---|---|---|
| Deepseek v4 Flash | 1.43m 52s | 1.13m 19s | 2.76m 9s | 0.82m 36s | 6.015m 56s |
| MiMo V2.5 Pro | 2.43m 51s | 5.214m 15s | 3.54m 35s | 6.111m 49s | 17.234m 30s |
| GPT-5.6 Luna | 3.01m 54s | 2.21m 20s | 7.92m 53s | 6.11m 36s | 19.27m 43s |
| Deepseek v4 Pro | 6.010m 20s | 6.37m 19s | 7.211m 47s | 4.88m 42s | 24.338m 8s |
| Gemini 3 Flash | 9.51m 10s | 30.42m 18s | 29.91m 48s | 27.71m 51s | 97.57m 7s |
| Gemini 3.7 Flash | 14.22m 38s | 25.91m 56s | 42.92m 31s | 64.23m | 147.210m 5s |
| Kimi k3 | 30.34m 51s | 52.26m 53s | 46.95m 48s | 50.25m 41s | 179.623m 13s |
The locked row, Kimi k3, is not yet available and sits here only as a ceiling reference. One note on credits: the Deepseek v4 Flash tier doubled in price after these runs, so its row reflects the older rate and would cost roughly twice as much today. Every figure here is what was actually billed at the time, not rescaled to current prices.
最省的一档,智力到一线水平。页面成型后的后续调整交给它。
Leaderboard version note
这个模型上游有过更新,榜单条目对不齐:Design Arena 同时挂着「DeepSeek-V4-Flash」和更新的「DeepSeek-V4-Flash-0731」(后者排第 38、Elo 1245),本页取的是前者;Artificial Analysis 只收录 0731 那一版。左边这一列请当作参考下限,本页的实跑产物才对应你实际会用到的版本。
设计能力第一梯队,Design Arena 网站生成榜排到第 21;缺点是慢,四道题跑了 34 分半。
四道题里有两道跑出全场最快(新增详情页 80 秒、SEO 96 秒);遇到会影响用户流程的取舍,它会先问一句再动手。
七个模型里只有它和 Gemini 3.7 Flash 把 robots、站点地图、给 AI 读的摘要三样都写了。代价是慢,四道题跑了 38 分钟,全场最慢。
Leaderboard version note
正式版 2026-08-13 才发布,两个榜单还没对齐:Artificial Analysis 收录的是当天的「DeepSeek V4 Pro 0813」,Design Arena 的条目没有版本标识、很可能仍是更早的预览版。左边这一列请当作参考下限。
See the site after all four steps从零生成首页
6.010m 20s新增详情页
6.37m 19s后端与登录
7.211m 47sSEO 与 GEO
4.88m 42s设计稳定、有个性,做新页面的首选。其余模型的花费都以它为参照。
第三方设计榜上七个里排第二(Website 榜 #7)。信息密度和收尾最完整,代价是比基准档贵一半。
See the site after all four steps从零生成首页
14.22m 38s新增详情页
25.91m 56s后端与登录
42.92m 31sSEO 与 GEO
64.23mDesign Arena 网站生成榜第 1,也是这里最贵的一档——同样四道题它花掉 179.6 积分,榜单方自己记录的平均生成耗时 627.9s。
暂未对外开放,放在这里只作天花板参照——它在 Design Arena 网站生成榜上排第 1,看过它的产物再回头看开放档位,才知道差距有多大、值不值那个价。
See the site after all four steps从零生成首页
30.34m 51s新增详情页
52.26m 53s后端与登录
46.95m 48sSEO 与 GEO
50.25m 41sThe link points to that step's version snapshot and does not change as the site is edited later
SUMMARY
Give Gemini 3 Flash the from-scratch build and the new pages, then switch to Deepseek v4 Flash once the design has settled — on the same four tasks the latter spent about one sixteenth of what Gemini did.
One identical brief, seven outputs with almost nothing in common in colour, type or layout. This is the step worth judging with your own eye.
No separate developer needed. In this run every model wired up the platform's auth and backend, and records really persist server-side — refresh, switch browsers, switch machines, still there. Easy check: submit, then reload and see if it survived.
One sentence is enough. Sitemap and robots.txt come with the platform; what the AI adds is structured data and an FAQ — exactly what an assistant quotes when someone asks it which local shop to trust.
Every open tier here is selectable in Creght right now — running one yourself is the fastest answer. For finer guidance by task type, plus tips on saving credits, see the docs.
Render diagnostics