Qwen 3.7 Max vs Kimi K3 vs DeepSeek V4: 34x Cost Gap (2026)

Same eval suite, same methodology: Kimi K3 scores 57 and costs $2,437 to run it. DeepSeek V4 Flash scores 50 and costs $72. Qwen 3.7 Max scores 46 for $1,604.

Qwen 3.7 Max vs Kimi K3 vs DeepSeek V4: 34x Cost Gap (2026)

Three Chinese flagship families, one benchmark suite, and a 34x spread in what it costs to finish it. The capability gap between the best and worst of them is 13 index points. The bill gap is $2,365.

  • Most capable: Kimi K3 (max), 57 on Artificial Analysis Intelligence Index v4.1
  • Best value by a distance: DeepSeek V4 Flash 0731, 50 for $72.02 per index run
  • Fastest: Qwen 3.7 Max, 206.4 output tokens/sec
  • The surprise: the cheapest model in the group outscores both flagships it undercuts
  • All four are on ofox today, so switching is a model-ID change

Three podium blocks of nearly identical height, each fronted by a stack of round tokens of wildly different heights, the tallest stack in bright orange, flat isometric on warm grey

TL;DR: Which One Should You Pick?

Your situationPickWhy
Hardest reasoning, budget is not the constraintKimi K3Top score of the four, 57 on the index
High-volume agent loops, cost dominatesDeepSeek V4 Flash 0731Scores 50, and the same eval run cost 34x less than K3
Latency matters more than the last few pointsQwen 3.7 Max206.4 tok/s against K3’s 35
You need image inputKimi K3 or Qwen 3.8 MaxThe other three are text-only in the catalog
You are on DeepSeek V4 Pro todayRe-evaluateFlash 0731 scores 6 points higher and costs 2.4x less to run
You are picking a default for a mixed workloadDeepSeek V4 Flash 0731Best score-per-dollar by an order of magnitude, and 1M context
You need the newest QwenQwen 3.8 MaxCheaper than 3.7 Max, but no third-party score yet

Two ways to stop reading here. If your monthly token spend is under about $200, pick DeepSeek V4 Flash and move on, because none of the optimisation below is worth the afternoon. If you are running production agent loops above $2,000 a month, go straight to the eval-cost section, because that is where the 34x lives.

Quick Specs Comparison

All figures below are from the ofox model catalog, pulled live on 2026-08-03.

Kimi K3DeepSeek V4 FlashDeepSeek V4 ProQwen 3.7 Max
Model IDmoonshotai/kimi-k3deepseek/deepseek-v4-flashdeepseek/deepseek-v4-probailian/qwen3.7-max
Input / 1M$3.00$0.14$0.45$2.50
Output / 1M$15.00$0.28$0.88$7.50
Cache read / 1M$0.30$0.0028$0.0037$0.50
Context window1,048,5761,000,0001,000,0001,064,000
Max output1,048,576384,000384,00064,000
Input modalitiesText, imageTextTextText
Open weightsYesYes (MIT)YesNo

Three things in that table are worth pausing on.

The cache-read column is not a rounding difference. DeepSeek charges $0.0028 per 1M cached input tokens. Qwen 3.7 Max charges $0.50 for the same thing, which is 179x more, and K3 charges $0.30, which is 107x more. For a workload with a large stable system prompt read many times, the cached-read rate is the number that decides your bill, not the headline input rate.

Qwen 3.7 Max’s 64K output ceiling is the tightest of the four by a factor of six against DeepSeek and sixteen against K3. If your task is “read a large repo and emit a large artefact,” that ceiling bites before anything else does. Qwen 3.8 Max lifts it to 131K, covered in the Qwen 3.8 Max launch breakdown.

Only K3 has open weights and a frontier score. DeepSeek’s weights are MIT-licensed and downloadable, but its top scorer here is the cheap model, not the flagship. Qwen 3.7 Max has no open weights at all, though Alibaba has announced a Max-class open-weight release for the week of 2026-08-10.

How Do They Score on the Same Benchmark?

Artificial Analysis Intelligence Index v4.1, snapshot 2026-08-03. This is the only source in this article that measures all four models the same way, which is why the whole comparison is anchored on it.

ModelIndex scoreRank in this group
Kimi K3 (max)571
DeepSeek V4 Flash 0731502
Qwen 3.7 Max463
DeepSeek V4 Pro (Reasoning, Max Effort)444

The index is a composite of nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR.

Read the ordering, not the absolute numbers. This is a rolling leaderboard that gets re-scored as the index version and the model builds change, so “K3 first, Flash second” will age better than “57 and 50”. The four-point gap between Flash and Qwen 3.7 Max is inside the range a re-run can move.

What does not move is the direction of the surprise: the cheapest model in the group sits above two flagships, one of which is from the same vendor.

Three things would change this ranking, and all three are plausible inside a quarter:

  • Qwen 3.8 Max gets scored. It is cheaper than 3.7 Max and Alibaba’s own table puts it well ahead of its predecessor. If it lands near K3 at $2.00 / $6.00, the value column shifts.
  • A new DeepSeek build lands. The 0731 retrain moved Flash 10 points with no architecture change. Nothing says that was the last one.
  • The index version changes. v4.1 added and reweighted evaluations; a v4.2 could reorder models separated by four points without any model changing at all.

That is the case for re-running your own evals quarterly rather than inheriting a ranking from an article, including this one.

Each vendor also publishes its own table, and those are not comparable with each other. Qwen’s launch table for 3.8 Max benchmarks against Claude and GPT with harnesses Qwen selected. Moonshot’s K3 launch table reports Terminal-Bench 2.1 at 88.3 against a leader of 88.8, which is a second place rather than a win. DeepSeek’s own Terminal-Bench figure for V4 Flash runs above the one Artificial Analysis measured for the same model. Those are different measurement setups, and mixing them into one table is how a comparison ends up meaningless.

What Does One Eval Run Actually Cost?

This is the number that reframes the whole comparison. Artificial Analysis publishes the total cost of running its own index on each model, so it is a real invoice for an identical workload rather than a sticker-price multiplication.

ModelIndex scoreOutput tokens generatedCost to run the index
Kimi K3 (max)57130M$2,437.41
Qwen 3.7 Max46100M$1,604.13
DeepSeek V4 Pro44180M$176.34
DeepSeek V4 Flash 073150210M$72.02

Derived from those four rows:

  • K3 costs 33.8x what V4 Flash costs, for +7 index points.
  • Qwen 3.7 Max costs 22.3x what V4 Flash costs, and loses by 4 points.
  • V4 Pro costs 2.4x what V4 Flash costs, and loses by 6 points.
  • K3 costs 1.5x what Qwen 3.7 Max costs, for +11 points.

One note on the price columns: Artificial Analysis lists DeepSeek V4 Pro at $0.435 / $0.87 while the ofox catalog rounds it to $0.45 / $0.88. The eval-cost figures come from Artificial Analysis and use its own numbers; the specs table above uses the catalog’s.

Only one of those four comparisons describes a trade. The other three describe a model that is worse and more expensive at the same time.

The K3 premium is at least defensible: you are paying 1.5x over Qwen 3.7 Max for a genuine 11-point lift, and if your work is bottlenecked on the hardest reasoning, that is a rational purchase. Paying 22x over V4 Flash to score four points lower is not a trade, it is an artefact of nobody re-running the numbers after 2026-07-31.

Why Is the Cheapest Model Beating Two Flagships?

Because the 0731 build was a retrain, not a new model, and it moved the score 10 points. DeepSeek V4 Flash 0731 kept the same 284B-total, 13B-active MoE architecture and came back scoring 50 against the previous build’s 40.

The second half of the answer is verbosity, and it cuts against DeepSeek.

  • V4 Flash generated 210M output tokens on the index, the most of the four.
  • Qwen 3.7 Max generated 100M, the least.
  • So Flash emits 2.1x the tokens Qwen does for the same suite.

That is why the realised gap is smaller than the sticker gap. Qwen 3.7 Max’s output rate is 26.8x Flash’s ($7.50 against $0.28), but the actual index bills came out 22.3x apart. Flash gives back roughly a fifth of its price advantage by talking more, and it still wins by a mile.

Thinking is on by default on DeepSeek and reasoning tokens bill as output, which is where the verbosity comes from. It can be turned off. If you do that, you also give up part of the score that put Flash above two flagships in the first place, so measure both ways on your own prompts before deciding.

The practical version of this section: sticker price per token is a bad proxy for cost. Two models with a 27x rate difference finished the same work 22x apart, and two models with a 1.5x rate difference finished 2.4x apart. Token counts vary between models more than prices do.

What Does This Cost You Per Month?

A benchmark suite is not your workload. Here is the same comparison on a concrete team, using the ofox catalog rates from the specs table.

Scenario: 5 developers running a coding agent.

  • 40 tasks per developer per day, 21 working days
  • Median task: 12,000 input tokens of repo context, 3,000 output tokens
  • 40% of input hits the prompt cache

Monthly volume per developer: 10.08M input tokens (4.03M cached, 6.05M uncached) and 2.52M output tokens.

Cached inputUncached inputOutputPer devTeam of 5
Kimi K3$1.21$18.14$37.80$57.15$285.77
Qwen 3.7 Max$2.02$15.12$18.90$36.04$180.18
DeepSeek V4 Pro$0.01$2.72$2.22$4.95$24.77
DeepSeek V4 Flash$0.01$0.85$0.71$1.56$7.82

That lands at 36.5x between K3 and Flash, close to the 33.8x the index run showed. Two completely different workloads producing the same ratio is a reasonable sign the ratio is real.

One correction to that table, in Flash’s disfavour. It assumes every model emits 3,000 output tokens per task, and the index run says that is wrong: Flash generated 2.1x the output tokens Qwen did for identical work. Scale Flash’s output line by 2.1x and its team cost goes from $7.82 to about $11.70 a month, which moves the gap against K3 from 36.5x down to roughly 24x.

Still 24x. The verbosity tax is real and it does not come close to closing the gap.

Where this scenario would mislead you:

  • Cache hit rate is doing a lot of work here. At 40% hits, DeepSeek’s near-free cached reads barely register because the absolute numbers are small. On a workload with a 200K-token system prompt read hundreds of times a day, that column becomes the entire bill and the gap widens further.
  • It prices tokens, not outcomes. If K3 finishes a task that Flash retries three times, the retry cost and the engineer time swamp everything above.
  • Rates move. Both DeepSeek builds have changed price and score inside the last quarter. Pull current numbers before you commit.

How Fast Are They?

Qwen 3.7 Max is 5.9x faster than Kimi K3 on output throughput.

ModelOutput tokens/sec
Qwen 3.7 Max206.4
DeepSeek V4 Pro55.0
Kimi K3 (max)35
DeepSeek V4 Flash 0731no published figure

Artificial Analysis flags K3 as notably slow and V4 Pro as below average, while Qwen 3.7 Max sits well above average for its price class. There was no speed measurement published for the 0731 Flash build at the time of writing, so that row stays empty rather than borrowing the previous build’s number.

Speed changes different things than score does. A 12-point index gap shows up as tasks the model can or cannot finish. A 5.9x throughput gap shows up as whether a 40-minute agent run becomes a four-hour one. For interactive tooling and long autonomous loops, the second is often the constraint people actually feel.

When Should You Pick Kimi K3?

When the task fails on the other three and the failure costs more than the tokens.

Concrete signals:

  • Your evals show the other models failing a class of task rather than doing it slightly worse.
  • You need image input alongside frontier reasoning. We tested this through the gateway on 2026-08-03: sending an image raised prompt tokens from 97 to 189 and returned an accurate description, so vision is live and not just listed.
  • You need a very large single response. K3’s 1,048,576 max output is 16x Qwen 3.7 Max’s ceiling.
  • You want open weights on the top scorer, which K3 is and Qwen 3.7 Max is not.

Signals it is the wrong pick: high request volume, latency-sensitive UX, or any workload where you have not actually measured a quality difference. At 35 tokens per second and $15 per 1M output, K3 punishes both volume and impatience. The Kimi K3 against GPT-5.5 and Opus 4.8 comparison covers how it holds up outside the Chinese-model bracket, and how to use Kimi K3 covers setup.

When Should You Pick DeepSeek V4 Flash?

As the default, until something makes it fail.

It scores second of four, costs least by an order of magnitude, has a 1M context window and a 384K output ceiling, and ships MIT-licensed open weights. The case against it is narrow:

  • Text only. No image input, so any pipeline touching screenshots or documents-as-images is out.
  • Verbose. 210M output tokens on the index, the most of the four. Budget for reasoning tokens, and cap max_tokens on long loops.
  • 7 points below K3. Real, and it will matter for the hardest slice of your work.

If you are currently on DeepSeek V4 Pro, this section is the one to act on. Flash 0731 scores 6 points higher and its index run cost 2.4x less. Any routing rule that sends hard work to Pro and easy work to Flash was written against a build that no longer exists. The V4 Pro against V4 Flash breakdown has the older framing and the update note, and V4 Flash against Gemini 3.6 Flash covers the cross-vendor budget bracket.

When Should You Pick Qwen?

When throughput is the constraint, or when you are already inside the Alibaba ecosystem.

Qwen 3.7 Max’s argument in this group is speed: 206.4 output tokens per second, nearly 4x V4 Pro and nearly 6x K3. For a chat product where users watch tokens appear, or an agent loop measured in wall-clock rather than dollars, that is a real product difference that no index score captures.

The argument against it, on these numbers, is that it is the middle option on capability and near the top on cost, which is an awkward place to be. It scores 4 points below a model that costs 22x less to run the same suite.

And it is no longer Alibaba’s newest. Qwen 3.8 Max shipped 2026-08-03 at $2.00/$6.00, cheaper than 3.7 Max on both sides, with a 131K output ceiling and image input. It is not in this comparison because Artificial Analysis had no page for it at the time of writing, so there is no like-for-like score to place it with. If you are choosing a Qwen today rather than comparing published numbers, 3.8 Max is the one to start from. The 3.7 Plus against 3.7 Max benchmark covers the cheaper sibling, and the 3.7 Max developer guide covers the verbosity tax on long agent sessions.

When Should You Pick None of These?

Three cases where the whole bracket is the wrong shelf:

  • You need video input. None of the four take it through the gateway. The Qwen 3.8 Max catalog entry lists text and image only, even though QwenCloud’s own page lists video.
  • You need a specific frontier capability these do not have. If your evals are calibrated against Claude or GPT behaviour, a 7-point index difference will not tell you whether the swap is safe. Run your own suite.
  • Your bottleneck is not the model. If a retrieval bug or a bad prompt is costing you more than the token bill, none of the numbers in this article will help.

Try All Three via ofox: A/B in 10 Lines

All four models sit in the same catalog, so the comparison is a loop over model IDs rather than four integrations.

Python: run the same prompt through all four

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.getenv("OFOX_API_KEY"),
    base_url="https://api.ofox.run/v1",
)

MODELS = [
    "moonshotai/kimi-k3",
    "deepseek/deepseek-v4-flash",
    "deepseek/deepseek-v4-pro",
    "bailian/qwen3.7-max",
]

prompt = "Refactor this module and explain the risk in the migration."

for model in MODELS:
    r = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": prompt}],
        max_tokens=2048,
    )
    u = r.usage
    print(f"{model:<32} out={u.completion_tokens:<7} total={u.total_tokens}")

Node: same shape

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.OFOX_API_KEY,
  baseURL: "https://api.ofox.run/v1",
});

for (const model of ["moonshotai/kimi-k3", "deepseek/deepseek-v4-flash", "bailian/qwen3.7-max"]) {
  const r = await client.chat.completions.create({
    model,
    messages: [{ role: "user", content: prompt }],
    max_tokens: 2048,
  });
  console.log(model, r.usage.completion_tokens);
}

Print completion_tokens, not just latency. The whole point of the eval-cost table above is that token counts differ more between these models than prices do, so a per-model output count on your own prompts is worth more than any published index.

One caveat on the catalog. K3’s entry lists /v1/chat/completions only, while the Qwen and DeepSeek entries also list /v1/responses. That field has been wrong in both directions before, so send one real request against the protocol you plan to use rather than planning around the metadata. Current per-model rates are on the model pages for Kimi K3, DeepSeek V4 Flash and Qwen 3.7 Max.

FAQ

Which is the best Chinese model in 2026? On the Artificial Analysis Intelligence Index v4.1 at the 2026-08-03 snapshot, Kimi K3 (max) at 57 is the highest scorer of this group. On score per dollar, DeepSeek V4 Flash 0731 wins by roughly an order of magnitude.

Is Kimi K3 worth the price over DeepSeek V4 Flash? It buys 7 index points for 33.8x the cost of the same eval run. That is worth it only when a specific class of task fails on Flash and succeeds on K3, which you have to measure on your own prompts.

Does DeepSeek V4 Flash have open weights? Yes. The 0731 build is on Hugging Face as deepseek-ai/DeepSeek-V4-Flash-0731 under MIT, in fp8, not gated.

Why is Qwen 3.7 Max in this comparison instead of 3.8 Max? Because 3.8 Max had no Artificial Analysis page when this was written, so there is no score measured the same way as the other three. Qwen 3.7 Max is the newest Qwen with a comparable third-party number.

Which model has the largest output ceiling? Kimi K3, at 1,048,576 tokens. DeepSeek’s two models allow 384,000 and Qwen 3.7 Max allows 64,000.

Do any of these support vision? Kimi K3 does, tested through the gateway on 2026-08-03. Qwen 3.8 Max does. Qwen 3.7 Max and both DeepSeek V4 models are text-only in the catalog.

Is the Artificial Analysis index cost the same as my bill will be? No. It is one specific 9-evaluation suite, so it tells you the ratio between models on a fixed workload, not your absolute spend. Use it to rank options, then measure your own prompts.

Sources Checked for This Refresh

Frequently Asked Questions

Which is better, Qwen 3.7 Max, Kimi K3, or DeepSeek V4?
On the Artificial Analysis Intelligence Index v4.1 (snapshot 2026-08-03), Kimi K3 (max) scores 57, DeepSeek V4 Flash 0731 scores 50, Qwen 3.7 Max scores 46 and DeepSeek V4 Pro scores 44. So K3 is the most capable of the four and DeepSeek V4 Flash is second, ahead of both flagships that cost far more per token. Capability order is not the same as value order: running the same index cost $2,437.41 on K3 and $72.02 on V4 Flash.
How much does it cost to run the same workload on each model?
Artificial Analysis publishes the total cost of its own Intelligence Index run per model, which is the cleanest like-for-like number available. On the 2026-08-03 snapshot: Kimi K3 (max) $2,437.41, Qwen 3.7 Max $1,604.13, DeepSeek V4 Pro $176.34, DeepSeek V4 Flash 0731 $72.02. That is a 33.8x spread between the most and least expensive, for a 7-point spread in index score.
Is DeepSeek V4 Flash really better than DeepSeek V4 Pro?
On this index, yes. V4 Flash 0731 scores 50 against V4 Pro's 44, and the index run cost $72.02 on Flash against $176.34 on Pro. The 0731 build was a retrain on the same architecture rather than a new model, and it moved the Flash score up 10 points. If your V4 Pro routing decision predates 2026-07-31, it is worth re-running your evals before renewing it.
Why is Kimi K3 so expensive?
Two things stack. Its token prices are the highest of the group at $3.00 input and $15.00 output per 1M, and it is slow, at 35 output tokens per second on the Artificial Analysis measurement. The token price is what shows up on the invoice; the speed is what shows up in your agent loop wall-clock. K3 was still the top scorer of the four, so the premium buys something real, just not cheaply.
What about Qwen 3.8 Max?
Qwen 3.8 Max launched on 2026-08-03 at $2.00 input and $6.00 output per 1M, cheaper than 3.7 Max on both sides, with a 131K output ceiling instead of 64K. It is deliberately not in the table above: Artificial Analysis had no page for it at the time of writing, so there is no score measured the same way as the others. Adding a vendor-reported number to a third-party table would break the comparison this article is built on.
Which of these models supports image input?
Kimi K3 does, and Qwen 3.8 Max does. Qwen 3.7 Max, DeepSeek V4 Pro and DeepSeek V4 Flash are listed as text-only in the ofox catalog. We tested K3 through the gateway on 2026-08-03: an image raised prompt tokens from 97 to 189 and came back with an accurate description, so image input is live rather than merely listed.
Can I run all three behind one API key?
Yes. All four models are in the ofox catalog as bailian/qwen3.7-max, moonshotai/kimi-k3, deepseek/deepseek-v4-pro and deepseek/deepseek-v4-flash, so switching between them is a model-ID change rather than a new integration. That is what makes an A/B on your own prompts cheap enough to actually run, which matters more than any published index.
Which model is fastest?
Qwen 3.7 Max, by a wide margin. Artificial Analysis measures it at 206.4 output tokens per second, against 55.0 for DeepSeek V4 Pro and 35 for Kimi K3, which its page describes as notably slow. There was no speed figure published for V4 Flash 0731 at the time of writing. For a long agent loop, a 5.9x speed difference changes the shape of the workday more than a few index points do.