All articles
AI News

Grok 4.6 vs DeepSeek V4 Pro vs Qwen 3.8: The Mid-Tier Compared

Six models landed in the week of 11 August, none of them flagships. Where the real competition moved, and why DeepSeek costs a quarter of its rivals with twice the context.

Travis Johnson

Travis Johnson

Founder, Deepest

August 20, 20264 min read

Six notable models landed in the week of 11 August: Grok 4.6, DeepSeek V4 Pro, Qwen 3.8 2.4T, two ByteDance Seed models and Gemini 3.7 Flash. None of them is a generational leap. Collectively they say something more useful than any single launch would, which is that the interesting competition has moved from the top of the market to the middle of it.

What shipped

  • Grok 4.6 (11 Aug): $2 in, $6 out. 500K context, and a 450,000-token maximum output, far beyond anything comparable.
  • DeepSeek V4 Pro 0813 (12 Aug): $0.46 in, $1.39 out. 1M context, 384K output.
  • Qwen 3.8 2.4T A95B (12 Aug): $2 in, $6 out. A 2.4-trillion-parameter mixture of experts with 95B active.
  • ByteDance Seed 2.1 Turbo (12 Aug): $0.50 in, $2.50 out, with image and video input.
  • ByteDance Seed 2.0 Code (12 Aug): $0.50 in, $3 out, aimed at coding.
  • Gemini 3.7 Flash (13 Aug): $0.75 in, $3.75 out, with audio and video input.
Key Finding: DeepSeek V4 Pro costs $0.46 per million input tokens against Grok 4.6 and Qwen 3.8's $2, roughly a quarter of the price, with twice the context window. The old assumption that price tracks capability in a straight line stopped being reliable some time ago, and this week makes the point plainly.

The middle is where the fight is now

Frontier launches get the coverage, and the frontier is genuinely crowded. But almost nothing released this week is trying to be the most capable model available. They are trying to be the best option at a particular price, and several succeed.

That is a healthier market than it sounds. A frontier model wins by being the best at everything, which means one winner. A mid-tier model wins by being the best at something specific at a specific price, which means many winners and a genuine reason to pick between them.

The cost is that choosing got harder. When there was one obviously best model, the decision was made for you. With six credible options in a band, the decision is yours and the tools for making it are poor.

Two specs worth noticing

Grok 4.6's output ceiling

450,000 tokens of maximum output against 128,000 for most Western flagships. Irrelevant for conversation, significant for bulk generation where the alternative is chunking and losing coherence at every seam.

DeepSeek's price-to-context ratio

A million-token context window at $0.46 per million input tokens is an unusual combination. Long-context work is normally where costs escalate fastest, because you are paying for every token of a very large prompt on every turn. At this rate, working with large documents stops being a budget decision.

How to keep up without drowning

Six models in a week is not a pace anyone can evaluate properly. The realistic approach is not to evaluate everything, but to have a cheap way of checking whether anything new beats your current default on your actual work.

What we do, and recommend: keep a small set of prompts that represent your real workload, ten or so, with known-good answers. When something new looks interesting on price or specs, run the set through it alongside your current model and read the results side by side. It takes minutes rather than an afternoon, and it answers the only question that matters, which is whether this one is better for you.

Most of the time the answer is no and you have spent five minutes. Occasionally it is yes, and you have found something a leaderboard would never have told you, because leaderboards measure average performance on standardised tasks and you do not have average tasks.

Frequently Asked Questions

Should I switch to whatever launched most recently?

No. Newer means newer, not better for you, and switching has real costs: re-running evaluations, adjusting prompts that were tuned to the old model's habits, and re-establishing trust in the output. Switch when you have measured an improvement on your own work, not when a launch post is convincing.

Is a 2.4-trillion-parameter model better than a smaller one?

Not reliably. Parameter counts describe the model, not its performance on your task, and mixture-of-experts architectures activate only a fraction of their parameters per token anyway. Qwen 3.8 2.4T activates 95B. Treat parameter counts as trivia rather than as a quality signal.

How much does it cost to test six models properly?

At current mid-tier prices, a ten-prompt comparison across six models costs a few cents of inference. The expensive part is your time, which is exactly what makes running them simultaneously rather than sequentially worth doing.

Grok 4.6DeepSeekQwenByteDanceGeminiAI newsmodel comparison

See it for yourself

Run any prompt across ChatGPT, Claude, Gemini, and 300+ other models simultaneously.30-day money-back guarantee.

Try Deepest free →

Related articles