Blank white background with no objects or features visible.

Découvrez TrueForge : l'infrastructure d'agents open-source et indépendante des fournisseurs. Réduisez vos coûts de 50%. Explorer maintenant→

We Checked Alibaba's Math: Qwen3.8-Max's Price Advantage Doesn't Survive Contact With Real Tasks

Par Amrutha Potluri

Published: August 6, 2026

On August 3, Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter model (95B active per inference) priced at $2/$6 per million input/output tokens, roughly 5x cheaper than GPT-5.6 Solon output. Alongside it, Alibaba published five benchmark scores, including a jump onFrontierSWE from 40.7 to 73.5 versus its predecessor.

We ran an independent check ourselves. Ten self-contained coding tasks, a separate hand-written suite (not a reproduction of FrontierSWE or DeepSWE) spanning easy string manipulation up to a dependency-order resolver and a regex-free glob matcher, through one TrueFoundry AI Gateway endpoint, against Qwen3.8-Max, GPT-5.6 Sol, and Kimi K3. We graded the generated code with real unit tests ourselves, not with another model, and ran the full suite three times per model to check whether any of it was just noise.

Qwen3.8-Max passed every task, and it still ends up the most expensive model per completed task of the three. All three models solved 10/10 tasks in most runs (Kimi K3 missed one, an empty-input edge case on the JSON-flattening task, in one of its three runs). On raw pass rate, Qwen3.8-Max looks like a clean win at a fraction of the price. Normalize by what it actually cost to get a correct answer, though, and the ranking flips. GPT-5.6 Sol landed at $0.0067 to $0.0070 per solved task across all three runs, tight and repeatable. Qwen3.8-Max landed at $0.0144 to $0.0284 per solved task: roughly2.8x GPT-5.6 Sol at the low end and over 4x at the high end, despite a 5x cheaper per-token price. Kimi K3 sat in between at $0.0084 to $0.0101.

Why? Qwen3.8-Max is dramatically more verbose on harder problems, and it occasionally spirals.Averaged across the four "hard" tasks in our suite, its completion length was 5,364 tokens per answer, against 729 for Kimi K3 and 282 for GPT-5.6 Sol: about 19x more tokens than GPT-5.6 Sol to solve the same problems, both correctly. The clearest example is the regex-free glob-matching task. One of Qwen3.8-Max's three runs generated 13,430 completion tokens and took298 seconds to answer a problem GPT-5.6 Sol solved in 3.5 seconds using roughly 200 tokens. Both got it right. That's not a rounding error, that's a five-minute coffee break versus a blink. On the JSON-flattening task, Qwen3.8-Max's completion length actually grew across our three repeated runs (6,104, then 7,752, then 13,965 tokens) while the answer stayed correct each time.The path to the right answer just kept getting longer and more expensive.

That's the piece the per-token pricing comparison misses. Six dollars per million output tokens sounds like a bargain next to GPT-5.6 Sol's thirty, until the model uses 20x more tokens to get there. Cost per correct answer, not cost per token, is the number that should drive a model-swap decision, and it isn't the number Alibaba's launch pricing leads with.

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

INSCRIVEZ-VOUS
Table des matières

Gouvernez, déployez et suivez l'IA dans votre propre infrastructure

Réservez un séjour de 30 minutes avec notre Expert en IA

Réservez une démo

Le moyen le plus rapide de créer, de gérer et de faire évoluer votre IA

Démo du livre
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Découvrez-en plus

July 20, 2023
|
5 min de lecture

LLMoPS CoE : la prochaine frontière dans le paysage MLOps

April 16, 2024
|
5 min de lecture

Cognita : Création d'applications RAG modulaires et open source pour la production

May 25, 2023
|
5 min de lecture

LLMs open source : Embrace or Perish

August 27, 2025
|
5 min de lecture

Cartographie du marché de l'IA sur site : des puces aux plans de contrôle

August 24, 2026
|
5 min de lecture

Claude Managed Agents vs. Vercel Eve: Which AI Agent Platform Should You Choose in 2026?

Aucun article n'a été trouvé.
TrueFoundry maps AI security frameworks to infrastructure enforcement controls
August 24, 2026
|
5 min de lecture

Cadres de sécurité de l'IA en 2026 : Lesquels s'appliquent et où chacun s'arrête

Aucun article n'a été trouvé.
OpenRouter pricing compared with TrueFoundry AI Gateway governance
August 24, 2026
|
5 min de lecture

OpenRouter Pricing in 2026: Full Breakdown of Plans, Costs, and Hidden Fees

Aucun article n'a été trouvé.
Top 8 Amazon Bedrock Alternative and Competitors for 2026
August 24, 2026
|
5 min de lecture

Les 8 meilleures alternatives et concurrents d'Amazon Bedrock pour 2026 [Examen détaillé]

Aucun article n'a été trouvé.
Aucun article n'a été trouvé.

Blogs récents

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.
Faites un rapide tour d'horizon des produits
Commencer la visite guidée du produit
Visite guidée du produit