Qwen3.8-Max: $2/$6 API, Open Weights Next Week
🧮 News

Qwen3.8-Max: $2/$6 API, Open Weights Next Week

Alibaba priced Qwen3.8-Max at $2/$6 per million tokens on August 3 and dated the open-weight release to next week. What that sequence means for builders.

The AI Dude · August 3, 2026 · 5 min read

Alibaba's Qwen team put a token rate on Qwen3.8-Max on August 3, 2026: $2.0 per million input tokens, $6.0 per million output, $0.25 per million on implicit cache hits, with open weights promised for next week.

One item in that announcement did not make our pricing write-up earlier today. Alongside the rates, Qwen linked a public GitHub repository it describes as a "complete project trace" of the autonomous build behind its headline claim, the one about ten-plus days of development from an empty folder. That artifact exists now. The weights do not. The order those two things arrived in is what the rest of this post is about.

"Next week, the open weights"

The phrase is verbatim from the August 3 post on X. The sequence that produced it fits inside a single morning.

1:46 AM, August 3. The @Alibaba_Qwen account posts a teaser video carrying one line of copy: "Meet Qwen3.8-Max: A New Bar for Coding and Cowork." One line, and no numbers in it. The post stood at 4.2 million views as of the August 3 fetch.

2:15 AM, August 3. Twenty-nine minutes later, the specification post lands with everything the video withheld. Parameter count at 2.4 trillion. The three price points above. Four capability claims, headed by autonomous coding. And the sentence this section is named for: "Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!" Two models, one drop, dated to the week of August 10. At fetch time it had drawn 600.7K views against the video's 4.2 million, which is the usual split between a trailer and a spec sheet.

Later that morning. Qwen's launch blog goes up, and the model appears in Qwen Studio and on the qwencloud.com API listing. The endpoint is live and billable. Our own pricing coverage followed a few hours after that.

Everything in that sequence is vendor-published. The specification post carries no benchmark figures, and no independent evaluation of Qwen3.8-Max appears in the sources behind this piece.

"from empty folder to production without hand-holding"

Read the announcement as a sequencing decision and it points somewhere fairly specific. Qwen priced the endpoint and dated the weights in the same post, about seven days apart. For a frontier-scale release that ordering is unusual. The hosted API is normally the entire product, and the weights are the subject a lab declines to bring up.

The gap is the thing worth planning around. When the weights land a week behind the endpoint, the hosted API stops being a dependency and starts being an evaluation surface. You wire an agent to qwencloud.com at $2/$6, run it against your own workload for a fortnight, and if the vendor raises rates or retires a checkpoint you have a copy of the model to fall back to. A closed frontier endpoint offers nothing comparable, and for a production team that hedge is worth more than the headline price.

The price still does work, though, and it does it in a place the launch copy skips over. Implicit caching at $0.25 per million is one-eighth of the input rate. Long-horizon agents replay the same context on every turn, and Qwen's own pitch is built on turn counts: 500-plus turns of chip design optimization, 365 days of e-commerce strategy. A workload shaped like that spends most of its input budget re-reading a prefix it has already sent. The cached rate is where the bill for a genuinely autonomous agent gets decided, and it is the number to model first if you are costing one out this week.

Now the deflating part. Two-point-four trillion parameters at 4-bit is roughly 1.2 terabytes of weights, which we worked through when the pricing landed this morning. Almost nobody reading this is going to serve that, on a home rig or on a single rented box. So the Max weights function mainly as a credible threat to the API price, and Qwen3.8-27B is the model most teams will actually download and run. Shipping both in one drop looks deliberate. The 2.4T number buys the attention, and the 27B buys the installed base.

One caution on the capability claims. Ten-plus days of self-evolving development, production-quality deliverables across hundreds of professions, native multimodal planning loops, system-level autonomous planning with closed-loop adaptive learning. All of that is prose in a launch post, and prose does not survive contact with a procurement spreadsheet. The GitHub trace matters more than any of those sentences because a trace has timestamps in it.

"complete project trace in the GitHub"

Publishing a repository as evidence is an unusual move for a launch of this size. Qwen says the repo holds the full record of the build that ran ten-plus days from an empty folder to production, which makes the loudest claim in the announcement checkable now, without the weights and without waiting for a leaderboard to catch up.

Three things next week would break the reading above. If the drop covers Qwen3.8-27B and the Max weights slip, then August 3 was a price move with a promise attached, and the case for treating the endpoint as a hedged position falls apart. If the accompanying license restricts commercial deployment or derivative work, the fallback-copy argument collapses, because a model you may not deploy hedges nothing. And if the published trace turns out to describe a heavily supervised build, the autonomous-coding claim carrying the entire announcement loses its only piece of evidence.

All three get settled the week of August 10, on the same repository host where Qwen has already put its evidence. The useful part of the order it chose is that the trace is public today, ahead of a single weight file, and anyone weighing whether to point a production agent at $2/$6 this month can go read it.

Qwen 3.8 Maxopen weightsAPI pricingAlibabaautonomous coding agents
Share 𝕏 / Twitter Reddit LinkedIn

Keep reading