Key Terms
- Bailian (Model Studio) - Alibaba Cloud's managed LLM platform. Provides API access, token billing, and subscription plans for Qwen and third-party models. Source: Aliyun – Models
- Token Plan - a Credits-based subscription available in Personal edition (Lite/Standard/Pro, individual developer) and Team edition (Standard/Pro/Max seats). Supports Claude Code, Cursor, Qwen Code, Qoder, Qoder CN, and OpenClaw as compatible tools. Beijing-only region. Source: Aliyun – Token Plan Overview
- Credits - Token Plan's billing unit. A qwen3.6-plus request with ~8,349 input tokens, ~40,794 cached tokens, and ~573 output tokens consumes approximately 3.18 Credits. Source: Aliyun – Token Plan Overview
- Coding Plan - a flat-rate monthly subscription (¥200/month Pro tier) providing per-request access to coding models through CLI tools. Being phased out: Lite stopped new purchases March 20, 2026, Pro is limited-stock. Grants Alibaba a license to use inputs and outputs; prohibits automated API use. Source: Aliyun – Token Plan Overview
- Context caching - reusing previously cached input tokens at a discount. Explicit cache creation costs 125% of standard input price, cache hits cost 10%. Supported on qwen3.7-max, qwen3.7-plus, qwen3.7-flash, and other models. Source: Aliyun – Context Cache
- Batch API - asynchronous inference at 50% of real-time pricing. Supported on qwen3.7-max, qwen3.7-plus, qwen3.7-flash, and other models. Source: Aliyun – Batch Interfaces Compatible With Openai
- Tiered pricing - input token cost rises with larger single-request context. For example, qwen3.7-flash costs ¥0.2/MTok input for 0-32K tokens, rising to ¥0.6/MTok for 32K-256K and ¥1.2/MTok for 256K-1M. Source: Aliyun – Billing For Model Studio
- 1折 Credits promo (qwen3.8-max-preview) - during the preview period, Credits consumption for qwen3.8-max-preview is reduced to 10% of the normal rate, effectively giving 10x usage. Personal edition also gets a night discount (22:00-08:00) reducing Credits further to 2% of normal (0.2折), effectively 50x usage. Source: Aliyun – Token Plan Overview
Latest Changes
Changes since the 2026-06 report.
- Verified (June watch-item): qwen3.7-max half-Credits Token Plan promo status. The Token Plan page no longer highlights the qwen3.7-max half-Credits promotional period that was extended to July 22. The Token Plan page now exclusively highlights qwen3.8-max-preview's 1折 (10%) Credits promo instead. However, the pay-as-you-go "限时5折" (50% off) discount for qwen3.7-max is still active on the billing page. Classification: partial (Token Plan Credits promo appears expired, pay-as-you-go promo continues). Source: Aliyun – Token Plan Overview , Aliyun – Billing For Model Studio
- Confirmed (June watch-item): New model release. qwen3.8-max-preview launched on 2026-07-19, announced via Alibaba_Qwen on X. Described as 2.4T parameters and "going open-weight soon." It is Token Plan exclusive (no pay-as-you-go access). This is Alibaba's answer to the frontier model competition. Source: X – Status , Aliyun – Token Plan Overview
- Confirmed (June watch-item): WAIC 2026 Agent Native Cloud announcements. At WAIC 2026 (mid-July), Alibaba showcased Qwen Office integrating agent products QoderWork, Wukong, MuleRun, plus Agent Native Cloud with multi-agent orchestration and DAMO Lingshu research agents. No direct pricing impact, but signals Alibaba's enterprise agent strategy beyond raw API access.
- New model: qwen3.7-flash launched 2026-07-21 (snapshot qwen3.7-flash-2026-07-15). A native vision-language Flash model that upgrades multimodal understanding, Agent execution (Search Agent, CI Agent), and vibe coding experience over the 3.6-Flash predecessor. Three-tier pricing: ¥0.2/¥0.8 (0-32K), ¥0.6/¥2.4 (32K-256K), ¥1.2/¥4.8 (256K-1M) per MTok. This is 83-89% cheaper than qwen3.6-flash at the lowest tier. Source: Aliyun – Newly Released Models , Aliyun – Billing For Model Studio
- Price change: Token Plan Team edition prices cut. Standard seat dropped from ¥198 to ¥150/month (24% cut), Pro (formerly Advanced) seat dropped from ¥698 to ¥550/month (21% cut). Max (formerly Premium) unchanged at ¥1,398/month. Credits allocations unchanged. Source: Aliyun – Token Plan Overview
- New product: Token Plan Personal edition. Three tiers: Lite (¥39/month, 700 Credits/5hr, 2,500/7days, 1-2 concurrent agents), Standard (¥139/month, 3,000 Credits/5hr, 10,000/7days, 3-4 agents), Pro (¥499/month, 12,000 Credits/5hr, 40,000/7days, 6-8 agents). Uses 5-hour and 7-day rolling window limits. All at limited-time promotional prices. Source: Aliyun – Token Plan Overview
- Plan change: Coding Plan being phased out. Coding Plan Lite stopped new purchases on 2026-03-20 and stopped renewals/upgrades on 2026-04-13. Coding Plan Pro is "限量抢购" (limited-stock flash sale) with no restock after inventory depletes. Alibaba explicitly recommends Token Plan as the replacement. Source: Aliyun – Token Plan Overview
- Price change: qwen3.7-plus pay-as-you-go now at 80% (限时8折). The billing page now shows a 20% discount on qwen3.7-plus input and output across all tiers, lowering the effective ≤256K rate to ¥1.6/¥6.4 per MTok. The original price of ¥2/¥8 remains listed as the base rate. Source: Aliyun – Billing For Model Studio
- Price change: qwen3.7-max pay-as-you-go 50% off still active (限时5折). The half-price promo on qwen3.7-max that was extended in June remains visible on the billing page, lowering effective pay-as-you-go rates to ¥6/¥18 per MTok. No explicit end date shown. Source: Aliyun – Billing For Model Studio
- New third-party model: kimi/kimi-k3 (added 2026-07-17). Kimi's most capable flagship model, 2.8T parameters, KDA hybrid linear attention, 1M context, vision support. The first open-source 3T-class model. Source: Aliyun – Newly Released Models
- New third-party model: glm-5.2-fast-preview (added 2026-07-09). High-speed variant of GLM-5.2 with 1.5-2x output TPS, 1M context. Source: Aliyun – Newly Released Models
- Catalog reshuffle: headline Qwen trio updated. The model catalog now features qwen3.8-max-preview (top, Token Plan only), qwen3.7-max, qwen3.7-plus, and qwen3.7-flash as the headline models. qwen3.6-flash is demoted to "more models." Source: Aliyun – Models
- HLD update needed: HLD.md lists the Latest LLM as "qwen3.7-max, qwen3.7-plus, qwen3.6-flash" but the current featured models are qwen3.8-max-preview, qwen3.7-max, qwen3.7-plus, and qwen3.7-flash. The Latest LLM column should be updated accordingly.
Plans
| Plan | Price | Billing | Usage Limits | Data Used for Training? | Key Models |
|---|---|---|---|---|---|
| Pay-as-you-go | Per-token | CNY per MTok | Undisclosed rate limits | No | All Qwen + third-party models |
| Coding Plan Lite | Stopped (new purchases 2026-03-20, renewals 2026-04-13) | Per-request | N/A | Yes (license to use inputs/outputs) | N/A (phased out) |
| Coding Plan Pro | ¥200/month (limited-stock, no restock) | Per-request | 6,000 req/5hr, 45,000/week, 90,000/month (per May report) | Yes (license to use inputs/outputs) | qwen3.6-plus, qwen3.5-plus, qwen3-coder-next, qwen3-coder-plus, and others |
| Token Plan Personal Lite | ¥39/month (was ¥60, limited-time) | Credits (rolling window) | 700 Credits/5hr, 2,500/7days, 1-2 concurrent agents | No (explicitly promised) | All models including qwen3.8-max-preview, qwen3.7-max, qwen3.7-plus, qwen3.7-flash, deepseek-v4-pro/flash, kimi-k3, glm-5.2 |
| Token Plan Personal Standard | ¥139/month (was ¥180, limited-time) | Credits (rolling window) | 3,000 Credits/5hr, 10,000/7days, 3-4 concurrent agents | No | Same as Personal Lite |
| Token Plan Personal Pro | ¥499/month (was ¥600, limited-time) | Credits (rolling window) | 12,000 Credits/5hr, 40,000/7days, 6-8 concurrent agents | No | Same as Personal Lite |
| Token Plan Team Standard | ¥150/seat/month (was ¥198, limited-time) | Credits (25,000/seat/month) | No 5hr/7day window, monthly total only | No | Same as Personal + team management |
| Token Plan Team Pro | ¥550/seat/month (was ¥698, limited-time) | Credits (100,000/seat/month) | No 5hr/7day window | No | Same as Team Standard |
| Token Plan Team Max | ¥1,398/seat/month (unchanged) | Credits (250,000/seat/month) | No 5hr/7day window | No | Same as Team Standard |
| Token Plan Shared Pack | ¥5,000/pack | Credits (625,000/pack) | 1-month expiry | No | Same as Team Standard |
| Free tier | Free | 1M input + 1M output tokens | One-time, 90-day expiry | No | Most Qwen models (qwen3.7-max, qwen3.7-plus, qwen3.7-flash, etc.) |
Terms explained:
- Credits - Token Plan's billing unit. A single qwen3.6-plus request with ~8K input tokens consumes roughly 3.18 Credits. Actual consumption varies by model, thinking mode, and tool usage. During the qwen3.8-max-preview period, Credits consumption is reduced to 10% of normal (Personal edition: further reduced to 2% during 22:00-08:00 night hours). Source: Aliyun – Token Plan Overview
- Rolling window limits (Personal edition) - Personal edition uses two fixed windows: a 5-hour window and a 7-day window. When either limit is hit, service pauses until that window expires. Unused Credits do not roll over. Team edition uses monthly totals with no 5hr/7day window. Source: Aliyun – Token Plan Overview
- Data used for training - Token Plan explicitly states it does not use conversation data to train models. The Coding Plan explicitly grants Alibaba a license to use inputs and outputs during the subscription. This distinction is critical for enterprise buyers. Source: Aliyun – Token Plan Overview
- Compatible tools - Token Plan supports Claude Code, Cursor, Qwen Code, Qoder, Qoder CN, and OpenClaw as compatible AI programming and agent tools, plus web search and code interpreter Harness tools. Source: Aliyun – Token Plan Overview
Source: Aliyun – Token Plan Overview
API Pricing
All prices in CNY per million tokens unless noted. China mainland region (Beijing) is the headline rate; international regions differ (noted below). Approximate USD at ~$1 = ¥7.2.
Qwen3.8 Max Preview (new, Token Plan exclusive)
No published per-token pay-as-you-go price. Available exclusively through Token Plan subscription with Credits billing. During the preview period, Credits consumption is at 1折 (10% of normal rate, effectively 10x usage). Personal edition also gets a night discount (22:00-08:00) reducing Credits further to 0.2折 (2% of normal). Alibaba announced weights will go open-source soon. Source: Aliyun – Token Plan Overview , X – Status
Qwen3.7 Max (flagship, 50% off promo still active)
| Model | Mode | Input Tier | Input (¥/MTok) | Output (¥/MTok) |
|---|---|---|---|---|
| qwen3.7-max (= 2026-05-20) | Non-thinking + Thinking | 0-1M | 12 (限时5折: effective 6) | 36 (限时5折: effective 18) |
| qwen3.7-max-2026-06-08 (vision) | Non-thinking + Thinking | 0-1M | 12 | 36 |
| qwen3.7-max-2026-05-20 | Non-thinking + Thinking | 0-1M | 12 | 36 |
| qwen3.7-max-preview (= 2026-05-17) | Thinking only | 0-1M | 12 | 36 |
Note: The "限时5折" (50% off) applies to the bare qwen3.7-max alias only. Dated snapshots and vision variants do not show the discount. Features: Batch API (50% discount), context caching. Free tier: 1M tokens input + 1M output, 90-day expiry. 1M context window, 64K max output, 256K thinking budget. Source: Aliyun – Billing For Model Studio
Qwen3.7 Plus (multimodal, 20% off promo active)
| Model | Input Tier | Input (¥/MTok) | Output: Non-thinking (¥/MTok) | Output: Thinking (¥/MTok) |
|---|---|---|---|---|
| qwen3.7-plus (= 2026-05-26) | 0-256K | 2 (限时8折: effective 1.6) | 8 (限时8折: effective 6.4) | 8 (限时8折: effective 6.4) |
| qwen3.7-plus (= 2026-05-26) | 256K-1M | 6 (限时8折: effective 4.8) | 24 (限时8折: effective 19.2) | 24 (限时8折: effective 19.2) |
| qwen3.7-plus-2026-05-26 (explicit ID) | 0-256K | 2 | 8 | 8 |
| qwen3.7-plus-2026-05-26 (explicit ID) | 256K-1M | 6 | 24 | 24 |
Note: The "限时8折" (20% off) applies to the bare qwen3.7-plus alias only. The explicit dated snapshot does not show the discount. Features: Batch API (50%), context caching. Free tier: 1M tokens each input/output, 90-day expiry. 1M context window. Multimodal: text, vision, GUI operation, screen reading. Source: Aliyun – Billing For Model Studio
Qwen3.7 Flash (new, launched 2026-07-21)
| Model | Mode | Input Tier | Input (¥/MTok) | Output (¥/MTok) |
|---|---|---|---|---|
| qwen3.7-flash (= 2026-07-15) | Non-thinking + Thinking | 0-32K | 0.2 | 0.8 |
| qwen3.7-flash (= 2026-07-15) | Non-thinking + Thinking | 32K-256K | 0.6 | 2.4 |
| qwen3.7-flash (= 2026-07-15) | Non-thinking + Thinking | 256K-1M | 1.2 | 4.8 |
Features: Batch API (50% discount), context caching. Free tier: 1M tokens input + 1M output, 90-day expiry. 1M context window. Multimodal: native vision-language model with multimodal coding, Search Agent, and CI Agent capabilities. This is the cheapest Qwen text model ever offered. Source: Aliyun – Billing For Model Studio , Aliyun – Newly Released Models
Qwen3.6 Flash (predecessor, still available)
| Model | Mode | Input Tier | Input (¥/MTok) | Output (¥/MTok) |
|---|---|---|---|---|
| qwen3.6-flash (= 2026-04-16) | Non-thinking + Thinking | 0-256K | 1.2 | 7.2 |
| qwen3.6-flash (= 2026-04-16) | Non-thinking + Thinking | 256K-1M | 4.8 | 28.8 |
Features: Batch API (50%), context caching. 1M context window. Source: Aliyun – Billing For Model Studio
Qwen3.6 Plus (predecessor, still available)
| Model | Input Tier | Input (¥/MTok) | Output: Non-thinking (¥/MTok) | Output: Thinking (¥/MTok) |
|---|---|---|---|---|
| qwen3.6-plus (= 2026-04-02) | 0-256K | 2 | 12 | 12 |
| qwen3.6-plus (= 2026-04-02) | 256K-1M | 8 | 48 | 48 |
Features: Batch API (50%). 1M context window. Source: Aliyun – Billing For Model Studio
Regional price differences (qwen3.7-max)
| Region | Input (¥/MTok) | Output (¥/MTok) | Note |
|---|---|---|---|
| Beijing (China mainland) | 12 (限时5折: 6) | 36 (限时5折: 18) | Headline rate with active promo |
| US (Virginia) | 12 (限时5折: 6) | 36 (限时5折: 18) | Global pricing, same as Beijing |
| Frankfurt (EU) | 12 (限时5折: 6) | 36 (限时5折: 18) | Global pricing, same as Beijing |
| Tokyo (Japan) | 12 (限时5折: 6) | 36 (限时5折: 18) | Global pricing, same as Beijing |
| Singapore | 18.736 (限时5折: 9.368) | 56.207 (限时5折: 28.104) | International rate, notably higher |
Source: Aliyun – Billing For Model Studio
Approximate USD Equivalent (Beijing rate, at $1 = ¥7.2)
| Model | Input ($/MTok) | Output ($/MTok) |
|---|---|---|
| qwen3.8-max-preview | undisclosed (Token Plan only) | undisclosed |
| qwen3.7-max (with 50% promo) | ~$0.83 | ~$2.50 |
| qwen3.7-max (base) | ~$1.67 | ~$5.00 |
| qwen3.7-plus (≤256K, with 20% promo) | ~$0.22 | ~$0.89 |
| qwen3.7-plus (≤256K, base) | ~$0.28 | ~$1.11 |
| qwen3.7-flash (0-32K) | ~$0.028 | ~$0.11 |
| qwen3.7-flash (32K-256K) | ~$0.083 | ~$0.33 |
| qwen3.6-flash (≤256K) | ~$0.17 | ~$1.00 |
Model Performance / Benchmarks
Alibaba has not published coding-specific benchmark scores (SWE-Bench Verified, TerminalBench, LiveCodeBench, etc.) for any closed-source model: qwen3.8-max-preview, qwen3.7-max, qwen3.7-plus, qwen3.7-flash, or qwen3.6-flash. The release notes describe capabilities qualitatively ("excels in programming, office productivity, and long-horizon autonomous execution") without numbers. Source: Aliyun – Newly Released Models
Community testing of qwen3.8-max-preview produced the following independent data points:
| Source | Finding |
|---|---|
| Fireworks.ai blog | Tested K3, Qwen 3.8 Max, Fable, and Sol; concluded Kimi K3 is competitive with Fable and both are SoTA. Qwen 3.8 Max not separately scored. Fireworks – Kimik3 Fable |
| senko (HN developer) | Vibecode-bench comparison: Qwen 3.8 roughly equal to previous-gen Opus 4.8 and GPT 5.5. For a web app task: Qwen 3.8 Max consumed 18M input tokens (17.8M cached), 114K output, cost $6.3 via API vs Fable $30 via Claude Code. News – Item |
| senko (HN, token cost) | Same task: Kimi K3 cost $5.5 (9.5M input), Qwen 3.8 cost $6.3 (18M input), Fable cost $30 (14M input). Chinese models "slow and token-inefficient" but "close to SOTA." News – Item |
The only verifiable Qwen benchmark scores for open-weight models remain from the Artificial Analysis intelligence index, via an independent blog post covering Qwen3.6 models:
| Model (open-weight) | Artificial Analysis index | Approx. frontier equivalence |
|---|---|---|
| Qwen3.6-27B (dense) | 37 | ~mid-2025 (GPT-5 / Claude Sonnet 4.5 level) |
| Qwen3.6-35B-A3B (MoE) | 32 | ~early-2025 (o3 / Claude 4 Sonnet level) |
Source: Quesma – Qwen 36 Is Awesome
Latest News
qwen3.8-max-preview Launched as Token Plan Exclusive (2026-07-19)
Alibaba announced qwen3.8-max-preview on July 19 via the official Alibaba_Qwen X/Twitter account, stating it is "going open-weight soon." The model is described as having 2.4T parameters. It is exclusively available through Token Plan subscription with promotional Credits pricing at 1折 (10% of normal rate), plus a night discount in the Personal edition. No pay-as-you-go pricing is published, and no coding benchmarks have been released. The HackerNews thread reached 961 upvotes and 731 comments, making it one of the most discussed AI model launches of the month. Source: X – Status , Hn – Items
qwen3.7-flash Launched as Budget Multimodal Model (2026-07-21)
qwen3.7-flash (snapshot 2026-07-15) was released on July 21 as the Flash tier of the Qwen3.7 native vision-language series. It improves multimodal understanding, Agent execution (Search Agent, CI Agent), and vibe coding over the 3.6-Flash predecessor. Three-tier pricing starting at ¥0.2/¥0.8 per MTok (0-32K) makes it the cheapest Qwen text model ever. Source: Aliyun – Newly Released Models
Token Plan Personal Edition Launched, Team Prices Cut (ongoing)
Alibaba introduced a Token Plan Personal edition with three tiers (Lite ¥39, Standard ¥139, Pro ¥499/month) using 5-hour and 7-day rolling window Credits limits. Simultaneously, Team edition prices were cut 21-24% (Standard ¥150, Pro ¥550 per seat/month). The Coding Plan is explicitly being phased out in favor of Token Plan. Source: Aliyun – Token Plan Overview
Kimi K3 Added to Bailian (2026-07-17)
Moonshot's kimi/kimi-k3 was added to the Bailian model catalog. Kimi K3 is described as 2.8T parameters with KDA hybrid linear attention and 1M context, claimed as the world's first open-source 3T-class model. It supports long-horizon programming, knowledge work, and reasoning. Source: Aliyun – Newly Released Models
"Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling" Analysis Published (2026-07-19)
An Emerging Trajectories analysis argued that Kimi K3 and Qwen 3.8 represent a strategic challenge to top-tier model developers because they prove the SOTA frontier is attainable with open models. The article claims Qwen 3.8 and Kimi K3 are "allegedly close to Anthropic's Fable 5 in performance" and that weights will be released publicly. It frames Alibaba's data center ownership as a structural margin advantage over model-only providers. The HackerNews thread reached 371 upvotes and 336 comments. Source: Emergingtrajectories – Frontier Lab Economics
WAIC 2026: Qwen Office and Agent Native Cloud (mid-July)
At the World Artificial Intelligence Conference 2026 (WAIC), Alibaba showcased Qwen Office integrating agent products QoderWork, Wukong, MuleRun, plus Agent Native Cloud with multi-agent orchestration and DAMO Lingshu research agents. This signals Alibaba's strategy to build full-stack agent products on top of its model infrastructure, not just sell API tokens.
Qwen Blog Still Redirects to qwen.ai
The old blog at qwenlm.github.io/blog/ continues to redirect to qwen.ai/research. The last substantive post on the old blog remains "Qwen3Guard" (September 23, 2025). The new qwen.ai location is heavily JavaScript-rendered and no July 2026 posts were extractable via standard fetch. Model-release detail is sourced from the Aliyun help docs rather than the Qwen blog. Source: Alibaba
Community Signals
"Qwen 3.8" dominated HackerNews with 961 upvotes and 731 comments (July 19)
The qwen3.8-max-preview announcement generated the largest Qwen-related HackerNews discussion to date, with 961 upvotes and 731 comments. The thread linked to both the Alibaba_Qwen X post and the qwencloud.com Token Plan pricing page. Discussion covered model quality, geopolitical implications, and pricing comparisons. Source: News – Item
Developer who cancelled Anthropic subscription for Qwen 3.8
nerdalytics reported that qwen3.8-max-preview was good enough to cancel their Anthropic subscription: "Qwen3.8-Max-Preview is already good enough that I cancelled my Anthropic subscription. There are other factors too, but Qwen is on a great path. As a European, this is choosing between the plague and cholera. Trusting US companies is as troubling as trusting Chinese companies. But I get more bang for the buck from the qwencloud.com Token Plan than from any US AI lab." News – Item
Developer comparison: Qwen 3.8 vs Fable vs Kimi K3 on coding tasks
senko tested all three models on a zero-shot web app from a detailed spec: "Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, I partially agree. Don't trust the benchmarks, and the Chinese models really are slow and token-inefficient. However they do seem very close to SOTA: I'd say roughly equal to previous gen (Opus 4.8, GPT 5.5)." Token cost for the same task: Kimi K3 at $5.5 (9.5M input), Qwen 3.8 Max at $6.3 (18M input, 17.8M cached), Fable at $30 (14M input). News – Item , News – Item
Negative coding experience with Qwen 3.8 in agentic tasks
isqueiros reported severe issues with qwen3.8-max-preview in agentic coding: "In my limited anecdotal experience, Kimi K3 is a bit better than Opus 4.8 and Qwen3.8 Max is disastrously bad. It can reason fine, but the moment it tries to do something it gets stuck into long second-guessing loops with no progress. It sometimes refuses to try and debug a problem even after multiple suggestions. I'm sticking to K3." News – Item
Verbose output and benchmark skepticism
Alifatisk raised concerns about Qwen model behavior and trust in benchmarks: "I remember when they released Qwen 3.7 Plus and Max. These models behaved way different from all prior models, it became too verbose. It wrote multiple paragraphs just to answer my prompt instead of the usual concise and direct way responding to me... Another thing I experience with the Qwen models is that I do not trust their benchmark scores at all." News – Item
Chinese models "slow and token-inefficient" but near-SOTA
paulddraper summarized the community sentiment: "Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, the Chinese models really are slow and token-inefficient. Those are all frontier-competitive models." This highlights the cost-per-task gap between Qwen/Kimi and Anthropic/OpenAI despite near-comparable capability. News – Item
Qwen 3.8 tested alongside GLM 5.2 and Mimo V2.5 Pro, found lacking
CapyToolkit tested multiple models in Claude Code's VS Code terminal extension: "I tested Mimo V2.5 Pro, Qwen 3.8 Max Preview, GLM 5.2, Macaron V1 Venti and many other, smaller model whenever hype on X appears. All behave dumb, don't follow instructions properly even when they write skills for themselves. They are literally on the level of Sonnet 3.5 which was released in 2024." News – Item
Frontier lab economics debate: Alibaba's structural advantage
The Emerging Trajectories analysis sparked debate about Alibaba's infrastructure advantage over model-only providers like Anthropic: the article argues Alibaba owns its data centers (the "Meta and Alibaba approach"), making inference costs fixed rather than variable, while model-only companies (Anthropic, DeepSeek, Moonshot, Knowledge Atlas/Zhipu) face a "constant race to the bottom on inference costs." Source: Emergingtrajectories – Frontier Lab Economics
Enterprise Readiness
| Feature | Available? | Details |
|---|---|---|
| SSO (SAML/OIDC) | Partial | Alibaba Cloud RAM (Resource Access Management) supports SAML SSO for console access. No native OIDC integration documented for API access. |
| SCIM | No | No SCIM-based user provisioning documented. Teams must manually add members via the Token Plan management console. |
| Audit logs | Yes | Billing and usage logs via Alibaba Cloud Billing Center. Token Plan provides per-member usage analytics in the team management console. Source: Aliyun – Token Plan Overview |
| IP indemnity | No | No IP indemnity commitment documented for Qwen models or the Bailian platform. |
| Data residency | Yes (pay-as-you-go) / Limited (plans) | Pay-as-you-go available in Beijing, Singapore, Tokyo, Frankfurt, and Virginia. Token Plan and Coding Plan are limited to Beijing (China mainland) only. Source: Aliyun – Models |
| HIPAA | No | No HIPAA compliance documented. |
| Air-gapped / on-prem | Partial | Model Studio supports importing and deploying custom models on dedicated instances, but no fully air-gapped deployment for hosted Qwen models is documented. Open-weight Qwen models can be self-hosted. |
| SLA | Undisclosed | No specific SLA for model availability documented in the public pricing or service pages. |
| Admin controls (RBAC) | Yes | Alibaba Cloud RAM supports role-based access control. Token Plan provides workspace-level permission management with admin and member roles, plus per-seat allocation and recall. Source: Aliyun – Token Plan Overview |
Terms explained:
- Data residency - pay-as-you-go is multi-region, but the subscription plans (Coding Plan, Token Plan) are Beijing-only. A team requiring data processing outside China cannot use the seat-based plans. Source: Aliyun – Token Plan Overview
- IP indemnity - no intellectual property indemnification is offered for Qwen outputs, unlike Microsoft's Copilot Copyright Commitment or comparable programs.
Transparency Gaps
| Gap | Details | Severity |
|---|---|---|
| qwen3.8-max-preview pay-as-you-go price undisclosed | The newest flagship model has no published per-token price. It is only accessible via Token Plan with Credits billing. Without a per-token rate, teams cannot calculate cost-per-task outside the promotional Credits framework. | High |
| Closed-model benchmarks missing | Alibaba publishes no SWE-Bench, TerminalBench, LiveCodeBench, or comparable coding scores for qwen3.8-max-preview, qwen3.7-max, qwen3.7-plus, qwen3.7-flash, or qwen3.6-flash. Capability is described only qualitatively. Community testing yields conflicting assessments (from "disastrously bad" to "cancelled my Anthropic subscription"). | High |
| Rate limits undisclosed | Pay-as-you-go API rate limits (RPM/TPM) are not published on the pricing page. Without published limits, teams cannot plan capacity. | High |
| Token Plan Credits formula not disclosed | Only one example is given (qwen3.6-plus at ~3.18 Credits for a specific request shape). Per-model Credit rates and the formula for thinking mode, tool calls, and cached tokens are not published. Users rely on the billing dashboard. | High |
| qwen3.8-max-preview promo end date unspecified | The 1折 (10%) Credits promo and night discount for qwen3.8-max-preview have no stated end date. Alibaba reserves the right to change or end the promo "based on operational conditions." | High |
| Coding Plan phase-out timeline unclear for Pro | Pro is "limited-stock" but Alibaba does not publish remaining inventory or a hard end date. Teams on existing Coding Plan subscriptions cannot plan migration timing. | Medium |
| Promotional pricing end dates unspecified | The "限时5折" (50% off) for qwen3.7-max and "限时8折" (20% off) for qwen3.7-plus on pay-as-you-go have no published end dates. Token Plan Personal/Team limited-time prices also lack end dates. | Medium |
| Coding Plan data-use scope | The Coding Plan grants Alibaba a license to use inputs and outputs for training, but the scope (which models it trains, retention period, whether it extends beyond Qwen) is not specified beyond a reference to the Bailian Service Agreement. | Medium |
| SLA not documented | No service-level agreement for model availability, latency, or error rates appears in the public documentation. | Medium |
| qwen.ai blog inaccessible | The new Qwen blog at qwen.ai is heavily JavaScript-rendered, so posts cannot be extracted via standard fetch. This reduces transparency for non-Chinese-speaking audiences tracking Qwen developments. | Medium |
| Singapore pricing premium unexplained | qwen3.7-max costs ¥18.736/¥56.207 per MTok in Singapore versus ¥12/¥36 in Beijing, US, Frankfurt, and Tokyo, a roughly 56% premium with no documented rationale. | Low |
| qwen3.8-max-preview open-weight timeline vague | Alibaba says "going open-weight soon" but provides no date or parameter count details beyond the 2.4T figure from X posts. | Low |