Key Terms
- Qwen3.8-Max - Alibaba's 2.4T-parameter MoE flagship (launched August 2, 2026, ~95B active), 1M-token context, native vision, described as capable of autonomous coding "for a dozen-plus days" to deliver complete projects. Pay-as-you-go: ¥12/¥36 per MTok (input/output, 0-1M tokens, untiered). Source: Aliyun – Newly Released Models , Aliyun – Billing For Model Studio
- qwen3.8-2.4t-a95b - the open-weight release of the 2.4T flagship (August 12), Apache-family open weights at the same ¥12/¥36 API price as the hosted max. Benchmarks: GPQA Diamond 92.6, PaperBench 93.0, OSWorld 86.1, CodeArena global #4. The largest open-weight release ever. Sources: Aliyun – Newly Released Models , Aliyun – Billing For Model Studio
- Prime mode (优速模式) - Alibaba's fast tier: qwen3.8-max-prime at ¥24/¥72 (2x standard), mirroring Anthropic's fast mode and OpenAI's Fast tier. Source: Aliyun – Billing For Model Studio
- qwen3.8-flash - multimodal 1M-context model (August 26) at ¥0.8/¥2.7 per MTok, explicitly OpenAI/Anthropic API-compatible and positioned for Claude Code and Codex integration. Source: Aliyun – Newly Released Models
- Token Plan - subscription replacing the Coding Plan: Personal (¥39/¥139/¥499 per month) and Team (¥150/¥550/¥1,398 per seat/month) editions with Credits and rolling-window limits. Source: Aliyun – Token Plan Overview
Latest Changes
Changes since the 2026-07 report.
July watch items verified:
- qwen3.8-max-preview open-weight release: confirmed. qwen3.8-2.4t-a95b, the open-weight version of the 2.4T flagship, was listed August 12. The preview's Token Plan exclusivity ended: qwen3.8-max is on pay-as-you-go at ¥12/¥36. Sources: Aliyun – Newly Released Models , Aliyun – Billing For Model Studio
- Coding Plan phase-out: confirmed (ongoing). Token Plan remains the recommended replacement; no reversal. Source: Aliyun – Token Plan Overview
- qwen3.7-max 50% pay-as-you-go promo: no longer listed. The billing page now publishes original prices only (¥12/¥36 for 3.7-max, same as 3.8-max) and directs to the console for active discounts; the 限时5折 marker is gone. Source: Aliyun – Billing For Model Studio
New August changes:
- New flagship (Aug 2): qwen3.8-max launched: 2.4T MoE, 1M context, native vision, "dozens of consecutive days" autonomous coding claim, ¥12/¥36 per MTok, 1M free tokens, batch at half price, context caching discounts. Sources: Aliyun – Newly Released Models , Aliyun – Billing For Model Studio
- [NEW CATEGORY] Largest open-weight release ever (Aug 12): qwen3.8-2.4t-a95b weights published (sparse MoE, 2.4T total/~95B active, hybrid attention, 1M context), served on Bailian at the same price as the hosted max. Source: Aliyun – Newly Released Models
- New models: qwen3.8-27b dense vision-language model (Aug 17) at ¥3/¥12 per MTok; qwen3.8-flash multimodal (Aug 26) at ¥0.8/¥2.7 with 1M context; qwen3.8-max-0902 snapshot (Sep 2) with further coding and agent gains at unchanged pricing; qwen3.7-text-rerank and text-embedding-flash (Sep 1). Source: Aliyun – Newly Released Models
- New speed tier: qwen3.8-max-prime at ¥24/¥72 (2x). Source: Aliyun – Billing For Model Studio
- Third-party catalog (Bailian): kimi-k3 re-listed (Aug 19), ZHIPU/GLM-5.3 (Aug 17) and GLM-5.3-Flash (Aug 31), MiniMax-H3 (Aug 17), deepseek-v4-pro-0813 (Aug 14) and v4-flash-0731 (Aug 1). Source: Aliyun – Newly Released Models
Plans
| Plan | Price | Usage | Notes |
|---|---|---|---|
| Pay-as-you-go | Per-token, per region | Undisclosed rate limits | Beijing cheapest; Singapore carries ~25% premium; US/Frankfurt/Tokyo global endpoints at Beijing price |
| Token Plan Personal Lite/Standard/Pro | ¥39/¥139/¥499 per month (promo) | 700/3,000/12,000 Credits per 5hr; 2,500/10,000/40,000 per 7 days; 1-2/3-4/6-8 concurrent agents | No training on data |
| Token Plan Team Standard/Pro/Max | ¥150/¥550/¥1,398 per seat/month (promo) | 25,000/100,000/250,000 Credits per seat/month | Team management |
| Coding Plan | Phasing out | - | No new Lite purchases; Pro limited stock |
Source: Aliyun – Token Plan Overview
API Pricing
| Model | Input ¥/MTok | Output ¥/MTok | Context | Notes |
|---|---|---|---|---|
| qwen3.8-max / max-0902 | ¥12 (~$1.67) | ¥36 (~$5.04) | 1M | Untiered; batch 50%; caching discounts; 1M free tokens |
| qwen3.8-max-prime | ¥24 | ¥72 | 1M | Fast tier |
| qwen3.8-2.4t-a95b (open weights) | ¥12 | ¥36 | 1M | Same price as hosted |
| qwen3.8-27b | ¥3 (~$0.42) | ¥12 (~$1.67) | 1M | Dense VLM |
| qwen3.8-flash | ¥0.8 (~$0.11) | ¥2.7 (~$0.38) | 1M | Cheapest 1M-context frontier-class model on Bailian |
| qwen3.7-max | ¥12 | ¥36 | 1M | Console promos may apply |
| qwen3.7-plus | ¥2/¥6 (tiered, 20% off listed) | ¥8/¥24 | 1M | 限时8折 marker persists |
| qwen3.7-flash | ¥0.2/¥0.6/¥1.2 | ¥0.8/¥2.4/¥4.8 | tiered 32K/256K/1M | Unchanged from July |
Singapore endpoints carry international pricing (e.g. qwen3.8-max ¥14.988/¥44.965). Source: Aliyun – Billing For Model Studio
Model Performance / Benchmarks
| Benchmark | qwen3.8-2.4t-a95b | Notes |
|---|---|---|
| GPQA Diamond | 92.6 | Source: Aliyun – Newly Released Models |
| PaperBench | 93.0 | Source: Aliyun – Newly Released Models |
| OSWorld | 86.1 | Source: Aliyun – Newly Released Models |
| BabyVision | 82.0 | Source: Aliyun – Newly Released Models |
| CodeArena | Global #4 | Source: Aliyun – Newly Released Models |
Third-party aggregator data: qwen3.8-27b posts Terminal-Bench 73.0 and DeepSWE 1.1 at 42.2, described as near-frontier agentic coding in a 27B body. Source: Local-Ai-Zone – Ai Updates August 2026
Latest News
The Qwen3.8 Wave (August 2026)
Alibaba shipped four new Qwen3.8 text models in one month: the 2.4T flagship max (Aug 2), its full open weights (Aug 12), a 27B dense VLM (Aug 17), and a ¥0.8-input flash with 1M context (Aug 26), plus a 0902 snapshot on Sep 2. The open-weight 2.4T release is the largest ever published, and qwen3.8-flash is explicitly positioned as a Claude Code/Codex-compatible drop-in via OpenAI/Anthropic protocol support. Sources: Aliyun – Newly Released Models , Aliyun – Billing For Model Studio
Bailian Becomes a Multi-Vendor Frontier Marketplace (August 2026)
Alibaba's Model Studio added Zhipu GLM-5.3/5.3-Flash, MiniMax H3, DeepSeek V4-Pro-0813/V4-Flash-0731, and Kimi K3 in August, making the platform a one-stop API for essentially every major Chinese frontier model. Source: Aliyun – Newly Released Models
Community Signals
Qwen's August releases were covered mainly through aggregator roundups rather than primary HN threads; the 2.4T open-weight release was called "the largest open-weight release ever" and the efficiency frontier (27B/flash class models running on consumer GPUs) drew the most commentary. Source: Local-Ai-Zone – Ai Updates August 2026 . The benchlm.ai September coding ranking places Qwen3.8-Max-tier models in the global top 10. Source: Benchlm – Coding
Enterprise Readiness
| Feature | Available? | Details |
|---|---|---|
| SSO (SAML) | Partial | Alibaba Cloud account-level IAM; no coding-agent-specific SSO documented. |
| SCIM | No | Not documented. |
| Audit logs | Partial | Alibaba Cloud ActionTrail platform-level; model-studio-specific audit not documented. |
| IP indemnity | No | Not documented. |
| Data residency | Yes | Beijing, Virginia, Singapore, Frankfurt, Tokyo regions with distinct pricing; Singapore "international" deployments and US/EU regional endpoints. Source: Aliyun – Billing For Model Studio |
| HIPAA | No | Not advertised. |
| Air-gapped / on-prem | Partial | Open weights (qwen3.8-2.4t-a95b, 27b) enable self-hosting; no enterprise on-prem product. |
| SLA | Yes | Alibaba Cloud platform SLAs apply. Source: Aliyun |
| Admin controls (RBAC) | Partial | RAM/IAM policies; Token Plan team management. Source: Aliyun – Token Plan Overview |
Transparency Gaps
| Gap | Details | Severity |
|---|---|---|
| Autonomous-coding claim | "Dozens of consecutive days" of autonomous coding has no published methodology, task set, or success criteria. | High |
| Promo pricing opacity | Active discounts moved to the console; the public billing page now shows only original prices, making effective rates unverifiable without an account. | Medium |
| Rate limits | Pay-as-you-go rate limits remain undisclosed. | Medium |
| 27B/flash agentic claims | Terminal-Bench/DeepSWE figures for qwen3.8-27b come from aggregator coverage, not an Alibaba-published benchmark page. | Medium |
| Open-weight license terms | The exact license (Apache 2.0 vs Qwen-style restrictions) for qwen3.8-2.4t-a95b is not stated on the release note. | Medium |