Alibaba

Executive Summary

What it is: Alibaba's Qwen family and Qwen Code run on Model Studio (Bailian) pay-as-you-go plus Token Plan subscriptions (Personal ¥39-¥499/mo, Team ¥150-¥1,398/seat/mo), and August delivered the Qwen3.8 wave: the 2.4T max flagship (¥12/¥36 per MTok), its open weights at the same price, a 27B dense model (¥3/¥12), and a ¥0.8/¥2.7 multimodal flash, all with 1M context. Source: https://help.aliyun.com/zh/model-studio/billing-for-model-studio

What to watch out for: The "dozens of consecutive days" autonomous-coding claim has no published methodology, active discounts moved to the console (public pages now show only original prices), and the qwen3.8-2.4t-a95b open-weight license terms are unstated on the release note. Source: https://help.aliyun.com/zh/model-studio/newly-released-models

Bottom line: Alibaba had the single biggest month of any supplier: the largest open-weight release ever, a full flagship-to-flash refresh in 25 days, and a multi-vendor marketplace (GLM-5.3, Kimi K3, DeepSeek V4, MiniMax H3) that makes Bailian the one-stop API for Chinese frontier models at a fraction of Western prices. Source: https://help.aliyun.com/zh/model-studio/newly-released-models

Key Terms

  • Qwen3.8-Max - Alibaba's 2.4T-parameter MoE flagship (launched August 2, 2026, ~95B active), 1M-token context, native vision, described as capable of autonomous coding "for a dozen-plus days" to deliver complete projects. Pay-as-you-go: ¥12/¥36 per MTok (input/output, 0-1M tokens, untiered). Source: Aliyun – Newly Released Models , Aliyun – Billing For Model Studio
  • qwen3.8-2.4t-a95b - the open-weight release of the 2.4T flagship (August 12), Apache-family open weights at the same ¥12/¥36 API price as the hosted max. Benchmarks: GPQA Diamond 92.6, PaperBench 93.0, OSWorld 86.1, CodeArena global #4. The largest open-weight release ever. Sources: Aliyun – Newly Released Models , Aliyun – Billing For Model Studio
  • Prime mode (优速模式) - Alibaba's fast tier: qwen3.8-max-prime at ¥24/¥72 (2x standard), mirroring Anthropic's fast mode and OpenAI's Fast tier. Source: Aliyun – Billing For Model Studio
  • qwen3.8-flash - multimodal 1M-context model (August 26) at ¥0.8/¥2.7 per MTok, explicitly OpenAI/Anthropic API-compatible and positioned for Claude Code and Codex integration. Source: Aliyun – Newly Released Models
  • Token Plan - subscription replacing the Coding Plan: Personal (¥39/¥139/¥499 per month) and Team (¥150/¥550/¥1,398 per seat/month) editions with Credits and rolling-window limits. Source: Aliyun – Token Plan Overview

Latest Changes

Changes since the 2026-07 report.

July watch items verified:

  • qwen3.8-max-preview open-weight release: confirmed. qwen3.8-2.4t-a95b, the open-weight version of the 2.4T flagship, was listed August 12. The preview's Token Plan exclusivity ended: qwen3.8-max is on pay-as-you-go at ¥12/¥36. Sources: Aliyun – Newly Released Models , Aliyun – Billing For Model Studio
  • Coding Plan phase-out: confirmed (ongoing). Token Plan remains the recommended replacement; no reversal. Source: Aliyun – Token Plan Overview
  • qwen3.7-max 50% pay-as-you-go promo: no longer listed. The billing page now publishes original prices only (¥12/¥36 for 3.7-max, same as 3.8-max) and directs to the console for active discounts; the 限时5折 marker is gone. Source: Aliyun – Billing For Model Studio

New August changes:

  • New flagship (Aug 2): qwen3.8-max launched: 2.4T MoE, 1M context, native vision, "dozens of consecutive days" autonomous coding claim, ¥12/¥36 per MTok, 1M free tokens, batch at half price, context caching discounts. Sources: Aliyun – Newly Released Models , Aliyun – Billing For Model Studio
  • [NEW CATEGORY] Largest open-weight release ever (Aug 12): qwen3.8-2.4t-a95b weights published (sparse MoE, 2.4T total/~95B active, hybrid attention, 1M context), served on Bailian at the same price as the hosted max. Source: Aliyun – Newly Released Models
  • New models: qwen3.8-27b dense vision-language model (Aug 17) at ¥3/¥12 per MTok; qwen3.8-flash multimodal (Aug 26) at ¥0.8/¥2.7 with 1M context; qwen3.8-max-0902 snapshot (Sep 2) with further coding and agent gains at unchanged pricing; qwen3.7-text-rerank and text-embedding-flash (Sep 1). Source: Aliyun – Newly Released Models
  • New speed tier: qwen3.8-max-prime at ¥24/¥72 (2x). Source: Aliyun – Billing For Model Studio
  • Third-party catalog (Bailian): kimi-k3 re-listed (Aug 19), ZHIPU/GLM-5.3 (Aug 17) and GLM-5.3-Flash (Aug 31), MiniMax-H3 (Aug 17), deepseek-v4-pro-0813 (Aug 14) and v4-flash-0731 (Aug 1). Source: Aliyun – Newly Released Models

Plans

Plan Price Usage Notes
Pay-as-you-go Per-token, per region Undisclosed rate limits Beijing cheapest; Singapore carries ~25% premium; US/Frankfurt/Tokyo global endpoints at Beijing price
Token Plan Personal Lite/Standard/Pro ¥39/¥139/¥499 per month (promo) 700/3,000/12,000 Credits per 5hr; 2,500/10,000/40,000 per 7 days; 1-2/3-4/6-8 concurrent agents No training on data
Token Plan Team Standard/Pro/Max ¥150/¥550/¥1,398 per seat/month (promo) 25,000/100,000/250,000 Credits per seat/month Team management
Coding Plan Phasing out - No new Lite purchases; Pro limited stock

Source: Aliyun – Token Plan Overview

API Pricing

Model Input ¥/MTok Output ¥/MTok Context Notes
qwen3.8-max / max-0902 ¥12 (~$1.67) ¥36 (~$5.04) 1M Untiered; batch 50%; caching discounts; 1M free tokens
qwen3.8-max-prime ¥24 ¥72 1M Fast tier
qwen3.8-2.4t-a95b (open weights) ¥12 ¥36 1M Same price as hosted
qwen3.8-27b ¥3 (~$0.42) ¥12 (~$1.67) 1M Dense VLM
qwen3.8-flash ¥0.8 (~$0.11) ¥2.7 (~$0.38) 1M Cheapest 1M-context frontier-class model on Bailian
qwen3.7-max ¥12 ¥36 1M Console promos may apply
qwen3.7-plus ¥2/¥6 (tiered, 20% off listed) ¥8/¥24 1M 限时8折 marker persists
qwen3.7-flash ¥0.2/¥0.6/¥1.2 ¥0.8/¥2.4/¥4.8 tiered 32K/256K/1M Unchanged from July

Singapore endpoints carry international pricing (e.g. qwen3.8-max ¥14.988/¥44.965). Source: Aliyun – Billing For Model Studio

Model Performance / Benchmarks

Benchmark qwen3.8-2.4t-a95b Notes
GPQA Diamond 92.6 Source: Aliyun – Newly Released Models
PaperBench 93.0 Source: Aliyun – Newly Released Models
OSWorld 86.1 Source: Aliyun – Newly Released Models
BabyVision 82.0 Source: Aliyun – Newly Released Models
CodeArena Global #4 Source: Aliyun – Newly Released Models

Third-party aggregator data: qwen3.8-27b posts Terminal-Bench 73.0 and DeepSWE 1.1 at 42.2, described as near-frontier agentic coding in a 27B body. Source: Local-Ai-Zone – Ai Updates August 2026

Latest News

The Qwen3.8 Wave (August 2026)

Alibaba shipped four new Qwen3.8 text models in one month: the 2.4T flagship max (Aug 2), its full open weights (Aug 12), a 27B dense VLM (Aug 17), and a ¥0.8-input flash with 1M context (Aug 26), plus a 0902 snapshot on Sep 2. The open-weight 2.4T release is the largest ever published, and qwen3.8-flash is explicitly positioned as a Claude Code/Codex-compatible drop-in via OpenAI/Anthropic protocol support. Sources: Aliyun – Newly Released Models , Aliyun – Billing For Model Studio

Bailian Becomes a Multi-Vendor Frontier Marketplace (August 2026)

Alibaba's Model Studio added Zhipu GLM-5.3/5.3-Flash, MiniMax H3, DeepSeek V4-Pro-0813/V4-Flash-0731, and Kimi K3 in August, making the platform a one-stop API for essentially every major Chinese frontier model. Source: Aliyun – Newly Released Models

Community Signals

Qwen's August releases were covered mainly through aggregator roundups rather than primary HN threads; the 2.4T open-weight release was called "the largest open-weight release ever" and the efficiency frontier (27B/flash class models running on consumer GPUs) drew the most commentary. Source: Local-Ai-Zone – Ai Updates August 2026 . The benchlm.ai September coding ranking places Qwen3.8-Max-tier models in the global top 10. Source: Benchlm – Coding

Enterprise Readiness

Feature Available? Details
SSO (SAML) Partial Alibaba Cloud account-level IAM; no coding-agent-specific SSO documented.
SCIM No Not documented.
Audit logs Partial Alibaba Cloud ActionTrail platform-level; model-studio-specific audit not documented.
IP indemnity No Not documented.
Data residency Yes Beijing, Virginia, Singapore, Frankfurt, Tokyo regions with distinct pricing; Singapore "international" deployments and US/EU regional endpoints. Source: Aliyun – Billing For Model Studio
HIPAA No Not advertised.
Air-gapped / on-prem Partial Open weights (qwen3.8-2.4t-a95b, 27b) enable self-hosting; no enterprise on-prem product.
SLA Yes Alibaba Cloud platform SLAs apply. Source: Aliyun
Admin controls (RBAC) Partial RAM/IAM policies; Token Plan team management. Source: Aliyun – Token Plan Overview

Transparency Gaps

Gap Details Severity
Autonomous-coding claim "Dozens of consecutive days" of autonomous coding has no published methodology, task set, or success criteria. High
Promo pricing opacity Active discounts moved to the console; the public billing page now shows only original prices, making effective rates unverifiable without an account. Medium
Rate limits Pay-as-you-go rate limits remain undisclosed. Medium
27B/flash agentic claims Terminal-Bench/DeepSWE figures for qwen3.8-27b come from aggregator coverage, not an Alibaba-published benchmark page. Medium
Open-weight license terms The exact license (Apache 2.0 vs Qwen-style restrictions) for qwen3.8-2.4t-a95b is not stated on the release note. Medium