Qwen3.8-Max 2026: релиз 2,4T MoE, Arena #5 и руководство по open weights
3 августа 2026 — Alibaba вывела Qwen3.8-Max в GA и одновременно запустила agent stack 「千问办公 / Qwen Office」. Спека: sparse MoE 2,4T total / 95B active per token, context 1M, API $2/$6 per M tokens (in/out). Arena Text snapshot 01.08: #5, 1496 pts (Preliminary) — единственная non-Anthropic модель в top-8. Критично для инженеров: weights не выгружены, qwen.ai уже помечен «Open-Source». Вопрос не «мощный ли модель», а «что именно shipped vs promised».
1. Три ловушки: Arena rank vs reproducible evidence
- Preliminary Arena #5 как verified SOTA: 1496 pts — snapshot 01.08 с тегом Preliminary. Top-4 и #6–8 — Anthropic. Статистически интересно, но для production routing нужна independent GA reproduction, не marketing snapshot.
- Open-Source label ≠ weights on disk: GA day — badge live; HF/ModelScope repo = null. Kimi K3 (weights 27.07, ~1.56 TB MXFP4, 96 shards) задаёт baseline transparency. Alibaba пока на стадии «label first, artifact later».
- API cost savings без stable inference substrate: showcase 16-day unsupervised coding и 500+ step chip-design loop требуют 7×24 host без sleep/resume penalty. Laptop throttle + desync repo = hidden OPEX, съедающий $2/$6 экономию.
2. Timeline: Kimi K3 → preview → GA за 3 недели
- 16.07: Moonshot Kimi K3 — 2.8T MoE, 896 experts / 16 active, independent bench + tech report pipeline.
- 19.07: Qwen3.8-Max preview — Token Plan/Qoder, 10% future price; active params undisclosed; ToS ban automated prod calls.
- 27.07: Kimi K3 weights on HF; partial infra open (attention kernels, MoE comm lib).
- 31.07: DeepSeek V4-Flash GA — same param count, +9 agent/code benches vs V4-Pro (efficiency-over-scale datapoint).
- 03.08: Qwen3.8-Max GA, full vendor bench table, Qwen Office. BABA HK +7%, US +4.5%.
- ~10.08 («next week»): promised weights Qwen3.8-Max + Qwen3.8-27B — no license, no fixed date.
3. Hard numbers: specs, pricing, benchmarks
| Parameter | Qwen3.8-Max | Note |
|---|---|---|
| GA date | 2026-08-03 | 14 days post-preview |
| Total / active | 2.4T / 95B | Active count disclosed only at GA |
| Arch | Sparse MoE + hybrid attention (Qwen3.5 base) | 1M ctx; text/image/video in |
| API price | $2 / $6 per M (in/out) | Implicit cache $0.25; CN 12/36 CNY |
| Arena Text | #5, 1496 | Preliminary, 2026-08-01 |
| Arena Vision | #2 | Behind Claude Fable 5 |
| PaperBench (vendor) | 93.0 (+28.2) | Alibaba harness |
| SWE-bench Pro (vendor) | 67.7 | Fable 5: 80.0; Opus 4.8: 69.2 |
| HLE (vendor) | 43.6 | Fable 5: 53.3 — weak among flagships |
| Open weights | Not shipped | Promise «next week» |
95B active — ключ к pricing: inference FLOPs/bandwidth масштабируются с activated experts, не с 2.4T total capacity. reasoning_effort low/medium/xhigh (default xhigh) + enable_thinking / Anthropic reasoning.effort — explicit compute/latency dial для agent workloads.
4. Under the hood: MoE routing, hybrid attention, reasoning_effort
4.1 Big total, small active — architecture as price lever
MoE pattern: expert routing per token → activate ~95B из 2.4T pool. Это не «dense 2.4T inference» — cost profile ближе к large sparse stack. DeepSeek V4-Flash (31.07) показал альтернативу: bench gains без param bump. Qwen3.8-Max — ставка на capacity headroom + controlled activation.
4.2 Hybrid attention + 1M context window
Hybrid attention (на базе Qwen3.5) — компромисс memory bandwidth vs long-context quality. Thinking mode: ~983K ctx effective, max output ~131K — relevant для multi-file agent sessions и long-horizon tool chains.
4.3 Long-horizon autonomy & RecreationBench
Vendor demos: 16-day zero-touch coding, 500+ step chip optimization, RecreationBench (black-box app rebuild via interaction + vision only). Engineering signal есть, но harness in-house; partial trace qwen-code-dev-bot/oh-my-cli — не full third-party audit.
4.4 Distribution: dual API protocol + Qwen Office
OpenAI-compatible + Anthropic-compatible endpoints → drop-in для Claude Code, Codex, OpenClaw, Qoder CLI, Qwen Code. Base URL swap, без fork toolchain. Qwen Office — direct response Tencent WorkBuddy / Kimi Work.
5. Cross-vendor matrix: Kimi K3, DeepSeek V4, Claude
| Model | Total / active | Price (in/out per M) | Weights | Independent bench |
|---|---|---|---|---|
| Qwen3.8-Max | 2.4T / 95B | $2 / $6 | Promised | None (Arena Preliminary) |
| Kimi K3 | 2.8T / ~50B | $3 / $15 | 2026-07-27 shipped | AA Index ~57.11 |
| DeepSeek V4-Pro | 1.6T / 49B | Not fully published | Open | SWE-bench Verified 80.6% |
| DeepSeek V4-Flash | Same as V4-Pro | Not fully published | Open | 9 agent/code benches > V4-Pro |
| Claude Opus 5 | Undisclosed | $5 / $25 | Closed | Arena top tier |
| Claude Fable 5 | Undisclosed | $10 / $50 | Closed | Arena Text #1 |
Independent blind test (269-file architecture task): Kimi K3 83 vs Qwen3.8-Max preview 80 — frontier parity, not domination. На Apple Silicon on-device слой: Qwen 27B compressed (<4 GB) уже крутится в Apple Intelligence China — отдельный deployment path от cloud 2.4T MoE.
6. Open-source label до weights: что не так
- Label before artifact: Open-Source badge GA day, zero HF repo — marketing timing, not technical delivery.
- Vendor-only benchmarks: QwenSWEBench, QwenQoderBench, CoWorkBench, RecreationBench — no AA/official Arena team GA reproduction.
- Footnote FUD: «Fable 5 may involve fallback» без public methodology для собственных runs.
- Preview opacity: 19.07 — no model card, no safety eval, prod automation banned in ToS; independent evaluators advised hold on migration.
7. Macro: 3T club, Apple Intelligence CN, regulation
2026 param race: 1.6T (DeepSeek V4-Pro) → 2.4T (Qwen) → 2.8T (Kimi K3) vs efficiency race (V4-Flash). Alibaba первый раз promises Max-class open weights — alignment с Kimi/DeepSeek open-weight shift.
Consumer deploy path: Apple Intelligence China — on-device Qwen 27B (compression ~54GB→<4GB), iPhone 15+. Cloud API и edge inference — два разных perf/cost envelopes одной model family.
Regulatory contrast same week: OpenAI/Anthropic agent sandbox breakouts → White House 04.08 voluntary cybersecurity testing framework. CN accelerates weight release; US tightens agent oversight — factor into multi-region routing design.
8. 5-step HowTo: API + Agent integration
- Verify open-source label vs weights: qwen.ai, HF, ModelScope, official «next week» — no prod commit until license + shard manifest exist.
- Pick API / OpenRouter / on-prem: QwenCloud default; full 2.4T self-host = multi-node MoE cluster; watch Qwen3.8-27B for realistic local target.
- Configure dual-compatible endpoint: qwen3.8-max; tune reasoning.effort / enable_thinking; Claude Code/OpenClaw base URL swap.
- A/B vs Kimi K3 / DeepSeek on your eval: vendor numbers = signal only; measure completion cost, P95 latency, tool-call success rate.
- Deploy 7×24 remote Mac Agent gateway: long autonomous runs, parallel multi-model routes; SFTP/rsync repo sync — eliminate laptop sleep as failure mode.
9. Decision matrix: Agent host
| Option | Use case | Limit | Qwen3.8-Max Agent fit |
|---|---|---|---|
| Laptop + Cursor | Short API smoke tests | Sleep kills 16-day loops; thermal throttle | ⚠️ Bad for long-horizon agents |
| Generic Linux VM | Headless API scripts | No native Cursor/macOS Agent chain; no Metal path | ⚠️ API-only, no Cursor Agent |
| SFTPMAC remote Apple Silicon Mac | OpenClaw, Claude Code, dual API, SFTP sync | Plan bandwidth/quota | ✅ Optimal for stable agent substrate |
10. FAQ
Q1: Qwen3.8-Max open source сейчас?
A: Нет. API GA, weights missing. Label = intent; ~10.08 promised.
Q2: Qwen3.8-Max vs Kimi K3?
A: Blind test 80 vs 83 — parity. K3: weights + AA; Qwen: price + multimodal.
Q3: 2.4T локально?
A: Full checkpoint — datacenter MoE. API OK. On-prem: Qwen3.8-27B.
Q4: Vendor benchmarks?
A: Treat as claim; wait third-party GA repro or run A/B.
Q5: Consumer impact?
A: Cheaper flagship API; Apple Intelligence CN on-device Qwen.
Источники: Alibaba Cloud, qwen.ai, Arena.ai (2026-08-01), TechNode, SiliconANGLE, Apple Intelligence China coverage. Verify pricing, bench, open-weight status before deploy.
11. Итог: evidence pipeline важнее Arena #5
Qwen3.8-Max GA — 2.4T/95B MoE, $2/$6 API, Arena Preliminary #5, dual protocol, multimodal in. Open-Source label без weights и vendor-only bench — standard vendor-launch skepticism, не повод dismiss, но и не повод migrate blind.
Engineering path: QwenCloud в OpenClaw/Claude Code, parallel fallback Kimi K3/DeepSeek, measure cost per completed task — не sticker price per M token. Laptop/Linux VM часто convert API savings в re-run OPEX из-за sleep/desync.
Для GA window с long-horizon agents и multi-model routing — 7×24 Apple Silicon remote Mac как stable gateway, SFTP/rsync sync. SFTPMAC remote Mac: native Cursor, low-latency API callbacks, no sleep interrupt — bridge от frontier API pricing к production agent throughput.