Alibaba Cloud · text · family comparison
Qwen API pricing by model and provider
Compare Qwen API pricing by exact model ID, provider channel, input and output token meter, context, and official source.
Exact-model monthly workload
Qwen3.8 Max price ranking status
Default workload: 100M uncached input + 20M output tokens. No provider has complete, fresh evidence for this exact workload. APIDir does not estimate missing meters.
Why there is no winner
No provider has complete, fresh evidence for this exact workload. APIDir does not estimate missing meters.
Model-level provider comparison
Qwen API price ranking status
Current answer: No current cheapest provider is published. Accepted evidence exists, but it does not yet form a named, like-for-like comparison group.
- Provider routes with evidence
- 4
- Accepted observations
- 13
- Billing meters
- 1M cached input tokens, 1M input tokens, 1M output tokens
- Latest accepted check
- Aug 22, 2026
Uncollapsed source records
All accepted offer evidence
These rows remain separate until their offered model ID, route, meter, tier, resolution, duration, audio, and region are proven comparable.
| Provider | Model | Configuration | Recorded price | Evidence |
|---|---|---|---|---|
| Alibaba Cloud Model Studioofficial | Qwen API | model: qwen3.8-max · chat-completions · Qwen3.8 Max global pay-as-you-go · not-applicable · no audio · source unit: CNY 1.5 per 1M implicit-cache input; normalized at 6.7412 CNY/USD (Federal Reserve H.10, 2026-08-14) | $0.222512312 / 1M cached input tokensstale | Provider sourceChecked Aug 22, 2026 |
| Alibaba Cloud Model Studioofficial | Qwen API | model: qwen3.8-max · chat-completions · Qwen3.8 Max global pay-as-you-go · not-applicable · no audio · source unit: CNY 12 per 1M; normalized at 6.7412 CNY/USD (Federal Reserve H.10, 2026-08-14) | $1.780098499 / 1M input tokensstale | Provider sourceChecked Aug 22, 2026 |
| Alibaba Cloud Model Studioofficial | Qwen API | model: qwen3.8-max · chat-completions · Qwen3.8 Max global pay-as-you-go · not-applicable · no audio · source unit: CNY 36 per 1M; normalized at 6.7412 CNY/USD (Federal Reserve H.10, 2026-08-14) | $5.340295496 / 1M output tokensstale | Provider sourceChecked Aug 22, 2026 |
| Novita AIaggregator | Qwen API | model: qwen/qwen3.8-max · chat-completions · Qwen3.8 Max serverless · not-applicable · no audio | $0.25 / 1M cached input tokensstale | Provider sourceChecked Aug 22, 2026 |
| Novita AIaggregator | Qwen API | model: qwen/qwen3.8-max · chat-completions · Qwen3.8 Max serverless · not-applicable · no audio | $2 / 1M input tokensstale | Provider sourceChecked Aug 22, 2026 |
| Novita AIaggregator | Qwen API | model: qwen/qwen3.8-max · chat-completions · Qwen3.8 Max serverless · not-applicable · no audio | $6 / 1M output tokensstale | Provider sourceChecked Aug 22, 2026 |
| OpenRouterrouter | Qwen API | model: qwen/qwen3.8-2.4t-a95b · chat-completions · standard · not-applicable · no audio | $2 / 1M input tokensstale | Provider sourceChecked Aug 22, 2026 |
| OpenRouterrouter | Qwen API | model: qwen/qwen3.8-2.4t-a95b · chat-completions · standard · not-applicable · no audio | $6 / 1M output tokensstale | Provider sourceChecked Aug 22, 2026 |
| OpenRouterrouter | Qwen API | model: qwen/qwen3.8-max · chat-completions · Qwen3.8 Max standard · not-applicable · no audio | $0.25 / 1M cached input tokensstale | Provider sourceChecked Aug 22, 2026 |
| OpenRouterrouter | Qwen API | model: qwen/qwen3.8-max · chat-completions · Qwen3.8 Max standard · not-applicable · no audio | $2 / 1M input tokensstale | Provider sourceChecked Aug 22, 2026 |
| OpenRouterrouter | Qwen API | model: qwen/qwen3.8-max · chat-completions · Qwen3.8 Max standard · not-applicable · no audio | $6 / 1M output tokensstale | Provider sourceChecked Aug 22, 2026 |
| Together AIaggregator | Qwen API | model: Qwen/Qwen3.8-2.4T-A95B · chat-completions · serverless · not-applicable · no audio | $2.5 / 1M input tokensstale | Provider sourceChecked Aug 22, 2026 |
| Together AIaggregator | Qwen API | model: Qwen/Qwen3.8-2.4T-A95B · chat-completions · serverless · not-applicable · no audio | $6.25 / 1M output tokensstale | Provider sourceChecked Aug 22, 2026 |
Checked evidence, not a ranking: fresh rows were checked against the linked first-party source on the shown date. APIDir does not name a cheapest provider unless exact configurations and units match. Recheck before purchasing.
Developer resources
Provider documentation and pricing links
Compare integration documentation beside the exact first-party source used for each accepted observation.
| Provider | Channel | Developer resources | Accepted evidence |
|---|---|---|---|
| Alibaba Cloud Model Studio | official | DocumentationPricing page | 3 observationsEvidence source 1 |
| Novita AI | aggregator | DocumentationPricing page | 3 observationsEvidence source 1Evidence source 2 |
| OpenRouter | router | DocumentationPricing page | 5 observationsEvidence source 1Evidence source 2 |
| Together AI | aggregator | DocumentationPricing page | 2 observationsEvidence source 1 |
Route-specific buying guide
How to budget Qwen3.8 Max on OpenRouter
Route ID: qwen/qwen3.8-max
The accepted rows describe OpenRouter's Qwen3.8 Max route. They are not Alibaba Cloud direct-list prices and should not be presented as a family-wide cheapest Qwen claim.
OpenRouter reported a 1,000,000-token context window when checked. A large context limit does not mean every prompt should use it: longer requests increase billable input and may change latency or upstream routing behavior.
Cost formula
Estimate one request as uncached input tokens × the input rate, cached input tokens × the cached-input rate, plus output tokens × the output rate. Divide every token count by one million before applying APIDir's displayed rates.
Costs that move the bill
- Prompt size, retrieved context, conversation history, and agent scratch data
- Cache-hit rate and any separately priced cache-write operation shown by the route
- Completion length, retry policy, tool loops, and rejected or abandoned outputs
- Router behavior, account limits, region, data terms, and availability requirements
Checks before purchase
- Record the exact qwen/qwen3.8-max route ID in every benchmark and invoice sample.
- Test a representative prompt-length distribution instead of multiplying one best-case request.
- Reconcile estimated prompt, cached-prompt, and completion tokens against returned usage fields.
- Compare the router offer with a direct Qwen route only after identity, meter, limits, and terms match.
Source checked Aug 22, 2026: OpenRouter public model API. Recheck the live route before committing spend.
Estimate this route with the LLM workload calculator →
Compare another model family
Decision guide
How to compare Qwen API providers
Qwen is a broad family covering many sizes, releases, modalities, and deployment choices. The current comparison uses Qwen3.8 2.4T A95B. APIDir stores each provider's route ID so a cheaper smaller Qwen model cannot be mistaken for the same product.
Input and output rates are only part of the decision. Context limits, caching, batching, quantization, throughput, rate limits, tool support, and hosted versus self-managed operations all affect total cost and suitability. Benchmark the exact route with representative prompts.
What changes Qwen API cost
- Exact Qwen release and size, context, input-output ratio, and cache behavior
- Shared serverless, dedicated, router, or self-hosted deployment channel
- Throughput, quantization, tool calls, retries, storage, and operations
Integration checks before production
- Pin the full provider model key rather than a Qwen family alias
- Validate context and output limits before submission
- Capture route, token counts, latency, retry count, and provider request ID
Workloads this comparison can inform
Multilingual application backends
Measure output quality and tokenization on your language mix beside the meter.
Large-context processing
Confirm the provider's supported context and any separate high-context rate.
Open-model deployment planning
Compare managed token pricing with realistic utilization and operations, not raw GPU price alone.
How to read the price result
Two accepted routes expose the exact reviewed Qwen model with separate input and output meters. APIDir displays both but does not publish a winner below the three-provider evidence threshold.
Related model price comparisons
Use this evidence safely
A recorded price is not a complete production quote. Open the source, reproduce the exact configuration, and include retries, failed work, storage, egress, support, rate limits, and downstream processing in total cost.
Questions developers ask
What makes two Qwen offers comparable?
The exact model and endpoint, provider route, input and output meters, cache policy, context-length price band, tools, region, and service tier must match.
Does APIDir name a cheapest provider on this page?
Only when at least three explicitly reviewed provider offers are fresh, publication-permitted, and within the same published comparison scope. Stale or mismatched evidence cannot support that claim.
How should I confirm a budget?
Open the linked provider source, verify the configuration, then calculate your own successful and failed workload volume.
APIDir Delta · controlled rollout
Join the source-linked pricing update waitlist.
Confirm your address to join the waitlist. Editorial delivery starts only after sender, suppression, and evidence checks are active.