Alibaba Cloud · text · family comparison
Qwen API pricing by model and provider
Compare Qwen API pricing by exact model ID, provider channel, input and output token meter, context, and official source.
Model-level provider comparison
Qwen API price ranking status
Current answer: No current cheapest provider is published. Accepted evidence exists, but it does not yet form a named, like-for-like comparison group.
- Provider routes with evidence
- 2
- Accepted observations
- 7
- Billing meters
- 1M cached input tokens, 1M input tokens, 1M output tokens
- Latest accepted check
- Aug 22, 2026
Why there is no winner
No reviewed comparison scope exists for this model yet. Accepted evidence can still be inspected below, but APIDir will not sort unlike offers or infer a cheapest provider.
Read the publication gateUncollapsed source records
All accepted offer evidence
These rows remain separate until their offered model ID, route, meter, tier, resolution, duration, audio, and region are proven comparable.
| Provider | Model | Configuration | Recorded price | Evidence |
|---|---|---|---|---|
| OpenRouterrouter | Qwen API | model: qwen/qwen3.8-2.4t-a95b · chat-completions · standard · not-applicable · no audio | $2 / 1M input tokensfresh | Provider sourceChecked Aug 22, 2026 |
| OpenRouterrouter | Qwen API | model: qwen/qwen3.8-2.4t-a95b · chat-completions · standard · not-applicable · no audio | $6 / 1M output tokensfresh | Provider sourceChecked Aug 22, 2026 |
| OpenRouterrouter | Qwen API | model: qwen/qwen3.8-max · chat-completions · Qwen3.8 Max standard · not-applicable · no audio | $0.25 / 1M cached input tokensfresh | Provider sourceChecked Aug 22, 2026 |
| OpenRouterrouter | Qwen API | model: qwen/qwen3.8-max · chat-completions · Qwen3.8 Max standard · not-applicable · no audio | $2 / 1M input tokensfresh | Provider sourceChecked Aug 22, 2026 |
| OpenRouterrouter | Qwen API | model: qwen/qwen3.8-max · chat-completions · Qwen3.8 Max standard · not-applicable · no audio | $6 / 1M output tokensfresh | Provider sourceChecked Aug 22, 2026 |
| Together AIaggregator | Qwen API | model: Qwen/Qwen3.8-2.4T-A95B · chat-completions · serverless · not-applicable · no audio | $2.5 / 1M input tokensfresh | Provider sourceChecked Aug 22, 2026 |
| Together AIaggregator | Qwen API | model: Qwen/Qwen3.8-2.4T-A95B · chat-completions · serverless · not-applicable · no audio | $6.25 / 1M output tokensfresh | Provider sourceChecked Aug 22, 2026 |
Checked evidence, not a ranking: fresh rows were checked against the linked first-party source on the shown date. APIDir does not name a cheapest provider unless exact configurations and units match. Recheck before purchasing.
Developer resources
Provider documentation and pricing links
Compare integration documentation beside the exact first-party source used for each accepted observation.
| Provider | Channel | Developer resources | Accepted evidence |
|---|---|---|---|
| OpenRouter | router | DocumentationPricing page | 5 observationsEvidence source 1Evidence source 2 |
| Together AI | aggregator | DocumentationPricing page | 2 observationsEvidence source 1 |
Route-specific buying guide
How to budget Qwen3.8 Max on OpenRouter
Route ID: qwen/qwen3.8-max
The accepted rows describe OpenRouter's Qwen3.8 Max route. They are not Alibaba Cloud direct-list prices and should not be presented as a family-wide cheapest Qwen claim.
OpenRouter reported a 1,000,000-token context window when checked. A large context limit does not mean every prompt should use it: longer requests increase billable input and may change latency or upstream routing behavior.
Cost formula
Estimate one request as uncached input tokens × the input rate, cached input tokens × the cached-input rate, plus output tokens × the output rate. Divide every token count by one million before applying APIDir's displayed rates.
Costs that move the bill
- Prompt size, retrieved context, conversation history, and agent scratch data
- Cache-hit rate and any separately priced cache-write operation shown by the route
- Completion length, retry policy, tool loops, and rejected or abandoned outputs
- Router behavior, account limits, region, data terms, and availability requirements
Checks before purchase
- Record the exact qwen/qwen3.8-max route ID in every benchmark and invoice sample.
- Test a representative prompt-length distribution instead of multiplying one best-case request.
- Reconcile estimated prompt, cached-prompt, and completion tokens against returned usage fields.
- Compare the router offer with a direct Qwen route only after identity, meter, limits, and terms match.
Source checked Aug 22, 2026: OpenRouter public model API. Recheck the live route before committing spend.
Estimate this route with the LLM workload calculator →
Compare another model family
Decision guide
How to compare Qwen API API providers
Qwen is a broad family covering many sizes, releases, modalities, and deployment choices. The current comparison uses Qwen3.8 2.4T A95B. APIDir stores each provider's route ID so a cheaper smaller Qwen model cannot be mistaken for the same product.
Input and output rates are only part of the decision. Context limits, caching, batching, quantization, throughput, rate limits, tool support, and hosted versus self-managed operations all affect total cost and suitability. Benchmark the exact route with representative prompts.
What changes the Qwen API API cost
- Exact Qwen release and size, context, input-output ratio, and cache behavior
- Shared serverless, dedicated, router, or self-hosted deployment channel
- Throughput, quantization, tool calls, retries, storage, and operations
Integration checks before production
- Pin the full provider model key rather than a Qwen family alias
- Validate context and output limits before submission
- Capture route, token counts, latency, retry count, and provider request ID
Workloads this comparison can inform
Multilingual application backends
Measure output quality and tokenization on your language mix beside the meter.
Large-context processing
Confirm the provider's supported context and any separate high-context rate.
Open-model deployment planning
Compare managed token pricing with realistic utilization and operations, not raw GPU price alone.
How to read the price result
Two accepted routes expose the exact reviewed Qwen model with separate input and output meters. APIDir displays both but does not publish a winner below the three-provider evidence threshold.
Related model price comparisons
Use this evidence safely
A recorded price is not a complete production quote. Open the source, reproduce the exact configuration, and include retries, failed work, storage, egress, support, rate limits, and downstream processing in total cost.
Questions developers ask
What makes two Qwen offers comparable?
The exact model and endpoint, provider route, input and output meters, cache policy, context-length price band, tools, region, and service tier must match.
Does APIDir name a cheapest provider on this page?
Only when at least three explicitly reviewed provider offers are fresh, publication-permitted, and within the same published comparison scope. Stale or mismatched evidence cannot support that claim.
How should I confirm a budget?
Open the linked provider source, verify the configuration, then calculate your own successful and failed workload volume.
APIDir Delta · controlled rollout
Reserve source-linked price research.
Confirm your address to join the rollout. Editorial delivery starts only after sender, suppression, and evidence checks are active.