OpenAI · text · exact comparison
GPT-OSS 120B API providers and route evidence
Compare source-linked GPT-OSS 120B input and output token prices across provider routes while keeping infrastructure compute meters separate.
Exact-model monthly workload
GPT-OSS 120B price ranking status
Default workload: 100M uncached input + 20M output tokens. No provider has complete, fresh evidence for this exact workload. APIDir does not estimate missing meters.
Why there is no winner
No provider has complete, fresh evidence for this exact workload. APIDir does not estimate missing meters.
Model-level provider comparison
GPT-OSS 120B price ranking status
Current answer: No current cheapest provider is published. The reviewed scopes do not have enough fresh, like-for-like offers to name a winner.
- Provider routes with evidence
- 4
- Accepted observations
- 9
- Billing meters
- 1M cached input tokens, 1M input tokens, 1M output tokens, compute second
- Latest accepted check
- Aug 27, 2026
GPT-OSS 120B API input token pricing
Input token price ranking
- Reviewed comparison scope
- GPT-OSS 120B · chat completions · text input · global route · USD per 1M input tokens
- Publication gate
- At least 3 unique providers with accepted, fresh, reviewed evidence
Why there is no winner
Ranking withheld: 0 of 3 required providers have fresh, matching evidence.
Provider tier names remain visible and may include different service features. This list ranks the published token meter only, not speed, quality, limits, or reliability.
GPT-OSS 120B API output token pricing
Output token price ranking
- Reviewed comparison scope
- GPT-OSS 120B · chat completions · text output · global route · USD per 1M output tokens
- Publication gate
- At least 3 unique providers with accepted, fresh, reviewed evidence
Why there is no winner
Ranking withheld: 0 of 3 required providers have fresh, matching evidence.
Provider tier names remain visible and may include different service features. This list ranks the published token meter only, not speed, quality, limits, or reliability.
Uncollapsed source records
All accepted offer evidence
These rows remain separate until their offered model ID, route, meter, tier, resolution, duration, audio, and region are proven comparable.
| Provider | Model | Configuration | Recorded price | Evidence |
|---|---|---|---|---|
| Fireworks AIaggregator | GPT-OSS 120B | model: accounts/fireworks/models/gpt-oss-120b · chat-completions · serverless standard · not-applicable · no audio | $0.015 / 1M cached input tokensstale | Provider sourceChecked Aug 27, 2026 |
| Fireworks AIaggregator | GPT-OSS 120B | model: accounts/fireworks/models/gpt-oss-120b · chat-completions · serverless standard · not-applicable · no audio | $0.15 / 1M input tokensstale | Provider sourceChecked Aug 27, 2026 |
| Fireworks AIaggregator | GPT-OSS 120B | model: accounts/fireworks/models/gpt-oss-120b · chat-completions · serverless standard · not-applicable · no audio | $0.6 / 1M output tokensstale | Provider sourceChecked Aug 27, 2026 |
| OpenRouterrouter | GPT-OSS 120B | model: openai/gpt-oss-120b · chat-completions · standard · not-applicable · no audio | $0.03 / 1M cached input tokensstale | Provider sourceChecked Aug 22, 2026 |
| OpenRouterrouter | GPT-OSS 120B | model: openai/gpt-oss-120b · chat-completions · standard · not-applicable · no audio | $0.037 / 1M input tokensstale | Provider sourceChecked Aug 27, 2026 |
| OpenRouterrouter | GPT-OSS 120B | model: openai/gpt-oss-120b · chat-completions · standard · not-applicable · no audio | $0.17 / 1M output tokensstale | Provider sourceChecked Aug 27, 2026 |
| RunPodinfrastructure | GPT-OSS 120B | custom-serverless-worker · A100 serverless · not-applicable · no audio · source unit: hour | $0.000755556 / compute secondstale | Provider sourceChecked Aug 22, 2026 |
| Together AIaggregator | GPT-OSS 120B | model: openai/gpt-oss-120b · chat-completions · serverless · not-applicable · no audio | $0.15 / 1M input tokensstale | Provider sourceChecked Aug 27, 2026 |
| Together AIaggregator | GPT-OSS 120B | model: openai/gpt-oss-120b · chat-completions · serverless · not-applicable · no audio | $0.6 / 1M output tokensstale | Provider sourceChecked Aug 27, 2026 |
Checked evidence, not a ranking: fresh rows were checked against the linked first-party source on the shown date. APIDir does not name a cheapest provider unless exact configurations and units match. Recheck before purchasing.
Developer resources
Provider documentation and pricing links
Compare integration documentation beside the exact first-party source used for each accepted observation.
| Provider | Channel | Developer resources | Accepted evidence |
|---|---|---|---|
| Fireworks AI | aggregator | DocumentationPricing page | 3 observationsEvidence source 1 |
| OpenRouter | router | DocumentationPricing page | 3 observationsEvidence source 1 |
| RunPod | infrastructure | DocumentationPricing page | 1 observationEvidence source 1 |
| Together AI | aggregator | DocumentationPricing page | 2 observationsEvidence source 1 |
Decision guide
How to compare GPT-OSS 120B API providers
GPT-OSS 120B can be consumed through shared inference APIs or deployed on rented infrastructure. Those are different products. Shared routes expose input and output token prices, while a GPU service may bill compute time and leave utilization, batching, serving software, and operations to the buyer. APIDir keeps the meters separate.
The published token rankings cover one exact model and one meter at a time. They do not score latency, rate limits, output quality, uptime, support, data terms, or the total cost of a self-hosted deployment. Use the ranking to shortlist routes, then run a representative prompt and concurrency benchmark.
What changes GPT-OSS 120B API cost
- Input-to-output token ratio, prompt reuse, and provider cache behavior
- Throughput, queueing, rate limits, and any priority or dedicated tier
- For self-hosting: GPU utilization, replicas, idle time, storage, networking, and operations
Integration checks before production
- Pin the exact GPT-OSS 120B provider key and API compatibility version
- Set token and timeout budgets at the caller boundary
- Log provider, route, input tokens, output tokens, retries, and terminal errors without storing sensitive prompts
Workloads this comparison can inform
Open-weight application backends
Compare managed routes before taking on deployment and utilization risk.
Batch text processing
Model the real input-output ratio and provider concurrency rather than comparing only input price.
Controlled self-hosting
Use compute rows for capacity planning, not as an all-in per-token quote.
How to read the price result
GPT-OSS 120B currently has three reviewed shared routes for both token meters, so APIDir can publish a narrow price ranking. Hardware-time evidence remains outside that order.
Related model price comparisons
Use this evidence safely
A recorded price is not a complete production quote. Open the source, reproduce the exact configuration, and include retries, failed work, storage, egress, support, rate limits, and downstream processing in total cost.
Questions developers ask
What makes two GPT-OSS offers comparable?
The exact model and endpoint, provider route, input and output meters, cache policy, context-length price band, tools, region, and service tier must match.
Does APIDir name a cheapest provider on this page?
Only when at least three explicitly reviewed provider offers are fresh, publication-permitted, and within the same published comparison scope. Stale or mismatched evidence cannot support that claim.
How should I confirm a budget?
Open the linked provider source, verify the configuration, then calculate your own successful and failed workload volume.
APIDir Delta · controlled rollout
Join the source-linked pricing update waitlist.
Confirm your address to join the waitlist. Editorial delivery starts only after sender, suppression, and evidence checks are active.