OpenAI · text · exact comparison

GPT-OSS 120B API providers and route evidence

Compare source-linked GPT-OSS 120B input and output token prices across provider routes while keeping infrastructure compute meters separate.

Current answer: 9 accepted GPT-OSS 120B observations; one current example is $0.015 / 1M cached input tokens, checked Aug 27, 2026. Open the evidence source, then compare an exact matching configuration below.

Exact-model monthly workload

GPT-OSS 120B price ranking status

No complete current comparison

Default workload: 100M uncached input + 20M output tokens. No provider has complete, fresh evidence for this exact workload. APIDir does not estimate missing meters.

Why there is no winner

No provider has complete, fresh evidence for this exact workload. APIDir does not estimate missing meters.

Enter your own token volume

Model-level provider comparison

GPT-OSS 120B price ranking status

Reviewed scopes, no current winner

Current answer: No current cheapest provider is published. The reviewed scopes do not have enough fresh, like-for-like offers to name a winner.

Provider routes with evidence
4
Accepted observations
9
Billing meters
1M cached input tokens, 1M input tokens, 1M output tokens, compute second
Latest accepted check
Aug 27, 2026

GPT-OSS 120B API input token pricing

Input token price ranking

Not currently rankable
Reviewed comparison scope
GPT-OSS 120B · chat completions · text input · global route · USD per 1M input tokens
Publication gate
At least 3 unique providers with accepted, fresh, reviewed evidence

Why there is no winner

Ranking withheld: 0 of 3 required providers have fresh, matching evidence.

Provider tier names remain visible and may include different service features. This list ranks the published token meter only, not speed, quality, limits, or reliability.

GPT-OSS 120B API output token pricing

Output token price ranking

Not currently rankable
Reviewed comparison scope
GPT-OSS 120B · chat completions · text output · global route · USD per 1M output tokens
Publication gate
At least 3 unique providers with accepted, fresh, reviewed evidence

Why there is no winner

Ranking withheld: 0 of 3 required providers have fresh, matching evidence.

Provider tier names remain visible and may include different service features. This list ranks the published token meter only, not speed, quality, limits, or reliability.

Uncollapsed source records

All accepted offer evidence

These rows remain separate until their offered model ID, route, meter, tier, resolution, duration, audio, and region are proven comparable.

GPT-OSS 120B provider offers · exact configuration scope
ProviderModelConfigurationRecorded priceEvidence
Fireworks AIaggregatorGPT-OSS 120Bmodel: accounts/fireworks/models/gpt-oss-120b · chat-completions · serverless standard · not-applicable · no audio$0.015 / 1M cached input tokensstaleProvider sourceChecked Aug 27, 2026
Fireworks AIaggregatorGPT-OSS 120Bmodel: accounts/fireworks/models/gpt-oss-120b · chat-completions · serverless standard · not-applicable · no audio$0.15 / 1M input tokensstaleProvider sourceChecked Aug 27, 2026
Fireworks AIaggregatorGPT-OSS 120Bmodel: accounts/fireworks/models/gpt-oss-120b · chat-completions · serverless standard · not-applicable · no audio$0.6 / 1M output tokensstaleProvider sourceChecked Aug 27, 2026
OpenRouterrouterGPT-OSS 120Bmodel: openai/gpt-oss-120b · chat-completions · standard · not-applicable · no audio$0.03 / 1M cached input tokensstaleProvider sourceChecked Aug 22, 2026
OpenRouterrouterGPT-OSS 120Bmodel: openai/gpt-oss-120b · chat-completions · standard · not-applicable · no audio$0.037 / 1M input tokensstaleProvider sourceChecked Aug 27, 2026
OpenRouterrouterGPT-OSS 120Bmodel: openai/gpt-oss-120b · chat-completions · standard · not-applicable · no audio$0.17 / 1M output tokensstaleProvider sourceChecked Aug 27, 2026
RunPodinfrastructureGPT-OSS 120Bcustom-serverless-worker · A100 serverless · not-applicable · no audio · source unit: hour$0.000755556 / compute secondstaleProvider sourceChecked Aug 22, 2026
Together AIaggregatorGPT-OSS 120Bmodel: openai/gpt-oss-120b · chat-completions · serverless · not-applicable · no audio$0.15 / 1M input tokensstaleProvider sourceChecked Aug 27, 2026
Together AIaggregatorGPT-OSS 120Bmodel: openai/gpt-oss-120b · chat-completions · serverless · not-applicable · no audio$0.6 / 1M output tokensstaleProvider sourceChecked Aug 27, 2026

Checked evidence, not a ranking: fresh rows were checked against the linked first-party source on the shown date. APIDir does not name a cheapest provider unless exact configurations and units match. Recheck before purchasing.

Developer resources

Provider documentation and pricing links

Compare integration documentation beside the exact first-party source used for each accepted observation.

Provider documentation, pricing, and accepted evidence for this model
ProviderChannelDeveloper resourcesAccepted evidence
Fireworks AIaggregatorDocumentationPricing page3 observationsEvidence source 1
OpenRouterrouterDocumentationPricing page3 observationsEvidence source 1
RunPodinfrastructureDocumentationPricing page1 observationEvidence source 1
Together AIaggregatorDocumentationPricing page2 observationsEvidence source 1

Decision guide

How to compare GPT-OSS 120B API providers

GPT-OSS 120B can be consumed through shared inference APIs or deployed on rented infrastructure. Those are different products. Shared routes expose input and output token prices, while a GPU service may bill compute time and leave utilization, batching, serving software, and operations to the buyer. APIDir keeps the meters separate.

The published token rankings cover one exact model and one meter at a time. They do not score latency, rate limits, output quality, uptime, support, data terms, or the total cost of a self-hosted deployment. Use the ranking to shortlist routes, then run a representative prompt and concurrency benchmark.

What changes GPT-OSS 120B API cost

  • Input-to-output token ratio, prompt reuse, and provider cache behavior
  • Throughput, queueing, rate limits, and any priority or dedicated tier
  • For self-hosting: GPU utilization, replicas, idle time, storage, networking, and operations

Integration checks before production

  • Pin the exact GPT-OSS 120B provider key and API compatibility version
  • Set token and timeout budgets at the caller boundary
  • Log provider, route, input tokens, output tokens, retries, and terminal errors without storing sensitive prompts

Workloads this comparison can inform

Open-weight application backends

Compare managed routes before taking on deployment and utilization risk.

Batch text processing

Model the real input-output ratio and provider concurrency rather than comparing only input price.

Controlled self-hosting

Use compute rows for capacity planning, not as an all-in per-token quote.

How to read the price result

GPT-OSS 120B currently has three reviewed shared routes for both token meters, so APIDir can publish a narrow price ranking. Hardware-time evidence remains outside that order.

Related model price comparisons

Use this evidence safely

A recorded price is not a complete production quote. Open the source, reproduce the exact configuration, and include retries, failed work, storage, egress, support, rate limits, and downstream processing in total cost.

Questions developers ask

What makes two GPT-OSS offers comparable?

The exact model and endpoint, provider route, input and output meters, cache policy, context-length price band, tools, region, and service tier must match.

Does APIDir name a cheapest provider on this page?

Only when at least three explicitly reviewed provider offers are fresh, publication-permitted, and within the same published comparison scope. Stale or mismatched evidence cannot support that claim.

How should I confirm a budget?

Open the linked provider source, verify the configuration, then calculate your own successful and failed workload volume.

APIDir Delta · controlled rollout

Join the source-linked pricing update waitlist.

Confirm your address to join the waitlist. Editorial delivery starts only after sender, suppression, and evidence checks are active.