aggregator channel · AI inference platform

Fireworks AI pricing by model

A hosted inference platform for generative AI models, fine-tuning, and compound AI systems.

Current answer: 10 accepted observations; an example row is $0.015 / 1M cached input tokens, checked Aug 27, 2026. Open the evidence source before purchase.

Where this provider fits

A hosted inference platform for generative AI models, fine-tuning, and compound AI systems.

What can change the bill

  • Serverless and dedicated costs are not directly interchangeable
  • Model-specific limits require documentation checks

Who should evaluate this route

  • Hosted open-model catalog
  • Fine-tuning and deployment controls
  • OpenAI-compatible API surface

Compare two nearby options

Use the same model, meter, configuration, and workload when comparing channels.

Recorded price evidence

Fireworks AI recorded offers · checked-source scope
ProviderModelConfigurationRecorded priceEvidence
Fireworks AIaggregatorGLM-5.2model: z-ai/glm-5.2 · chat-completions · serverless-standard · not-applicable · no audio$1.4 / 1M input tokensstaleProvider sourceChecked Aug 22, 2026
Fireworks AIaggregatorGLM-5.2model: z-ai/glm-5.2 · chat-completions · serverless-standard · not-applicable · no audio$4.4 / 1M output tokensstaleProvider sourceChecked Aug 22, 2026
Fireworks AIaggregatorGPT-OSS 120Bmodel: accounts/fireworks/models/gpt-oss-120b · chat-completions · serverless standard · not-applicable · no audio$0.015 / 1M cached input tokensstaleProvider sourceChecked Aug 27, 2026
Fireworks AIaggregatorGPT-OSS 120Bmodel: accounts/fireworks/models/gpt-oss-120b · chat-completions · serverless standard · not-applicable · no audio$0.15 / 1M input tokensstaleProvider sourceChecked Aug 27, 2026
Fireworks AIaggregatorGPT-OSS 120Bmodel: accounts/fireworks/models/gpt-oss-120b · chat-completions · serverless standard · not-applicable · no audio$0.6 / 1M output tokensstaleProvider sourceChecked Aug 27, 2026
Fireworks AIaggregatorKimi APImodel: fireworks/kimi-k3 · chat-completions · serverless standard · not-applicable · no audio$0.3 / 1M cached input tokensstaleProvider sourceChecked Aug 22, 2026
Fireworks AIaggregatorKimi APImodel: fireworks/kimi-k3 · chat-completions · serverless standard · not-applicable · no audio$3 / 1M input tokensstaleProvider sourceChecked Aug 22, 2026
Fireworks AIaggregatorKimi APImodel: fireworks/kimi-k3 · chat-completions · serverless-standard · not-applicable · no audio$3 / 1M input tokensstaleProvider sourceChecked Aug 22, 2026
Fireworks AIaggregatorKimi APImodel: fireworks/kimi-k3 · chat-completions · serverless-standard · not-applicable · no audio$15 / 1M output tokensstaleProvider sourceChecked Aug 22, 2026
Fireworks AIaggregatorKimi APImodel: fireworks/kimi-k3 · chat-completions · serverless standard · not-applicable · no audio$15 / 1M output tokensstaleProvider sourceChecked Aug 22, 2026

Checked evidence, not a ranking: fresh rows were checked against the linked first-party source on the shown date. APIDir does not name a cheapest provider unless exact configurations and units match. Recheck before purchasing.

Model price comparisons for this provider

Follow each model page to compare this provider route with other accepted, source-linked offers on matching billing meters.

Questions developers ask

What does APIDir verify about Fireworks AI?

APIDir separates the service channel from the underlying model owner, records exact offer configuration, links the source, and exposes the check date.

Does an entry mean APIDir recommends this provider?

No. A directory record is a research starting point. Review current sources, legal terms, limits, and a workload-specific test.

Why can a displayed observation be stale?

Historical observations remain visible for audit, but stale records do not support current cheapest or savings claims.

APIDir Delta · controlled rollout

Join the source-linked pricing update waitlist.

Confirm your address to join the waitlist. Editorial delivery starts only after sender, suppression, and evidence checks are active.