Multi-model API

Select GPT-5.5, GPT-5.6 sol or terra, and GPT-6 astra through one InferGate endpoint without changing SDK configuration.

Keep your client and endpoint fixed, and select a published identifier per request. This lets an application compare models without creating a separate SDK connection for each model.

Explicit selection instead of hidden replacement

Store the chosen model alongside your request ID and output. If you implement fallback in your own application, make that decision explicit and retain both requested and returned identities. An unavailable model must not appear as a successful call to another model.

Evaluate a model switch

Prepaid usage and costs

API usage consumes available account credit. Review current input and output rates in application pricing before running a workload. This page does not promise free requests or a fixed discount. See how prepaid pricing works and compare actual usage logs after a small test.

One client, separate model decisions

DecisionKeep fixedChange deliberately
Client setupBase URL and server-side keyAccount credentials only when intended
EvaluationPrompt set and pass criteriaExact model identifier
StreamingCompletion and cancellation handlingModel-specific optional parameters after testing
FallbackOutcome loggingAn explicit model change disclosed by your application

Runnable integration example

curl --fail-with-body https://api.useinfergate.com/v1/responses \
  -H "Authorization: Bearer $INFERGATE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-6-astra","input":"Return a short release note for an API client update.","max_output_tokens":256}'

Frequently asked questions

Is a model family name a routing alias?

No. Use an exact published identifier. InferGate does not silently turn an unavailable model into a different model identity.

Does every API key see all four models?

Not necessarily. The public catalog describes the available service identities; account permissions and key restrictions can narrow the authenticated list.

Should I automatically retry with another model?

That is an application decision. Record the original failure, new requested model and actual response identity. Consider duplicate output, cost and user expectations before adding fallback.

Continue your integration

Review prepaid pricing · Read API documentation · Get Started

Start with a small request

Create an InferGate account Read API documentation

Explore the API

OpenAI-compatible APIMulti-model APIAI API gateway