Stream API responses with InferGate

Handle Responses API streaming text, completion events and interrupted output with an OpenAI Python SDK example for InferGate.

Streaming emits events as a response is produced. A successful connection does not mean generation has completed. Keep completion and output checks separate from HTTP status.

Consume deltas and require completion

import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["INFERGATE_API_KEY"],
                base_url="https://api.useinfergate.com/v1")
events = client.responses.create(model="gpt-5.5",
    input="Explain streaming in two sentences.", stream=True,
    max_output_tokens=256)
completed = False
output = []
try:
    for event in events:
        if event.type == "response.output_text.delta":
            output.append(event.delta)
            print(event.delta, end="", flush=True)
        elif event.type == "response.completed":
            completed = True
        elif event.type in ("response.failed", "response.incomplete", "error"):
            raise RuntimeError(str(event))
finally:
    events.close()
if not completed or not "".join(output).strip():
    raise RuntimeError("Stream did not complete with text output")

Deploy behind a proxy

Debug an incomplete stream

Record the request ID, model, start time, first-text time and termination reason. Compare client, proxy and application logs. Zero useful text without a completed event is incomplete, even when the original response status is 200.

Classify a stream outcome

Layer or outcomeExpected value or observationAction
Completed with textCompletion event and useful textAccept and record usage
Client canceledCaller intentionally stopped readingRecord cancellation, not successful completion
Upstream failureFailure event or upstream errorPreserve error and request ID
Incomplete outputConnection ends without completionTreat as incomplete even after HTTP 200

Frequently asked questions

Why is HTTP 200 not sufficient for streaming success?

The status is sent before generation necessarily finishes. A later interruption can leave no text or only partial text. Require a completion signal and useful output.

Can I automatically retry after a disconnect?

A retry may duplicate both output and usage. Decide based on whether the original request completed, whether output was already consumed and your application workflow.

Should I hide client cancellation errors?

No. Separate intentional client cancellation from upstream failure and server interruption so reliability metrics remain meaningful.

Continue your integration

Review prepaid pricing · Read API documentation · Get Started

Start with a small request

Create an InferGate account Read API documentation

Explore the API

OpenAI-compatible APIMulti-model APIAI API gateway