Stream API responses with InferGate
Handle Responses API streaming text, completion events and interrupted output with an OpenAI Python SDK example for InferGate.
Streaming emits events as a response is produced. A successful connection does not mean generation has completed. Keep completion and output checks separate from HTTP status.
Consume deltas and require completion
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["INFERGATE_API_KEY"],
base_url="https://api.useinfergate.com/v1")
events = client.responses.create(model="gpt-5.5",
input="Explain streaming in two sentences.", stream=True,
max_output_tokens=256)
completed = False
output = []
try:
for event in events:
if event.type == "response.output_text.delta":
output.append(event.delta)
print(event.delta, end="", flush=True)
elif event.type == "response.completed":
completed = True
elif event.type in ("response.failed", "response.incomplete", "error"):
raise RuntimeError(str(event))
finally:
events.close()
if not completed or not "".join(output).strip():
raise RuntimeError("Stream did not complete with text output")Deploy behind a proxy
- Use an appropriate read timeout for your workload and upstream latency.
- Disable response buffering for SSE where your proxy supports it.
- Flush generated text as it arrives and preserve the request context.
- Classify client cancellation separately from upstream or server interruption.
- Do not blindly retry a partially consumed stream: it may duplicate output and usage.
Debug an incomplete stream
Record the request ID, model, start time, first-text time and termination reason. Compare client, proxy and application logs. Zero useful text without a completed event is incomplete, even when the original response status is 200.
Classify a stream outcome
| Layer or outcome | Expected value or observation | Action |
|---|---|---|
| Completed with text | Completion event and useful text | Accept and record usage |
| Client canceled | Caller intentionally stopped reading | Record cancellation, not successful completion |
| Upstream failure | Failure event or upstream error | Preserve error and request ID |
| Incomplete output | Connection ends without completion | Treat as incomplete even after HTTP 200 |
Frequently asked questions
Why is HTTP 200 not sufficient for streaming success?
The status is sent before generation necessarily finishes. A later interruption can leave no text or only partial text. Require a completion signal and useful output.
Can I automatically retry after a disconnect?
A retry may duplicate both output and usage. Decide based on whether the original request completed, whether output was already consumed and your application workflow.
Should I hide client cancellation errors?
No. Separate intentional client cancellation from upstream failure and server interruption so reliability metrics remain meaningful.
Continue your integration
Review prepaid pricing · Read API documentation · Get Started