axio-transport-openai

OpenAI and OpenAI-compatible completion transports.

api selects the endpoint. OpenAITransport defaults to "responses" and posts to /v1/responses, reading it through axio-responses. NebiusTransport, OpenRouterTransport and OpenAICompatibleTransport set "chat" and post to /v1/chat/completions. That endpoint refuses function tools beside any reasoning effort other than "none" — see the troubleshooting entry.

extra_params is folded into the request. Its tools are merged with the agent’s function declarations rather than replacing them, so adding a hosted tool does not take away the calls the agent is there to dispatch.

class axio_transport_openai.OpenAITransport(name: 'str' = 'OpenAI', base_url: 'str' = <factory>, api_key: 'str' = <factory>, model: 'ModelSpec' = <factory>, models: 'ModelRegistry' = <factory>, session: 'aiohttp.ClientSession | None' = None, max_retries: 'int' = 10, retry_base_delay: 'float' = 5.0, extra_params: 'Mapping[str, Any]'=mappingproxy({}), *, api: "Literal['responses', 'chat'] | None"=None, reasoning_summary: "Literal['auto', 'concise', 'detailed']"='auto')[source]
api: Literal['responses', 'chat'] | None

Which endpoint this server speaks. "responses" is the one that takes function tools and reasoning together, and OpenAI recommends it for new work. Left unset it follows base_url: only OpenAI’s own host is assumed to implement it.

build_responses_payload(messages: list[Message], tools: list[Tool[Any]], system: str) → dict[str, Any][source]

The request /v1/responses takes.

The system prompt goes in instructions rather than in a message. Tool calls and their outputs are items beside the messages rather than blocks inside them. This endpoint takes tools and reasoning together, which is the reason to prefer it. /v1/chat/completions refuses that pair outright for a model that reasons.

async embed(texts: list[str]) → list[list[float]][source]

Call the OpenAI-compatible /v1/embeddings endpoint.

reasoning_summary: Literal['auto', 'concise', 'detailed']

How much of its reasoning the model summarises for a reader on /v1/responses. The raw chain is never returned; without a summary there is nothing to show.

class axio_transport_openai.nebius.NebiusTransport(name: 'str' = 'Nebius AI Studio', base_url: 'str' = 'https://api.tokenfactory.nebius.com/v1', api_key: 'str' = <factory>, model: 'ModelSpec' = <factory>, models: 'ModelRegistry' = <factory>, session: 'aiohttp.ClientSession | None' = None, max_retries: 'int' = 10, retry_base_delay: 'float' = 5.0, extra_params: 'Mapping[str, Any]'=mappingproxy({}), thinking: 'bool' = False, *, api: "Literal['responses', 'chat']"='chat', reasoning_summary: "Literal['auto', 'concise', 'detailed']"='auto')[source]
async fetch_models() → None[source]

Fetch available models from Nebius /v1/models?verbose=true.

class axio_transport_openai.openrouter.OpenRouterTransport(name: 'str' = 'OpenRouter', base_url: 'str' = 'https://openrouter.ai/api/v1', api_key: 'str' = <factory>, model: 'ModelSpec' = ModelSpec(id='google/gemini-2.5-pro-preview', capabilities=frozenset(), max_output_tokens=8192, context_window=128000, input_cost=0.0, output_cost=0.0), models: 'ModelRegistry' = <factory>, session: 'aiohttp.ClientSession | None' = None, max_retries: 'int' = 10, retry_base_delay: 'float' = 5.0, extra_params: 'Mapping[str, Any]'=mappingproxy({}), thinking: 'bool' = False, *, api: "Literal['responses', 'chat']"='chat', reasoning_summary: "Literal['auto', 'concise', 'detailed']"='auto')[source]
async fetch_models() → None[source]

Fetch available models from OpenRouter /v1/models.

class axio_transport_openai.custom.OpenAICompatibleTransport(name: str = 'OpenAI', base_url: str = '', api_key: str = <factory>, model: ModelSpec = <factory>, models: ModelRegistry = <factory>, session: aiohttp.ClientSession | None = None, max_retries: int = 10, retry_base_delay: float = 5.0, extra_params: Mapping[str, Any]=mappingproxy({}), *, api: Literal['responses', 'chat']='chat', reasoning_summary: Literal['auto', 'concise', 'detailed']='auto')[source]

OpenAI-compatible transport for a single user-defined provider.

Instances are created by CustomHubScreen with name, base_url, api_key, and models populated from the JSON config. Supports JSON round-trip via to_dict() / from_dict().

api: Literal['responses', 'chat']

Which endpoint this server speaks. "responses" is the one that takes function tools and reasoning together, and OpenAI recommends it for new work. Left unset it follows base_url: only OpenAI’s own host is assumed to implement it.

classmethod from_dict(data: dict[str, Any], *, session: ClientSession | None = None) → Self[source]

As OpenAITransport.from_dict(), but base_url/api_key come verbatim from data.

The base implementation reads a value saved empty as empty, and falls back to the OPENAI_BASE_URL/OPENAI_API_KEY env vars only where the key is absent altogether. That is right for the built-in OpenAI provider, whose settings dict is written by hand and omits what it wants the default for.

A custom provider’s config is a full round-trip of to_dict(), so an absent key means the same thing as an empty one: this server takes no credential. Reading it through the environment would point a local server at an unrelated real endpoint.

reasoning_summary: Literal['auto', 'concise', 'detailed']

How much of its reasoning the model summarises for a reader on /v1/responses. The raw chain is never returned; without a summary there is nothing to show.

Realtime

class axio_transport_openai.OpenAIRealtimeTransport(api_key: str = <factory>, base_url: str = <factory>, model: str = 'gpt-realtime', http_session: ClientSession | None = None, vad_threshold: float = 0.5, vad_prefix_padding_ms: int = 300, vad_silence_duration_ms: int = 500, vad_create_response: bool = True, vad_interrupt_response: bool = True, input_noise_reduction: str | None = 'near_field')[source]

RealtimeTransport for OpenAI’s Realtime WebSocket API.

Server-VAD knobs live here so callers can tune barge-in sensitivity without touching the transport internals. When AEC is imperfect or the room is loud, the defaults can cause the model’s own audio to trip interrupt_response and cancel its response. Bump vad_threshold higher, lengthen vad_silence_duration_ms, or set vad_interrupt_response=False to require an explicit agent.interrupt().

input_noise_reduction: str | None

Server-side noise reduction profile applied to mic audio. Options: "near_field" (close-talking mic / headset), "far_field" (laptop or conference mic), or None to disable.

vad_create_response: bool

Whether the server auto-creates a response after each user turn.

vad_interrupt_response: bool

Whether incoming user speech cancels the model’s in-flight response. Setting this to False makes interruption explicit (call agent.interrupt()) — useful when AEC isn’t strong enough and the model’s own voice would otherwise loop back through the mic and interrupt itself.

vad_prefix_padding_ms: int

Audio retained before a detected speech start.

vad_silence_duration_ms: int

Silence required before the server considers a turn complete.

vad_threshold: float

Server VAD activation threshold (0..1). Higher → less sensitive to quiet sounds and bleed-through. Default 0.5.

class axio_transport_openai.OpenAIRealtimeSession(ws: Any, output_audio_media_type: str = 'audio/pcm;rate=24000', http_session: ClientSession | None = None, own_http_session: bool = False)[source]

RealtimeSession backed by an aiohttp WebSocket.

ws may be any object that quacks like aiohttp.ClientWebSocketResponse (send_str, async iteration yielding aiohttp.WSMessage-like objects, close, closed). Tests inject a stub.

async close() → None[source]

Tear down the session and release resources.

async commit() → None[source]

Signal end-of-utterance for manual VAD; no-op with server VAD.

async events() → AsyncIterator[StreamEvent][source]

Async iterator over server events for the lifetime of this session.

async interrupt() → None[source]

Abort in-flight assistant generation (e.g. user interrupted).

async send(content: ContentBlock | list[ContentBlock]) → None[source]

Append user content (audio chunk, text, image) to the input buffer.

async send_tool_result(tool_use_id: ToolCallID, name: ToolName, content: str | list[ContentBlock]) → None[source]

Deliver a tool’s result to the provider so generation can resume.

name is included because some providers (e.g. Gemini Live) require the tool name alongside the call id. OpenAI realtime can ignore it.