axio-transport-google

Google Gemini completion transports for the Developer API and Vertex AI.

Gemini’s stream carries no per-event discriminator, so this transport reads it with axio_sse.payloads() and Wire payload shapes rather than with a Reader. Anything it has no typed event for — grounding and citation metadata, executable code, parts added later — is emitted as ProviderEvent(provider="google").

Its usage counts are converted into the axio rule. toolUsePromptTokenCount is not inside promptTokenCount, and thoughtsTokenCount is not inside candidatesTokenCount, so both are added. cachedContentTokenCount already is inside its total, and is not added.

class axio_transport_google.GoogleTransport(name: 'str' = 'Google GenAI', api_key: 'str' = '', vertexai: 'bool | None' = None, project: 'str' = '', location: 'str' = '', model: 'ModelSpec' = <factory>, models: 'ModelRegistry' = <factory>, session: 'aiohttp.ClientSession | None' = None, max_retries: 'int' = 5, retry_base_delay: 'float' = 5.0, temperature: 'float | None' = None, top_p: 'float | None' = None, top_k: 'float | None' = None, seed: 'int | None' = None, safety_settings: 'list[SafetySettingDict] | None' = None, debug: 'bool' = False, nudge_on_media_tool_result: 'bool' = True, max_output_tokens: 'int | None' = None, thinking_budget: 'int | None' = None, thinking_level: 'str | None' = None, service_tier: 'str | None' = None, media_resolution: 'str | None' = None, _thought_signatures: 'dict[str, str]'=<factory>, _streams: 'int' = 0, last_usage: 'Usage | None' = None, _credentials: 'Any' = None)[source]
async fetch_models() → None[source]

Fetch available Gemini models.

Developer API: GET /v1beta/models?key=… Vertex AI: GET /v1beta1/publishers/google/models (no project prefix)

async generate_images(prompt: str, *, model: str | None = None, n: int = 1) → list[bytes][source]

Generate images via Gemini Nano Banana (generateContent with IMAGE response modality).

async generate_videos(prompt: str, *, model: str | None = None, n: int = 1, image: bytes | None = None, duration_seconds: int | None = None, aspect_ratio: str | None = None) → list[bytes][source]

Generate videos using Veo models. Polls until the operation completes.

get_thinking_options() → tuple[str, ...] | None[source]

Valid thinkingLevel values for the current model, or None if budget-based (2.5).

retry_base_delay: float

Seconds before the first retry, doubling after that. The other transports name it too.

class axio_transport_google.VertexAITransport(name: str = 'Google Vertex AI', api_key: str = '', vertexai: bool | None = True, project: str = '', location: str = '', model: ModelSpec = <factory>, models: ModelRegistry = <factory>, session: ClientSession | None = None, max_retries: int = 5, retry_base_delay: float = 5.0, temperature: float | None = None, top_p: float | None = None, top_k: float | None = None, seed: int | None = None, safety_settings: list[SafetySetting] | None = None, debug: bool = False, nudge_on_media_tool_result: bool = True, max_output_tokens: int | None = None, thinking_budget: int | None = None, thinking_level: str | None = None, service_tier: str | None = None, media_resolution: str | None = None, _thought_signatures: dict[str, str]=<factory>, _streams: int = 0, last_usage: Usage | None = None, _credentials: Any = None)[source]

GoogleTransport pre-configured for Vertex AI (includes Anthropic models).

Realtime

class axio_transport_google.realtime.GeminiLiveTransport(api_key: str = <factory>, base_url: str | None = None, model: str | None = None, http_session: ClientSession | None = None, vertexai: bool = <factory>, project: str | None = <factory>, location: str | None = <factory>, language_code: str | None = None, auto_region: bool = False)[source]

RealtimeTransport for Gemini Live.

Supports both backends:

  • AI Studio / developer API (default) — auth via GEMINI_API_KEY query param, model id like "models/gemini-live-2.5-flash-native-audio".

  • Vertex AI — set vertexai=True (or env GOOGLE_GENAI_USE_VERTEXAI=1). Requires project and location. Auth uses google.auth.default (gcloud user creds, service account, GCE metadata, etc.). The full model resource path is built automatically.

language_code is forwarded to generationConfig.speechConfig.languageCode so the model picks the right TTS voice for the user’s locale.

auto_region: bool

When True (Vertex only), probe every supported Live region at connect time and pick the lowest-latency one for THIS network — overrides location. Adds ~1 s of startup latency. The result is cached for the lifetime of the transport instance.

language_code: str | None

BCP-47 code passed to generationConfig.speechConfig.languageCode. NOTE: this only steers the TTS voice the server uses for the audio output — it does not tell the model which language to think / write in. Most callers will also want to inject a language instruction into the system prompt (chat.py does this automatically).

class axio_transport_google.realtime.VertexLiveTransport(api_key: str = <factory>, base_url: str | None = None, model: str | None = None, http_session: ClientSession | None = None, vertexai: bool = True, project: str | None = <factory>, location: str | None = <factory>, language_code: str | None = None, auto_region: bool = False)[source]

GeminiLiveTransport pre-configured for Vertex AI (mirrors VertexAITransport).

class axio_transport_google.realtime.GeminiLiveSession(ws: Any, output_audio_media_type: str = 'audio/pcm;rate=24000', http_session: ClientSession | None = None, own_http_session: bool = False)[source]

RealtimeSession backed by an aiohttp WebSocket against Gemini Live.

ws is any object that quacks like aiohttp.ClientWebSocketResponse; tests inject a stub.

async close() → None[source]

Tear down the session and release resources.

async commit() → None[source]

Signal end-of-utterance for manual VAD; no-op with server VAD.

async events() → AsyncIterator[StreamEvent][source]

Async iterator over server events for the lifetime of this session.

async interrupt() → None[source]

Abort in-flight assistant generation (e.g. user interrupted).

async send(content: ContentBlock | list[ContentBlock]) → None[source]

Append user content (audio chunk, text, image) to the input buffer.

async send_tool_result(tool_use_id: ToolCallID, name: ToolName, content: str | list[ContentBlock]) → None[source]

Deliver a tool’s result to the provider so generation can resume.

name is included because some providers (e.g. Gemini Live) require the tool name alongside the call id. OpenAI realtime can ignore it.

Media tools

async axio_transport_google.tools.generate_image(prompt: ~typing.Annotated[str, ~axio.field.FieldInfo(description=, default=MISSING, ge=None, le=None, strict=True)], model: ~typing.Annotated[str, ~axio.field.FieldInfo(description=, default=MISSING, ge=None, le=None, strict=True)] = 'gemini-3-pro-image-preview', n: int = 1) → list[TextBlock | ImageBlock][source]

Generate images from a text prompt using Google Gemini image models (Nano Banana family). Only use when the user explicitly asks to generate an image — never for screenshots, UI testing, or verifying application output. Use a descriptive, detailed prompt for best results.

async axio_transport_google.tools.generate_video(prompt: ~typing.Annotated[str, ~axio.field.FieldInfo(description=, default=MISSING, ge=None, le=None, strict=True)], model: ~typing.Annotated[str, ~axio.field.FieldInfo(description=, default=MISSING, ge=None, le=None, strict=True)] = 'veo-3.1-fast-generate-001', duration_seconds: int = 6, aspect_ratio: ~typing.Annotated[str, ~axio.field.FieldInfo(description=, default=MISSING, ge=None, le=None, strict=True)] = '16:9') → list[TextBlock | VideoBlock][source]

Generate a video from a text prompt using Google Veo models. Only use when the user explicitly asks to generate a video — never for screen recordings, UI testing, or verifying application output. This is an async operation that may take 1-3 minutes.

Thinking levels

axio_transport_google.valid_thinking_levels(model_id: str) → tuple[str, ...] | None[source]

Return valid thinkingLevel values for a Gemini 3+ model, or None for budget-based (2.5) models.

A thinking_level the model does not accept is replaced by that family’s highest level, so an invented value buys maximum thinking rather than none.