axio-audio

Microphone capture and speaker playback for realtime agents.

class axio_audio.Microphone(sample_rate: int = 24000, chunk_ms: int = 50, device: int | str | None = None, queue_maxsize: int = 100)[source]

Async-iterable wrapper around a sounddevice RawInputStream.

Captures PCM16 mono audio at sample_rate and yields AudioBlock chunks of approximately chunk_ms milliseconds each. The defaults match the OpenAI Realtime API (24 kHz PCM16).

Usage:

async with Microphone() as mic:
    async for chunk in mic:
        await agent.send(chunk)
class axio_audio.Speaker(sample_rate: int = 24000, device: int | str | None = None, playback_tap: Callable[[bytes], None] | None = None)[source]

Async-friendly wrapper around a sounddevice RawOutputStream.

Holds an internal byte buffer; feed() appends, the audio callback drains as the device asks for samples. stop() clears the buffer (use it to honour user interruptions — drops everything still queued so the assistant goes silent immediately).

Usage:

async with Speaker() as spk:
    async for ev in agent.events():
        if isinstance(ev, AudioOutputDelta):
            await spk.feed(ev.data)
async feed(pcm: bytes) → None[source]

Append PCM16 mono bytes to the playback buffer.

pending_bytes() → int[source]

Bytes still waiting to be played — useful for back-pressure decisions.

playback_tap: Callable[[bytes], None] | None = None

Optional callback invoked from the audio thread with each chunk that is actually being played. Useful as the far-end reference for an echo canceller — the timing here matches what hits the speaker driver, not when the application called feed().

async stop() → None[source]

Drop everything queued for playback (user interrupted).

class axio_audio.DuplexAudio(sample_rate: int = 48000, chunk_ms: int = 20, channels: int = 1, device: int | str | tuple[int | str | None, int | str | None] | None = None, mono_io: bool = True, queue_maxsize: int = 100, playback_tap: Callable[[bytes], None] | None = None)[source]

One sd.RawStream, both directions, sample-aligned by PortAudio.

The callback runs on PortAudio’s audio thread for every block of chunk_ms worth of frames; it (a) pulls the next chunk of speaker bytes out of an internal buffer into outdata and (b) forwards indata to a thread-safe queue that the asyncio side drains via mic_chunks().

channels is how many channels we open the device with. On a stereo-only device (e.g. PipeWire’s Echo-Cancel Sink/Source, or most laptop default outputs) you must set channels=2 to drive both speakers correctly. The feed_speaker / mic_chunks API itself is mono — feed_speaker upmixes by duplicating samples across channels and mic_chunks downmixes by averaging. Set mono_io=False to skip the conversions and pass interleaved PCM through unchanged.

async feed_speaker(pcm: bytes) → None[source]

Append PCM16 to the speaker buffer.

Input is mono when mono_io is True (samples are duplicated across channels for the device), otherwise must be in the device’s exact channel layout.

mic_chunks() → AsyncIterator[AudioBlock][source]

Async-iterate captured mic data as AudioBlock chunks.

Chunk size is whatever PortAudio hands the callback (typically chunk_ms worth of frames). Yields mono PCM16 when mono_io is True, otherwise interleaved as opened.

playback_tap: Callable[[bytes], None] | None = None

Optional sync callback fired from the audio thread with the mono bytes we just handed to the speaker (silence-padded when the buffer underruns). Useful as a far-end reference for an external AEC, or for level meters that need real playback timing rather than enqueue timing. Receives the same bytes the mono feed_speaker API accepts, so it’s symmetric with mic_chunks.

speaker_pending_bytes() → int[source]

Mono-equivalent bytes still waiting to be played.

async stop_speaker() → None[source]

Drop everything queued for playback (use on user interruption).