axio-audio¶
Microphone capture and speaker playback for realtime agents.
- class axio_audio.Microphone(sample_rate: int = 24000, chunk_ms: int = 50, device: int | str | None = None, queue_maxsize: int = 100)[source]¶
Async-iterable wrapper around a sounddevice
RawInputStream.Captures PCM16 mono audio at
sample_rateand yieldsAudioBlockchunks of approximatelychunk_msmilliseconds each. The defaults match the OpenAI Realtime API (24 kHz PCM16).Usage:
async with Microphone() as mic: async for chunk in mic: await agent.send(chunk)
- class axio_audio.Speaker(sample_rate: int = 24000, device: int | str | None = None, playback_tap: Callable[[bytes], None] | None = None)[source]¶
Async-friendly wrapper around a sounddevice
RawOutputStream.Holds an internal byte buffer;
feed()appends, the audio callback drains as the device asks for samples.stop()clears the buffer (use it to honour user interruptions — drops everything still queued so the assistant goes silent immediately).Usage:
async with Speaker() as spk: async for ev in agent.events(): if isinstance(ev, AudioOutputDelta): await spk.feed(ev.data)
- class axio_audio.DuplexAudio(sample_rate: int = 48000, chunk_ms: int = 20, channels: int = 1, device: int | str | tuple[int | str | None, int | str | None] | None = None, mono_io: bool = True, queue_maxsize: int = 100, playback_tap: Callable[[bytes], None] | None = None)[source]¶
One sd.RawStream, both directions, sample-aligned by PortAudio.
The callback runs on PortAudio’s audio thread for every block of
chunk_msworth of frames; it (a) pulls the next chunk of speaker bytes out of an internal buffer intooutdataand (b) forwardsindatato a thread-safe queue that the asyncio side drains viamic_chunks().channelsis how many channels we open the device with. On a stereo-only device (e.g. PipeWire’sEcho-Cancel Sink/Source, or most laptop default outputs) you must setchannels=2to drive both speakers correctly. Thefeed_speaker/mic_chunksAPI itself is mono —feed_speakerupmixes by duplicating samples across channels andmic_chunksdownmixes by averaging. Setmono_io=Falseto skip the conversions and pass interleaved PCM through unchanged.- async feed_speaker(pcm: bytes) None[source]¶
Append PCM16 to the speaker buffer.
Input is mono when
mono_iois True (samples are duplicated across channels for the device), otherwise must be in the device’s exact channel layout.
- mic_chunks() AsyncIterator[AudioBlock][source]¶
Async-iterate captured mic data as
AudioBlockchunks.Chunk size is whatever PortAudio hands the callback (typically
chunk_msworth of frames). Yields mono PCM16 whenmono_iois True, otherwise interleaved as opened.
- playback_tap: Callable[[bytes], None] | None = None¶
Optional sync callback fired from the audio thread with the mono bytes we just handed to the speaker (silence-padded when the buffer underruns). Useful as a far-end reference for an external AEC, or for level meters that need real playback timing rather than enqueue timing. Receives the same bytes the mono
feed_speakerAPI accepts, so it’s symmetric withmic_chunks.