Overview
This plugin allows you to use Microsoft AI (MAI) transcription models as an STT provider for your voice agents. The plugin streams audio to the Azure realtime transcription WebSocket API. Batch recognize() isn't supported.
Installation
Install the plugin from PyPI:
uv add "livekit-agents[microsoft-ai]~=1.8"
Authentication
The plugin requires the endpoint URL, API key, and model deployment name of your transcription resource. Set the following environment variables in your .env file:
MICROSOFT_AI_STT_URL=wss://<your-resource>.services.ai.azure.com/mai/v1/realtime?intent=transcriptionMICROSOFT_AI_STT_API_KEY=<your-api-key>MICROSOFT_AI_STT_MODEL=<your-deployment-name>MICROSOFT_AI_STT_AUTH_HEADER=api-key
To authenticate with an Azure resource key, set MICROSOFT_AI_STT_AUTH_HEADER to api-key. If you don't set this variable, the plugin sends the key in an Authorization: Bearer header.
Usage
Use Microsoft AI STT in an AgentSession or as a standalone transcription service. For example, you can use this STT in the Voice AI quickstart.
from livekit.agents import AgentSession, inferencefrom livekit.plugins import microsoft_aivad = inference.VAD(model="silero")session = AgentSession(vad=vad,stt=microsoft_ai.STT(vad=vad, language="en"),# ... llm, tts, etc.)
The plugin requires its own VAD to detect the end of each utterance. The AgentSession VAD doesn't commit audio to the STT stream. Pass the same VAD to the STT and to the AgentSession.
Parameters
This section describes some of the available parameters. See the plugin reference for a complete list of all available parameters.
vadVADVAD that detects the end of each utterance and commits the audio for transcription. To commit audio manually with flush() or end_input(), pass None.
urlstringEnv: MICROSOFT_AI_STT_URLFull WebSocket endpoint URL, including the path and query string. The plugin doesn't add any path or query parameters.
modelstringEnv: MICROSOFT_AI_STT_MODELDeployment name of the transcription model.
auth_headerstringDefault: AuthorizationEnv: MICROSOFT_AI_STT_AUTH_HEADERHeader that carries the API key. Authorization sends Bearer <key>. api-key sends the raw key in an api-key header.
languagestringEnv: MICROSOFT_AI_STT_LANGUAGELanguage code of the input audio. The model uses this value as a hint. To override it per stream, use stream(language=...).
Additional resources
The following resources provide more information about using Microsoft AI with LiveKit Agents.