Skip to main content

Microsoft AI STT plugin guide

How to use the Microsoft AI STT plugin for LiveKit Agents.

Available inPython

Overview

This plugin allows you to use Microsoft AI (MAI)  transcription models as an STT provider for your voice agents. The plugin streams audio to the Azure realtime transcription WebSocket API. Batch recognize() isn't supported.

Installation

Install the plugin from PyPI:

uv add "livekit-agents[microsoft-ai]~=1.8"

Authentication

The plugin requires the endpoint URL, API key, and model deployment name of your transcription resource. Set the following environment variables in your .env file:

MICROSOFT_AI_STT_URL=wss://<your-resource>.services.ai.azure.com/mai/v1/realtime?intent=transcription
MICROSOFT_AI_STT_API_KEY=<your-api-key>
MICROSOFT_AI_STT_MODEL=<your-deployment-name>
MICROSOFT_AI_STT_AUTH_HEADER=api-key

To authenticate with an Azure resource key, set MICROSOFT_AI_STT_AUTH_HEADER to api-key. If you don't set this variable, the plugin sends the key in an Authorization: Bearer header.

Usage

Use Microsoft AI STT in an AgentSession or as a standalone transcription service. For example, you can use this STT in the Voice AI quickstart.

from livekit.agents import AgentSession, inference
from livekit.plugins import microsoft_ai
vad = inference.VAD(model="silero")
session = AgentSession(
vad=vad,
stt=microsoft_ai.STT(vad=vad, language="en"),
# ... llm, tts, etc.
)
VAD requirement

The plugin requires its own VAD to detect the end of each utterance. The AgentSession VAD doesn't commit audio to the STT stream. Pass the same VAD to the STT and to the AgentSession.

Parameters

This section describes some of the available parameters. See the plugin reference for a complete list of all available parameters.

vad
Required
VAD

VAD that detects the end of each utterance and commits the audio for transcription. To commit audio manually with flush() or end_input(), pass None.

url
Required
stringEnv: MICROSOFT_AI_STT_URL

Full WebSocket endpoint URL, including the path and query string. The plugin doesn't add any path or query parameters.

model
Required
stringEnv: MICROSOFT_AI_STT_MODEL

Deployment name of the transcription model.

auth_headerstringDefault: AuthorizationEnv: MICROSOFT_AI_STT_AUTH_HEADER

Header that carries the API key. Authorization sends Bearer <key>. api-key sends the raw key in an api-key header.

languagestringEnv: MICROSOFT_AI_STT_LANGUAGE

Language code of the input audio. The model uses this value as a hint. To override it per stream, use stream(language=...).

Additional resources

The following resources provide more information about using Microsoft AI with LiveKit Agents.