Overview
Speaker diarization identifies which speaker is talking at each point in a single audio track. A diarization-capable STT groups speech by voice and labels each transcript with a speaker identifier, such as S1 or A.
In apps like video conferencing, each person joins from their own device and publishes their own audio track. You already know who's speaking from the track. Diarization matters when several people share one microphone. This is common in physical AI, where a robot or device hears everyone in the room through its own microphones. Examples include the following:
- A humanoid or service robot that people talk to in a shared space.
- An in-store kiosk, drive-thru, or reception device with bystanders nearby.
- An in-car, smart home, or conference room assistant shared by a group.
With diarization, your agent can tell these voices apart, address the right person, and ignore background chatter.
A speaker label tells you that two utterances came from the same voice, not who that person is. The provider assigns labels in the order it first hears each voice, and they apply only within the current session. S1 in one session has no relation to S1 in another.
Enable diarization
Diarization is off by default for most providers. To turn it on, set the provider's diarization option. Option names vary by provider. For each provider's option, see Supported providers.
With LiveKit Inference, pass the option in extra_kwargs (modelOptions in Node.js). The following example enables diarization for Deepgram:
from livekit.agents import AgentSession, inferencesession = AgentSession(stt=inference.STT(model="deepgram/nova-3",language="en",extra_kwargs={"diarize": True},),# ... llm, tts, etc.)
import { AgentSession, inference } from '@livekit/agents';const session = new AgentSession({stt: new inference.STT({model: 'deepgram/nova-3',language: 'en',modelOptions: { diarize: true },}),// ... llm, tts, etc.});
With a plugin, pass the diarization parameter to the STT constructor instead. For example, deepgram.STT(enable_diarization=True).
When diarization is on, the STT reports capabilities.diarization as True.
Read speaker labels
The speaker label appears as speaker_id (speakerId in Node.js) on each transcript's SpeechData, on individual words when the provider labels them, and on the user_input_transcribed session event. It's empty when the provider can't attribute an utterance. This is common for interim transcripts and very short utterances.
The following example logs each final transcript with its speaker:
@session.on("user_input_transcribed")def on_user_input_transcribed(event):if event.is_final:print(f"[{event.speaker_id}] {event.transcript}")
import { voice } from '@livekit/agents';session.on(voice.AgentSessionEventTypes.UserInputTranscribed, (event) => {if (event.isFinal) {console.log(`[${event.speakerId}] ${event.transcript}`);}});
Speaker labels don't reach the LLM on their own. To act on them in the conversation, use primary speaker detection or format the transcript yourself in the STT node.
Primary speaker detection
Most agents talk to one person at a time, even when others are audible. MultiSpeakerAdapter wraps a diarization-enabled STT and identifies the primary speaker based on audio level: the voice that's consistently loudest, usually the person closest to the microphone. A different speaker becomes primary only when they're clearly louder, or after the current primary speaker has been silent for about a minute.
You can label background speech so the LLM knows who said what, or suppress it so only the primary speaker reaches the LLM. The following example labels background speech:
from livekit.agents import AgentSession, inference, sttsession = AgentSession(stt=stt.MultiSpeakerAdapter(stt=inference.STT(model="deepgram/nova-3",language="en",extra_kwargs={"diarize": True},),background_format="[Background speaker {speaker_id}] {text}",),# ... llm, tts, etc.)
When you label background speech, explain the labels in your agent's instructions. For example: "Messages that start with [Background speaker] come from other people nearby. Don't respond to them unless the user refers to them."
MultiSpeakerAdapter accepts the following parameters:
sttSTTThe STT to wrap. Diarization must be enabled on it, or the adapter raises a ValueError.
detect_primary_speakerboolDefault: TrueWhether to detect the primary speaker. When False, transcripts pass through unchanged.
suppress_background_speakerboolDefault: FalseWhether to drop transcripts from speakers other than the primary speaker.
primary_formatstrDefault: {text}Format for the primary speaker's transcripts. Supports the {text} and {speaker_id} placeholders.
background_formatstrDefault: {text}Format for other speakers' transcripts. Supports the {text} and {speaker_id} placeholders.
primary_detection_optionsPrimarySpeakerDetectionOptionsTuning for primary speaker detection, such as how much louder a new speaker must be to take over. The defaults suit most environments. To learn more, see the Python reference.
Supported providers
The following STT providers support diarization in LiveKit Agents. For more diarization options, such as a maximum speaker count, see each provider's page.
LiveKit Inference
Speaker labels from LiveKit Inference are available in Python and Node.js:
| Provider | Models | Option |
|---|---|---|
| AssemblyAI | All models | speaker_labels: True |
| Deepgram | Nova-3 and Nova-2 | diarize: True |
| Speechmatics | All models | diarization: "speaker" |
| SpaceXAI | xai/stt-1 | diarize: True |
Plugins
Pass the parameter to the STT constructor. Node.js parameters use camelCase:
| Plugin | Parameter | Speaker labels |
|---|---|---|
| AssemblyAI | speaker_labels=True | Python, Node.js |
| Deepgram | enable_diarization=True | Python |
| NVIDIA Riva | enable_diarization=True | Python |
| Smallest AI | diarize=True | Python |
| Soniox | enable_speaker_diarization=True | Python, Node.js |
| SpaceXAI | enable_diarization=True | Python |
| Speechmatics | enable_diarization=True (default) | Python |
To get speaker labels from Deepgram or SpaceXAI in Node.js, use LiveKit Inference.
Limitations
Keep the following limitations in mind:
- Accuracy: Similar voices, overlapping speech, and noisy audio can merge or split speakers.
- Streaming only: Diarization requires a streaming STT. It isn't available through
StreamAdapter. - Fallback: A
FallbackAdaptersupports diarization only if every STT in it has diarization enabled. - Interruptions: Suppressing background speakers filters transcripts, not audio. Background voices can still interrupt the agent.