Create a new agent in your browser using this model
Overview
Speechmatics speech-to-text is available in LiveKit Agents through LiveKit Inference and the Speechmatics plugin. With LiveKit Inference, your agent runs on LiveKit's infrastructure to minimize latency. No separate provider API key is required, and usage and rate limits are managed through LiveKit Cloud. Use the plugin instead if you want to manage your own billing and rate limits. Pricing for LiveKit Inference is available on the pricing page .
LiveKit Inference
Use LiveKit Inference to access Speechmatics STT without a separate Speechmatics API key.
A checkmark in the EU endpoint column means the model has a dedicated EU endpoint, in addition to the global endpoint. To learn more, see Data residency.
| Model name | Model ID | Languages | EU endpoint |
|---|---|---|---|
Speechmatics Linden-1 | speechmatics/linden-1 | arar_enbabebgbncacmncmn_encmn_en_ms_tacscydadeelenen_msen_taeoeseteufafifrgaglhehihrhuiaiditjakoltlvmnmrmsmtnlnoplptroruskslsvswtathtltrugukurviyue | EU endpoint |
Speechmatics Enhanced Deprecated | speechmatics/enhanced | arar_enbabebgbncacmncmn_encmn_en_ms_tacscydadeelenen_msen_taeoeseteufafifrgaglhehihrhuiaiditjakoltlvmnmrmsmtnlnoplptroruskslsvswtathtltrugukurviyue | EU endpoint |
Speechmatics Standard Deprecated | speechmatics/standard | arar_enbabebgbncacmncmn_encmn_en_ms_tacscydadeelenen_msen_taeoeseteufafifrgaglhehihrhuiaiditjakoltlvmnmrmsmtnlnoplptroruskslsvswtathtltrugukurviyue | EU endpoint |
Usage
To use Speechmatics, use the STT class from the inference module:
from livekit.agents import AgentSession, inferencesession = AgentSession(stt=inference.STT(model="speechmatics/linden-1",language="en"),# ... llm, tts, vad, turn_handling, etc.)
import { AgentSession, inference } from '@livekit/agents';const session = new AgentSession({stt: new inference.STT({model: "speechmatics/linden-1",language: "en"}),// ... llm, tts, vad, turnHandling, etc.});
Voice activity detection
The speechmatics/enhanced and speechmatics/standard models are deprecated. They are scheduled to be retired on October 5, 2026. Use the speechmatics/linden-1 model instead.
This section applies to LiveKit Inference. With the plugin, turn detection is set by turn_detection_mode, which defaults to EXTERNAL and needs a VAD running in your agent. See Turn detection.
Used through LiveKit Inference, speechmatics/linden-1 detects the end of speech server-side, so your agent doesn't need a VAD and the vad parameter is ignored.
The speechmatics/enhanced and speechmatics/standard models don't detect end of speech server-side, so LiveKit Inference relies on a VAD running in your agent. The SDK runs the VAD on incoming audio locally, and when the speaker stops, it signals the LiveKit Inference gateway to flush the final transcript.
For those models, the framework loads a VAD for you by default. To tune it, or to use a different VAD, pass one with the vad parameter:
from livekit.agents import AgentSession, inferencesession = AgentSession(stt=inference.STT(model="speechmatics/enhanced",# optional; the framework loads one if omittedvad=inference.VAD(min_silence_duration=0.4),),# ... llm, tts, etc.)
import { AgentSession, inference } from '@livekit/agents';const session = new AgentSession({stt: new inference.STT({model: "speechmatics/enhanced",// optional; the framework loads one if omittedvad: new inference.VAD({ minSilenceDuration: 0.4 }),}),// ... llm, tts, etc.});
Parameters
modelstringThe model to use for the STT. See model IDs for available models.
languageLanguageCodeLanguage code for the transcription. If not set, the provider default applies.
vadVADVoice activity detector used to detect end of speech. Set this parameter to override the default VAD.
This parameter is ignored when used with speechmatics/linden-1, which detects the end of speech server-side. See Voice activity detection.
extra_kwargsdictAdditional parameters to pass to the Speechmatics STT API. See model parameters for supported fields.
In Node.js this parameter is called modelOptions.
Model parameters
Pass the following parameters inside extra_kwargs (Python) or modelOptions (Node.js):
| Parameter | Type | Default | Notes |
|---|---|---|---|
domain | string | Domain-specific language pack for improved accuracy in a vertical, for example finance. | |
output_locale | string | BCP-47 locale that controls output formatting conventions, such as spelling and number formats. Locales are per language, for example en-GB, en-US, and en-AU for English. See Languages . A locale the language doesn't offer fails the session, and the error names the language it was rejected for. | |
max_delay | float | 1.0 | Maximum delay in seconds between the end of a spoken word and the final transcript. Valid range 0.7–4.0. Lower values reduce latency but can reduce accuracy. Not supported by speechmatics/linden-1. |
max_delay_mode | string | How max_delay is applied. flexible lets the model exceed the delay to finish recognizing entities such as numbers; fixed enforces the delay strictly. Not supported by speechmatics/linden-1. | |
diarization | string | none | Speaker diarization mode. speaker attributes each transcript segment to a speaker; none disables it. These are the only two values speechmatics/linden-1 accepts, and any other value returns an error. See Speaker diarization. |
speaker_sensitivity | float | Sensitivity of speaker detection, from 0.0 to 1.0. Higher values help distinguish speakers with similar voices. Applies when diarization is set to speaker. | |
max_speakers | int | Maximum number of speakers to detect when diarization is enabled. | |
prefer_current_speaker | bool | false | When diarization is enabled, reduces the likelihood of switching between similar-sounding speakers. |
enable_partials | bool | true | Emit interim transcript results before the final transcript. |
enable_entities | bool | Enable entity detection and formatting, such as numbers, dates, and currencies. See the Speechmatics entities docs . Entity formatting is always on for speechmatics/linden-1, which doesn't support this parameter. | |
punctuation_overrides | dict | Override default punctuation behavior, such as the set of permitted punctuation marks. | |
additional_vocab | list[dict] | Custom vocabulary to boost recognition of domain-specific words. Each entry is a dict with a content field and optional sounds_like pronunciation variants. See Keyterms. | |
end_of_utterance_silence_trigger | float | Duration of silence in seconds that triggers the end of an utterance and emits a final transcript. Not supported by speechmatics/linden-1. | |
audio_filtering_config | dict | Advanced audio filtering configuration passed through to Speechmatics, such as a volume threshold. Not supported by speechmatics/linden-1. | |
transcript_filtering_config | dict | Advanced transcript filtering configuration passed through to Speechmatics. Not supported by speechmatics/linden-1. |
For more details on these options, see the Agent STT API reference for speechmatics/linden-1, or the real-time API reference for the deprecated models.
Plugin
The Speechmatics plugin connects directly to the Speechmatics Agent STT API with your own API key.
Installation
Install the plugin from PyPI:
uv add "livekit-agents[speechmatics,silero]~=1.8"
Authentication
The Speechmatics plugin requires an API key .
Set SPEECHMATICS_API_KEY in your .env file.
Endpoint
The plugin connects to wss://eu2.rt.speechmatics.com/v2/agent by default. To use a different region or a self-hosted deployment, set the base_url parameter:
stt = speechmatics.STT(base_url="wss://global.rt.speechmatics.com/v2/agent",)
Speechmatics automatically routes each connection to its nearest regional server through the global endpoint. If you have data residency requirements, pin to a region instead:
| Region | Endpoint |
|---|---|
| All regions | wss://global.rt.speechmatics.com/v2/agent |
| Europe | wss://eu.rt.speechmatics.com/v2/agent |
| US | wss://us.rt.speechmatics.com/v2/agent |
| Australia | wss://au.rt.speechmatics.com/v2/agent |
For the full list, see Supported endpoints .
Usage
Use Speechmatics STT in an AgentSession or as a standalone transcription service. For example, you can use this STT in the Voice AI quickstart.
from livekit.plugins import speechmaticssession = AgentSession(stt=speechmatics.STT(),# ... llm, tts, etc.)
Turn detection
The turn_detection_mode parameter sets which component detects the end of a turn. It supports two modes:
| Mode | Detection method |
|---|---|
EXTERNAL (default) | A VAD running in your agent. Speechmatics doesn't detect the end of speech in this mode, so the VAD signals it. Works with LiveKit's turn detection. |
VAD | Speechmatics' own server-side voice activity detection. |
The default EXTERNAL mode needs no extra configuration, as shown in the Usage example.
In EXTERNAL mode, the plugin loads Silero VAD automatically when you don't pass a vad parameter. To use a different VAD, pass it as vad. To provide the end-of-turn signal yourself, pass vad=None and call finalize() from your own logic.
To have Speechmatics detect the end of a turn instead, set turn_detection_mode to VAD, and set turn_detection="stt" in the turn handling options. Without turn_detection="stt", the session ignores the end-of-speech events the plugin emits:
from livekit.agents import AgentSession, TurnHandlingOptionsfrom livekit.plugins import speechmaticssession = AgentSession(stt=speechmatics.STT(turn_detection_mode=speechmatics.TurnDetectionMode.VAD,),turn_handling=TurnHandlingOptions(turn_detection="stt",),# ... llm, tts, etc.)
Parameters
This section describes the key parameters for the Speechmatics STT plugin. See the plugin reference for a complete list of all available parameters.
api_keystringSpeechmatics API key. Defaults to the SPEECHMATICS_API_KEY environment variable. See Authentication.
base_urlstringDefault: wss://eu2.rt.speechmatics.com/v2/agentThe Speechmatics endpoint to connect to. See Endpoint.
modelModel | stringDefault: linden-1The transcription model. linden-1 is currently the only model the Speechmatics Agent STT API accepts. A name it doesn't accept falls back to linden-1 with a warning, rather than failing the session.
languageLanguageCodeDefault: enLanguage code for the input audio. All languages are global, meaning that regardless of which language you select, the system can recognize different dialects and accents. For the full list, see Supported Languages .
include_partialsboolDefault: falseInclude partial transcript segments in the output. Partial transcript segments arrive as AddPartialSegment messages and allow you to receive preliminary results that update as more context is available, until the higher-accuracy final transcript segment arrives as an AddSegment message. Partial transcript segments are returned faster but without any post-processing such as formatting. When enabled, the STT service emits INTERIM_TRANSCRIPT events.
enable_diarizationboolDefault: trueEnable speaker diarization. When enabled, each transcript segment is attributed to a speaker. You can use the speaker_sensitivity parameter to adjust the sensitivity of diarization. To learn more, see Speaker diarization.
turn_detection_modeTurnDetectionModeDefault: TurnDetectionMode.EXTERNALSets which component detects the end of a turn. See Turn detection for the available modes and examples.
vadVADVoice activity detector that signals the end of a turn. This parameter applies only to EXTERNAL turn detection mode. It's ignored in VAD mode, where Speechmatics closes turns itself.
In EXTERNAL mode, the plugin loads Silero VAD automatically when this parameter isn't set. Pass vad=None to opt out and call finalize() from your own logic. See Turn detection.
speaker_formatstringFormatter for speaker identification in transcription output. The following attributes are available:
{speaker_id}: The ID of the speaker.{text}: The text spoken by the speaker.
By default, if speaker diarization is enabled and this parameter is not set, the transcription output is not formatted for speaker identification.
The system instructions for the language model might need to include any necessary instructions to handle the formatting. To learn more, see Speaker diarization.
speaker_sensitivityfloatDefault: 0.5Sensitivity of speaker detection. Valid values are between 0 and 1. Higher values increase sensitivity and can help when two or more speakers have similar voices. To learn more, see Speaker sensitivity .
This parameter applies only when enable_diarization is True, which is the default.
prefer_current_speakerboolDefault: falseWhen speaker diarization is enabled and this is set to True, it reduces the likelihood of switching between similar sounding speakers. To learn more, see Prefer current speaker .
max_speakersintThe maximum number of speakers to detect in the session. Valid values are between 2 and 100. If not set, the number of speakers is unrestricted.
This parameter applies only when enable_diarization is True, which is the default.
known_speakerslist[SpeakerIdentifier]Speakers identified in an earlier session, so the same labels are applied to them again. Retrieve them from a running session with get_speaker_ids(). To learn more, see Reusing speaker labels across sessions.
This parameter applies only when enable_diarization is True, which is the default.
additional_vocablist[AdditionalVocabEntry]Custom vocabulary that biases recognition towards domain-specific words. Each entry has a content field and optional sounds_like pronunciation variants. See Keyterms.
domainstringDomain-specific language pack for improved accuracy in a vertical, for example finance.
output_localestringBCP-47 locale that controls output formatting conventions, such as spelling and number formats. For example, en-GB.
Keyterms
Speechmatics supports keyterms through the additional_vocab parameter, on all models including speechmatics/linden-1. Speechmatics calls this feature the custom dictionary . You can only set keyterms when the session starts. Speechmatics doesn't accept changes to additional_vocab mid-session.
With the plugin, pass it to the STT constructor. Each entry takes a content value and optional sounds_like pronunciation variants, for when the spelling doesn't match how the word is said:
from livekit.plugins import speechmaticsfrom livekit.plugins.speechmatics import AdditionalVocabEntrystt = speechmatics.STT(additional_vocab=[AdditionalVocabEntry(content="Speechmatics"),AdditionalVocabEntry(content="LiveKit", sounds_like=["live kit"]),],)
With LiveKit Inference, pass the same entries as plain dicts inside extra_kwargs:
from livekit.agents import inferencestt = inference.STT(model="speechmatics/linden-1",extra_kwargs={"additional_vocab": [{"content": "Speechmatics"},{"content": "LiveKit", "sounds_like": ["live kit"]},],},)
To manage keyterms across providers, or detect them automatically from the conversation, see Keyterm management.
Speaker diarization
Speaker diarization attributes speech to individual speakers. When enabled, the STT reports capabilities.diarization = True, and the transcription output can include speaker information through the speaker_id and text attributes.
Agent STT labels one speaker per transcript segment, not per word. A segment is the unit the service emits: AddSegment for a final transcript and AddPartialSegment for an interim one. Each segment has at most one speaker, so overlapping speech from two people is never labeled within the same segment. The plugin surfaces each segment as a LiveKit transcript event, which means the speaker label changes between events, never inside one.
Speakers are labeled in the order they are first heard: S1, S2, and so on. Labels hold for the session but aren't stable across sessions unless you carry them over with known_speakers.
With diarization enabled, wrap the Speechmatics STT with MultiSpeakerAdapter to detect the primary speaker and format transcripts by speaker.
Enable diarization in the STT constructor. LiveKit Inference uses a string diarization mode, while the plugin uses the boolean enable_diarization:
stt = inference.STT(model="speechmatics/linden-1",extra_kwargs={# "none" or "speaker""diarization": "speaker",},)
stt = speechmatics.STT(enable_diarization=True,speaker_format="<{speaker_id}>{text}</{speaker_id}>",)
With the plugin, use speaker_format to format speaker output. For example:
<{speaker_id}>{text}</{speaker_id}>produces<S1>Hello</S1>.[Speaker {speaker_id}] {text}produces[Speaker S1] Hello.
Segments the service can't attribute to a speaker are labeled UU, so a format string renders them as <UU>Hello</UU>. With diarization disabled, every segment is labeled UU.
Include the format in your agent instructions so the language model reads the labels as speaker attribution rather than as part of what was said. Two settings help when speakers are hard to separate: raise speaker_sensitivity to split similar-sounding voices more aggressively, and set prefer_current_speaker to reduce switching mid-turn. To learn more, see Speaker diarization .
Reusing speaker labels across sessions
Speaker labels are assigned per session, so the same person is usually S1 in one session and something else in the next. To keep labels stable, call get_speaker_ids() on the STT once each speaker has said at least five words, then pass the result as known_speakers when you create the STT for the next session:
speakers = await stt.get_speaker_ids()# ... later, in a new sessionstt = speechmatics.STT(enable_diarization=True,known_speakers=speakers,)
get_speaker_ids() returns the speakers for the STT's open stream. If the STT has more than one stream open, it returns a list of results, one per stream.
Additional resources
The following resources provide more information about using Speechmatics with LiveKit Agents.
Python package
The livekit-plugins-speechmatics package on PyPI.
Plugin reference
Reference for the Speechmatics STT plugin.
GitHub repo
View the source or contribute to the LiveKit Speechmatics STT plugin.
Speechmatics plugin guide
Speechmatics' own guide to this plugin.
Agent STT API reference
The WebSocket API the plugin speaks.
Voice AI quickstart
Get started with LiveKit Agents and Speechmatics STT.