From ebe650c2682f554d1352b2c1f31cd3dd04976452 Mon Sep 17 00:00:00 2001 From: karamouche Date: Mon, 28 Sep 2026 15:22:37 -0400 Subject: [PATCH 1/3] docs(speaker-diarization): optimize page for GEO Clarify pre-recorded scope, fix number_of_speakers naming, and document speaker-count hints as non-guarantees with evaluation links. --- .../speaker-diarization.mdx | 24 ++++++++++++------- 1 file changed, 15 insertions(+), 9 deletions(-) diff --git a/chapters/audio-intelligence/speaker-diarization.mdx b/chapters/audio-intelligence/speaker-diarization.mdx index 6e1a302..e83bec7 100644 --- a/chapters/audio-intelligence/speaker-diarization.mdx +++ b/chapters/audio-intelligence/speaker-diarization.mdx @@ -1,7 +1,7 @@ --- -title: 'Speaker Diarization' +title: 'Speaker diarization API: configuration and output' sidebarTitle: Diarization -description: 'Detect speakers and understand who said what.' +description: 'Enable speaker diarization for pre-recorded audio in Gladia. Configure speaker-count hints and read speaker labels, timestamps and limitations.' --- import PrerecordedBadge from "/snippets/badges/prerecorded.mdx" @@ -9,7 +9,7 @@ import PrerecordedBadge from "/snippets/badges/prerecorded.mdx" -Speaker diarization is the process of detecting multiple speakers in an audio, and understanding which parts of the transcription each speaker said. +Speaker diarization assigns speaker labels to segments of pre-recorded audio so you can identify who spoke when. In Gladia, enable `diarization` on the pre-recorded transcription request. The response associates each utterance with a speaker index in order of first appearance. These labels distinguish speakers within the recording and do not establish a person's identity. ## Enabling diarization @@ -49,12 +49,18 @@ Speakers will be assigned indexes by **order of appearance** (i.e. the 1st speak ## Improving diarization accuracy -You can improve the accuracy of the diarization by providing the model with hints regarding the expected number or lower/upper bounds ofspeakers using the `diarization_config.num_of_speakers`, `diarization_config.min_speakers` and `diarization_config.max_speakers` parameters respectively. - -**Important:** These parameters are hints, not hard constraints. The actual number of speakers detected by the model may not comply with the provided parameters. +Provide speaker-count hints with `diarization_config.number_of_speakers`, `diarization_config.min_speakers` and `diarization_config.max_speakers`. These specify the expected count, lower hint and upper hint respectively. They are hints, not hard constraints; the detected count may differ. | Key | Type | Description | | --- | --- | --- | -| `diarization_config.number_of_speakers` | number | Guiding number of speakers - instructs the model to detect an exact number of speakers in the audio. | -| `diarization_config.min_speakers` | number | Instructs the model to detect no less than this number of speakers in the audio. | -| `diarization_config.max_speakers` | number | Causes the model to detect no more than this number of speakers in the audio. | +| `diarization_config.number_of_speakers` | number | Expected speaker-count hint. It does not guarantee that the detected count matches this value. | +| `diarization_config.min_speakers` | number | Lower speaker-count hint, not an enforced minimum. | +| `diarization_config.max_speakers` | number | Upper speaker-count hint, not an enforced maximum. | + +## Diarization scope and evaluation + +This guide covers speaker diarization for **pre-recorded audio** only. + +Speaker diarization, transcription accuracy, and channel identification are different capabilities. For concept definitions, see [speaker diarization concepts](https://www.gladia.io/blog/what-is-diarization). + +Async accuracy comparisons use the model and dataset scope described in the [async benchmark methodology](https://www.gladia.io/competitors/benchmarks). Use the [blind API comparison](https://www.gladia.io/compare-stt-apis) tool and check [pricing](https://www.gladia.io/pricing) for current plans. Speaker-count hints do not guarantee a specific detected speaker count. From 55f784dfd83be5a945d66a030f08d514b3a0ae34 Mon Sep 17 00:00:00 2001 From: karamouche Date: Thu, 1 Oct 2026 14:18:24 -0400 Subject: [PATCH 2/3] docs(speaker-diarization): use H1 wording for page title --- chapters/audio-intelligence/speaker-diarization.mdx | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/chapters/audio-intelligence/speaker-diarization.mdx b/chapters/audio-intelligence/speaker-diarization.mdx index e83bec7..fdaf332 100644 --- a/chapters/audio-intelligence/speaker-diarization.mdx +++ b/chapters/audio-intelligence/speaker-diarization.mdx @@ -1,5 +1,5 @@ --- -title: 'Speaker diarization API: configuration and output' +title: 'Speaker diarization for pre-recorded audio' sidebarTitle: Diarization description: 'Enable speaker diarization for pre-recorded audio in Gladia. Configure speaker-count hints and read speaker labels, timestamps and limitations.' --- From c96a070513c2a459822a9a56f4a63e9760e4a63e Mon Sep 17 00:00:00 2001 From: karamouche Date: Thu, 1 Oct 2026 14:34:53 -0400 Subject: [PATCH 3/3] docs(speaker-diarization): clarify scope vs channels, fix JSON sample --- chapters/audio-intelligence/speaker-diarization.mdx | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/chapters/audio-intelligence/speaker-diarization.mdx b/chapters/audio-intelligence/speaker-diarization.mdx index fdaf332..b299f69 100644 --- a/chapters/audio-intelligence/speaker-diarization.mdx +++ b/chapters/audio-intelligence/speaker-diarization.mdx @@ -38,7 +38,7 @@ Speakers will be assigned indexes by **order of appearance** (i.e. the 1st speak "end": 2.364, "confidence": 0.8914285714285715, "channel": 0, - "speaker": 0, + "speaker": 0 }, ... ] @@ -61,6 +61,6 @@ Provide speaker-count hints with `diarization_config.number_of_speakers`, `diari This guide covers speaker diarization for **pre-recorded audio** only. -Speaker diarization, transcription accuracy, and channel identification are different capabilities. For concept definitions, see [speaker diarization concepts](https://www.gladia.io/blog/what-is-diarization). +Speaker diarization labels who spoke when within a single mixed audio track, while transcription accuracy measures word errors in the transcribed text. Channel identification instead relies on separate audio channels, reported in each utterance's `channel` field, rather than telling speakers apart within one track (see [Multiple channels](/chapters/limits-and-specifications/multiple-channels)). For concept definitions, see [speaker diarization concepts](https://www.gladia.io/blog/what-is-diarization). Async accuracy comparisons use the model and dataset scope described in the [async benchmark methodology](https://www.gladia.io/competitors/benchmarks). Use the [blind API comparison](https://www.gladia.io/compare-stt-apis) tool and check [pricing](https://www.gladia.io/pricing) for current plans. Speaker-count hints do not guarantee a specific detected speaker count.