Our method · step 04 of 10

Transcribe

Speakers are told apart without being identified, speech is transcribed in the language spoken, then the text is cut into sentences.

  1. 1The voice is isolated from music and noise.
  2. 2Speakers are told apart, never identified: A, B, C.
  3. 3Speech is transcribed in the language spoken, turn by turn.
  4. 4The text is cut into sentences, ready for the classifiers.
Models
pyannote 3.1, Whisper large-v3-turbo (MLX), SaT
Guard
language detected by windows; segmentation anomalies re-cut
Writes
speaking turns, transcript, numbered sentences
Processing

From sound to sentences, in three operations

The three operations run on the computers of the project’s network, one after the other, with the models in service in October 2026. The resulting sentences then go to the classifiers (step 05).

01

Speaking turns

The pyannote model detects who speaks and when. Each speaking turn is attributed to an anonymous speaker (A, B…). Speakers are never identified: the record speaks of “the host”, “the guest”.

pyannote 3.1
02

Transcription

Whisper large-v3, in its turbo version since 2 October 2026, transcribes each passage of speech in its own language, detected by windows: a French video stays in French, an American one is transcribed in English, a mixed one passage by passage. A video without speech is marked as such.

WhisperMLX
03

Segmentation

Speaking turns are cut into sentences by a segmentation model, with a ceiling of 510 tokens per unit, so that each classifier reads a whole sentence and never a fragment. Segmentation anomalies are detected and re-cut.

SaT

This method has been in use since 3 October 2026. Before that, the voice was first separated from the music (Demucs), then the whole video was transcribed before speaking turns were attributed. Most videos in the corpus were transcribed this way.

Project paper