Transcribe
Voice isolated, speakers told apart, speech transcribed in the language spoken, then cut into sentences.
- 1The voice is isolated from music and noise.
- 2Speakers are told apart, never identified: A, B, C.
- 3Speech is transcribed in the language spoken, turn by turn.
- 4The text is cut into sentences, ready for the classifiers.
Six processing steps
The method has six steps. The first five run on the computers of the project’s network, the last on the central server. The models shown are those in service in October 2026.
Speaking turns
The pyannote model detects who speaks and when. Each speaking turn is attributed to an anonymous speaker (A, B…). Speakers are not identified by name.
Transcription
Whisper large-v3, in its turbo version since 2 October 2026, transcribes each passage of speech in its own language: an English-language video is transcribed in English, a mixed one passage by passage.
Segmentation
Speaking turns are split into sentences no longer than the classifiers’ maximum input.
Politicisation
Each transcript sentence from a French-language channel is classified as political or not. Comment sentences go through a second classifier, pre-trained on social-media posts.
Themes
Each political sentence then goes through binary classifiers, one per category of the project’s coding schemes (Boursier and Lemor, 2025). Each indicates whether the sentence deals with its theme, whatever the position expressed. Nine of them are published.
Meaning-based search
On the central server, each sentence is converted into a numerical vector. The data explorer uses these vectors to find sentences with a similar meaning.
This method has been in use since 3 October 2026. Before that, the voice was first separated from the music (Demucs), then the whole video was transcribed before speaking turns were attributed. Most videos in the corpus were transcribed this way.