Our method · 03

Transcribe

Voice isolated, speakers told apart, speech transcribed in the language spoken, then cut into sentences.

  1. 1The voice is isolated from music and noise.
  2. 2Speakers are told apart, never identified: A, B, C.
  3. 3Speech is transcribed in the language spoken, turn by turn.
  4. 4The text is cut into sentences, ready for the classifiers.
Processing

Six processing steps

The method has six steps. The first five run on the computers of the project’s network, the last on the central server. The models shown are those in service in October 2026.

01

Speaking turns

The pyannote model detects who speaks and when. Each speaking turn is attributed to an anonymous speaker (A, B…). Speakers are not identified by name.

pyannote 3.1
02

Transcription

Whisper large-v3, in its turbo version since 2 October 2026, transcribes each passage of speech in its own language: an English-language video is transcribed in English, a mixed one passage by passage.

WhisperMLX
03

Segmentation

Speaking turns are split into sentences no longer than the classifiers’ maximum input.

SaT
04

Politicisation

Each transcript sentence from a French-language channel is classified as political or not. Comment sentences go through a second classifier, pre-trained on social-media posts.

CamemBERTav2XLM-RoBERTa
05

Themes

Each political sentence then goes through binary classifiers, one per category of the project’s coding schemes (Boursier and Lemor, 2025). Each indicates whether the sentence deals with its theme, whatever the position expressed. Nine of them are published.

CamemBERTLLM_Tool
06

Meaning-based search

On the central server, each sentence is converted into a numerical vector. The data explorer uses these vectors to find sentences with a similar meaning.

Qwen3-Embedding-8B

This method has been in use since 3 October 2026. Before that, the voice was first separated from the music (Demucs), then the whole video was transcribed before speaking turns were attributed. Most videos in the corpus were transcribed this way.

Project paper