The eleven steps
The full diagram, from the task queue to publication.
The eleven steps of the method
Each step runs on the project’s computers, linked by a distributed-computing application. Click a step: what it does, with what, what it writes, and what is known of its reliability.
01 · Detect
The RSS feeds of the YouTube channels are read every six hours; every week, the full list of each channel’s videos is reread, for YouTube and TikTok alike. A detected video gets its publication date and enters the queue; those of the last seven days go ahead of everything else.
- Cadence
- RSS every 6 h, full listing every week
- Priority
- last seven days first, then since April 2026, then the backlog
- Writes
- detected videos, date, task queue
02 · Extract
The video, its metadata (title, length, views, likes, announced comment count) and its comments are extracted. The audience is measured again at day 1, 7 and 30, then during the monthly passes. A removal is recorded with its reason when the platform gives one.
- Tools
- yt-dlp, RSS feeds, the platforms’ public APIs
- Snapshots
- day 1, 7, 30, then monthly
- Writes
- metadata, audience history, comments, deletions
03 · Isolate the voice
The audio track is separated from music and noise (Demucs), so that speech recognition only has speech to read. On small computers this step runs alone, one model at a time, to fit in memory.
- Model
- Demucs (htdemucs)
- Writes
- temporary voice track, deleted after transcription
04 · Tell speakers apart
Speaking turns are attributed to anonymous speakers (A, B, C…) by diarisation (pyannote). Speakers are never identified: the record speaks of “the host”, “the guest”.
- Model
- pyannote 3.1
- Writes
- speaking turns, number of speakers, share of the main one
05 · Transcribe
Whisper (large-v3-turbo, on MLX) transcribes each speaking turn in the spoken language, detected by windows: a French video stays in French, an American one in English, a mixed one segment by segment. A video without speech is marked as such.
- Model
- Whisper large-v3-turbo (MLX)
- Quality
- the automatic record rates the transcript’s reliability (good, medium, poor)
- Writes
- raw and cleaned transcript, per speaking turn
06 · Cut into sentences
Speaking turns are cut into sentences by a segmentation model (SaT), with a ceiling of 510 tokens per unit, so that each classifier reads a whole sentence and never a fragment.
- Model
- SaT (wtpsplit)
- Guard
- segmentation anomalies detected and re-segmented
- Writes
- numbered sentences, linked to their speaking turn
07 · Political or not
Each sentence is classified as political or not by a CamemBERTav2 classifier trained on human annotations, then assigned to nine themes by nine binary classifiers. The classifiers are trained on French: English-language channels do not go through them.
- Models
- CamemBERTav2, one classifier per label
- Reliability
- F1 and checked precision published below
- Writes
- label and probability per sentence and per classifier
08 · The people named
Named entities are extracted (GLiNER, multilingual); they serve as a guard. The published counts come from a closed list of public figures, parties, organisations, media and countries, recognised in the sentences by checked patterns. A private person is never counted.
- Model
- GLiNER (entities), closed list (counts)
- List
- 240 entries
- Writes
- mentions per video and per entry
09 · Summarise
A local language model (Gemma 4, 31 billion parameters) writes a record for each video of the week: title, summary, claims with their topic and speaker, public figures named, transcript reliability, in French and in English. The record is checked against the transcript; a name that could be a private person’s holds it back.
- Model
- Gemma 4 31B, on the project’s computers
- Guard
- names checked against the closed list; record held in case of doubt
- Writes
- record per video, syntheses of the day and the week, records of deleted videos
10 · Extract the arguments
For each theme and period, a sample of political sentences is drawn with quotas by orientation and a cap per channel; the language model extracts the arguments, separates those put forward from those contested, and each argument is linked to the sentences carrying it.
- Sample
- 120 sentences per draw, two draws reconciled
- Writes
- arguments, quotes, weighted shares by orientation
11 · Publish
Every hour, the counts per video are recomputed and the site is refreshed: videos of the week, themes, public figures, trends, channels, deletions, in French and English. The same data are served as JSON and, with a key, through the explorer and the API.
- Cadence
- hourly
- Measures
- comparisons between orientations as means per video
- Writes
- this site, the JSON and CSV files, the explorer and the API