Our method

Ten steps, run on the project’s computing infrastructure, turn each video into an analysable format. Follow each step below to see how it works in detail, the models used, the decision thresholds and the reliability indicators.

The methods paper (PDF)

Steps 01 to 03

Build and collect

Channels chosen and coded, whose videos are detected, then extracted with their comments and their audience.

  1. 01

    Build the corpus

    Three countries, seven orientations: YouTube and TikTok channels chosen and coded by the team. The English-language corpus was extended in October 2026; its classifiers will soon be trained.

    The detail
  2. 02

    Detect

    The channels’ feeds are read every six hours, their full list every week. The videos of the last seven days go ahead of everything else.

    The detail
  3. 03

    Extract

    The video, its metadata and its comments are extracted. The audience is measured again at day 1, 7 and 30, then every month.

    The detail
Steps 04 to 06

Transcribe and classify

From sound to text, then to classified sentences, in which public figures are counted.

  1. 04

    Transcribe

    Speakers are told apart without being identified, speech is transcribed in the language spoken, then the text is cut into sentences.

    The detail
  2. 05

    Classify

    Each sentence is classified as political or not, then assigned to nine themes by nine classifiers whose reliability is published.

    The detail
  3. 06

    Name the people cited

    The public figures, parties and media of a closed list are recognised in the sentences and counted video by video; a private person is never counted.

    The detail
Steps 07 to 10

Write and publish

The records and syntheses, the arguments by theme, the tracking of removals, and the site recomputed every hour.

  1. 07

    Summarise

    Our local language model writes the record of each video of the week, checked against the transcript, then the syntheses of the day and of the week.

    The detail
  2. 08

    Draw out the arguments

    For each theme, the model reads blind a sample of sentences from several channels and draws out the arguments; each links back to the sentences that carry it.

    The detail
  3. 09

    Track removals

    The removal of a video is noticed, dated and explained when the platform says why. Its record stays, with its audience figures.

    The detail
  4. 10

    Publish

    Every hour, the counts per video are recomputed and the site refreshed, with the corpus coverage, the known limits and the corrections log.

    The detail
Project paper