Our method · step 08 of 10

Draw out the arguments

For each theme, the model reads blind a sample of sentences from several channels and draws out the arguments; each links back to the sentences that carry it.

  1. 1A sample of at most 120 sentences of the theme is drawn, from several channels and orientations.
  2. 2The model reads them blind, without the channels’ names or orientation.
  3. 3It draws out three to six arguments; each must rest on at least three sentences from two videos.
  4. 4Each argument links to its exact quotes; the shares by orientation are weighted.
Sample
up to 120 sentences per draw, quotas by orientation, a cap per channel
Guard
at least three sentences from two videos per argument; two draws reconciled when possible
Writes
arguments, quotes, weighted shares by orientation
Method in detail

From text to arguments

Each argument links back to the sentences supporting it. Open a step to read the calculation rules.

01Periods

Sentences are grouped by period, the most recent months, each year since 2018 (the current one included) and the years 2009 to 2017 together. The recent period covers three, six or twelve months, the shortest span in which every theme has at least 600 sentences from at least eight channels, with at least 50 in each of two orientations.

02Sample

The draw covers political sentences of 60 to 400 characters that the classifier assigns to the theme with high probability. A sample holds at most 120 sentences. Each orientation gets a quota proportional to the square root of its number of sentences, with at least eight sentences where possible. A channel supplies at most eight videos (sixteen for 2009 to 2017) and 15% of the sample; a video supplies at most two sentences.

03Filters

Before the draw, filters set aside duplicates, speech-recognition loops (the same phrase repeated) and sentences with almost the same meaning. They also set aside sentences containing a person’s name that is not on the predefined list of public figures. A period is processed only if it has at least 150 candidate sentences from at least five channels.

04Blind reading

The open language model used for the summaries (Gemma 4 31B), run on the central server, reads the sentences without knowing their channel or orientation. It draws out three to six arguments. For each, it returns the numbers of the sentences that put it forward and of those that criticise it or defend the opposite position. It also flags the sentences that do not belong to the theme.

05Two draws

When the period has at least 240 candidate sentences, two separate samples are read independently, then a third pass reconciles the two lists of arguments. Arguments found in both readings come first. An argument is kept only if it rests on at least three sentences from at least two videos.

06Quotes and shares

Quotes reproduce the transcribed sentences unedited, cut beyond 280 characters. The number of sentences, the channels and the breakdown by orientation are computed from the sample, without the model. Orientation shares are weighted to correct for the quotas; each sentence counts for the number of sentences of its orientation that it represents.

Publication, updates and reading the results

A result that seems to name a private individual is not published. A period is marked provisional when less than half of its YouTube videos are transcribed. The analysis of the recent period is redone at most every four weeks, when its number of sentences has changed by at least 15% or its length has changed. A year is analysed again when it ends, when its number of sentences grows by at least 15% or when its share of transcribed videos passes 50%, 75% or 90%.

The arguments describe a sample. They are ordered by their frequency in that sample; this frequency gives an order of magnitude for all the sentences of the theme. Orientation shares depend on the make-up of the corpus, where the far right is over-represented, and on the share of videos transcribed in each orientation. The classifier sometimes assigns a sentence to the theme wrongly; the model flags these sentences, which are left out of the arguments. The model may also misread an ironic or ambiguous sentence. Transcription errors carry over into the quotes. The texts produced are not reviewed by the team. See the arguments by theme →

Frequencies describe sentences in the sample; they do not measure public opinion. Quotes let you return to the statements analysed.

Read the arguments →

Project paper