Our method · step 04 of 11

Classify

Political or not, then nine themes, by classifiers whose reliability is published.

  1. 1Each sentence is read by a classifier trained on human annotations.
  2. 2It is classified as political or not.
  3. 3A political sentence is assigned to nine themes, each with its probability.
  4. 4The reliability of each classifier is measured and published.
Themes

Nine themes, nine classifiers

The classifiers detect the theme of a sentence, whatever the position expressed. They were trained on annotations designed to measure far-right and neo-reactionary ideas, but each category was defined by its subject: a criticism of migration policy counts as a sentence about immigration.

These categories come from the project’s coding schemes (Boursier and Lemor, 2025). In the annotation protocol, a sentence fell under immigration as soon as it mentioned foreigners, immigrants or migrants. The reactionary ways of treating that subject, such as immigration described as a threat to national identity, were coded separately, as subcategories. The published classifiers learned the general category only.

In the training data, the share of sentences annotated as belonging to a theme without any reactionary subcategory varies widely: 75% for the environment, 46% for immigration, 33% for democracy, 31% for distrust of the state, 29% for technology. It stays below 6% for authority, tradition, progress and equality, whose training examples almost all treated the theme from a reactionary angle. These four classifiers may miss more of the sentences that approach their theme from another angle; this has not been measured.

The shares published on this site show what political sentences talk about. The positions taken on each theme are analysed separately, on samples of sentences: see the arguments by theme.

Wherever one orientation is compared with another, the share of a theme is a mean per video: for each video, the share of its political sentences that the classifier assigns to the theme, then the mean of these shares over the videos of the orientation, each video counting once. A video enters this mean from 20 political sentences. Adding up sentences would weigh the most prolific channels, not the orientation: in 2026, left-wing channels produce close to three quarters of the transcribed political sentences. “All orientations” is the mean of the orientations, each counting equally. The split of a theme’s sentences between orientations remains shown on the themes page, as a description of the corpus.

The nine published themes
Immigration
Political sentences about immigration: foreigners, immigrants, migrants, asylum.
Democracy
Political sentences about democracy, as an ideal or as a regime: elections, institutions, representation.
Authority
Political sentences about authority: obedience, order, the use of force against those who break norms.
Tradition
Political sentences about tradition: values, customs, family models, heritage.
Progress
Political sentences about progress in a broad sense: technical, human or social progress.
Equality and inequality
Political sentences about equality or inequality between human beings: sexes, origins, social groups.
Environment
Political sentences about the environment or ecology: climate, energy, nature, green parties.
Technology
Political sentences about technology, technique or innovation, or about tech companies and their leaders.
Distrust of the state
Political sentences expressing libertarian ideas or distrust of the state, public services, taxation or norms.
Validation

Reliability of the classifiers

The classifiers are trained on sentences annotated by a large language model using the LLM_Tool software (technical paper). The reliability of this model was first measured on 1,000 transcript sentences also annotated, independently, by two researchers: for politicisation, Light’s κ is 0.787 and the macro F1 is 0.922 (project paper, table 2). This human check covers the politicisation of transcripts only. Each classifier is then evaluated on a validation set kept apart from training. According to the project’s model registry, the macro F1 is 0.92 for the politicisation of transcripts (2,000 sentences) and 0.89 for that of comments (4,339 sentences).

This score measures each classifier’s agreement with the language model’s annotations. For themes, a further check was carried out in June 2026: 30 to 40 sentences that each classifier assigned to its theme with high probability were annotated again by a different language model. Precision is the share of these sentences that do deal with the theme, with borderline cases counting as half. This check was not split by orientation. The share of sentences on a theme that the classifier misses (recall) has not been measured.

ThemeMacro F1 (validation)SentencesPrecision (check)Sentences
Immigration0.952,2200.7640
Democracy0.882,1340.9030
Authority0.872,5150.7040
Tradition0.911,5100.7030
Progress0.891,0420.6230
Equality and inequality0.892,3810.8530
Environment0.961,8090.9030
Technology0.955420.6330
Distrust of the state0.905130.6530

For themes whose precision is close to 0.6, the published shares are orders of magnitude. Two classifiers are not published: nationalism, which picks up almost any mention of the nation or of France (precision of 0.39 in the check), and fictional metaphors, whose validation set has only 77 sentences.

Project paper