Classify
Each sentence is classified as political or not, then assigned to nine themes by nine classifiers whose reliability is published.
- 1Each sentence is read by a classifier trained on human annotations.
- 2It is classified as political or not.
- 3A political sentence is assigned to nine themes, each with its probability.
- 4The reliability of each classifier is measured and published.
- Models
- CamemBERTav2, one classifier per label; XLM-RoBERTa for comments
- Reliability
- validation F1 and checked precision published for each classifier
- Writes
- label and probability per sentence and per classifier
Political or not, then the themes
Two families of classifiers read the sentences one after the other, on the computers of the network. They are trained on French. The English-language corpus, extended in October 2026, does not go through them yet; their English counterparts will soon be trained and deployed.
Politicisation
Each transcript sentence is classified as political or not by a classifier trained on human annotations. Comment sentences go through a second classifier, pre-trained on social-media posts.
Themes
Each political sentence then goes through binary classifiers, one per category of the project’s coding schemes (Boursier and Lemor, 2025). Each indicates whether the sentence deals with its theme, whatever the position expressed, with a probability. Nine of them are published.
English-language channels
The sentences of English-language channels count neither in the shares of political sentences nor in the theme shares. Their videos are transcribed in the spoken language, summarised in French and in English, and the public figures they name are counted like the others.
Nine themes, nine classifiers
The classifiers detect the theme of a sentence, whatever the position expressed. They were trained on annotations designed to measure far-right and neo-reactionary ideas, but each category was defined by its subject, so that a criticism of migration policy counts as a sentence about immigration.
These categories come from the project’s coding schemes (Boursier and Lemor, 2025). In the annotation protocol, a sentence fell under immigration as soon as it mentioned foreigners, immigrants or migrants. The reactionary ways of treating that subject, such as immigration described as a threat to national identity, were coded separately, as subcategories. The published classifiers learned the general category only.
In the training data, the share of sentences annotated as belonging to a theme without any reactionary subcategory varies widely, 75% for the environment, 46% for immigration, 33% for democracy, 31% for distrust of the state, 29% for technology. It stays below 6% for authority, tradition, progress and equality, whose training examples almost all treated the theme from a reactionary angle. These four classifiers may miss more of the sentences that approach their theme from another angle; this has not been measured.
The shares published on this site show what political sentences talk about. The positions taken on each theme are analysed separately, on samples of sentences; see the arguments by theme.
Wherever one orientation is compared with another, the share of a theme is a mean per video. For each video, the share of its political sentences that the classifier assigns to the theme is computed, then the mean of these shares over the videos of the orientation, each video counting once. A video enters this mean from 20 political sentences. Adding up sentences would weigh the most prolific channels, not the orientation; in 2026, left-wing channels produce close to three quarters of the transcribed political sentences. “All orientations” is the mean of the orientations, each counting equally. The split of a theme’s sentences between orientations remains shown on the themes page, as a description of the corpus.
- Immigration
- Political sentences about immigration, foreigners, immigrants, migrants or asylum.
- Democracy
- Political sentences about democracy, as an ideal or as a regime, about elections, institutions or representation.
- Authority
- Political sentences about authority, obedience, order or the use of force against those who break norms.
- Tradition
- Political sentences about tradition, values, customs, family models or heritage.
- Progress
- Political sentences about progress in a broad sense, technical, human or social.
- Equality and inequality
- Political sentences about equality or inequality between human beings, by sex, origin or social group.
- Environment
- Political sentences about the environment or ecology, climate, energy, nature or green parties.
- Technology
- Political sentences about technology, technique or innovation, or about tech companies and their leaders.
- Distrust of the state
- Political sentences expressing libertarian ideas or distrust of the state, public services, taxation or norms.
Reliability of the classifiers
The classifiers are trained on sentences annotated by a large language model using the LLM_Tool software (technical paper). The reliability of this model was first measured on 1,000 transcript sentences also annotated, independently, by two researchers. For politicisation, Light’s κ is 0.787 and the macro F1 is 0.922 (project paper, table 2). This human check covers the politicisation of transcripts only. Each classifier is then evaluated on a validation set kept apart from training. According to the project’s model registry, the macro F1 is 0.92 for the politicisation of transcripts (2,000 sentences) and 0.89 for that of comments (4,339 sentences).
This score measures each classifier’s agreement with the language model’s annotations. For themes, a further check was carried out in June 2026. Thirty to forty sentences that each classifier assigned to its theme with high probability were annotated again by a different language model. Precision is the share of these sentences that do deal with the theme, with borderline cases counting as half. This check was not split by orientation. The share of sentences on a theme that the classifier misses (recall) has not been measured.
| Theme | Macro F1 (validation) | Sentences | Precision (check) | Sentences |
|---|---|---|---|---|
| Immigration | 0.95 | 2,220 | 0.76 | 40 |
| Democracy | 0.88 | 2,134 | 0.90 | 30 |
| Authority | 0.87 | 2,515 | 0.70 | 40 |
| Tradition | 0.91 | 1,510 | 0.70 | 30 |
| Progress | 0.89 | 1,042 | 0.62 | 30 |
| Equality and inequality | 0.89 | 2,381 | 0.85 | 30 |
| Environment | 0.96 | 1,809 | 0.90 | 30 |
| Technology | 0.95 | 542 | 0.63 | 30 |
| Distrust of the state | 0.90 | 513 | 0.65 | 30 |
For themes whose precision is close to 0.6, the published shares are orders of magnitude. Two classifiers are not published, nationalism, which picks up almost any mention of the nation or of France (precision of 0.39 in the check), and fictional metaphors, whose validation set has only 77 sentences.