Our method · step 03 of 10

Extract

The video, its metadata and its comments are extracted. The audience is measured again at day 1, 7 and 30, then every month.

  1. 1The video is downloaded by the project’s network of computers.
  2. 2Its metadata is recorded: title, length, views, likes.
  3. 3Its comments are extracted and kept with their date.
  4. 4The audience is measured again at day 1, 7 and 30, then every month.
Tools
yt-dlp, RSS feeds, the platforms’ public APIs
Snapshots
day 1, 7, 30, then monthly
Writes
metadata, audience history, dated comments
Extraction

The video, its comments, its audience

Within an hour of detection, a video less than nine days old moves to the front of the extraction queue. Older videos are extracted in batches. No video is redistributed: the site’s player points to the platform.

  1. within the hour

    The metadata

    The title, length, date, view count, likes and announced comment count are recorded and dated. The video file is downloaded for transcription only, then deleted.

  2. +7 d · +30 d

    The comments

    Comments are extracted with their date and their author’s pseudonym, most recent first. Those of recent videos are extracted again after one week and after one month, to follow arrivals and disappearances.

  3. day 1 · 7 · 30 · monthly

    The audience figures

    Views, likes and comment counts of recent videos are measured again after one day, one week and one month. The whole corpus is measured once a month. The figures are timestamped and kept, even after the video is removed.

The text of comments stays under controlled access, under fair dealing for research; it is neither published on this site nor redistributed as such. The like count is not shown when it is below 0.05% of the views of a video viewed at least 10,000 times: such values come from a reading error during collection, logged in the corrections log (step 10).

Project paper