yt-research-skill
An agent skill that batch-researches a YouTube channel's latest N videos — captions first, local ASR fallback — and writes one cross-video report instead of N summaries. Works in DSH, Codex, Claude Code, Hermes and WorkBuddy.
- Stars
- 0
- Language
- Python
- Created
- Sep 11, 2026
- Updated
- Sep 11, 2026
Introduction
youtube-channel-research
[!IMPORTANT] This repository is superseded by multiplat-research. That project keeps the same reports and traps table but splits the code into swappable data sources (YouTube · Bilibili · local files · Douyin own-account exports) and a platform-agnostic analysis core. It also fixes two things this version still has: the ffmpeg dependency was removed (faster-whisper decodes audio through PyAV) and a client fallback chain works around YouTube returning
HTTP 403on media streams. Use the newer repository unless you specifically want the single-platform version.
Turn "what has this channel been posting lately?" into a cross-video research report. Give it a channel and N — it enumerates the latest N uploads, pulls the text of each one (captions first, local ASR as fallback), then writes a single comparative report. Not N disconnected summaries.
A pure-Python, zero-external-service Agent Skill that works in any host supporting the standard skill layout: DSH / Codex / Claude Code / Hermes / WorkBuddy.
The problem it solves
Researching a channel by hand means: find the videos → figure out which have captions → download captions, or download audio and transcribe the rest → read N long transcripts → compare them. The first four steps are pure labour, and they come with traps that bite every single time (client-fingerprint blocking, corrupted HuggingFace model caches on Windows, zero-width characters polluting caption text, runaway ASR cost).
This skill scripts steps 1–4 and writes down every trap it hit, so the agent can spend its effort on step 5.
| Capability | Detail |
|---|---|
| Two-pass enumeration | Pass 1: one request for the video list. Pass 2: per-video publish date / views / caption availability — without downloading any media |
| Cost estimate up front | Tells you before downloading how many of the N videos need ASR, and roughly how many minutes that will take |
| Captions first | Manual and auto captions, dual json3 / vtt parsing, automatic de-duplication of scrolling caption repeats |
| ASR fallback | faster-whisper (CPU / int8), automatic language detection, robust multi-candidate model resolution |
| Text cleaning | Strips zero-width characters and HTML tags, collapses whitespace, removes spurious spaces between CJK characters |
| Cost guardrail | Refuses batches whose estimated ASR time exceeds --max-asr-min (default 30 min) unless explicitly confirmed |
| Consistent dating | Publish dates come from YouTube's own upload_date (US Pacific), so files never disagree by a day |
Quick start
# 1) Dependencies (a skill-local venv keeps your system Python clean)
python -m venv .venv
.venv/bin/python -m pip install -r requirements.txt # Windows: .venv\Scripts\python.exe
# 2) Self-check: verify interpreter + dependency resolution
python scripts/python_env.py
# 3) Enumerate the latest 5 videos -> channel.json + index.md
python scripts/channel_list.py "@MrBeast" -n 5 --out ./yt-research
# 4) Fetch text (captions first, ASR only when there are none)
python scripts/fetch_transcript.py --index ./yt-research/channel.json --out ./yt-research/transcripts
# 5) Let the agent write research.md from index.md + transcripts/*.plain
Then have your agent write the report following the six-section format in
SKILL.md, or just use that structure yourself.
Installing into an agent host
The layout is a plain folder, so "installing" simply means making this folder visible inside a host's skills directory.
| Host | Skills directory | Notes |
|---|---|---|
| DSH | ~/.dsh/skills/<name> | also auto-discovers ~/.skillshub/* |
| Codex | ~/.codex/skills/<name> | same; on our test machine both share ~/.skillshub |
| Claude Code | ~/.claude/skills/<name> | project-level .claude/skills/ also works |
| Hermes | %LOCALAPPDATA%\hermes\skills (Windows) / ~/.hermes/skills | also hermes skills install <repo> and hermes skills tap add <repo> |
| WorkBuddy | ~/.workbuddy/skills/<name> | real directory, not a link, on our machine |
One-command install (recommended)
python scripts/install_to_agents.py # detect existing host dirs, write nothing
python scripts/install_to_agents.py --install # actually install
python scripts/install_to_agents.py --agents dsh,codex --mode copy # pick hosts and copy instead
The script creates a directory link where supported (Windows junction / POSIX symlink),
so a later git pull updates every host at once. It falls back to copying when linking
is unavailable or denied.
Manual install
Windows (PowerShell, DSH as the example)
New-Item -ItemType Junction -Path "$env:USERPROFILE\.dsh\skills\youtube-channel-research" `
-Target "C:\path\to\yt-research-skill"
macOS / Linux
ln -s /path/to/yt-research-skill ~/.claude/skills/youtube-channel-research
Hermes (install straight from the repo)
hermes skills install <your-github-user>/yt-research-skill
# or register it as a skill source, then use hermes skills browse/search
hermes skills tap add <your-github-user>/yt-research-skill
Environment variables
No machine-specific paths are hard-coded anywhere. Two variables exist for overrides:
| Variable | Purpose | Default |
|---|---|---|
YTDLP_PY | Python interpreter that has yt_dlp | auto-resolved (below) |
YT_ASR_MODEL_DIR | faster-whisper model root | first writable candidate |
Interpreter resolution order (scripts/python_env.py):
$YTDLP_PY / $YT_ASR_PY → skill-local .venv/ → current interpreter
→ common third-party agent Pythons → python3 / python
There are deliberately two resolution tiers: resolve_downloader() only needs
yt_dlp, while resolve_asr() additionally requires faster_whisper and is called
before any audio is downloaded. A single-tier check picks interpreters that have
yt_dlp but not faster_whisper, and the failure then surfaces only after a full audio
download — the worst possible moment.
Output
| File | Content |
|---|---|
channel.json | Channel info + full metadata for N videos + ASR cost estimate |
index.md | Human-readable index table (with a "source" column: caption language or ASR) |
transcripts/<NN>-<id>.txt | Transcript with [mm:ss] timestamps |
transcripts/<NN>-<id>.plain | Plain text (for reading / feeding to a model) |
transcripts/<NN>-<id>.json | Structured segments + metadata + content source |
transcripts/<NN>-<id>.meta.json | Metadata quick-reference |
transcripts/transcripts.json | Batch manifest (success / failure / elapsed) |
research.md | The deliverable, written by the agent in the six-section format |
Example output
A real, unedited run lives in examples/:
python scripts/channel_list.py "@MrBeast" -n 3 --out ./smoke
# [done] channel=MrBeast videos=3 total=1:41:46 needASR=0 est. transcribe=0.0 min
The index produced by step one (examples/index-mrbeast.md) —
note how the "source" column already tells you the cost:
| # | Title | Length | Views | Published | Source |
|---|---|---|---|---|---|
| 1 | I Survived The Most Extreme Places On Earth | 23:27 | 77,909,716 | 2026-09-06 | captions manual:zh-Hans |
| 2 | Escape 100 Cops, Win $500,000 | 20:28 | 99,800,365 | 2026-08-23 | captions manual:zh-Hans |
| 3 | Last To Leave Mansion, Keeps It | 57:51 | 111,328,601 | 2026-08-09 | captions manual:zh-Hans |
Step two fetched three transcripts (30,874 characters total) and step three produced a
≈24 KB report (examples/research-mrbeast.md, in Chinese)
with this structure:
1. Channel overview — subscribers, sample range, publish-date span, source mix (captions 3 / ASR 0)
2. Per-video cards — one-line thesis + 3-6 key points + source and confidence, for each video
3. Comparison table — 16 dimensions side by side (theme / mechanic / prize / structural formula / sponsors / emotional payoff …)
4. Themes and drift — recurring narrative devices, series patterns, what changed over time
5. Conflicts and doubts — where videos disagree; where caption translation may be off
6. Method and limits — captions vs ASR, timestamp error, "latest N" ≠ the whole channel
The two findings that best show why this is cross-video research rather than N summaries:
- Section 3 isolates a structural formula shared by all three videos (cold open with the stakes → one-sentence rules → numeric/timestamp progress calls → one mid-video rule twist → at least two brand integrations → a self-referential closing line);
- Section 4 finds a narrative device drifting: videos 1 and 2 both use a uniform "win money for your mom" frame, which by video 3 has decayed into individual contestants' backstories.
Cost and runtime (measured)
- Caption path: essentially free — roughly one request per video.
- ASR path: ≈0.35× realtime on CPU with medium/int8 and 8 threads (13 min of audio ≈ 4.5 min to transcribe).
- Because of that, the script refuses batches estimated above 30 minutes of ASR by default;
narrow the range with
--ordersor confirm explicitly with--yes.
Traps worth knowing first
The full list of 12 measured traps is in SKILL.md. The four that matter most:
- Do not use
tempfile.mkdtemp()as yt-dlp's working directory. In a sandboxed environment the directory it creates cannot even beos.listdir-ed (WinError 5): Python can write there, the yt-dlp subprocess cannot. The scripts use.work/under the output directory instead. - Do not use curl for YouTube metadata. Some networks gate on TLS/client fingerprint and
will block curl (HTTP 000). Always go through
yt-dlp --dump-json. - HuggingFace model caches do break.
snapshots/can hold 0-byte placeholders while the real files sit inblobs/. Runpython scripts/setup_asr_model.py --repair-hf-cacheto repair in place (one-time, uses hard links so it costs no extra disk). - Auto-generated CJK captions carry zero-width characters (used for karaoke timing).
Unclean text poisons everything downstream;
clean_text()in the scripts handles it.
License
MIT.
This skill was generated by an AI agent in a working session and verified end to end on a
real channel. Every entry in the "traps" table comes from an actual error and its fix,
not from speculation. See ATTRIBUTION.md for details.