yt-briefing 0.4.0 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/yt-search/SKILL.md +53 -0
- package/README.md +31 -12
- package/dist/bootstrap.js +5 -4
- package/dist/cli.js +3 -1
- package/dist/install-skill.js +6 -5
- package/dist/lib/paths.js +4 -0
- package/dist/lib/skill-install.js +12 -2
- package/dist/lib/yt-api.js +29 -0
- package/dist/yt-search.js +294 -0
- package/package.json +1 -1
|
@@ -0,0 +1,53 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: yt-search
|
|
3
|
+
description: Research a topic across YouTube — describe an intent, the engine expands it into search queries, ranks results against your intent, then lazily yields ONE video at a time with a rich summary. You keep or skip each; at the end it synthesizes a comparison from everything you kept. Same transcript engine + proxy as /yt; lazy on purpose (no transcript bursts → no IP block). Summaries and prompts use the language chosen at onboarding.
|
|
4
|
+
argument-hint: A descriptive intent in quotes, e.g. "which terminal for coding with Claude Code". Optional --max N, --since YYYY-MM-DD, --queries 1..3.
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## How it works
|
|
8
|
+
|
|
9
|
+
`src/yt-search.ts` is the whole engine: intent → query expansion (LLM) → `search.list` → re-rank against the intent on metadata only (no transcript) → **lazy** one-candidate-at-a-time yield with a rich summary → record keep/skip → on demand synthesize a comparison from everything kept. Matching is **descriptive, not exact-keyword** — YouTube ranks by relevance and the LLM bridges intent→query and filters noise. This skill is a thin loop: paste the summary, collect keep/skip, show the final comparison. It runs no filters and formats nothing itself.
|
|
10
|
+
|
|
11
|
+
**Lazy on purpose:** one transcript per step, never a burst — a burst looks like scraping and gets the IP blocked (same reason `/yt` is lazy). Run the engine bare — stdout is a single JSON line, stderr empty; never redirect.
|
|
12
|
+
|
|
13
|
+
## Language
|
|
14
|
+
|
|
15
|
+
Read `data/config.json` → `output_lang` once at the start. Phrase the question text and option descriptions in that language. The two button labels stay the short English words `Keep` / `Skip`. Summaries and the final comparison are already written in `output_lang` by the engine — paste them verbatim.
|
|
16
|
+
|
|
17
|
+
## Loop
|
|
18
|
+
|
|
19
|
+
```
|
|
20
|
+
out = JSON.parse(`bun run src/yt-search.ts "<intent from the user>" --reset`) // first call
|
|
21
|
+
while true:
|
|
22
|
+
out.status:
|
|
23
|
+
"error" → show out.error verbatim, stop
|
|
24
|
+
"no_results" → tell the user nothing relevant was found, stop
|
|
25
|
+
"rate_limited" → transcript fetch blocked (datacenter IP) — tell the user, stop; recovery in README.md → Running on a VPS
|
|
26
|
+
"decision_needed" → steps A–C
|
|
27
|
+
"done" → step D
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
**On `decision_needed`:**
|
|
31
|
+
|
|
32
|
+
- **A.** Take `out.summary` (markdown) and `out.pending` (`{videoId,title,channelTitle,publishedAt,position,total}`).
|
|
33
|
+
- **B.** In the SAME turn, as your chat text (NOT command output — the UI doesn't show it), paste `summary` **verbatim** — no paraphrase, no shortening. Optionally prefix one line like `Wynik {position}/{total}`. The user must see it before the popup.
|
|
34
|
+
- **C.** In the same message call `AskUserQuestion` — 1 call, 1 question, phrased in `output_lang`:
|
|
35
|
+
- Question e.g. "Brać pod uwagę w porównaniu?" with two options: **Keep** = include this video in the final comparison; **Skip** = drop it.
|
|
36
|
+
- **Other** is the stop channel: if the user types `stop` (case-insensitive, trimmed) or dismisses the popup (✕) → **end the loop early** and go to step D (compare what's kept so far).
|
|
37
|
+
- Then act:
|
|
38
|
+
- Keep → `bun run src/yt-search.ts --keep`
|
|
39
|
+
- Skip → `bun run src/yt-search.ts --skip`
|
|
40
|
+
- stop/dismissed → `bun run src/yt-search.ts --compare` (skip straight to D)
|
|
41
|
+
- The script reads the pending candidate from cache; pass only `--keep` / `--skip` / `--compare`. Its JSON becomes the next `out` — back to the top of the loop.
|
|
42
|
+
|
|
43
|
+
**On `done` (step D):**
|
|
44
|
+
|
|
45
|
+
- If `out.kept > 0` → run `out = JSON.parse(\`bun run src/yt-search.ts --compare\`)`; when it returns `status:"compare"`, paste `out.comparison` **verbatim** as your chat text (it's the artifact — a decision-grade comparison in `output_lang`). Stop.
|
|
46
|
+
- If `out.kept == 0` → tell the user nothing was kept, so there's nothing to compare. Stop.
|
|
47
|
+
|
|
48
|
+
## Rules
|
|
49
|
+
|
|
50
|
+
- **Verbatim:** paste `summary` and `comparison` exactly as returned; never paste a raw transcript.
|
|
51
|
+
- **Language:** question text + option descriptions follow `output_lang`; button labels stay `Keep` / `Skip`.
|
|
52
|
+
- **Cost awareness:** each search runs `search.list` (100 quota units/query, up to `--queries`). Don't silently re-run `--reset` in a loop. A bare resume (no `--reset`) continues the same ranked queue without new searches.
|
|
53
|
+
- **Stateless triage:** this is independent of `/yt` (no channels, no ratings). For the recurring channel briefing use `/yt`; for one video use `/yt-transcribe`.
|
package/README.md
CHANGED
|
@@ -24,7 +24,8 @@ the last one left off.
|
|
|
24
24
|
You'll need Node 18+ or Bun, a YouTube Data API v3 key, an LLM key (a
|
|
25
25
|
[free Gemini key](https://aistudio.google.com/apikey) works, see [Providers](#providers)), and
|
|
26
26
|
a tool that runs skills: [Claude Code](https://claude.com/claude-code),
|
|
27
|
-
[Cursor](https://cursor.com), or anything else
|
|
27
|
+
[Cursor](https://cursor.com), [Codex](https://developers.openai.com/codex), or anything else
|
|
28
|
+
that loads the standard `SKILL.md` (Agent Skills — 30+ agents).
|
|
28
29
|
|
|
29
30
|
1. Install yt-dlp (it pulls the subtitles):
|
|
30
31
|
|
|
@@ -61,12 +62,6 @@ npx yt-briefing remove @handle # also deletes its l
|
|
|
61
62
|
npx yt-briefing list # show the current list
|
|
62
63
|
```
|
|
63
64
|
|
|
64
|
-
## Run it
|
|
65
|
-
|
|
66
|
-
Open your project in Claude Code or Cursor and run `/yt`. If it's not listed, start a fresh
|
|
67
|
-
session. To install the skills again for another tool or project, run
|
|
68
|
-
`npx yt-briefing install-skill` (it installs both `/yt` and `/yt-transcribe`).
|
|
69
|
-
|
|
70
65
|
## One-off: transcribe a single video
|
|
71
66
|
|
|
72
67
|
Just want one video summarized — no channels, no queue, no rating? Run `/yt-transcribe` and
|
|
@@ -74,8 +69,7 @@ paste a URL or video ID. It pulls that video's transcript and writes a journalis
|
|
|
74
69
|
summary in the language you chose at setup (the same `output_lang` as `/yt`). Want a one-off in
|
|
75
70
|
another language? Just say so when you run it (e.g. `/yt-transcribe <url> in German`) — it
|
|
76
71
|
won't change your setup. `--lang pl|en` is separate — it picks which caption track to fetch,
|
|
77
|
-
not the summary language.
|
|
78
|
-
server it benefits from the same [WARP proxy](#running-on-a-vps).
|
|
72
|
+
not the summary language.
|
|
79
73
|
|
|
80
74
|
The skill is installed alongside `/yt` by `init` / `install-skill`. From the plain CLI:
|
|
81
75
|
|
|
@@ -83,6 +77,31 @@ The skill is installed alongside `/yt` by `init` / `install-skill`. From the pla
|
|
|
83
77
|
npx yt-briefing transcribe <url-or-id> --lang auto # prints the transcript to stdout
|
|
84
78
|
```
|
|
85
79
|
|
|
80
|
+
## Research a topic across YouTube
|
|
81
|
+
|
|
82
|
+
Want to know what YouTube says about something — and get an actual comparison, not a pile of
|
|
83
|
+
links? Run `/yt-search` and describe the intent, e.g. `which terminal for coding with Claude
|
|
84
|
+
Code`. The matching is **descriptive, not exact keywords**: an LLM turns your intent into search
|
|
85
|
+
queries, then ranks the results against what you actually meant (on titles/descriptions — no
|
|
86
|
+
transcripts yet). Then it hands you **one video at a time** with a rich summary; you **Keep** or
|
|
87
|
+
**Skip** each. At the end it synthesizes a **comparison** from everything you kept.
|
|
88
|
+
|
|
89
|
+
It's lazy on purpose — one transcript per step, never a burst (a burst gets your IP blocked, same
|
|
90
|
+
as `/yt`). Summaries and the comparison use the language you chose at setup.
|
|
91
|
+
|
|
92
|
+
```bash
|
|
93
|
+
yt-briefing search "<intent>" [--max N] [--queries 1..3] [--since YYYY-MM-DD] # JSON status line
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
> Each search calls the YouTube `search.list` endpoint, which costs **100 quota units per query**
|
|
97
|
+
> (plain reads cost 1; the daily free quota is 10,000). `--queries` (default 2) caps how many.
|
|
98
|
+
|
|
99
|
+
## Run it
|
|
100
|
+
|
|
101
|
+
Open your project in Claude Code or Cursor and run `/yt`. If it's not listed, start a fresh
|
|
102
|
+
session. To install the skills again for another tool or project, run
|
|
103
|
+
`npx yt-briefing install-skill` (it installs `/yt`, `/yt-transcribe`, and `/yt-search`).
|
|
104
|
+
|
|
86
105
|
## Providers
|
|
87
106
|
|
|
88
107
|
Any OpenAI-compatible endpoint works. Gemini 2.5 Flash is the easy default. It's fast, cheap,
|
|
@@ -110,9 +129,9 @@ summarizing the next video while you rate the current one, so the following step
|
|
|
110
129
|
ready with no wait. An agent's turn-by-turn loop cannot prefetch like that, and every step pays
|
|
111
130
|
its own cold start, which adds up across a whole queue.
|
|
112
131
|
|
|
113
|
-
Compatibility. A standard API plus a
|
|
114
|
-
Cursor, any other tool
|
|
115
|
-
to that one tool and one model.
|
|
132
|
+
Compatibility. A standard API plus a standard `SKILL.md` means one engine runs everywhere: Claude
|
|
133
|
+
Code, Cursor, Codex, any other Agent-Skills-compatible tool, or the plain CLI. A tool-native
|
|
134
|
+
approach would tie it to that one tool and one model.
|
|
116
135
|
|
|
117
136
|
## Why one transcript at a time
|
|
118
137
|
|
package/dist/bootstrap.js
CHANGED
|
@@ -108,10 +108,11 @@ function main() {
|
|
|
108
108
|
// 6. Coding agent ------------------------------------------------------------
|
|
109
109
|
// Place the skill INTO THIS PROJECT (the package folder you open in the agent) — never a
|
|
110
110
|
// home-global dir (that's the npm -g antipattern: machine-wide, invisible, easy to forget).
|
|
111
|
-
//
|
|
112
|
-
//
|
|
113
|
-
|
|
114
|
-
console.log('
|
|
111
|
+
// SKILL.md is the cross-agent standard, so the shipped skill runs in any compatible agent —
|
|
112
|
+
// we just install it into that agent's skills dir (.claude/skills, .cursor/skills, .codex/skills).
|
|
113
|
+
// 1/2/3 = known agents; 4 = any other compatible agent (a project folder you name).
|
|
114
|
+
console.log('\n 6) Which agent will you run /yt in? (it ships a standard Agent Skill — any compatible agent works)');
|
|
115
|
+
console.log(' 1) Claude Code 2) Cursor 3) Codex 4) Custom folder (any other agent)\n');
|
|
115
116
|
const agentKey = ask(' Your agent', '1');
|
|
116
117
|
// For a custom target, ask the folder now (keeps all prompts in the interactive block).
|
|
117
118
|
const customDir = AGENTS[agentKey] ? '' : ask(' Skills folder to install into', customSkillsRootDefault());
|
package/dist/cli.js
CHANGED
|
@@ -12,6 +12,7 @@
|
|
|
12
12
|
* yt-briefing sweep [--reset] advance one step; prints a JSON status line
|
|
13
13
|
* yt-briefing rate --rating 0|1 [...] record a rating for the pending video
|
|
14
14
|
* yt-briefing transcribe <url|id> print a single video's transcript
|
|
15
|
+
* yt-briefing search "<intent>" [...] topic search → lazy triage → compare (JSON status line)
|
|
15
16
|
*/
|
|
16
17
|
import { spawnSync } from 'node:child_process';
|
|
17
18
|
import { script } from "./lib/paths.js";
|
|
@@ -24,6 +25,7 @@ const TARGETS = {
|
|
|
24
25
|
sweep: 'yt-sweep',
|
|
25
26
|
rate: 'yt-rating',
|
|
26
27
|
transcribe: 'yt-transcript',
|
|
28
|
+
search: 'yt-search',
|
|
27
29
|
};
|
|
28
30
|
const CHANNEL_ACTIONS = new Set(['add', 'remove', 'list']);
|
|
29
31
|
let argv = null;
|
|
@@ -32,7 +34,7 @@ if (cmd && CHANNEL_ACTIONS.has(cmd))
|
|
|
32
34
|
else if (cmd && TARGETS[cmd])
|
|
33
35
|
argv = [script(TARGETS[cmd]), ...rest];
|
|
34
36
|
if (!argv) {
|
|
35
|
-
console.error('Usage: yt-briefing <init|install-skill|add|remove|list|sweep|rate|transcribe> [args...]');
|
|
37
|
+
console.error('Usage: yt-briefing <init|install-skill|add|remove|list|sweep|rate|transcribe|search> [args...]');
|
|
36
38
|
process.exit(cmd ? 1 : 0);
|
|
37
39
|
}
|
|
38
40
|
const res = spawnSync(process.execPath, argv, { stdio: 'inherit' });
|
package/dist/install-skill.js
CHANGED
|
@@ -12,9 +12,9 @@
|
|
|
12
12
|
* standalone command is for re-installing, a different project, or a second agent. There is
|
|
13
13
|
* deliberately no home-global install — the skills live with the project that uses them.
|
|
14
14
|
*
|
|
15
|
-
*
|
|
16
|
-
*
|
|
17
|
-
*
|
|
15
|
+
* SKILL.md is the cross-agent Agent Skills standard, so the shipped skills run in any compatible
|
|
16
|
+
* agent (Claude Code, Cursor, Codex, and 30+ others); this command just (re)places them in the
|
|
17
|
+
* skills dir of whichever agent you pick.
|
|
18
18
|
*/
|
|
19
19
|
import { AGENTS, installSkills, projectSkillsRoot, customSkillsRootDefault, isPackageDevCwd } from "./lib/skill-install.js";
|
|
20
20
|
import { question } from "./lib/prompt.js";
|
|
@@ -28,8 +28,9 @@ function done(targets) {
|
|
|
28
28
|
// 1) which agent → which skills subdir
|
|
29
29
|
console.log('\n Install the /yt + /yt-transcribe skills — which agent?\n');
|
|
30
30
|
console.log(' 1) Claude Code');
|
|
31
|
-
console.log(' 2) Cursor
|
|
32
|
-
console.log(' 3)
|
|
31
|
+
console.log(' 2) Cursor');
|
|
32
|
+
console.log(' 3) Codex');
|
|
33
|
+
console.log(' 4) Custom folder (any other compatible agent)\n');
|
|
33
34
|
const agentKey = ask(' Agent', '1');
|
|
34
35
|
const agent = AGENTS[agentKey];
|
|
35
36
|
// 3) Custom — write the skills straight into a skills root the user names (their agent's dir).
|
package/dist/lib/paths.js
CHANGED
|
@@ -47,6 +47,10 @@ export const REST_FILE = join(CACHE_DIR, 'queue-rest.json');
|
|
|
47
47
|
export const PENDING_FILE = join(CACHE_DIR, 'pending.json');
|
|
48
48
|
export const PREFETCH_FILE = join(CACHE_DIR, 'prefetch.json');
|
|
49
49
|
export const LOG_FILE = join(CACHE_DIR, 'sweep.log');
|
|
50
|
+
// /yt-search (ad-hoc topic search → lazy triage → compare). All throwaway, under .cache/.
|
|
51
|
+
export const SEARCH_QUEUE_FILE = join(CACHE_DIR, 'search-queue.json'); // ranked candidates + cursor
|
|
52
|
+
export const SEARCH_PENDING_FILE = join(CACHE_DIR, 'search-pending.json'); // current candidate awaiting keep/skip
|
|
53
|
+
export const SEARCH_KEPT_FILE = join(CACHE_DIR, 'search-kept.json'); // kept mega-summaries → compare corpus
|
|
50
54
|
/** Absolute path to a channel profile from its slug. */
|
|
51
55
|
export const profilePath = (slug) => join(CHANNELS_DIR, `${slug}.md`);
|
|
52
56
|
/**
|
|
@@ -26,7 +26,7 @@ import { PKG_ROOT, DATA_DIR } from "./paths.js";
|
|
|
26
26
|
/** Compiled output dir — what a baked (dist) skill command points the runtime at. */
|
|
27
27
|
const DIST_DIR = join(PKG_ROOT, 'dist');
|
|
28
28
|
/** The skills this package ships — each lives at `.claude/skills/<name>/SKILL.md`. */
|
|
29
|
-
export const SKILLS = ['yt', 'yt-transcribe'];
|
|
29
|
+
export const SKILLS = ['yt', 'yt-transcribe', 'yt-search'];
|
|
30
30
|
/** True when the installer itself is running under Bun (vs plain Node). */
|
|
31
31
|
export const isBun = process.versions.bun != null;
|
|
32
32
|
/**
|
|
@@ -39,10 +39,19 @@ export const isBun = process.versions.bun != null;
|
|
|
39
39
|
export const isPackageDevCwd = () => isBun && resolve(process.cwd()) === PKG_ROOT;
|
|
40
40
|
/** Source path of a shipped skill's SKILL.md, by skill name. */
|
|
41
41
|
export const skillSource = (name) => join(PKG_ROOT, '.claude', 'skills', name, 'SKILL.md');
|
|
42
|
-
/**
|
|
42
|
+
/**
|
|
43
|
+
* Agent key → display name + the skills ROOT directory it scans (skills install under it).
|
|
44
|
+
*
|
|
45
|
+
* SKILL.md is the cross-agent Agent Skills standard (Anthropic, Dec 2025), now read by 30+
|
|
46
|
+
* tools that each scan their own `<agent-home>/skills/` dir. We only need the right directory
|
|
47
|
+
* per agent — the shipped SKILL.md works unmodified in all of them. The "custom folder" picker
|
|
48
|
+
* option (no AGENTS entry) covers every other compatible agent (Gemini CLI, Copilot, Windsurf…)
|
|
49
|
+
* and defaults to the neutral `.agents/skills/` location.
|
|
50
|
+
*/
|
|
43
51
|
export const AGENTS = {
|
|
44
52
|
'1': { name: 'Claude Code', sub: join('.claude', 'skills') },
|
|
45
53
|
'2': { name: 'Cursor', sub: join('.cursor', 'skills') },
|
|
54
|
+
'3': { name: 'Codex', sub: join('.codex', 'skills') },
|
|
46
55
|
};
|
|
47
56
|
/**
|
|
48
57
|
* One shipped skill's SKILL.md. `dist=false` (default) returns it verbatim — the
|
|
@@ -64,6 +73,7 @@ export function skillBody(name, dist = false) {
|
|
|
64
73
|
.replace(/bun run src\/yt-sweep\.ts/g, cmd('yt-sweep'))
|
|
65
74
|
.replace(/bun run src\/yt-rating\.ts/g, cmd('yt-rating'))
|
|
66
75
|
.replace(/bun run src\/yt-transcript\.ts/g, cmd('yt-transcript'))
|
|
76
|
+
.replace(/bun run src\/yt-search\.ts/g, cmd('yt-search'))
|
|
67
77
|
.replace(/data\//g, DATA_DIR + '/');
|
|
68
78
|
}
|
|
69
79
|
/**
|
package/dist/lib/yt-api.js
CHANGED
|
@@ -110,6 +110,35 @@ async function enrichWithTypes(videos) {
|
|
|
110
110
|
}
|
|
111
111
|
return out;
|
|
112
112
|
}
|
|
113
|
+
/**
|
|
114
|
+
* Free-text video search via `search.list`. YouTube ranks by relevance, so the query can be
|
|
115
|
+
* descriptive — no exact keyword match required. NOTE: search.list costs 100 quota units per
|
|
116
|
+
* call (plain reads cost 1), so callers should keep the number of queries small.
|
|
117
|
+
*
|
|
118
|
+
* `since` (ISO date) maps to publishedAfter — useful to cut stale results on fast-moving topics.
|
|
119
|
+
*/
|
|
120
|
+
export async function searchVideos(query, opts = {}) {
|
|
121
|
+
const { maxResults = 10, since = null } = opts;
|
|
122
|
+
const params = {
|
|
123
|
+
part: 'snippet',
|
|
124
|
+
q: query,
|
|
125
|
+
type: 'video',
|
|
126
|
+
order: 'relevance',
|
|
127
|
+
maxResults: String(Math.min(Math.max(maxResults, 1), 50)),
|
|
128
|
+
};
|
|
129
|
+
if (since)
|
|
130
|
+
params.publishedAfter = new Date(since).toISOString();
|
|
131
|
+
const data = await get('search', params);
|
|
132
|
+
return (data.items || [])
|
|
133
|
+
.filter((it) => it.id?.videoId)
|
|
134
|
+
.map((it) => ({
|
|
135
|
+
videoId: it.id.videoId,
|
|
136
|
+
title: it.snippet?.title ?? '',
|
|
137
|
+
channelTitle: it.snippet?.channelTitle ?? '',
|
|
138
|
+
publishedAt: it.snippet?.publishedAt ?? '',
|
|
139
|
+
description: it.snippet?.description ?? '',
|
|
140
|
+
}));
|
|
141
|
+
}
|
|
113
142
|
/**
|
|
114
143
|
* List a channel's uploads (newest first). With `enrich` (default true) each video is
|
|
115
144
|
* typed (short/live/longform) and current/upcoming live broadcasts are filtered out.
|
|
@@ -0,0 +1,294 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
/**
|
|
3
|
+
* yt-search.ts — ad-hoc topic search → lazy triage → comparison. The third yt-briefing mode,
|
|
4
|
+
* sibling to the channel briefing (yt-sweep) and one-shot transcribe (yt-transcript).
|
|
5
|
+
*
|
|
6
|
+
* You describe an intent ("which terminal for coding with Claude Code"); the engine:
|
|
7
|
+
* 1. expands it into a few YouTube search queries (LLM),
|
|
8
|
+
* 2. runs search.list and merges candidates,
|
|
9
|
+
* 3. re-ranks candidates against your intent on metadata only — title/channel/description,
|
|
10
|
+
* NO transcript yet (cheap; protects the expensive/rate-limited transcript step),
|
|
11
|
+
* 4. yields ONE candidate at a time with a rich summary, lazily — never a burst of transcript
|
|
12
|
+
* fetches (a burst looks like scraping and gets the IP blocked, same reason yt-sweep is lazy),
|
|
13
|
+
* 5. records your keep/skip decision; kept summaries accumulate in a cache,
|
|
14
|
+
* 6. on demand synthesizes a comparison across everything you kept.
|
|
15
|
+
*
|
|
16
|
+
* Matching is descriptive, not exact-keyword: search.list already ranks by relevance, and the
|
|
17
|
+
* LLM bridges intent→query (step 1) and filters noise (step 3).
|
|
18
|
+
*
|
|
19
|
+
* Usage (the skill / CLI drives these; one JSON line per call):
|
|
20
|
+
* yt-search "<intent>" [--reset] [--max N] [--queries N] [--since DATE] [--lang auto]
|
|
21
|
+
* yt-search --keep record the pending candidate, advance, yield next
|
|
22
|
+
* yt-search --skip drop the pending candidate, advance, yield next
|
|
23
|
+
* yt-search --compare synthesize a comparison from everything kept
|
|
24
|
+
*
|
|
25
|
+
* Output statuses:
|
|
26
|
+
* {"status":"decision_needed","summary":"<md>","pending":{videoId,title,channelTitle,publishedAt,position,total}}
|
|
27
|
+
* {"status":"done","kept":N} queue exhausted — caller runs --compare if kept>0
|
|
28
|
+
* {"status":"compare","comparison":"<md>"}
|
|
29
|
+
* {"status":"no_results"} search returned nothing for the intent
|
|
30
|
+
* {"status":"rate_limited"} transcript fetch blocked (datacenter IP — see docs/warp-proxy.md)
|
|
31
|
+
* {"status":"error","error":"<msg>"} setup/config problem (missing key, etc.)
|
|
32
|
+
*
|
|
33
|
+
* Cache (all throwaway, under DATA_DIR/.cache): search-queue.json (ranked candidates + cursor),
|
|
34
|
+
* search-pending.json (current candidate), search-kept.json (kept summaries = compare corpus).
|
|
35
|
+
*/
|
|
36
|
+
import { readFileSync, writeFileSync, existsSync, mkdirSync } from 'node:fs';
|
|
37
|
+
import { spawn } from 'node:child_process';
|
|
38
|
+
import dotenv from 'dotenv';
|
|
39
|
+
import { chat, getModel } from "./lib/llm.js";
|
|
40
|
+
import { outputLang } from "./lib/config.js";
|
|
41
|
+
import { searchVideos } from "./lib/yt-api.js";
|
|
42
|
+
import { PKG_ROOT, ENV_PATH, CACHE_DIR, SEARCH_QUEUE_FILE, SEARCH_PENDING_FILE, SEARCH_KEPT_FILE, script, } from "./lib/paths.js";
|
|
43
|
+
dotenv.config({ path: ENV_PATH });
|
|
44
|
+
mkdirSync(CACHE_DIR, { recursive: true });
|
|
45
|
+
const RUNTIME = process.execPath;
|
|
46
|
+
const LANG = outputLang();
|
|
47
|
+
const argv = process.argv.slice(2);
|
|
48
|
+
const has = (f) => argv.includes(f);
|
|
49
|
+
const flagVal = (f) => {
|
|
50
|
+
const i = argv.indexOf(f);
|
|
51
|
+
return i !== -1 && argv[i + 1] ? argv[i + 1] : null;
|
|
52
|
+
};
|
|
53
|
+
// First non-flag token is the intent (only on the initial / --reset call).
|
|
54
|
+
const positional = argv.filter((a, i) => !a.startsWith('--') && !(i > 0 && argv[i - 1].startsWith('--') && argv[i - 1] !== '--reset' && argv[i - 1] !== '--keep' && argv[i - 1] !== '--skip' && argv[i - 1] !== '--compare'));
|
|
55
|
+
const intentArg = positional[0] ?? null;
|
|
56
|
+
const RESET = has('--reset');
|
|
57
|
+
const KEEP = has('--keep');
|
|
58
|
+
const SKIP = has('--skip');
|
|
59
|
+
const COMPARE = has('--compare');
|
|
60
|
+
const MAX = Math.max(1, parseInt(flagVal('--max') || '8', 10));
|
|
61
|
+
const QUERIES = Math.max(1, Math.min(3, parseInt(flagVal('--queries') || '2', 10)));
|
|
62
|
+
const SINCE = flagVal('--since');
|
|
63
|
+
const LANGTRACK = flagVal('--lang') || 'auto';
|
|
64
|
+
function emit(obj) {
|
|
65
|
+
process.stdout.write(JSON.stringify(obj));
|
|
66
|
+
process.exit(0);
|
|
67
|
+
}
|
|
68
|
+
const readJSON = (p, fallback) => {
|
|
69
|
+
if (!existsSync(p))
|
|
70
|
+
return fallback;
|
|
71
|
+
try {
|
|
72
|
+
return JSON.parse(readFileSync(p, 'utf8'));
|
|
73
|
+
}
|
|
74
|
+
catch {
|
|
75
|
+
return fallback;
|
|
76
|
+
}
|
|
77
|
+
};
|
|
78
|
+
const writeJSON = (p, v) => writeFileSync(p, JSON.stringify(v));
|
|
79
|
+
const loadQueue = () => readJSON(SEARCH_QUEUE_FILE, null);
|
|
80
|
+
const loadKept = () => readJSON(SEARCH_KEPT_FILE, []);
|
|
81
|
+
/** Fetch a transcript via the sibling script with the launching runtime. Mirrors yt-sweep. */
|
|
82
|
+
function fetchTranscript(videoId) {
|
|
83
|
+
return new Promise((resolve, reject) => {
|
|
84
|
+
const p = spawn(RUNTIME, [script('yt-transcript'), videoId, '--lang', LANGTRACK], { cwd: PKG_ROOT, env: { ...process.env } });
|
|
85
|
+
let stdout = '';
|
|
86
|
+
p.stdout.on('data', d => { stdout += d.toString(); });
|
|
87
|
+
p.stderr.resume();
|
|
88
|
+
p.on('close', code => resolve({ stdout, code: code ?? 1 }));
|
|
89
|
+
p.on('error', reject);
|
|
90
|
+
});
|
|
91
|
+
}
|
|
92
|
+
/** Strip ```fences``` and slice the outermost JSON array from an LLM reply. */
|
|
93
|
+
function parseJsonArray(out) {
|
|
94
|
+
const start = out.indexOf('['), end = out.lastIndexOf(']');
|
|
95
|
+
if (start === -1 || end === -1 || end < start)
|
|
96
|
+
return null;
|
|
97
|
+
try {
|
|
98
|
+
return JSON.parse(out.slice(start, end + 1));
|
|
99
|
+
}
|
|
100
|
+
catch {
|
|
101
|
+
return null;
|
|
102
|
+
}
|
|
103
|
+
}
|
|
104
|
+
// ---------- stage 1: intent → search queries ----------
|
|
105
|
+
async function expandQueries(intent) {
|
|
106
|
+
const prompt = `A user wants to research a topic on YouTube. Turn their intent into up to ${QUERIES} effective YouTube search queries (short, keyword-rich, the way people actually search). Cover slightly different angles if useful. Use the language the topic is most discussed in (usually English for tech).
|
|
107
|
+
|
|
108
|
+
Intent: "${intent}"
|
|
109
|
+
|
|
110
|
+
Output ONLY a raw JSON array of strings, no fences, no commentary. Example: ["query one","query two"]`;
|
|
111
|
+
try {
|
|
112
|
+
const out = await chat(prompt, { system: 'You output ONLY a raw JSON array of search-query strings.', temperature: 0.4 });
|
|
113
|
+
const arr = parseJsonArray(out);
|
|
114
|
+
const qs = (arr ?? []).filter((s) => typeof s === 'string' && s.trim().length > 0).slice(0, QUERIES);
|
|
115
|
+
return qs.length ? qs : [intent];
|
|
116
|
+
}
|
|
117
|
+
catch {
|
|
118
|
+
return [intent]; // expansion is best-effort; fall back to the raw intent
|
|
119
|
+
}
|
|
120
|
+
}
|
|
121
|
+
// ---------- stage 3: re-rank candidates against intent (metadata only) ----------
|
|
122
|
+
async function rerank(intent, hits) {
|
|
123
|
+
if (hits.length === 0)
|
|
124
|
+
return [];
|
|
125
|
+
const compact = hits.map(h => ({ id: h.videoId, title: h.title, channel: h.channelTitle, published: h.publishedAt, desc: (h.description || '').slice(0, 280) }));
|
|
126
|
+
const prompt = `Rank these YouTube videos by how well they serve the user's intent. Judge on title + channel + description only (no transcripts). Drop clearly off-topic, clickbait, or duplicate-angle results.
|
|
127
|
+
|
|
128
|
+
Intent: "${intent}"
|
|
129
|
+
|
|
130
|
+
Candidates:
|
|
131
|
+
${JSON.stringify(compact)}
|
|
132
|
+
|
|
133
|
+
Output ONLY a raw JSON array, best first, no fences:
|
|
134
|
+
[{"id":"VIDEO_ID","keep":true,"score":0-100,"reason":"max 12 words"},...]
|
|
135
|
+
Set keep=false for anything not worth the user's time.`;
|
|
136
|
+
try {
|
|
137
|
+
const out = await chat(prompt, { system: 'You output ONLY a raw JSON array as instructed.', temperature: 0 });
|
|
138
|
+
const arr = parseJsonArray(out);
|
|
139
|
+
if (!arr)
|
|
140
|
+
return hits.map(h => ({ ...h })); // fall back: keep all, original order
|
|
141
|
+
const byId = new Map(hits.map(h => [h.videoId, h]));
|
|
142
|
+
const ranked = [];
|
|
143
|
+
for (const r of arr) {
|
|
144
|
+
if (!r || r.keep === false)
|
|
145
|
+
continue;
|
|
146
|
+
const h = byId.get(r.id);
|
|
147
|
+
if (h)
|
|
148
|
+
ranked.push({ ...h, score: typeof r.score === 'number' ? r.score : undefined, reason: r.reason });
|
|
149
|
+
}
|
|
150
|
+
return ranked.length ? ranked : hits.map(h => ({ ...h }));
|
|
151
|
+
}
|
|
152
|
+
catch {
|
|
153
|
+
return hits.map(h => ({ ...h }));
|
|
154
|
+
}
|
|
155
|
+
}
|
|
156
|
+
// ---------- mega-summary for one candidate (the triage artifact) ----------
|
|
157
|
+
async function megaSummary(c, transcript, intent) {
|
|
158
|
+
const prompt = `Summarize this YouTube video in ${LANG} for a user researching: "${intent}". Make it a RICH, standalone summary they can decide on and that will later feed a cross-video comparison — OR return 'OFFTOPIC: <reason>' if the transcript clearly doesn't serve the intent.
|
|
159
|
+
|
|
160
|
+
Video:
|
|
161
|
+
- title: ${c.title}
|
|
162
|
+
- channel: ${c.channelTitle}
|
|
163
|
+
- published: ${c.publishedAt}
|
|
164
|
+
- url: https://youtube.com/watch?v=${c.videoId}
|
|
165
|
+
|
|
166
|
+
Transcript:
|
|
167
|
+
${transcript}
|
|
168
|
+
|
|
169
|
+
If on-topic, write:
|
|
170
|
+
- Header: ### ${c.channelTitle} — "${c.title}"
|
|
171
|
+
- Subtitle: _${c.publishedAt} · https://youtube.com/watch?v=${c.videoId}_
|
|
172
|
+
- One sentence on relevance to the intent
|
|
173
|
+
- 3-6 numbered thematic sections × 2-4 sentences, concrete: which options/tools are discussed, the criteria, the author's verdict and reasoning. Pull out anything directly comparable (names, pros/cons, recommendations).
|
|
174
|
+
- A short "Bottom line for the intent" line
|
|
175
|
+
- At most 5-8 short quotes from the transcript. No timestamps.
|
|
176
|
+
|
|
177
|
+
Language: natural ${LANG}; foreign words only for proper nouns or established technical terms.
|
|
178
|
+
Output ONLY the summary OR 'OFFTOPIC: <reason>'. No preamble.`;
|
|
179
|
+
return chat(prompt, {
|
|
180
|
+
system: `You are a research-grade video summarizer writing in ${LANG}. Output only the summary or 'OFFTOPIC: <reason>'.`,
|
|
181
|
+
model: getModel(),
|
|
182
|
+
});
|
|
183
|
+
}
|
|
184
|
+
// ---------- stage 6: comparison across kept summaries ----------
|
|
185
|
+
async function synthesizeComparison(intent, kept) {
|
|
186
|
+
const corpus = kept.map((k, i) => `--- VIDEO ${i + 1}: ${k.channelTitle} — "${k.title}" (${k.publishedAt})\nhttps://youtube.com/watch?v=${k.videoId}\n${k.summary}`).join('\n\n');
|
|
187
|
+
const prompt = `The user researched "${intent}" and kept ${kept.length} YouTube video summaries below. Synthesize a single comparison in ${LANG} that actually helps them decide.
|
|
188
|
+
|
|
189
|
+
${corpus}
|
|
190
|
+
|
|
191
|
+
Write:
|
|
192
|
+
- One-paragraph bottom line answering the intent directly.
|
|
193
|
+
- A comparison of the concrete options/tools across the videos (a Markdown table when it fits: option · who recommends it · pros · cons · best for).
|
|
194
|
+
- Consensus vs disagreements between the sources.
|
|
195
|
+
- A final recommendation with the reasoning, and who it's for.
|
|
196
|
+
|
|
197
|
+
Language: natural ${LANG}; foreign words only for proper nouns or established technical terms. Cite videos as [1], [2]… matching the order above. Output only the comparison.`;
|
|
198
|
+
return chat(prompt, { system: `You synthesize a decision-grade comparison in ${LANG}. No preamble.`, model: getModel() });
|
|
199
|
+
}
|
|
200
|
+
// ---------- lazy yield: advance to the next candidate that has a transcript ----------
|
|
201
|
+
async function yieldNext(queue) {
|
|
202
|
+
while (queue.cursor < queue.candidates.length) {
|
|
203
|
+
const c = queue.candidates[queue.cursor];
|
|
204
|
+
const t = await fetchTranscript(c.videoId);
|
|
205
|
+
if (t.code === 2)
|
|
206
|
+
emit({ status: 'rate_limited' }); // blocked IP — stop, don't advance
|
|
207
|
+
if (t.code !== 0 || !t.stdout.trim()) { // no transcript → auto-skip
|
|
208
|
+
queue.cursor++;
|
|
209
|
+
writeJSON(SEARCH_QUEUE_FILE, queue);
|
|
210
|
+
continue;
|
|
211
|
+
}
|
|
212
|
+
const summary = await megaSummary(c, t.stdout, queue.intent);
|
|
213
|
+
if (summary.startsWith('OFFTOPIC:')) { // re-rank missed it → auto-skip
|
|
214
|
+
queue.cursor++;
|
|
215
|
+
writeJSON(SEARCH_QUEUE_FILE, queue);
|
|
216
|
+
continue;
|
|
217
|
+
}
|
|
218
|
+
const pending = { videoId: c.videoId, title: c.title, channelTitle: c.channelTitle, publishedAt: c.publishedAt, summary };
|
|
219
|
+
writeJSON(SEARCH_PENDING_FILE, pending);
|
|
220
|
+
emit({
|
|
221
|
+
status: 'decision_needed',
|
|
222
|
+
summary,
|
|
223
|
+
pending: { videoId: c.videoId, title: c.title, channelTitle: c.channelTitle, publishedAt: c.publishedAt, position: queue.cursor + 1, total: queue.candidates.length },
|
|
224
|
+
});
|
|
225
|
+
}
|
|
226
|
+
emit({ status: 'done', kept: loadKept().length });
|
|
227
|
+
}
|
|
228
|
+
// ---------- main ----------
|
|
229
|
+
async function main() {
|
|
230
|
+
if (!process.env.YT_BRIEFING_YOUTUBE_API_KEY && (RESET || (intentArg && !loadQueue()))) {
|
|
231
|
+
emit({ status: 'error', error: 'YT_BRIEFING_YOUTUBE_API_KEY is not set — add a YouTube Data API v3 key to .yt-briefing/.env (see README → setup / .env.example).' });
|
|
232
|
+
}
|
|
233
|
+
// --compare: synthesize from kept summaries.
|
|
234
|
+
if (COMPARE) {
|
|
235
|
+
const queue = loadQueue();
|
|
236
|
+
const kept = loadKept();
|
|
237
|
+
if (kept.length === 0)
|
|
238
|
+
emit({ status: 'done', kept: 0 });
|
|
239
|
+
const comparison = await synthesizeComparison(queue?.intent ?? '', kept);
|
|
240
|
+
emit({ status: 'compare', comparison });
|
|
241
|
+
}
|
|
242
|
+
// --keep / --skip: record decision on the pending candidate, then advance + yield next.
|
|
243
|
+
if (KEEP || SKIP) {
|
|
244
|
+
const queue = loadQueue();
|
|
245
|
+
if (!queue)
|
|
246
|
+
emit({ status: 'error', error: 'No active search. Start one: yt-search "<intent>".' });
|
|
247
|
+
if (KEEP) {
|
|
248
|
+
const pending = readJSON(SEARCH_PENDING_FILE, null);
|
|
249
|
+
if (pending) {
|
|
250
|
+
const kept = loadKept();
|
|
251
|
+
kept.push(pending);
|
|
252
|
+
writeJSON(SEARCH_KEPT_FILE, kept);
|
|
253
|
+
}
|
|
254
|
+
}
|
|
255
|
+
queue.cursor++;
|
|
256
|
+
writeJSON(SEARCH_QUEUE_FILE, queue);
|
|
257
|
+
await yieldNext(queue);
|
|
258
|
+
}
|
|
259
|
+
// Resume an in-progress search (bare call, same intent) without rebuilding.
|
|
260
|
+
const existing = loadQueue();
|
|
261
|
+
if (existing && !RESET && (!intentArg || intentArg === existing.intent)) {
|
|
262
|
+
await yieldNext(existing);
|
|
263
|
+
}
|
|
264
|
+
// Fresh search (new intent or --reset).
|
|
265
|
+
if (!intentArg)
|
|
266
|
+
emit({ status: 'error', error: 'Provide an intent: yt-search "<what to research>".' });
|
|
267
|
+
const queries = await expandQueries(intentArg);
|
|
268
|
+
const seen = new Set();
|
|
269
|
+
const hits = [];
|
|
270
|
+
for (const q of queries) {
|
|
271
|
+
let batch = [];
|
|
272
|
+
try {
|
|
273
|
+
batch = await searchVideos(q, { maxResults: 10, since: SINCE });
|
|
274
|
+
}
|
|
275
|
+
catch (e) {
|
|
276
|
+
emit({ status: 'error', error: e.message });
|
|
277
|
+
}
|
|
278
|
+
for (const h of batch)
|
|
279
|
+
if (!seen.has(h.videoId)) {
|
|
280
|
+
seen.add(h.videoId);
|
|
281
|
+
hits.push(h);
|
|
282
|
+
}
|
|
283
|
+
}
|
|
284
|
+
if (hits.length === 0)
|
|
285
|
+
emit({ status: 'no_results' });
|
|
286
|
+
const ranked = (await rerank(intentArg, hits)).slice(0, MAX);
|
|
287
|
+
if (ranked.length === 0)
|
|
288
|
+
emit({ status: 'no_results' });
|
|
289
|
+
const queue = { built_at: new Date().toISOString(), intent: intentArg, candidates: ranked, cursor: 0 };
|
|
290
|
+
writeJSON(SEARCH_QUEUE_FILE, queue);
|
|
291
|
+
writeJSON(SEARCH_KEPT_FILE, []); // fresh search → fresh compare corpus
|
|
292
|
+
await yieldNext(queue);
|
|
293
|
+
}
|
|
294
|
+
main().catch(err => emit({ status: 'error', error: err.message }));
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "yt-briefing",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.5.0",
|
|
4
4
|
"description": "A self-learning YouTube briefing engine: it sweeps the channels you follow, filters noise in two stages (title, then transcript), summarizes the rest in your language, and adapts to your ratings — one video at a time.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|