yt-briefing 0.4.0 → 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/yt-search/SKILL.md +57 -0
- package/README.md +33 -12
- package/dist/bootstrap.js +5 -4
- package/dist/cli.js +3 -1
- package/dist/install-skill.js +6 -5
- package/dist/lib/paths.js +4 -0
- package/dist/lib/skill-install.js +12 -2
- package/dist/lib/yt-api.js +1 -1
- package/dist/yt-search.js +287 -0
- package/package.json +1 -1
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: yt-search
|
|
3
|
+
description: Search WITHIN one YouTube channel by intent — name a channel and what you're after; the engine lists that channel's uploads, ranks them against your intent (metadata only, no transcript yet), then lazily yields ONE matching video at a time with a rich summary. You keep or skip each; at the end it synthesizes a comparison from everything you kept. Channel-scoped, not whole-YouTube. Same transcript engine + proxy as /yt; lazy on purpose (no transcript bursts → no IP block). Summaries and prompts use the language chosen at onboarding.
|
|
4
|
+
argument-hint: A channel (@handle or URL) and a descriptive intent, e.g. "@t3dotgg which terminal for AI coding". Optional --max N, --scan N, --since YYYY-MM-DD.
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## How it works
|
|
8
|
+
|
|
9
|
+
`src/yt-search.ts` is the whole engine: list one channel's uploads (cheap — `playlistItems`, ~1 quota unit/page; NOT `search.list`) → re-rank them against your intent on metadata only (title/description, no transcript) → **lazy** one-candidate-at-a-time yield with a rich summary → record keep/skip → on demand synthesize a comparison from everything kept. **Channel-scoped on purpose** — you choose where to look; it does NOT search all of YouTube. Matching is descriptive: the LLM filters the channel's videos by intent. This skill is a thin loop — paste the summary, collect keep/skip, show the final comparison.
|
|
10
|
+
|
|
11
|
+
**Lazy on purpose:** one transcript per step, never a burst — a burst looks like scraping and gets the IP blocked (same reason `/yt` is lazy). Run the engine bare — stdout is a single JSON line, stderr empty; never redirect.
|
|
12
|
+
|
|
13
|
+
## Inputs
|
|
14
|
+
|
|
15
|
+
The user gives a **channel** (`@handle` or a channel URL) and an **intent** (what to look for). Pass the channel via `--channel` and the intent as the quoted positional. If the user names a channel but no clear intent (or vice versa), ask for the missing half before running.
|
|
16
|
+
|
|
17
|
+
## Language
|
|
18
|
+
|
|
19
|
+
Read `data/config.json` → `output_lang` once at the start. Phrase the question text and option descriptions in that language. The two button labels stay the short English words `Keep` / `Skip`. Summaries and the final comparison are already written in `output_lang` by the engine — paste them verbatim.
|
|
20
|
+
|
|
21
|
+
## Loop
|
|
22
|
+
|
|
23
|
+
```
|
|
24
|
+
out = JSON.parse(`bun run src/yt-search.ts "<intent>" --channel <@handle|url> --reset`) // first call
|
|
25
|
+
while true:
|
|
26
|
+
out.status:
|
|
27
|
+
"error" → show out.error verbatim, stop
|
|
28
|
+
"no_results" → tell the user nothing in that channel matched, stop
|
|
29
|
+
"rate_limited" → transcript fetch blocked (datacenter IP) — tell the user, stop; recovery in README.md → Running on a VPS
|
|
30
|
+
"decision_needed" → steps A–C
|
|
31
|
+
"done" → step D
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
**On `decision_needed`:**
|
|
35
|
+
|
|
36
|
+
- **A.** Take `out.summary` (markdown) and `out.pending` (`{videoId,title,channelTitle,publishedAt,position,total}`).
|
|
37
|
+
- **B.** In the SAME turn, as your chat text (NOT command output — the UI doesn't show it), paste `summary` **verbatim** — no paraphrase, no shortening. Optionally prefix one line like `Wynik {position}/{total}`. The user must see it before the popup.
|
|
38
|
+
- **C.** In the same message call `AskUserQuestion` — 1 call, 1 question, phrased in `output_lang`:
|
|
39
|
+
- Question e.g. "Brać pod uwagę w porównaniu?" with two options: **Keep** = include this video in the final comparison; **Skip** = drop it.
|
|
40
|
+
- **Other** is the stop channel: if the user types `stop` (case-insensitive, trimmed) or dismisses the popup (✕) → **end the loop early** and go to step D (compare what's kept so far).
|
|
41
|
+
- Then act:
|
|
42
|
+
- Keep → `bun run src/yt-search.ts --keep`
|
|
43
|
+
- Skip → `bun run src/yt-search.ts --skip`
|
|
44
|
+
- stop/dismissed → `bun run src/yt-search.ts --compare` (skip straight to D)
|
|
45
|
+
- The script reads the pending candidate from cache; pass only `--keep` / `--skip` / `--compare` (no channel/intent again). Its JSON becomes the next `out` — back to the top of the loop.
|
|
46
|
+
|
|
47
|
+
**On `done` (step D):**
|
|
48
|
+
|
|
49
|
+
- If `out.kept > 0` → run `out = JSON.parse(\`bun run src/yt-search.ts --compare\`)`; when it returns `status:"compare"`, paste `out.comparison` **verbatim** as your chat text (the artifact — a decision-grade comparison in `output_lang`). Stop.
|
|
50
|
+
- If `out.kept == 0` → tell the user nothing was kept, so there's nothing to compare. Stop.
|
|
51
|
+
|
|
52
|
+
## Rules
|
|
53
|
+
|
|
54
|
+
- **Verbatim:** paste `summary` and `comparison` exactly as returned; never paste a raw transcript.
|
|
55
|
+
- **Language:** question text + option descriptions follow `output_lang`; button labels stay `Keep` / `Skip`.
|
|
56
|
+
- **Scope:** one channel per search. Listing is cheap; the cost is the lazy transcript fetches, so let the user keep/skip rather than pulling everything. A bare resume (no `--reset`) continues the same ranked queue. `--scan N` (default 50) caps how many recent uploads are considered; `--since` widens by date.
|
|
57
|
+
- **Stateless triage:** independent of `/yt` (no channel profiles, no ratings written). For the recurring multi-channel briefing use `/yt`; for one known video use `/yt-transcribe`.
|
package/README.md
CHANGED
|
@@ -24,7 +24,8 @@ the last one left off.
|
|
|
24
24
|
You'll need Node 18+ or Bun, a YouTube Data API v3 key, an LLM key (a
|
|
25
25
|
[free Gemini key](https://aistudio.google.com/apikey) works, see [Providers](#providers)), and
|
|
26
26
|
a tool that runs skills: [Claude Code](https://claude.com/claude-code),
|
|
27
|
-
[Cursor](https://cursor.com), or anything else
|
|
27
|
+
[Cursor](https://cursor.com), [Codex](https://developers.openai.com/codex), or anything else
|
|
28
|
+
that loads the standard `SKILL.md` (Agent Skills — 30+ agents).
|
|
28
29
|
|
|
29
30
|
1. Install yt-dlp (it pulls the subtitles):
|
|
30
31
|
|
|
@@ -61,12 +62,6 @@ npx yt-briefing remove @handle # also deletes its l
|
|
|
61
62
|
npx yt-briefing list # show the current list
|
|
62
63
|
```
|
|
63
64
|
|
|
64
|
-
## Run it
|
|
65
|
-
|
|
66
|
-
Open your project in Claude Code or Cursor and run `/yt`. If it's not listed, start a fresh
|
|
67
|
-
session. To install the skills again for another tool or project, run
|
|
68
|
-
`npx yt-briefing install-skill` (it installs both `/yt` and `/yt-transcribe`).
|
|
69
|
-
|
|
70
65
|
## One-off: transcribe a single video
|
|
71
66
|
|
|
72
67
|
Just want one video summarized — no channels, no queue, no rating? Run `/yt-transcribe` and
|
|
@@ -74,8 +69,7 @@ paste a URL or video ID. It pulls that video's transcript and writes a journalis
|
|
|
74
69
|
summary in the language you chose at setup (the same `output_lang` as `/yt`). Want a one-off in
|
|
75
70
|
another language? Just say so when you run it (e.g. `/yt-transcribe <url> in German`) — it
|
|
76
71
|
won't change your setup. `--lang pl|en` is separate — it picks which caption track to fetch,
|
|
77
|
-
not the summary language.
|
|
78
|
-
server it benefits from the same [WARP proxy](#running-on-a-vps).
|
|
72
|
+
not the summary language.
|
|
79
73
|
|
|
80
74
|
The skill is installed alongside `/yt` by `init` / `install-skill`. From the plain CLI:
|
|
81
75
|
|
|
@@ -83,6 +77,33 @@ The skill is installed alongside `/yt` by `init` / `install-skill`. From the pla
|
|
|
83
77
|
npx yt-briefing transcribe <url-or-id> --lang auto # prints the transcript to stdout
|
|
84
78
|
```
|
|
85
79
|
|
|
80
|
+
## Search within a channel
|
|
81
|
+
|
|
82
|
+
Following a creator and want to mine *their* videos for something specific — and get an actual
|
|
83
|
+
comparison, not a pile of links? Run `/yt-search`, name a **channel** and your **intent**, e.g.
|
|
84
|
+
`@t3dotgg which terminal for AI coding`. It lists that channel's uploads, ranks them against your
|
|
85
|
+
intent (titles/descriptions — no transcripts yet, matching is **descriptive, not exact keywords**),
|
|
86
|
+
then hands you **one matching video at a time** with a rich summary; you **Keep** or **Skip** each.
|
|
87
|
+
At the end it synthesizes a **comparison** from everything you kept.
|
|
88
|
+
|
|
89
|
+
It's channel-scoped on purpose (you choose where to look) and lazy — one transcript per step, never
|
|
90
|
+
a burst (a burst gets your IP blocked, same as `/yt`). Summaries and the comparison use the
|
|
91
|
+
language you chose at setup.
|
|
92
|
+
|
|
93
|
+
```bash
|
|
94
|
+
yt-briefing search "<intent>" --channel <@handle|url> [--max N] [--scan N] [--since YYYY-MM-DD]
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
> Listing a channel's uploads is cheap (`playlistItems`, ~1 quota unit per 50 videos — not the
|
|
98
|
+
> 100-unit `search.list`). `--scan` (default 50) caps how many recent uploads are considered;
|
|
99
|
+
> `--since` widens by date.
|
|
100
|
+
|
|
101
|
+
## Run it
|
|
102
|
+
|
|
103
|
+
Open your project in Claude Code or Cursor and run `/yt`. If it's not listed, start a fresh
|
|
104
|
+
session. To install the skills again for another tool or project, run
|
|
105
|
+
`npx yt-briefing install-skill` (it installs `/yt`, `/yt-transcribe`, and `/yt-search`).
|
|
106
|
+
|
|
86
107
|
## Providers
|
|
87
108
|
|
|
88
109
|
Any OpenAI-compatible endpoint works. Gemini 2.5 Flash is the easy default. It's fast, cheap,
|
|
@@ -110,9 +131,9 @@ summarizing the next video while you rate the current one, so the following step
|
|
|
110
131
|
ready with no wait. An agent's turn-by-turn loop cannot prefetch like that, and every step pays
|
|
111
132
|
its own cold start, which adds up across a whole queue.
|
|
112
133
|
|
|
113
|
-
Compatibility. A standard API plus a
|
|
114
|
-
Cursor, any other tool
|
|
115
|
-
to that one tool and one model.
|
|
134
|
+
Compatibility. A standard API plus a standard `SKILL.md` means one engine runs everywhere: Claude
|
|
135
|
+
Code, Cursor, Codex, any other Agent-Skills-compatible tool, or the plain CLI. A tool-native
|
|
136
|
+
approach would tie it to that one tool and one model.
|
|
116
137
|
|
|
117
138
|
## Why one transcript at a time
|
|
118
139
|
|
package/dist/bootstrap.js
CHANGED
|
@@ -108,10 +108,11 @@ function main() {
|
|
|
108
108
|
// 6. Coding agent ------------------------------------------------------------
|
|
109
109
|
// Place the skill INTO THIS PROJECT (the package folder you open in the agent) — never a
|
|
110
110
|
// home-global dir (that's the npm -g antipattern: machine-wide, invisible, easy to forget).
|
|
111
|
-
//
|
|
112
|
-
//
|
|
113
|
-
|
|
114
|
-
console.log('
|
|
111
|
+
// SKILL.md is the cross-agent standard, so the shipped skill runs in any compatible agent —
|
|
112
|
+
// we just install it into that agent's skills dir (.claude/skills, .cursor/skills, .codex/skills).
|
|
113
|
+
// 1/2/3 = known agents; 4 = any other compatible agent (a project folder you name).
|
|
114
|
+
console.log('\n 6) Which agent will you run /yt in? (it ships a standard Agent Skill — any compatible agent works)');
|
|
115
|
+
console.log(' 1) Claude Code 2) Cursor 3) Codex 4) Custom folder (any other agent)\n');
|
|
115
116
|
const agentKey = ask(' Your agent', '1');
|
|
116
117
|
// For a custom target, ask the folder now (keeps all prompts in the interactive block).
|
|
117
118
|
const customDir = AGENTS[agentKey] ? '' : ask(' Skills folder to install into', customSkillsRootDefault());
|
package/dist/cli.js
CHANGED
|
@@ -12,6 +12,7 @@
|
|
|
12
12
|
* yt-briefing sweep [--reset] advance one step; prints a JSON status line
|
|
13
13
|
* yt-briefing rate --rating 0|1 [...] record a rating for the pending video
|
|
14
14
|
* yt-briefing transcribe <url|id> print a single video's transcript
|
|
15
|
+
* yt-briefing search "<intent>" --channel <@handle|url> search within one channel → triage → compare
|
|
15
16
|
*/
|
|
16
17
|
import { spawnSync } from 'node:child_process';
|
|
17
18
|
import { script } from "./lib/paths.js";
|
|
@@ -24,6 +25,7 @@ const TARGETS = {
|
|
|
24
25
|
sweep: 'yt-sweep',
|
|
25
26
|
rate: 'yt-rating',
|
|
26
27
|
transcribe: 'yt-transcript',
|
|
28
|
+
search: 'yt-search',
|
|
27
29
|
};
|
|
28
30
|
const CHANNEL_ACTIONS = new Set(['add', 'remove', 'list']);
|
|
29
31
|
let argv = null;
|
|
@@ -32,7 +34,7 @@ if (cmd && CHANNEL_ACTIONS.has(cmd))
|
|
|
32
34
|
else if (cmd && TARGETS[cmd])
|
|
33
35
|
argv = [script(TARGETS[cmd]), ...rest];
|
|
34
36
|
if (!argv) {
|
|
35
|
-
console.error('Usage: yt-briefing <init|install-skill|add|remove|list|sweep|rate|transcribe> [args...]');
|
|
37
|
+
console.error('Usage: yt-briefing <init|install-skill|add|remove|list|sweep|rate|transcribe|search> [args...]');
|
|
36
38
|
process.exit(cmd ? 1 : 0);
|
|
37
39
|
}
|
|
38
40
|
const res = spawnSync(process.execPath, argv, { stdio: 'inherit' });
|
package/dist/install-skill.js
CHANGED
|
@@ -12,9 +12,9 @@
|
|
|
12
12
|
* standalone command is for re-installing, a different project, or a second agent. There is
|
|
13
13
|
* deliberately no home-global install — the skills live with the project that uses them.
|
|
14
14
|
*
|
|
15
|
-
*
|
|
16
|
-
*
|
|
17
|
-
*
|
|
15
|
+
* SKILL.md is the cross-agent Agent Skills standard, so the shipped skills run in any compatible
|
|
16
|
+
* agent (Claude Code, Cursor, Codex, and 30+ others); this command just (re)places them in the
|
|
17
|
+
* skills dir of whichever agent you pick.
|
|
18
18
|
*/
|
|
19
19
|
import { AGENTS, installSkills, projectSkillsRoot, customSkillsRootDefault, isPackageDevCwd } from "./lib/skill-install.js";
|
|
20
20
|
import { question } from "./lib/prompt.js";
|
|
@@ -28,8 +28,9 @@ function done(targets) {
|
|
|
28
28
|
// 1) which agent → which skills subdir
|
|
29
29
|
console.log('\n Install the /yt + /yt-transcribe skills — which agent?\n');
|
|
30
30
|
console.log(' 1) Claude Code');
|
|
31
|
-
console.log(' 2) Cursor
|
|
32
|
-
console.log(' 3)
|
|
31
|
+
console.log(' 2) Cursor');
|
|
32
|
+
console.log(' 3) Codex');
|
|
33
|
+
console.log(' 4) Custom folder (any other compatible agent)\n');
|
|
33
34
|
const agentKey = ask(' Agent', '1');
|
|
34
35
|
const agent = AGENTS[agentKey];
|
|
35
36
|
// 3) Custom — write the skills straight into a skills root the user names (their agent's dir).
|
package/dist/lib/paths.js
CHANGED
|
@@ -47,6 +47,10 @@ export const REST_FILE = join(CACHE_DIR, 'queue-rest.json');
|
|
|
47
47
|
export const PENDING_FILE = join(CACHE_DIR, 'pending.json');
|
|
48
48
|
export const PREFETCH_FILE = join(CACHE_DIR, 'prefetch.json');
|
|
49
49
|
export const LOG_FILE = join(CACHE_DIR, 'sweep.log');
|
|
50
|
+
// /yt-search (ad-hoc topic search → lazy triage → compare). All throwaway, under .cache/.
|
|
51
|
+
export const SEARCH_QUEUE_FILE = join(CACHE_DIR, 'search-queue.json'); // ranked candidates + cursor
|
|
52
|
+
export const SEARCH_PENDING_FILE = join(CACHE_DIR, 'search-pending.json'); // current candidate awaiting keep/skip
|
|
53
|
+
export const SEARCH_KEPT_FILE = join(CACHE_DIR, 'search-kept.json'); // kept mega-summaries → compare corpus
|
|
50
54
|
/** Absolute path to a channel profile from its slug. */
|
|
51
55
|
export const profilePath = (slug) => join(CHANNELS_DIR, `${slug}.md`);
|
|
52
56
|
/**
|
|
@@ -26,7 +26,7 @@ import { PKG_ROOT, DATA_DIR } from "./paths.js";
|
|
|
26
26
|
/** Compiled output dir — what a baked (dist) skill command points the runtime at. */
|
|
27
27
|
const DIST_DIR = join(PKG_ROOT, 'dist');
|
|
28
28
|
/** The skills this package ships — each lives at `.claude/skills/<name>/SKILL.md`. */
|
|
29
|
-
export const SKILLS = ['yt', 'yt-transcribe'];
|
|
29
|
+
export const SKILLS = ['yt', 'yt-transcribe', 'yt-search'];
|
|
30
30
|
/** True when the installer itself is running under Bun (vs plain Node). */
|
|
31
31
|
export const isBun = process.versions.bun != null;
|
|
32
32
|
/**
|
|
@@ -39,10 +39,19 @@ export const isBun = process.versions.bun != null;
|
|
|
39
39
|
export const isPackageDevCwd = () => isBun && resolve(process.cwd()) === PKG_ROOT;
|
|
40
40
|
/** Source path of a shipped skill's SKILL.md, by skill name. */
|
|
41
41
|
export const skillSource = (name) => join(PKG_ROOT, '.claude', 'skills', name, 'SKILL.md');
|
|
42
|
-
/**
|
|
42
|
+
/**
|
|
43
|
+
* Agent key → display name + the skills ROOT directory it scans (skills install under it).
|
|
44
|
+
*
|
|
45
|
+
* SKILL.md is the cross-agent Agent Skills standard (Anthropic, Dec 2025), now read by 30+
|
|
46
|
+
* tools that each scan their own `<agent-home>/skills/` dir. We only need the right directory
|
|
47
|
+
* per agent — the shipped SKILL.md works unmodified in all of them. The "custom folder" picker
|
|
48
|
+
* option (no AGENTS entry) covers every other compatible agent (Gemini CLI, Copilot, Windsurf…)
|
|
49
|
+
* and defaults to the neutral `.agents/skills/` location.
|
|
50
|
+
*/
|
|
43
51
|
export const AGENTS = {
|
|
44
52
|
'1': { name: 'Claude Code', sub: join('.claude', 'skills') },
|
|
45
53
|
'2': { name: 'Cursor', sub: join('.cursor', 'skills') },
|
|
54
|
+
'3': { name: 'Codex', sub: join('.codex', 'skills') },
|
|
46
55
|
};
|
|
47
56
|
/**
|
|
48
57
|
* One shipped skill's SKILL.md. `dist=false` (default) returns it verbatim — the
|
|
@@ -64,6 +73,7 @@ export function skillBody(name, dist = false) {
|
|
|
64
73
|
.replace(/bun run src\/yt-sweep\.ts/g, cmd('yt-sweep'))
|
|
65
74
|
.replace(/bun run src\/yt-rating\.ts/g, cmd('yt-rating'))
|
|
66
75
|
.replace(/bun run src\/yt-transcript\.ts/g, cmd('yt-transcript'))
|
|
76
|
+
.replace(/bun run src\/yt-search\.ts/g, cmd('yt-search'))
|
|
67
77
|
.replace(/data\//g, DATA_DIR + '/');
|
|
68
78
|
}
|
|
69
79
|
/**
|
package/dist/lib/yt-api.js
CHANGED
|
@@ -58,7 +58,7 @@ async function listUploads(playlistId, since, maxCount) {
|
|
|
58
58
|
hitOld = true;
|
|
59
59
|
break;
|
|
60
60
|
}
|
|
61
|
-
videos.push({ videoId: item.snippet.resourceId.videoId, title: item.snippet.title, publishedAt });
|
|
61
|
+
videos.push({ videoId: item.snippet.resourceId.videoId, title: item.snippet.title, publishedAt, description: item.snippet.description ?? '' });
|
|
62
62
|
if (maxCount !== null && videos.length >= maxCount)
|
|
63
63
|
return videos;
|
|
64
64
|
}
|
|
@@ -0,0 +1,287 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
/**
|
|
3
|
+
* yt-search.ts — search WITHIN one channel by intent, then lazy triage → comparison. The third
|
|
4
|
+
* yt-briefing mode, sibling to the channel briefing (yt-sweep) and one-shot transcribe.
|
|
5
|
+
*
|
|
6
|
+
* You point it at a channel and describe what you're after ("which terminal does he recommend
|
|
7
|
+
* for AI coding"); the engine:
|
|
8
|
+
* 1. lists that channel's uploads (cheap — playlistItems, 1 quota unit/page; NOT search.list),
|
|
9
|
+
* 2. re-ranks them against your intent on metadata only — title/description, NO transcript yet,
|
|
10
|
+
* 3. yields ONE matching video at a time with a rich summary, lazily — never a burst of
|
|
11
|
+
* transcript fetches (a burst looks like scraping and gets the IP blocked),
|
|
12
|
+
* 4. records your keep/skip decision; kept summaries accumulate in a cache,
|
|
13
|
+
* 5. on demand synthesizes a comparison across everything you kept.
|
|
14
|
+
*
|
|
15
|
+
* Channel-scoped on purpose: you choose where to look. Matching is descriptive — the LLM filters
|
|
16
|
+
* the channel's videos by intent (no exact-keyword needed).
|
|
17
|
+
*
|
|
18
|
+
* Usage (the skill / CLI drives these; one JSON line per call):
|
|
19
|
+
* yt-search "<intent>" --channel <@handle|url> [--reset] [--max N] [--scan N] [--since DATE] [--lang auto]
|
|
20
|
+
* yt-search --keep record the pending candidate, advance, yield next
|
|
21
|
+
* yt-search --skip drop the pending candidate, advance, yield next
|
|
22
|
+
* yt-search --compare synthesize a comparison from everything kept
|
|
23
|
+
*
|
|
24
|
+
* Output statuses:
|
|
25
|
+
* {"status":"decision_needed","summary":"<md>","pending":{videoId,title,channelTitle,publishedAt,position,total}}
|
|
26
|
+
* {"status":"done","kept":N} queue exhausted — caller runs --compare if kept>0
|
|
27
|
+
* {"status":"compare","comparison":"<md>"}
|
|
28
|
+
* {"status":"no_results"} the channel has no videos matching the intent
|
|
29
|
+
* {"status":"rate_limited"} transcript fetch blocked (datacenter IP — see docs/warp-proxy.md)
|
|
30
|
+
* {"status":"error","error":"<msg>"} setup/config problem (missing key, missing channel, …)
|
|
31
|
+
*
|
|
32
|
+
* Cache (all throwaway, under DATA_DIR/.cache): search-queue.json (ranked candidates + cursor),
|
|
33
|
+
* search-pending.json (current candidate), search-kept.json (kept summaries = compare corpus).
|
|
34
|
+
*/
|
|
35
|
+
import { readFileSync, writeFileSync, existsSync, mkdirSync } from 'node:fs';
|
|
36
|
+
import { spawn } from 'node:child_process';
|
|
37
|
+
import dotenv from 'dotenv';
|
|
38
|
+
import { chat, getModel } from "./lib/llm.js";
|
|
39
|
+
import { outputLang } from "./lib/config.js";
|
|
40
|
+
import { fetchChannelVideos } from "./lib/yt-api.js";
|
|
41
|
+
import { normalizeHandle } from "./lib/channels.js";
|
|
42
|
+
import { PKG_ROOT, ENV_PATH, CACHE_DIR, SEARCH_QUEUE_FILE, SEARCH_PENDING_FILE, SEARCH_KEPT_FILE, script, } from "./lib/paths.js";
|
|
43
|
+
dotenv.config({ path: ENV_PATH });
|
|
44
|
+
mkdirSync(CACHE_DIR, { recursive: true });
|
|
45
|
+
const RUNTIME = process.execPath;
|
|
46
|
+
const LANG = outputLang();
|
|
47
|
+
const argv = process.argv.slice(2);
|
|
48
|
+
const VALUE_FLAGS = new Set(['--channel', '--max', '--scan', '--since', '--lang']);
|
|
49
|
+
const has = (f) => argv.includes(f);
|
|
50
|
+
const flagVal = (f) => {
|
|
51
|
+
const i = argv.indexOf(f);
|
|
52
|
+
return i !== -1 && argv[i + 1] != null ? argv[i + 1] : null;
|
|
53
|
+
};
|
|
54
|
+
// First token that is neither a flag nor a flag's value is the intent.
|
|
55
|
+
function positionalIntent() {
|
|
56
|
+
for (let i = 0; i < argv.length; i++) {
|
|
57
|
+
const a = argv[i];
|
|
58
|
+
if (a.startsWith('--')) {
|
|
59
|
+
if (VALUE_FLAGS.has(a))
|
|
60
|
+
i++;
|
|
61
|
+
continue;
|
|
62
|
+
}
|
|
63
|
+
if (i > 0 && VALUE_FLAGS.has(argv[i - 1]))
|
|
64
|
+
continue;
|
|
65
|
+
return a;
|
|
66
|
+
}
|
|
67
|
+
return null;
|
|
68
|
+
}
|
|
69
|
+
const intentArg = positionalIntent();
|
|
70
|
+
const RESET = has('--reset');
|
|
71
|
+
const KEEP = has('--keep');
|
|
72
|
+
const SKIP = has('--skip');
|
|
73
|
+
const COMPARE = has('--compare');
|
|
74
|
+
const CHANNEL = flagVal('--channel');
|
|
75
|
+
const MAX = Math.max(1, parseInt(flagVal('--max') || '8', 10));
|
|
76
|
+
const SCAN = Math.max(1, parseInt(flagVal('--scan') || '50', 10)); // recent uploads to consider when no --since
|
|
77
|
+
const SINCE = flagVal('--since');
|
|
78
|
+
const LANGTRACK = flagVal('--lang') || 'auto';
|
|
79
|
+
function emit(obj) {
|
|
80
|
+
process.stdout.write(JSON.stringify(obj));
|
|
81
|
+
process.exit(0);
|
|
82
|
+
}
|
|
83
|
+
const readJSON = (p, fallback) => {
|
|
84
|
+
if (!existsSync(p))
|
|
85
|
+
return fallback;
|
|
86
|
+
try {
|
|
87
|
+
return JSON.parse(readFileSync(p, 'utf8'));
|
|
88
|
+
}
|
|
89
|
+
catch {
|
|
90
|
+
return fallback;
|
|
91
|
+
}
|
|
92
|
+
};
|
|
93
|
+
const writeJSON = (p, v) => writeFileSync(p, JSON.stringify(v));
|
|
94
|
+
const loadQueue = () => readJSON(SEARCH_QUEUE_FILE, null);
|
|
95
|
+
const loadKept = () => readJSON(SEARCH_KEPT_FILE, []);
|
|
96
|
+
/** Fetch a transcript via the sibling script with the launching runtime. Mirrors yt-sweep. */
|
|
97
|
+
function fetchTranscript(videoId) {
|
|
98
|
+
return new Promise((resolve, reject) => {
|
|
99
|
+
const p = spawn(RUNTIME, [script('yt-transcript'), videoId, '--lang', LANGTRACK], { cwd: PKG_ROOT, env: { ...process.env } });
|
|
100
|
+
let stdout = '';
|
|
101
|
+
p.stdout.on('data', d => { stdout += d.toString(); });
|
|
102
|
+
p.stderr.resume();
|
|
103
|
+
p.on('close', code => resolve({ stdout, code: code ?? 1 }));
|
|
104
|
+
p.on('error', reject);
|
|
105
|
+
});
|
|
106
|
+
}
|
|
107
|
+
/** Slice the outermost JSON array from an LLM reply (tolerates stray prose / fences). */
|
|
108
|
+
function parseJsonArray(out) {
|
|
109
|
+
const start = out.indexOf('['), end = out.lastIndexOf(']');
|
|
110
|
+
if (start === -1 || end === -1 || end < start)
|
|
111
|
+
return null;
|
|
112
|
+
try {
|
|
113
|
+
return JSON.parse(out.slice(start, end + 1));
|
|
114
|
+
}
|
|
115
|
+
catch {
|
|
116
|
+
return null;
|
|
117
|
+
}
|
|
118
|
+
}
|
|
119
|
+
/** Re-rank a channel's videos against the intent (metadata only — no transcript). */
|
|
120
|
+
async function rerank(intent, items) {
|
|
121
|
+
if (items.length === 0)
|
|
122
|
+
return [];
|
|
123
|
+
const compact = items.map(h => ({ id: h.videoId, title: h.title, published: h.publishedAt, desc: (h.description || '').slice(0, 280) }));
|
|
124
|
+
const prompt = `From this YouTube channel's videos, pick the ones that serve the user's intent and rank them. Judge on title + description only (no transcripts). Drop anything off-topic.
|
|
125
|
+
|
|
126
|
+
Intent: "${intent}"
|
|
127
|
+
|
|
128
|
+
Videos:
|
|
129
|
+
${JSON.stringify(compact)}
|
|
130
|
+
|
|
131
|
+
Output ONLY a raw JSON array, best first, no fences:
|
|
132
|
+
[{"id":"VIDEO_ID","keep":true,"score":0-100,"reason":"max 12 words"},...]
|
|
133
|
+
Set keep=false for anything not relevant to the intent.`;
|
|
134
|
+
try {
|
|
135
|
+
const out = await chat(prompt, { system: 'You output ONLY a raw JSON array as instructed.', temperature: 0 });
|
|
136
|
+
const arr = parseJsonArray(out);
|
|
137
|
+
if (!arr)
|
|
138
|
+
return items; // fall back: keep all, original (newest-first) order
|
|
139
|
+
const byId = new Map(items.map(h => [h.videoId, h]));
|
|
140
|
+
const ranked = [];
|
|
141
|
+
for (const r of arr) {
|
|
142
|
+
if (!r || r.keep === false)
|
|
143
|
+
continue;
|
|
144
|
+
const h = byId.get(r.id);
|
|
145
|
+
if (h)
|
|
146
|
+
ranked.push({ ...h, score: typeof r.score === 'number' ? r.score : undefined, reason: r.reason });
|
|
147
|
+
}
|
|
148
|
+
return ranked.length ? ranked : items;
|
|
149
|
+
}
|
|
150
|
+
catch {
|
|
151
|
+
return items;
|
|
152
|
+
}
|
|
153
|
+
}
|
|
154
|
+
/** Rich, standalone summary for one candidate — the triage artifact + compare input. */
|
|
155
|
+
async function megaSummary(c, transcript, intent) {
|
|
156
|
+
const prompt = `Summarize this YouTube video in ${LANG} for a user researching: "${intent}". Make it a RICH, standalone summary they can decide on and that will later feed a cross-video comparison — OR return 'OFFTOPIC: <reason>' if the transcript clearly doesn't serve the intent.
|
|
157
|
+
|
|
158
|
+
Video:
|
|
159
|
+
- title: ${c.title}
|
|
160
|
+
- channel: ${c.channelTitle}
|
|
161
|
+
- published: ${c.publishedAt}
|
|
162
|
+
- url: https://youtube.com/watch?v=${c.videoId}
|
|
163
|
+
|
|
164
|
+
Transcript:
|
|
165
|
+
${transcript}
|
|
166
|
+
|
|
167
|
+
If on-topic, write:
|
|
168
|
+
- Header: ### ${c.channelTitle} — "${c.title}"
|
|
169
|
+
- Subtitle: _${c.publishedAt} · https://youtube.com/watch?v=${c.videoId}_
|
|
170
|
+
- One sentence on relevance to the intent
|
|
171
|
+
- 3-6 numbered thematic sections × 2-4 sentences, concrete: which options/tools are discussed, the criteria, the author's verdict and reasoning. Pull out anything directly comparable (names, pros/cons, recommendations).
|
|
172
|
+
- A short "Bottom line for the intent" line
|
|
173
|
+
- At most 5-8 short quotes from the transcript. No timestamps.
|
|
174
|
+
|
|
175
|
+
Language: natural ${LANG}; foreign words only for proper nouns or established technical terms.
|
|
176
|
+
Output ONLY the summary OR 'OFFTOPIC: <reason>'. No preamble.`;
|
|
177
|
+
return chat(prompt, {
|
|
178
|
+
system: `You are a research-grade video summarizer writing in ${LANG}. Output only the summary or 'OFFTOPIC: <reason>'.`,
|
|
179
|
+
model: getModel(),
|
|
180
|
+
});
|
|
181
|
+
}
|
|
182
|
+
/** Synthesize a comparison across everything kept. */
|
|
183
|
+
async function synthesizeComparison(intent, kept) {
|
|
184
|
+
const corpus = kept.map((k, i) => `--- VIDEO ${i + 1}: ${k.channelTitle} — "${k.title}" (${k.publishedAt})\nhttps://youtube.com/watch?v=${k.videoId}\n${k.summary}`).join('\n\n');
|
|
185
|
+
const prompt = `The user researched "${intent}" and kept ${kept.length} YouTube video summaries below. Synthesize a single comparison in ${LANG} that actually helps them decide.
|
|
186
|
+
|
|
187
|
+
${corpus}
|
|
188
|
+
|
|
189
|
+
Write:
|
|
190
|
+
- One-paragraph bottom line answering the intent directly.
|
|
191
|
+
- A comparison of the concrete options/tools across the videos (a Markdown table when it fits: option · who recommends it · pros · cons · best for).
|
|
192
|
+
- Consensus vs disagreements between the sources.
|
|
193
|
+
- A final recommendation with the reasoning, and who it's for.
|
|
194
|
+
|
|
195
|
+
Language: natural ${LANG}; foreign words only for proper nouns or established technical terms. Cite videos as [1], [2]… matching the order above. Output only the comparison.`;
|
|
196
|
+
return chat(prompt, { system: `You synthesize a decision-grade comparison in ${LANG}. No preamble.`, model: getModel() });
|
|
197
|
+
}
|
|
198
|
+
/** Lazy yield: advance to the next candidate that has a transcript, summarize, emit. */
|
|
199
|
+
async function yieldNext(queue) {
|
|
200
|
+
while (queue.cursor < queue.candidates.length) {
|
|
201
|
+
const c = queue.candidates[queue.cursor];
|
|
202
|
+
const t = await fetchTranscript(c.videoId);
|
|
203
|
+
if (t.code === 2)
|
|
204
|
+
emit({ status: 'rate_limited' }); // blocked IP — stop, don't advance
|
|
205
|
+
if (t.code !== 0 || !t.stdout.trim()) { // no transcript → auto-skip
|
|
206
|
+
queue.cursor++;
|
|
207
|
+
writeJSON(SEARCH_QUEUE_FILE, queue);
|
|
208
|
+
continue;
|
|
209
|
+
}
|
|
210
|
+
const summary = await megaSummary(c, t.stdout, queue.intent);
|
|
211
|
+
if (summary.startsWith('OFFTOPIC:')) { // re-rank missed it → auto-skip
|
|
212
|
+
queue.cursor++;
|
|
213
|
+
writeJSON(SEARCH_QUEUE_FILE, queue);
|
|
214
|
+
continue;
|
|
215
|
+
}
|
|
216
|
+
const pending = { videoId: c.videoId, title: c.title, channelTitle: c.channelTitle, publishedAt: c.publishedAt, summary };
|
|
217
|
+
writeJSON(SEARCH_PENDING_FILE, pending);
|
|
218
|
+
emit({
|
|
219
|
+
status: 'decision_needed',
|
|
220
|
+
summary,
|
|
221
|
+
pending: { videoId: c.videoId, title: c.title, channelTitle: c.channelTitle, publishedAt: c.publishedAt, position: queue.cursor + 1, total: queue.candidates.length },
|
|
222
|
+
});
|
|
223
|
+
}
|
|
224
|
+
emit({ status: 'done', kept: loadKept().length });
|
|
225
|
+
}
|
|
226
|
+
async function main() {
|
|
227
|
+
// --compare: synthesize from kept summaries.
|
|
228
|
+
if (COMPARE) {
|
|
229
|
+
const queue = loadQueue();
|
|
230
|
+
const kept = loadKept();
|
|
231
|
+
if (kept.length === 0)
|
|
232
|
+
emit({ status: 'done', kept: 0 });
|
|
233
|
+
const comparison = await synthesizeComparison(queue?.intent ?? '', kept);
|
|
234
|
+
emit({ status: 'compare', comparison });
|
|
235
|
+
}
|
|
236
|
+
// --keep / --skip: record decision on the pending candidate, then advance + yield next.
|
|
237
|
+
if (KEEP || SKIP) {
|
|
238
|
+
const queue = loadQueue();
|
|
239
|
+
if (!queue)
|
|
240
|
+
emit({ status: 'error', error: 'No active search. Start one: yt-search "<intent>" --channel <@handle>.' });
|
|
241
|
+
if (KEEP) {
|
|
242
|
+
const pending = readJSON(SEARCH_PENDING_FILE, null);
|
|
243
|
+
if (pending) {
|
|
244
|
+
const kept = loadKept();
|
|
245
|
+
kept.push(pending);
|
|
246
|
+
writeJSON(SEARCH_KEPT_FILE, kept);
|
|
247
|
+
}
|
|
248
|
+
}
|
|
249
|
+
queue.cursor++;
|
|
250
|
+
writeJSON(SEARCH_QUEUE_FILE, queue);
|
|
251
|
+
await yieldNext(queue);
|
|
252
|
+
}
|
|
253
|
+
// Resume an in-progress search (bare call, same intent + channel) without rebuilding.
|
|
254
|
+
const existing = loadQueue();
|
|
255
|
+
if (existing && !RESET && (!intentArg || intentArg === existing.intent) && (!CHANNEL || normalizeHandle(CHANNEL) === existing.channel)) {
|
|
256
|
+
await yieldNext(existing);
|
|
257
|
+
}
|
|
258
|
+
// Fresh search: needs an intent AND a channel.
|
|
259
|
+
if (!process.env.YT_BRIEFING_YOUTUBE_API_KEY) {
|
|
260
|
+
emit({ status: 'error', error: 'YT_BRIEFING_YOUTUBE_API_KEY is not set — add a YouTube Data API v3 key to .yt-briefing/.env (see README → setup / .env.example).' });
|
|
261
|
+
}
|
|
262
|
+
if (!intentArg)
|
|
263
|
+
emit({ status: 'error', error: 'Provide an intent: yt-search "<what to look for>" --channel <@handle|url>.' });
|
|
264
|
+
if (!CHANNEL)
|
|
265
|
+
emit({ status: 'error', error: 'Provide a channel: --channel <@handle|url>. /yt-search searches within one channel, not all of YouTube.' });
|
|
266
|
+
const handle = normalizeHandle(CHANNEL);
|
|
267
|
+
if (!handle)
|
|
268
|
+
emit({ status: 'error', error: `Could not read a channel handle from "${CHANNEL}" — use @name or the channel URL.` });
|
|
269
|
+
let videos;
|
|
270
|
+
try {
|
|
271
|
+
videos = await fetchChannelVideos(handle, { since: SINCE, limit: SINCE ? null : SCAN, enrich: false });
|
|
272
|
+
}
|
|
273
|
+
catch (e) {
|
|
274
|
+
emit({ status: 'error', error: e.message });
|
|
275
|
+
}
|
|
276
|
+
if (videos.length === 0)
|
|
277
|
+
emit({ status: 'no_results' });
|
|
278
|
+
const pool = videos.map(v => ({ videoId: v.videoId, title: v.title, channelTitle: handle, publishedAt: v.publishedAt, description: v.description }));
|
|
279
|
+
const ranked = (await rerank(intentArg, pool)).slice(0, MAX);
|
|
280
|
+
if (ranked.length === 0)
|
|
281
|
+
emit({ status: 'no_results' });
|
|
282
|
+
const queue = { built_at: new Date().toISOString(), intent: intentArg, channel: handle, candidates: ranked, cursor: 0 };
|
|
283
|
+
writeJSON(SEARCH_QUEUE_FILE, queue);
|
|
284
|
+
writeJSON(SEARCH_KEPT_FILE, []); // fresh search → fresh compare corpus
|
|
285
|
+
await yieldNext(queue);
|
|
286
|
+
}
|
|
287
|
+
main().catch(err => emit({ status: 'error', error: err.message }));
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "yt-briefing",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.6.0",
|
|
4
4
|
"description": "A self-learning YouTube briefing engine: it sweeps the channels you follow, filters noise in two stages (title, then transcript), summarizes the rest in your language, and adapts to your ratings — one video at a time.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|