yt-briefing 0.9.2 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,12 +1,12 @@
1
1
  ---
2
2
  name: yt-search
3
3
  description: Search WITHIN one YouTube channel by intent — name a channel and what you're after; the engine lists that channel's uploads, ranks them against your intent (metadata only, no transcript yet), then lazily yields ONE matching video at a time with a rich summary. You keep or skip each; at the end it synthesizes a comparison from everything you kept. Channel-scoped, not whole-YouTube. Same transcript engine + proxy as /yt; lazy on purpose (no transcript bursts → no IP block). Summaries and prompts use the language chosen at onboarding.
4
- argument-hint: A channel (@handle or URL) and a descriptive intent, e.g. "@t3dotgg which terminal for AI coding". Optional --top N (default 10), --scan N, --since YYYY-MM-DD.
4
+ argument-hint: A channel (@handle or URL) and a descriptive intent, e.g. "@t3dotgg which terminal for AI coding". Searches the channel's whole upload history. Optional --top N (default 10).
5
5
  ---
6
6
 
7
7
  ## How it works
8
8
 
9
- `src/yt-search.ts` is the whole engine: list one channel's uploads (cheap — `playlistItems`, ~1 quota unit/page; NOT `search.list`) → re-rank them against your intent on metadata only (title/description, no transcript) → **lazy** one-candidate-at-a-time yield with a rich summary → record keep/skip → on demand synthesize a comparison from everything kept. **Channel-scoped on purpose** — you choose where to look; it does NOT search all of YouTube. Matching is descriptive: the LLM filters the channel's videos by intent. This skill is a thin loop — paste the summary, collect keep/skip, show the final comparison.
9
+ `src/yt-search.ts` is the whole engine: list the channel's **full upload history** (cheap — `playlistItems`, ~1 quota unit/page; NOT `search.list`) → re-rank it against your intent on metadata only (title/description, no transcript; ranked in chunks and merged so a 1000+ video channel never overflows the LLM) → **lazy** one-candidate-at-a-time yield with a rich summary → record keep/skip → on demand synthesize a comparison from everything kept. **Channel-scoped on purpose** — you choose where to look; it does NOT search all of YouTube. Matching is descriptive: the LLM filters the channel's videos by intent. This skill is a thin loop — paste the summary, collect keep/skip, show the final comparison.
10
10
 
11
11
  **Lazy on purpose:** one transcript per step, never a burst — a burst looks like scraping and gets the IP blocked (same reason `/yt` is lazy). Run the engine bare — stdout is a single JSON line, stderr empty; never redirect.
12
12
 
@@ -53,5 +53,5 @@ while true:
53
53
 
54
54
  - **Verbatim:** paste `summary` and `comparison` exactly as returned; never paste a raw transcript.
55
55
  - **Language:** question text + option descriptions follow `output_lang`; button labels stay `Keep` / `Skip`.
56
- - **Scope:** one channel per search. Listing is cheap; the cost is the lazy transcript fetches, so let the user keep/skip rather than pulling everything. A bare resume (no `--reset`) continues the same ranked queue. `--scan N` (default 50) caps how many recent uploads are considered; `--since` widens by date.
56
+ - **Scope:** one channel per search, ranked across its **whole upload history**. Listing is cheap; the cost is the lazy transcript fetches, so let the user keep/skip rather than pulling everything. A bare resume (no `--reset`) continues the same ranked queue. `--top N` (default 10) caps how many of the top re-ranked matches get triaged — it is the only flag.
57
57
  - **Stateless triage:** independent of `/yt` (no channel profiles, no ratings written). For the recurring multi-channel briefing use `/yt`; for one known video use `/yt-transcribe`.
package/README.md CHANGED
@@ -1,5 +1,12 @@
1
1
  # yt-briefing
2
2
 
3
+ [![npm](https://img.shields.io/npm/v/yt-briefing)](https://www.npmjs.com/package/yt-briefing)
4
+ [![license](https://img.shields.io/npm/l/yt-briefing)](./LICENSE)
5
+ [![node](https://img.shields.io/node/v/yt-briefing)](https://www.npmjs.com/package/yt-briefing)
6
+ [![mac](https://github.com/michal90r/yt-briefing/actions/workflows/ci-mac.yml/badge.svg)](https://github.com/michal90r/yt-briefing/actions/workflows/ci-mac.yml)
7
+ [![ubuntu](https://github.com/michal90r/yt-briefing/actions/workflows/ci-ubuntu.yml/badge.svg)](https://github.com/michal90r/yt-briefing/actions/workflows/ci-ubuntu.yml)
8
+ [![windows](https://github.com/michal90r/yt-briefing/actions/workflows/ci-windows.yml/badge.svg)](https://github.com/michal90r/yt-briefing/actions/workflows/ci-windows.yml)
9
+
3
10
  Save hours on YouTube. yt-briefing watches the channels you follow so you don't have to. For
4
11
  each new video it gives you a short briefing in your own language — every point that matters, with
5
12
  only the filler cut, so nothing important is lost. Reading it takes a fraction of the time the
@@ -13,18 +13,31 @@
13
13
  * agent's cwd IS the package folder AND the runtime is Bun (which runs TypeScript directly).
14
14
  * That's true for the publisher's own day-to-day use, so it stays the default.
15
15
  *
16
- * For everyone else — a Node user, or any install whose cwd won't be the package — we bake an
17
- * absolute, runtime-correct command instead: `"<this runtime>" "<abs>/dist/X.js"`. The runtime
18
- * is `process.execPath` of whoever runs the installer (so `init` under Node bakes node, under
19
- * Bun bakes bun), and it points at the compiled build, so it runs from any cwd. The engine
20
- * resolves data/.env from its own location regardless. (Requires `dist/` — build once with
21
- * `bun run build` / `npm run build`; that's exactly how a Node user already got here.)
16
+ * For everyone else — a Node user, or any install whose cwd won't be the package — we rewrite to
17
+ * a PORTABLE command: `<runtime> "<project-relative>/dist/X.js"`. The runtime is the bare name
18
+ * (`node`/`bun`) resolved from PATH, never an absolute binary; the script and `data/` paths are
19
+ * relative to the PROJECT ROOT, never machine-absolute. The invariant this rests on is the same
20
+ * one the whole package already relies on (paths.ts derives BASE_DIR/DATA_DIR from
21
+ * `process.cwd()` when consumed): the agent runs from the project root. So the rewritten skill is
22
+ * machine-independent — it survives being committed to git and shared across machines (e.g. a
23
+ * Mac dev box and a Linux VPS), which an absolute `process.execPath`/`<abs>/dist` baking did not.
24
+ * (Requires `dist/` — build once with `bun run build` / `npm run build`.)
22
25
  */
23
26
  import { readFileSync, writeFileSync, mkdirSync } from 'node:fs';
24
- import { join, resolve } from 'node:path';
25
- import { PKG_ROOT, DATA_DIR } from "./paths.js";
26
- /** Compiled output dir — what a baked (dist) skill command points the runtime at. */
27
+ import { join, resolve, relative, dirname, sep } from 'node:path';
28
+ import { PKG_ROOT, BASE_DIR, DATA_DIR } from "./paths.js";
29
+ /** Compiled output dir — what a rewritten (dist) skill command points the runtime at. */
27
30
  const DIST_DIR = join(PKG_ROOT, 'dist');
31
+ /** Consumed as a dependency? Then PKG_ROOT lives under node_modules (mirrors paths.ts). */
32
+ const CONSUMED = PKG_ROOT.split(sep).includes('node_modules');
33
+ /**
34
+ * The project root the agent runs from — the cwd against which the rewritten skill's relative
35
+ * paths resolve. When consumed, that's the user's project (parent of `<project>/.yt-briefing`);
36
+ * in a clone it's the package itself. Matches how paths.ts picks BASE_DIR.
37
+ */
38
+ const PROJECT_ROOT = CONSUMED ? dirname(BASE_DIR) : PKG_ROOT;
39
+ /** An absolute path expressed relative to PROJECT_ROOT, with POSIX `/` (portable on Windows too). */
40
+ const toProjectRel = (abs) => relative(PROJECT_ROOT, abs).split(sep).join('/');
28
41
  /** The skills this package ships — each lives at `.claude/skills/<name>/SKILL.md`. */
29
42
  export const SKILLS = ['yt', 'yt-transcribe', 'yt-search'];
30
43
  /** True when the installer itself is running under Bun (vs plain Node). */
@@ -56,25 +69,26 @@ export const AGENTS = {
56
69
  /**
57
70
  * One shipped skill's SKILL.md. `dist=false` (default) returns it verbatim — the
58
71
  * `bun run src/X.ts` dev form, correct only when cwd is the package AND the runtime is Bun.
59
- * `dist=true` rewrites for the consumed case: engine commands become
60
- * `"<process.execPath>" "<abs>/dist/X.js"` (this machine's runtime, Node or Bun, against the
61
- * compiled build, so they run from any cwd), and the bare `data/…` paths the agent reads (e.g.
62
- * `data/config.json`) become the absolute `DATA_DIR`. In dev the agent's cwd IS the package so
63
- * `data/` resolves; when consumed, DATA_DIR moves to `<project>/.yt-briefing/data`, so the
64
- * dev-relative paths would miss — hence the rewrite.
72
+ * `dist=true` rewrites for the consumed case into PORTABLE, project-relative form: engine
73
+ * commands become `<node|bun> "<rel>/dist/X.js"` (bare runtime from PATH + a path relative to the
74
+ * project root, so they run on any machine from the project cwd), and the bare `data/…` paths the
75
+ * agent reads (e.g. `data/config.json`) become the project-relative DATA_DIR (`.yt-briefing/data/`
76
+ * when consumed). Nothing machine-absolute is baked, so the rewritten skill can be committed and
77
+ * shared across machines. The runtime name follows whoever runs the installer (Node→`node`,
78
+ * Bun→`bun`); the compiled `dist/` build runs under either.
65
79
  */
66
80
  export function skillBody(name, dist = false) {
67
81
  const raw = readFileSync(skillSource(name), 'utf8');
68
82
  if (!dist)
69
83
  return raw;
70
- const exe = process.execPath;
71
- const cmd = (base) => `"${exe}" "${join(DIST_DIR, base + '.js')}"`;
84
+ const runtime = isBun ? 'bun' : 'node';
85
+ const cmd = (base) => `${runtime} "${toProjectRel(join(DIST_DIR, base + '.js'))}"`;
72
86
  return raw
73
87
  .replace(/bun run src\/yt-sweep\.ts/g, cmd('yt-sweep'))
74
88
  .replace(/bun run src\/yt-rating\.ts/g, cmd('yt-rating'))
75
89
  .replace(/bun run src\/yt-transcript\.ts/g, cmd('yt-transcript'))
76
90
  .replace(/bun run src\/yt-search\.ts/g, cmd('yt-search'))
77
- .replace(/data\//g, DATA_DIR + '/');
91
+ .replace(/data\//g, toProjectRel(DATA_DIR) + '/');
78
92
  }
79
93
  /**
80
94
  * Write every shipped skill into `root`, each under its own `<name>/SKILL.md` subdir
package/dist/yt-search.js CHANGED
@@ -5,8 +5,9 @@
5
5
  *
6
6
  * You point it at a channel and describe what you're after ("which terminal does he recommend
7
7
  * for AI coding"); the engine:
8
- * 1. lists that channel's uploads (cheap — playlistItems, 1 quota unit/page; NOT search.list),
9
- * 2. re-ranks them against your intent on metadata only — title/description, NO transcript yet,
8
+ * 1. lists that channel's FULL upload history (cheap — playlistItems, 1 quota unit/page; NOT search.list),
9
+ * 2. re-ranks it against your intent on metadata only — title/description, NO transcript yet
10
+ * (the history is ranked in chunks and merged, so a 1000+ video channel never overflows the LLM),
10
11
  * 3. yields ONE matching video at a time with a rich summary, lazily — never a burst of
11
12
  * transcript fetches (a burst looks like scraping and gets the IP blocked),
12
13
  * 4. records your keep/skip decision; kept summaries accumulate in a cache,
@@ -16,7 +17,7 @@
16
17
  * the channel's videos by intent (no exact-keyword needed).
17
18
  *
18
19
  * Usage (the skill / CLI drives these; one JSON line per call):
19
- * yt-search "<intent>" --channel <@handle|url> [--reset] [--top N] [--scan N] [--since DATE] [--lang auto]
20
+ * yt-search "<intent>" --channel <@handle|url> [--reset] [--top N] [--lang auto]
20
21
  * yt-search --keep record the pending candidate, advance, yield next
21
22
  * yt-search --skip drop the pending candidate, advance, yield next
22
23
  * yt-search --compare synthesize a comparison from everything kept
@@ -45,7 +46,7 @@ mkdirSync(CACHE_DIR, { recursive: true });
45
46
  const RUNTIME = process.execPath;
46
47
  const LANG = outputLang();
47
48
  const argv = process.argv.slice(2);
48
- const VALUE_FLAGS = new Set(['--channel', '--top', '--scan', '--since', '--lang']);
49
+ const VALUE_FLAGS = new Set(['--channel', '--top', '--lang']);
49
50
  const has = (f) => argv.includes(f);
50
51
  const flagVal = (f) => {
51
52
  const i = argv.indexOf(f);
@@ -73,8 +74,6 @@ const SKIP = has('--skip');
73
74
  const COMPARE = has('--compare');
74
75
  const CHANNEL = flagVal('--channel');
75
76
  const TOP = Math.max(1, parseInt(flagVal('--top') || '10', 10));
76
- const SCAN = Math.max(1, parseInt(flagVal('--scan') || '50', 10)); // recent uploads to consider when no --since
77
- const SINCE = flagVal('--since');
78
77
  const LANGTRACK = flagVal('--lang') || 'auto';
79
78
  function emit(obj) {
80
79
  process.stdout.write(JSON.stringify(obj));
@@ -116,8 +115,12 @@ function parseJsonArray(out) {
116
115
  return null;
117
116
  }
118
117
  }
119
- /** Re-rank a channel's videos against the intent (metadata only — no transcript). */
120
- async function rerank(intent, items) {
118
+ // The whole channel history is ranked, so the pool can be large (a years-old channel is 1000+
119
+ // uploads). One LLM prompt would overflow context, so the pool is ranked in chunks and merged.
120
+ const RERANK_CHUNK = 200;
121
+ /** Re-rank ONE chunk against the intent (metadata only). Returns [] on a parse/LLM failure so a
122
+ * bad chunk is dropped rather than flooding the queue with unranked videos. */
123
+ async function rerankBatch(intent, items) {
121
124
  if (items.length === 0)
122
125
  return [];
123
126
  const compact = items.map(h => ({ id: h.videoId, title: h.title, published: h.publishedAt, desc: (h.description || '').slice(0, 280) }));
@@ -135,7 +138,7 @@ Set keep=false for anything not relevant to the intent.`;
135
138
  const out = await chat(prompt, { system: 'You output ONLY a raw JSON array as instructed.', temperature: 0 });
136
139
  const arr = parseJsonArray(out);
137
140
  if (!arr)
138
- return items; // fall back: keep all, original (newest-first) order
141
+ return [];
139
142
  const byId = new Map(items.map(h => [h.videoId, h]));
140
143
  const ranked = [];
141
144
  for (const r of arr) {
@@ -145,11 +148,26 @@ Set keep=false for anything not relevant to the intent.`;
145
148
  if (h)
146
149
  ranked.push({ ...h, score: typeof r.score === 'number' ? r.score : undefined, reason: r.reason });
147
150
  }
148
- return ranked.length ? ranked : items;
151
+ return ranked;
149
152
  }
150
153
  catch {
151
- return items;
154
+ return [];
155
+ }
156
+ }
157
+ /** Re-rank the channel's full history against the intent: chunk the pool, rank each chunk, merge
158
+ * the keepers globally by score (highest first), dedup by id. */
159
+ async function rerank(intent, items) {
160
+ if (items.length === 0)
161
+ return [];
162
+ const merged = [];
163
+ for (let i = 0; i < items.length; i += RERANK_CHUNK) {
164
+ merged.push(...await rerankBatch(intent, items.slice(i, i + RERANK_CHUNK)));
152
165
  }
166
+ const byId = new Map();
167
+ for (const c of merged)
168
+ if (!byId.has(c.videoId))
169
+ byId.set(c.videoId, c);
170
+ return [...byId.values()].sort((a, b) => (b.score ?? -1) - (a.score ?? -1));
153
171
  }
154
172
  /** Rich, standalone summary for one candidate — the triage artifact + compare input. */
155
173
  async function megaSummary(c, transcript, intent) {
@@ -278,7 +296,7 @@ async function main() {
278
296
  emit({ status: 'error', error: `Could not read a channel handle from "${CHANNEL}" — use @name or the channel URL.` });
279
297
  let videos;
280
298
  try {
281
- videos = await fetchChannelVideos(handle, { since: SINCE, limit: SINCE ? null : SCAN, enrich: false });
299
+ videos = await fetchChannelVideos(handle, { limit: null, enrich: false }); // full channel history
282
300
  }
283
301
  catch (e) {
284
302
  emit({ status: 'error', error: e.message });
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "yt-briefing",
3
- "version": "0.9.2",
3
+ "version": "0.11.0",
4
4
  "description": "A self-learning YouTube briefing engine: it sweeps the channels you follow, filters noise in two stages (title, then transcript), summarizes the rest in your language, and adapts to your ratings — one video at a time.",
5
5
  "type": "module",
6
6
  "bin": {
@@ -22,7 +22,8 @@
22
22
  "sweep": "bun run src/yt-sweep.ts",
23
23
  "rate": "bun run src/yt-rating.ts",
24
24
  "transcribe": "bun run src/yt-transcript.ts",
25
- "typecheck": "tsc --noEmit"
25
+ "typecheck": "tsc --noEmit",
26
+ "test": "vitest run"
26
27
  },
27
28
  "engines": {
28
29
  "node": ">=18",
@@ -51,6 +52,7 @@
51
52
  "devDependencies": {
52
53
  "@types/node": "^22",
53
54
  "typescript": "^5.7",
54
- "@types/bun": "latest"
55
+ "@types/bun": "latest",
56
+ "vitest": "^3"
55
57
  }
56
58
  }