@tekyzinc/gsd-t 5.14.11 → 5.16.11
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +176 -0
- package/README.md +2 -1
- package/commands/gsd-t-demo-videos.md +581 -0
- package/commands/gsd-t-gap-analysis.md +124 -10
- package/commands/gsd-t-help.md +6 -0
- package/package.json +1 -1
- package/templates/demo-videos/README.md +90 -0
- package/templates/demo-videos/e2e/cast.mjs +104 -0
- package/templates/demo-videos/e2e/example.lines.mjs +79 -0
- package/templates/demo-videos/e2e/example.spec.ts +144 -0
- package/templates/demo-videos/e2e/manifest.ts +30 -0
- package/templates/demo-videos/e2e/preflight.spec.ts +109 -0
- package/templates/demo-videos/e2e/runtime.ts +293 -0
- package/templates/demo-videos/e2e/signin.ts +42 -0
- package/templates/demo-videos/scripts/seed-lib.mjs +106 -0
- package/templates/demo-videos/scripts/walkthrough-mux.mjs +177 -0
- package/templates/demo-videos/scripts/walkthrough-normalise.mjs +123 -0
- package/templates/demo-videos/scripts/walkthrough-trim.mjs +86 -0
- package/templates/demo-videos/scripts/walkthrough-voice-check.mjs +197 -0
- package/templates/demo-videos/scripts/walkthrough-voice-ensure.mjs +72 -0
- package/templates/demo-videos/scripts/walkthrough-voice.mjs +469 -0
|
@@ -0,0 +1,581 @@
|
|
|
1
|
+
# GSD-T: Demo Videos — Narrated Walkthrough Videos of a Running App
|
|
2
|
+
|
|
3
|
+
Produce narrated screen-recording walkthroughs of an application, from a coverage
|
|
4
|
+
plan through to finished MP4s. `$ARGUMENTS` names what to do: a walkthrough name
|
|
5
|
+
(`fleet`), `--all`, `--plan`, `--seed`, `--audit`, or nothing (plan then build
|
|
6
|
+
everything).
|
|
7
|
+
|
|
8
|
+
Distilled from a real 200-turn production run (HILO ATOS, 12 walkthroughs,
|
|
9
|
+
2026-08-19 → 2026-08-27). **Every gate below exists because its absence cost a
|
|
10
|
+
re-record, a re-render, or a shipped video the user had to catch.** Do not
|
|
11
|
+
re-derive them.
|
|
12
|
+
|
|
13
|
+
---
|
|
14
|
+
|
|
15
|
+
## The shape of the thing
|
|
16
|
+
|
|
17
|
+
**Narration is the master clock. The picture obeys it.**
|
|
18
|
+
|
|
19
|
+
A step is one spoken sentence plus one thing happening on screen. The step lasts
|
|
20
|
+
exactly as long as the sentence takes to say — **measured from the rendered
|
|
21
|
+
audio file**, never estimated, never a constant. The action fires when the
|
|
22
|
+
sentence starts, then the pointer rests wherever it landed for the remainder.
|
|
23
|
+
That is what reads on camera as "pointing at a thing while explaining it".
|
|
24
|
+
|
|
25
|
+
The first version of this pipeline gave every step a fixed 5 seconds. That single
|
|
26
|
+
constant is why narration drifted out of sync with the screen, and it is the
|
|
27
|
+
defect the whole design exists to remove.
|
|
28
|
+
|
|
29
|
+
**One continuous recording per walkthrough**, clipped afterward — never
|
|
30
|
+
per-screen clips assembled later, and never stills. Stills cannot show a
|
|
31
|
+
transition, and a slideshow does not read as software being used.
|
|
32
|
+
|
|
33
|
+
---
|
|
34
|
+
|
|
35
|
+
## Prerequisites — check these before Step 1
|
|
36
|
+
|
|
37
|
+
| Need | Why | Halt if missing |
|
|
38
|
+
|---|---|---|
|
|
39
|
+
| A **running app with real data** — deployed/preview URL, not a local build with an empty database | An empty tenant on video is indistinguishable from a feature that was never built | Yes — ask for the URL, a login, and the name of a tenant that actually has data |
|
|
40
|
+
| **Playwright** installed | The recorder | Yes — `gsd-t setup-playwright` |
|
|
41
|
+
| **ffmpeg** | Every audio operation | Yes |
|
|
42
|
+
| **auto-editor** (Python venv, bundled into the project, not PATH) | Silence removal, Stage 4 | Yes |
|
|
43
|
+
| A **text-to-speech service with FIXED voices** (e.g. Google Cloud TTS) | The narrator. A language-model TTS re-reads a style hint per request and the voice drifts | Yes — see Stage 1 |
|
|
44
|
+
|
|
45
|
+
**Ask for the tenant by name.** In the source run, two videos were filmed
|
|
46
|
+
against the wrong location before the right ID was pinned down; the numbering
|
|
47
|
+
was counter-intuitive and no amount of code reading would have revealed it.
|
|
48
|
+
Record the answer in the project's walkthrough `signin` module as a named
|
|
49
|
+
constant with a comment, so it is never re-derived.
|
|
50
|
+
|
|
51
|
+
---
|
|
52
|
+
|
|
53
|
+
## Step 1 — PLAN: the coverage map
|
|
54
|
+
|
|
55
|
+
Do not start from the code. **Walk the running app** and enumerate what a viewer
|
|
56
|
+
could actually be shown. The source run's first nine videos covered roughly 25 of
|
|
57
|
+
~55 populated screens, and the gaps were only found by an audit against the live
|
|
58
|
+
site — a whole top-level navigation section, the app's actual landing page, and
|
|
59
|
+
the richest populated feature in the product were all missed.
|
|
60
|
+
|
|
61
|
+
Produce a plan table, one row per walkthrough:
|
|
62
|
+
|
|
63
|
+
| Column | Content |
|
|
64
|
+
|---|---|
|
|
65
|
+
| Name | kebab-case, becomes every filename (`fleet`, `flight-risk`) |
|
|
66
|
+
| Workflow | the job a real user is doing, not the menu name |
|
|
67
|
+
| Screens | every route it visits |
|
|
68
|
+
| Tabs/sub-surfaces | **count them on screen** — a record with 7 tabs narrated as 3 is a shipped error |
|
|
69
|
+
| Data state | populated / partly empty / empty |
|
|
70
|
+
| Target length | under 90 seconds preferred; note if genuinely longer |
|
|
71
|
+
|
|
72
|
+
**Group by workflow, not by menu.** Give the viewer the dependency chain: *before
|
|
73
|
+
you can schedule a student for a course, a program must exist and be linked to a
|
|
74
|
+
course.* That context is what makes a walkthrough useful rather than a tour.
|
|
75
|
+
|
|
76
|
+
Write the plan to `docs/demo-videos/PLAN.md`. If the user supplies a spreadsheet,
|
|
77
|
+
mirror it there too.
|
|
78
|
+
|
|
79
|
+
---
|
|
80
|
+
|
|
81
|
+
## Step 2 — PROBE the empty screens (never assume)
|
|
82
|
+
|
|
83
|
+
For every screen the plan marks empty or doubtful, **open it in the running app
|
|
84
|
+
and look**. In the source run the empty-screen table was written from the seed
|
|
85
|
+
scripts, and probing the live app **reversed four of five verdicts**:
|
|
86
|
+
|
|
87
|
+
- one screen looked empty only because it opened on a period with no data — the
|
|
88
|
+
fix was a click, not a seeder
|
|
89
|
+
- three were already populated
|
|
90
|
+
- built-but-empty is not the same as a stub: an empty screen with working
|
|
91
|
+
controls is worth filming
|
|
92
|
+
|
|
93
|
+
Classify each: **already populated** / **needs seeding** / **empty by design**
|
|
94
|
+
(say so in narration) / **blocked by an app bug** (report it, do not film it).
|
|
95
|
+
|
|
96
|
+
---
|
|
97
|
+
|
|
98
|
+
## Step 3 — SEED, through the real UI only
|
|
99
|
+
|
|
100
|
+
Seeders drive the actual forms as a real user, so every guardrail the app
|
|
101
|
+
enforces still applies. A record created this way is a record a person could have
|
|
102
|
+
created, which is the only reason the screen it fills is showing something true.
|
|
103
|
+
|
|
104
|
+
**The rule every seeder follows: when the app refuses, report what it said and
|
|
105
|
+
stop.** Never retry past a refusal, never reach around the form into the
|
|
106
|
+
database. A refusal is information — usually that the screen is empty for a
|
|
107
|
+
reason worth knowing. (In the source run, one seeder's refusal turned out to be a
|
|
108
|
+
genuine app bug: an enabled button with an inert click handler.)
|
|
109
|
+
|
|
110
|
+
Two silent-failure traps seen in practice, worth checking for in any form:
|
|
111
|
+
|
|
112
|
+
- **Labels not wired to inputs** — `getByLabel` matches nothing, every required
|
|
113
|
+
field stays empty, and the dialog just sits there. Fall back to filling by
|
|
114
|
+
input position, and say so in a comment.
|
|
115
|
+
- **Save navigates elsewhere** — the control you need is gone on the next
|
|
116
|
+
iteration, so each item must start from a fresh page load.
|
|
117
|
+
|
|
118
|
+
Also: several elements can carry `role=dialog` (a sidebar, an assistant panel).
|
|
119
|
+
Match a dialog by its title, never `.first()`.
|
|
120
|
+
|
|
121
|
+
**Every cast member must have a script that recreates it.** A shared demo site
|
|
122
|
+
usually resets (nightly is common), so anything the walkthrough names aloud is
|
|
123
|
+
gone by morning — and a walkthrough that references a course which no longer
|
|
124
|
+
exists fails at its first selector. Treat the reset as normal and make the data
|
|
125
|
+
reproducible: one seeder per cast member, re-runnable, and idempotent where the
|
|
126
|
+
app allows it (check whether the record is already there and say so, rather than
|
|
127
|
+
creating a duplicate).
|
|
128
|
+
|
|
129
|
+
`--seed` re-runs the whole cast, so the day's first recording starts from a known
|
|
130
|
+
state. State plainly in the handoff which parts of the cast are seeded ahead and
|
|
131
|
+
which are created live on camera.
|
|
132
|
+
|
|
133
|
+
---
|
|
134
|
+
|
|
135
|
+
## Step 4 — CAST the demo (do this BEFORE writing a word)
|
|
136
|
+
|
|
137
|
+
**A demo that describes what a form is for, while the form sits empty, teaches
|
|
138
|
+
nothing.** The viewer learns that a button exists — not what the software does.
|
|
139
|
+
That is the clinical failure, and it is what this step removes.
|
|
140
|
+
|
|
141
|
+
**Two stories, woven.** The operator's story is what happens on screen, told in
|
|
142
|
+
their own voice — *"I'm setting fifty-five hours."* The customer's story is why
|
|
143
|
+
every value they type is that value. Neither works alone: the customer alone is a
|
|
144
|
+
bio, the operator alone is a person explaining a form.
|
|
145
|
+
|
|
146
|
+
| The reason | The decision | On screen |
|
|
147
|
+
|---|---|---|
|
|
148
|
+
| Maya works night shifts | Tyler picks Part 61, not Part 141 | dropdown → **Part 61** |
|
|
149
|
+
| Maya has never flown | Tyler sets 55 hours, not the FAA's 40 | types **55** |
|
|
150
|
+
| Maya can only fly mornings | Tyler assigns James, who flies mornings | dropdown → **James Rivera** |
|
|
151
|
+
|
|
152
|
+
**The reason comes BEFORE the value, in the same breath** — *"She's on nights,
|
|
153
|
+
so — Part 61."* Value-then-justification is the teacher voice creeping back.
|
|
154
|
+
|
|
155
|
+
**The tell that you have slipped back into explaining:** a "because" clause
|
|
156
|
+
pointing at the software. *"so it's required rather than optional"*, *"which is
|
|
157
|
+
what the invoice uses later"*, *"the schedule refuses it otherwise"*. Every
|
|
158
|
+
reason must point at the customer, never at the mechanism. Read the finished
|
|
159
|
+
narration aloud: **a sentence that would survive with the names removed is
|
|
160
|
+
explaining the software, and it is wrong.**
|
|
161
|
+
|
|
162
|
+
Pick real, named specifics and write them down as a cast list before any
|
|
163
|
+
narration is drafted. Not "a course" — a course with a name someone could say out
|
|
164
|
+
loud. Not "a student" — a person with a name.
|
|
165
|
+
|
|
166
|
+
Put the cast in ONE constants block that both the seeder and the spec import, so
|
|
167
|
+
a name can be changed in a single edit and can never drift between the narration
|
|
168
|
+
and the screen:
|
|
169
|
+
|
|
170
|
+
```js
|
|
171
|
+
// e2e/walkthrough/cast.mjs — the demo's cast. One edit changes it everywhere.
|
|
172
|
+
export const CAST = {
|
|
173
|
+
course: 'Private Pilot Certificate — Part 61',
|
|
174
|
+
student: { first: 'Maya', last: 'Ellison', email: 'maya.ellison@example.com' },
|
|
175
|
+
// …aircraft, instructor, dates — everything the walkthrough names aloud
|
|
176
|
+
};
|
|
177
|
+
```
|
|
178
|
+
|
|
179
|
+
**Three rules, and the third is the one a spec silently loses:**
|
|
180
|
+
|
|
181
|
+
1. **Every named thing is real and specific.** A syllabus with actual stage
|
|
182
|
+
names, a certificate someone actually earns, a rate someone actually pays.
|
|
183
|
+
Generic placeholders (`Test Course 1`, `Student A`) read as fake and make the
|
|
184
|
+
whole demo read as fake with them.
|
|
185
|
+
|
|
186
|
+
2. **The data is ENTERED on camera, not described — and EVERY DROPDOWN IS
|
|
187
|
+
OPENED AND PICKED.** The walkthrough types the values and saves. A sentence
|
|
188
|
+
explaining what a field is for, over an empty field, is the defect.
|
|
189
|
+
|
|
190
|
+
A dropdown that is merely highlighted shows nothing: the viewer cannot see
|
|
191
|
+
what the alternatives were, or that a choice happened at all. **The choice is
|
|
192
|
+
the most informative moment in a create-flow** — it is where the customer's
|
|
193
|
+
situation becomes the operator's decision. Use `choose()`, never
|
|
194
|
+
`highlight()`, on a select.
|
|
195
|
+
|
|
196
|
+
If the flow creates something, the demo creates it — that also proves the
|
|
197
|
+
create-flow works, which describing it never does.
|
|
198
|
+
|
|
199
|
+
3. **The names CARRY FORWARD.** Once the course is created, every later sentence
|
|
200
|
+
says that course BY NAME. Once the student is enrolled, they are referred to
|
|
201
|
+
by name for the rest of the video — "Maya's next lesson", not "the student's
|
|
202
|
+
next lesson". This is what makes the walkthrough one story instead of a tour
|
|
203
|
+
of screens. It is easy to lose because each step is written independently, so
|
|
204
|
+
check it as a pass over the finished narration: **a sentence that says "the
|
|
205
|
+
student" or "the course" after the cast has been introduced is a bug.**
|
|
206
|
+
|
|
207
|
+
**Give any value with a symbol or abbreviation a spoken twin.** The narrator
|
|
208
|
+
reads text literally, so `$185/hr` comes out as "dollar one eight five slash h
|
|
209
|
+
r" and `9:00 AM` as "nine colon zero zero A M". Keep the typed value for the
|
|
210
|
+
form field and a said-aloud version for the sentence, both in the cast block.
|
|
211
|
+
|
|
212
|
+
**Price it with the real billable parts.** A course, a plan or a subscription is
|
|
213
|
+
the SUM of the things a charge attaches to, so build those things on camera and
|
|
214
|
+
attach them — not one summary price. In the flight-school example that is four
|
|
215
|
+
products: the airplane per hour, the instructor per hour, ground instruction per
|
|
216
|
+
hour, and the materials kit once. The payoff line is the one that makes the whole
|
|
217
|
+
section land: *"an hour of dual bills Maya two-sixty — a hundred and eighty-five
|
|
218
|
+
for the airplane, seventy-five for James"* is arithmetic the viewer just watched
|
|
219
|
+
being set up.
|
|
220
|
+
|
|
221
|
+
**Order the walkthrough as the real-life sequence**, so each screen is visited
|
|
222
|
+
because the previous one made it necessary: build the course → enrol the named
|
|
223
|
+
student → schedule their first lesson → fly it → bill it. That ordering is what
|
|
224
|
+
makes the dependency context land ("before you can schedule a student for a
|
|
225
|
+
course, a program must exist and be linked to a course") instead of being
|
|
226
|
+
asserted.
|
|
227
|
+
|
|
228
|
+
**When creation hits a guardrail**, that is information, not a blocker — the app
|
|
229
|
+
refusing an incomplete enrolment is worth showing. But do not fight it on camera:
|
|
230
|
+
fall back to an existing named record, and say in the handoff which parts of the
|
|
231
|
+
cast are created live and which are pre-seeded.
|
|
232
|
+
|
|
233
|
+
---
|
|
234
|
+
|
|
235
|
+
## Step 5 — WRITE the narration as beats
|
|
236
|
+
|
|
237
|
+
Two files per walkthrough, and the separation is load-bearing:
|
|
238
|
+
|
|
239
|
+
- `<name>.lines.mjs` — an ordered array of sentences. Nothing else.
|
|
240
|
+
- `<name>.spec.ts` — what happens on screen for each sentence, in the same order.
|
|
241
|
+
|
|
242
|
+
**A beat is one idea being explained, NOT one screen.** A beat may dwell on three
|
|
243
|
+
things within a screen, or carry across a navigation. Building around screens is
|
|
244
|
+
what produced fixed-length steps and the drift that followed.
|
|
245
|
+
|
|
246
|
+
Narration rules, each from a user correction:
|
|
247
|
+
|
|
248
|
+
- **Never name something the viewer cannot see.** "Groups", "the tabs", "the
|
|
249
|
+
address bar" — if the sentence names a thing, the step must point at that
|
|
250
|
+
thing. Do not invent jargon; say "the left sidebar's top-level menus, which
|
|
251
|
+
expand to show…".
|
|
252
|
+
- **Explain, do not sell.** No "exciting", no "powerful", no enthusiasm. A
|
|
253
|
+
colleague showing you how the job is done.
|
|
254
|
+
- **Give the dependency context.** Why this screen exists, and what downstream
|
|
255
|
+
reads from it.
|
|
256
|
+
- **Count what is on screen before writing about it.** "Eight-step wizard" shipped
|
|
257
|
+
in a video where the UI says *Step 1 of 9*.
|
|
258
|
+
- **Say the cast's names, every time.** After Step 4's cast is introduced, "the
|
|
259
|
+
student" and "the course" are bugs — it is *Maya Ellison* and the *Private
|
|
260
|
+
Pilot Certificate — Part 61*. Read the finished narration once looking only
|
|
261
|
+
for this.
|
|
262
|
+
- **Narrate the value being typed, not the field's purpose.** "Her first lesson
|
|
263
|
+
is Tuesday at nine, with James in the Cessna 172" — not "you would select a
|
|
264
|
+
date, an instructor and an aircraft here".
|
|
265
|
+
- **Naming a whole strip highlights the strip; making a point about one control
|
|
266
|
+
highlights that control; navigating by it moves the mouse and clicks it.**
|
|
267
|
+
|
|
268
|
+
---
|
|
269
|
+
|
|
270
|
+
## Step 6 — the five-stage build
|
|
271
|
+
|
|
272
|
+
Run in this order, every time. Nothing here is optional.
|
|
273
|
+
|
|
274
|
+
```
|
|
275
|
+
0. PREFLIGHT playwright test --grep preflight — check every target, no recording
|
|
276
|
+
1. VOICE walkthrough-voice-ensure.mjs <name> — render, measure, re-render if it drifts
|
|
277
|
+
2. RECORD playwright test --grep "<name>" — one continuous run
|
|
278
|
+
3. MUX walkthrough-mux.mjs <name> — lay audio on the recording
|
|
279
|
+
4. TRIM walkthrough-trim.mjs <name> — CUT THE SILENCE
|
|
280
|
+
```
|
|
281
|
+
|
|
282
|
+
### Stage 0 — Preflight
|
|
283
|
+
|
|
284
|
+
A single wrong selector fails an entire 5–8 minute recording, at the first
|
|
285
|
+
sentence whose target is missing. Three bad selectors therefore cost three full
|
|
286
|
+
recordings to find. Preflight walks the same screens without recording and
|
|
287
|
+
reports **every** missing target in one pass. It never asserts; it prints a
|
|
288
|
+
report. The real gate is still the recording itself.
|
|
289
|
+
|
|
290
|
+
### Stage 1 — Voice
|
|
291
|
+
|
|
292
|
+
**Pick a text-to-speech service whose voice is a FIXED TRAINED SPEAKER, not a
|
|
293
|
+
language model reading a style hint.** This one choice decides whether the
|
|
294
|
+
narrator can drift at all, and everything else in this stage follows from it.
|
|
295
|
+
|
|
296
|
+
The source run learned it the expensive way. It started on a language model's
|
|
297
|
+
audio output (Gemini `generateContent`), where a voice name is a *style hint the
|
|
298
|
+
model re-interprets on every request* — so the narrator audibly changed
|
|
299
|
+
part-way through a video. Enormous effort went into mitigation: batching 8
|
|
300
|
+
sentences per request so one request meant one performance, splitting the
|
|
301
|
+
returned audio back apart on silences, verifying every split, and re-rendering
|
|
302
|
+
whole takes that measured as drifted. It reduced how OFTEN the voice changed and
|
|
303
|
+
could never stop it, because **a batch boundary is still a boundary between two
|
|
304
|
+
different readings**.
|
|
305
|
+
|
|
306
|
+
Moving to a dedicated speech service (Google Cloud Text-to-Speech) ended it in
|
|
307
|
+
one change. A voice id there is a fixed trained speaker: the same id returns the
|
|
308
|
+
same speaker every time, forever. That removed batch boundaries, silence
|
|
309
|
+
splitting, miscut clips, per-model daily quotas, and drift — all at once — and
|
|
310
|
+
made **one sentence per request** both the simple thing and the correct thing.
|
|
311
|
+
|
|
312
|
+
> **If you are on a language-model TTS and cannot switch**, the mitigation is
|
|
313
|
+
> batching: 8 sentences per request (measured — an 18-line single take spread
|
|
314
|
+
> 46 Hz of pitch, 8-line batches 16 Hz), a verified split that retries rather
|
|
315
|
+
> than guessing where a sentence ended, and a re-render loop. Treat it as a
|
|
316
|
+
> workaround, not a design.
|
|
317
|
+
|
|
318
|
+
**Volume is forced, on every service.** Two-pass ffmpeg `loudnorm` to a fixed
|
|
319
|
+
target (EBU R128, −18 LUFS). Measured on the shipped set: ±0.3–0.7 dB within a
|
|
320
|
+
video, every video landing on the same target so they match each other too. This
|
|
321
|
+
is worth doing even with a fixed speaker — it is the one number a service will
|
|
322
|
+
not hold steady for you.
|
|
323
|
+
|
|
324
|
+
**Every narration line must be a full sentence — 8 words or more.** Integrated
|
|
325
|
+
loudness needs enough audio to measure against; a two- or three-word clip
|
|
326
|
+
("Create Curriculum.") lands off target and blows the video's volume spread,
|
|
327
|
+
and also trips miscut and speaking-rate checks. If a step needs an action the
|
|
328
|
+
narration does not describe, run it untimed rather than inventing a stub line.
|
|
329
|
+
|
|
330
|
+
### The gate — measure IDENTITY, not expressiveness
|
|
331
|
+
|
|
332
|
+
Measure per clip and report across the video: loudness (LUFS), pitch (median
|
|
333
|
+
fundamental, Hz), speaking rate (energy peaks/sec), and **miscut clips** — a
|
|
334
|
+
clip whose length is far from what its sentence should take to say.
|
|
335
|
+
|
|
336
|
+
**Gate on the MEAN pitch, not the within-video spread.** This is the correction
|
|
337
|
+
that matters, and it was found by running the gate against twelve videos a human
|
|
338
|
+
had already confirmed sounded perfect: **a ±35 Hz spread threshold failed nine
|
|
339
|
+
of them.** Within-video pitch spread is ordinary sentence intonation — a
|
|
340
|
+
question rising, a list falling, a short line sitting higher — and flattening it
|
|
341
|
+
would make the narration robotic. The speaker-identity signal is the mean: across
|
|
342
|
+
those same twelve videos it sat in a 7 Hz band (102–109 Hz), which is what a
|
|
343
|
+
fixed speaker looks like.
|
|
344
|
+
|
|
345
|
+
Working thresholds: volume ±1.5 dB, mean pitch inside a band calibrated from
|
|
346
|
+
known-good output, rate ±5.5/s, zero miscuts.
|
|
347
|
+
|
|
348
|
+
> **Calibrate a gate against output a human has approved, before trusting it.**
|
|
349
|
+
> A gate that fails most of your known-good work is measuring the wrong thing,
|
|
350
|
+
> and the cost of believing it is re-rendering audio that was already correct.
|
|
351
|
+
|
|
352
|
+
**Ensure, don't just check.** Render → measure → re-render if it fails (3
|
|
353
|
+
attempts, then halt rather than ship). Even with a fixed speaker this catches a
|
|
354
|
+
bad take: `build-a-course` failed its first attempt on volume and passed the
|
|
355
|
+
second.
|
|
356
|
+
|
|
357
|
+
### Stage 2 — Record
|
|
358
|
+
|
|
359
|
+
One continuous run, real browser, real mouse movement, real clicks, real
|
|
360
|
+
transitions. Pin the video size in the Playwright project config (an unpinned
|
|
361
|
+
size gets downscaled then upscaled to a blurry result).
|
|
362
|
+
|
|
363
|
+
Two assertions the runtime must enforce, both from shipped errors:
|
|
364
|
+
|
|
365
|
+
- **A narrated step with no visible target throws.** Mark genuinely abstract
|
|
366
|
+
lines explicitly; everything else must point at something.
|
|
367
|
+
- **`step()` also asserts the page.** A failed click used to produce confident
|
|
368
|
+
narration about a screen that was never on camera.
|
|
369
|
+
|
|
370
|
+
Give pages **6–8 seconds to populate, not 5.** Reading at 5s reported a populated
|
|
371
|
+
page as empty and produced a wrong "this module is empty" call.
|
|
372
|
+
|
|
373
|
+
The run writes a **step log** — index, start/end ms, narration, measured audio
|
|
374
|
+
ms, route, focus, action. That log is the proof of what was on screen when, and
|
|
375
|
+
the mux reads it. Copy the recording out of `test-results/` immediately;
|
|
376
|
+
Playwright wipes that directory each run, and a kept copy means a re-mux never
|
|
377
|
+
needs a re-record.
|
|
378
|
+
|
|
379
|
+
**When a recording fails, look at `test-results/<name>/test-failed-1.png`
|
|
380
|
+
FIRST.** Every wrong theory in the source run came from reasoning about the DOM
|
|
381
|
+
instead of looking at the picture the run had already saved. One three-recording
|
|
382
|
+
failure was diagnosed twice-wrongly (virtualised rows, then a timing race) before
|
|
383
|
+
the failure screenshot showed the truth immediately.
|
|
384
|
+
|
|
385
|
+
> **Read UI state before clicking it.** Filter chips that are already ON look
|
|
386
|
+
> identical to buttons that turn something on. Clicking "Aircraft" removed every
|
|
387
|
+
> aircraft row and the next narrated sentence pointed at nothing. Read the state
|
|
388
|
+
> and click only to *change* it.
|
|
389
|
+
|
|
390
|
+
### Stage 3 — Mux
|
|
391
|
+
|
|
392
|
+
Place each clip at the timestamp the step log recorded. Both sides come from the
|
|
393
|
+
same measured timeline, so nothing needs aligning afterward.
|
|
394
|
+
|
|
395
|
+
Three things it must handle:
|
|
396
|
+
|
|
397
|
+
- **Wait for the recording to settle.** Playwright finishes writing the `.webm`
|
|
398
|
+
*after* the test function returns; a copy taken instantly can be short, and the
|
|
399
|
+
tail of the narration then plays over black.
|
|
400
|
+
- **Trim the unnarrated head** — sign-in, gates, first navigation are all on tape
|
|
401
|
+
before the first word.
|
|
402
|
+
- **Never let two sentences overlap.** If a clip is still playing when the next is
|
|
403
|
+
due, start the next after it ends. Re-rendering narration at a different speed
|
|
404
|
+
or voice makes clips no longer fit the slots the recording left for them; a
|
|
405
|
+
slightly late sentence is far less noticeable than two voices at once. (The
|
|
406
|
+
user caught this as "it's almost overspeaking at the transitions" and as a
|
|
407
|
+
clipped sentence end.)
|
|
408
|
+
- **Say so loudly if narration outruns the recording** — that means the capture
|
|
409
|
+
was cut short and the video must be re-recorded, not shipped quietly.
|
|
410
|
+
|
|
411
|
+
### Stage 4 — Trim (never skip)
|
|
412
|
+
|
|
413
|
+
The recording pauses on every page load, which lands **20–60 seconds of dead air**
|
|
414
|
+
in each finished video and makes a 3-minute walkthrough feel far longer.
|
|
415
|
+
`auto-editor <file> --margin 0.2s` cuts every stretch where nobody is speaking.
|
|
416
|
+
Across twelve videos this removed **7.4 minutes** total.
|
|
417
|
+
|
|
418
|
+
Run it after **every** mux. It rewrites the MP4 in place and is safe to re-run.
|
|
419
|
+
Bundle auto-editor in a project-local venv so it does not depend on PATH.
|
|
420
|
+
|
|
421
|
+
The user's verdict on this stage was "It's perfect. Run this after every video."
|
|
422
|
+
|
|
423
|
+
---
|
|
424
|
+
|
|
425
|
+
## Step 7 — VERIFY before showing the user
|
|
426
|
+
|
|
427
|
+
The user should never be the one who finds these. Check, per video:
|
|
428
|
+
|
|
429
|
+
1. **Every planned screen and tab was actually visited** — read the step log's
|
|
430
|
+
routes, not the spec source.
|
|
431
|
+
2. **Every narrated claim points at something** — the runtime enforces it, but
|
|
432
|
+
confirm no line is wrongly marked abstract.
|
|
433
|
+
3. **Voice gate passes** — one narrator, one volume, zero miscuts.
|
|
434
|
+
4. **No narration past the end of the recording.**
|
|
435
|
+
5. **No implied click that did not happen** — a `highlight` step immediately
|
|
436
|
+
followed by a `goto` reads on camera as "they clicked that and it took us
|
|
437
|
+
here". It did not. Either click the real navigation, or park the cursor
|
|
438
|
+
somewhere neutral before navigating. Find them by scanning the step log for
|
|
439
|
+
`highlight` → `goto` adjacency. (Audited at 22 instances across 6 videos in
|
|
440
|
+
the source run.)
|
|
441
|
+
6. **Counts in the narration match the UI** — tabs, wizard steps, row counts.
|
|
442
|
+
7. **The cast is named throughout** — grep the finished narration for "the
|
|
443
|
+
student", "the course", "a user", "the aircraft". After the cast is
|
|
444
|
+
introduced, each of those is a line that should say a name instead.
|
|
445
|
+
8. **Data was entered, not described** — any step whose sentence explains what a
|
|
446
|
+
field is for should be typing into that field. A form that stays empty while
|
|
447
|
+
the narration explains it is the clinical failure this exists to prevent.
|
|
448
|
+
9. **Every dropdown was opened and picked** — grep the spec for `highlight(`
|
|
449
|
+
on a select; each one should be `choose()`. The step log records `choose`
|
|
450
|
+
as its own action, so count them against the number of selects in the flow.
|
|
451
|
+
|
|
452
|
+
Then show the user each video as it finishes, not in a batch at the end.
|
|
453
|
+
|
|
454
|
+
---
|
|
455
|
+
|
|
456
|
+
## Quota and auth — only if you are on a language-model TTS
|
|
457
|
+
|
|
458
|
+
A dedicated speech service typically bills per character with no per-model daily
|
|
459
|
+
cap, and authenticates with a cloud login rather than an API key (Google Cloud
|
|
460
|
+
TTS is OAuth-only — it refuses an API key, and needs a quota project). If that is
|
|
461
|
+
what you are on, this section does not apply.
|
|
462
|
+
|
|
463
|
+
On a language-model TTS the quota rules bite hard:
|
|
464
|
+
|
|
465
|
+
- A free tier can be as low as **10 requests per day per model**, and the error
|
|
466
|
+
names it (`…PerDayPerProjectPerModel-FreeTier`). It dies almost immediately.
|
|
467
|
+
- A paid project's quota is separate; a **new key on a new project** is the
|
|
468
|
+
reliable way past a spent one.
|
|
469
|
+
- Because quota counts **per model**, switching model is also a way past a spent
|
|
470
|
+
allowance. Order the models in a list and skip a spent one for the rest of the
|
|
471
|
+
run rather than sleeping on a ~22-hour reset.
|
|
472
|
+
- A billing page showing a balance does not mean the key works.
|
|
473
|
+
|
|
474
|
+
Batching pays for itself here: 188 sentences across 12 videos cost roughly **28
|
|
475
|
+
requests** total.
|
|
476
|
+
|
|
477
|
+
Pace requests deliberately (several seconds apart) — the per-minute limit bites
|
|
478
|
+
before the per-day one.
|
|
479
|
+
|
|
480
|
+
---
|
|
481
|
+
|
|
482
|
+
## Files this command creates
|
|
483
|
+
|
|
484
|
+
```
|
|
485
|
+
docs/demo-videos/PLAN.md the coverage map
|
|
486
|
+
docs/demo-videos/HANDOFF.md hard-won facts, bugs found, what is open
|
|
487
|
+
docs/demo-videos/walkthrough-<name>.mp4 output (gitignore it)
|
|
488
|
+
e2e/walkthrough/<name>.lines.mjs narration, one sentence per entry
|
|
489
|
+
e2e/walkthrough/<name>.spec.ts what happens on screen per sentence
|
|
490
|
+
e2e/walkthrough/runtime.ts step(), highlight(), click(), enter(), choose(), act(), goTo()
|
|
491
|
+
e2e/walkthrough/cast.mjs the demo's cast — every name said aloud
|
|
492
|
+
e2e/walkthrough/signin.ts shared sign-in + the tenant constant
|
|
493
|
+
e2e/walkthrough/manifest.ts loads narration; skips if audio is missing
|
|
494
|
+
e2e/walkthrough/preflight.spec.ts checks every target without recording
|
|
495
|
+
scripts/walkthrough-voice-gcloud.mjs the narrator — fixed-voice TTS, one sentence/request
|
|
496
|
+
scripts/walkthrough-voice.mjs SUPERSEDED — batched language-model TTS
|
|
497
|
+
scripts/walkthrough-voice-check.mjs volume spread + MEAN-pitch identity + miscuts; exit 4
|
|
498
|
+
scripts/walkthrough-voice-ensure.mjs render → measure → re-render → halt
|
|
499
|
+
scripts/walkthrough-normalise.mjs force existing clips to one loudness
|
|
500
|
+
scripts/walkthrough-mux.mjs lay audio on the recording
|
|
501
|
+
scripts/walkthrough-trim.mjs remove silent gaps
|
|
502
|
+
scripts/demo-data/seed-lib.mjs shared seeder sign-in + refusal reporting
|
|
503
|
+
.demo-build/ audio, recordings, step logs (gitignore)
|
|
504
|
+
```
|
|
505
|
+
|
|
506
|
+
**Working templates for all of these ship with the GSD-T package. Copy them —
|
|
507
|
+
do not re-derive the pipeline.** Resolve the package directory first:
|
|
508
|
+
|
|
509
|
+
```bash
|
|
510
|
+
GSD_T_DIR=$(npm root -g 2>/dev/null)/@tekyzinc/gsd-t
|
|
511
|
+
TPL="$GSD_T_DIR/templates/demo-videos"
|
|
512
|
+
[ -d "$TPL" ] || { echo "demo-video templates not found at $TPL"; exit 1; }
|
|
513
|
+
|
|
514
|
+
mkdir -p e2e/walkthrough scripts scripts/demo-data docs/demo-videos
|
|
515
|
+
cp "$TPL"/e2e/* e2e/walkthrough/
|
|
516
|
+
cp "$TPL"/scripts/walkthrough-*.mjs scripts/
|
|
517
|
+
cp "$TPL"/scripts/seed-lib.mjs scripts/demo-data/
|
|
518
|
+
python3 -m venv .venv-tools/auto-editor
|
|
519
|
+
.venv-tools/auto-editor/bin/pip install auto-editor
|
|
520
|
+
printf '\n.demo-build/\n.tts-cache-v2/\ndocs/demo-videos/*.mp4\n' >> .gitignore
|
|
521
|
+
```
|
|
522
|
+
|
|
523
|
+
**Halt if that directory is absent** — an older installed package does not carry
|
|
524
|
+
it, and re-deriving the pipeline is exactly what this command exists to prevent.
|
|
525
|
+
Run `/gsd-t-version-update` and try again. `$TPL/README.md` carries the wiring
|
|
526
|
+
details, including the Playwright project config.
|
|
527
|
+
|
|
528
|
+
**A spec whose audio has not been rendered yet must SKIP, not throw.** Playwright
|
|
529
|
+
imports every spec before applying `--grep`, so an import-time throw takes down
|
|
530
|
+
the whole run including the walkthroughs that were ready.
|
|
531
|
+
|
|
532
|
+
---
|
|
533
|
+
|
|
534
|
+
## Captions
|
|
535
|
+
|
|
536
|
+
**Off by default.** They were built, then removed at the user's request — "the
|
|
537
|
+
captions just get in the way and are not needed, just the voice over". If a
|
|
538
|
+
project wants them: bottom of frame, ~50% width, not full width, drop shadow
|
|
539
|
+
(they are unreadable against white), and scroll any narrated target clear of the
|
|
540
|
+
caption strip.
|
|
541
|
+
|
|
542
|
+
---
|
|
543
|
+
|
|
544
|
+
## Handoff document
|
|
545
|
+
|
|
546
|
+
Maintain `docs/demo-videos/HANDOFF.md` throughout, holding everything that would
|
|
547
|
+
otherwise be re-derived: the tenant ID and why it is counter-intuitive, settle
|
|
548
|
+
times, quota state per key, the pipeline stages, **which app bugs were found
|
|
549
|
+
while probing**, and what is still open. The source run's handoff is what made a
|
|
550
|
+
context-clear-and-resume possible at all.
|
|
551
|
+
|
|
552
|
+
App bugs found while filming are a real deliverable — they go to the team, not
|
|
553
|
+
into the video.
|
|
554
|
+
|
|
555
|
+
---
|
|
556
|
+
|
|
557
|
+
## Document Ripple
|
|
558
|
+
|
|
559
|
+
| Trigger | Update |
|
|
560
|
+
|---|---|
|
|
561
|
+
| Walkthrough added/removed | `docs/demo-videos/PLAN.md` + `HANDOFF.md` |
|
|
562
|
+
| Pipeline stage changed | `HANDOFF.md` pipeline block + this command file |
|
|
563
|
+
| Voice settings changed | `HANDOFF.md` (and re-render — settings are in the cache key) |
|
|
564
|
+
| Seeder added | `HANDOFF.md` § Seeding, with what it refuses and why |
|
|
565
|
+
| App bug found while probing | `HANDOFF.md` § Bugs + `.gsd-t/techdebt.md` |
|
|
566
|
+
| New test projects in `playwright.config.ts` | note it — this command must not modify application code |
|
|
567
|
+
| Any file changed | `.gsd-t/progress.md` Decision Log entry |
|
|
568
|
+
|
|
569
|
+
**This command modifies no application code.** Its footprint is the walkthrough
|
|
570
|
+
directory, the scripts, two Playwright test projects, and `.gitignore`.
|
|
571
|
+
|
|
572
|
+
---
|
|
573
|
+
|
|
574
|
+
## ▶ Next Up
|
|
575
|
+
|
|
576
|
+
**Verify** — check coverage, voice gate, and implied clicks before shipping.
|
|
577
|
+
|
|
578
|
+
`/gsd-t-verify`
|
|
579
|
+
|
|
580
|
+
**Also available:**
|
|
581
|
+
- `/gsd-t-backlog-add` — record any app bugs the probing turned up
|