@bigsteele/the-big-sean 0.3.0 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -26,6 +26,21 @@ npx @bigsteele/the-big-sean --check # after the audit: re-compute the report'
26
26
  npx @bigsteele/the-big-sean --stdout # print the protocol
27
27
  ```
28
28
 
29
+ ## It grades against your North Star, not the rubric's
30
+
31
+ A rubric with 140 checks will always find things missing. Before any grading, the audit
32
+ establishes what the software is actually for, from evidence in the repository: the
33
+ pricing page, the money path, the schema, the code that got the most effort. That becomes
34
+ `NORTH-STAR.md` (the one sentence, the value moment, the money path, the stage, the three
35
+ to five workflows the business dies without, and what is deliberately out of scope), and
36
+ every finding afterwards has to earn its place against it.
37
+
38
+ So each recommendation carries one line saying which workflow or money path it protects,
39
+ the order follows the business rather than the rubric's numbering, and anything that
40
+ cannot make that connection lands in a section called **"Rubric items that do not serve
41
+ your North Star"** with a reason it waits. The audit is not allowed to tell you to build
42
+ something your product does not need just because a check exists for it.
43
+
29
44
  ## What it targets
30
45
 
31
46
  The audit reads **this repository**, and any live system it reaches must be one this
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@bigsteele/the-big-sean",
3
- "version": "0.3.0",
3
+ "version": "0.4.0",
4
4
  "description": "The Big Sean: the Launch Report Card, by Big Steele and LaSean Pickens. Runs the Big Steele AI Audit, drops a 140-check launch-readiness and autonomy audit protocol into a repo for Claude Code to run, names the deliverable after the app, and independently re-computes any report's math from its grade.json so the number on the card can be proven, not trusted.",
5
5
  "type": "module",
6
6
  "license": "UNLICENSED",
@@ -47,8 +47,21 @@ The access ladder. For every system, climb every rung before you write UNKNOWN.
47
47
  Never ask me to connect something. If a rung needs a login prompt, skip it. If a rung would write, skip it.
48
48
  Write a table SOURCES.md: every system you found, the identifier you targeted and the file and line in this repository that named it, whether that identifier was confirmed from config or only claimed in prose, the rung that got you in (or every rung that failed), whether you could reach it live or only in code, and what you could not reach. A row with a live read and no confirmed identifier is a defect in the audit, not a finding about the app. This table is the first thing on the report card. If you reached nothing live, the report is titled CODE-ONLY REPORT CARD and every live-dependent category is marked UNKNOWN with the reason. Never pretend a code read was a live check.
49
49
  PART A. LAUNCH READINESS
50
- Step 2. Learn the app
51
- Read the planning docs, README, specs, and any state files first. Then read the real code: every route, page, API handler, server action, edge or serverless function, worker, cron, webhook, database migration, policy, trigger, function, storage rule, and integration. Trace the main user journeys end to end: sign up, first value, the core workflow of this product, pay, cancel, delete account, get support. Trace the owner's journeys: onboarding a customer, seeing what happened, handling a failure, getting paid. Write WHAT-THIS-APP-DOES.md: what it is, who uses it, the money flows, the external services, the background jobs, and the workflows that must work for a customer to pay and stay.
50
+ Step 2. Learn the app, then fix the North Star before you grade anything
51
+ NORTH STAR FIRST. A rubric with one hundred forty checks will always find things missing. Without an anchor, the report becomes a list of what the rubric wanted instead of a plan for this business, and the owner is handed work that has nothing to do with what they are building. Establish what this software is for BEFORE Part A, and judge every recommendation against it afterward.
52
+ Derive it from evidence, not from a question to me. Read, in this order, and cite the file and line for each claim: the README and landing or marketing copy (what it promises and to whom); the pricing page, plans table, or products in the payments provider (what someone actually pays for, which is the sharpest statement of value in the repository); planning docs, PRDs, roadmaps, and issue titles (what is being built next); the schema (the entities the business is actually about, and which tables carry the money); the code paths that are most built out and most defended (where the effort went); and analytics or event names (what the team decided to measure).
53
+ Write NORTH-STAR.md, and keep every line traceable:
54
+ • The one sentence. What this software is for, in the owner's terms, not the rubric's. Who it serves, what it replaces, and what it promises.
55
+ • The value moment. The single step where a user first gets the thing they came for, named as a route, handler, or job. Everything upstream is acquisition; everything downstream is retention.
56
+ • The money path. The exact sequence from intent to paid, and what breaks the business if it fails. If there is provably no money yet, say so and name what stands in its place (activation, a waitlist, an internal cost saved).
57
+ • The stage. Pre-launch, private beta, paying customers, or scaling, with the evidence that says so (customers in the database, live keys, a deploy history, a support surface). A pre-launch app and a platform with five thousand tenants deserve different advice from the same rubric.
58
+ • The critical few. The three to five workflows this business dies without. These are the ones whose failures are P0 no matter what the score says.
59
+ • Out of scope by design. What this app deliberately does not do, and where the repo says so. A deliberate omission is not a gap.
60
+ • Confidence. High where a pricing page and a schema agree; low where you are inferring from code shape alone. Say which.
61
+ • Anything you could not determine. Write UNKNOWN and say what would settle it. Never invent a mission.
62
+ If the evidence contradicts itself, that is a finding, not a puzzle to resolve quietly: a README promising one thing while the code and the money path serve another means the team does not agree on what it is building, and it goes in the report as its own finding with both citations.
63
+ THE DRIFT RULE, which governs the rest of this audit. Every FAIL, every UNKNOWN, and every recommendation must connect to the North Star. In the action plan and the path to 100, each item carries one line, "why this matters here", that names the workflow, the money path, or the critical few it protects. An item that cannot make that connection is not dropped and not silently downgraded: it is grouped under "Rubric items that do not serve your North Star" with a one-line reason, so the owner sees you considered it and why it waits. Ordering follows the North Star, not the rubric's numbering: two weight-5 failures are not equal when one breaks the money path and the other breaks a surface nobody has shipped yet. Never recommend building a capability the North Star does not need just because a check exists for it. If the audit's honest conclusion is that the app should do less, say that.
64
+ Now learn the app. Read the planning docs, README, specs, and any state files first. Then read the real code: every route, page, API handler, server action, edge or serverless function, worker, cron, webhook, database migration, policy, trigger, function, storage rule, and integration. Trace the main user journeys end to end: sign up, first value, the core workflow of this product, pay, cancel, delete account, get support. Trace the owner's journeys: onboarding a customer, seeing what happened, handling a failure, getting paid. Write WHAT-THIS-APP-DOES.md: what it is, who uses it, the money flows, the external services, the background jobs, and the workflows that must work for a customer to pay and stay. Where it disagrees with NORTH-STAR.md, the code wins for what the app DOES and the North Star governs what it is FOR; note the gap.
52
65
  Step 3. Check the live system where you can (all read-only)
53
66
  Where the database is reachable:
54
67
  • Every table with its row count. Flag tables that are empty but should have content (an academy with no lessons, a pricing table with no rows, a policies table with no current version).
@@ -277,7 +290,7 @@ Verification and unattended outcomes
277
290
  Lock the product's expected workflows, control applicability, test cases, and measurable acceptance thresholds before running checks. Derive thresholds from verified business contracts or explicitly approved design targets, never invent industry averages. Application-specific questions must be made concrete: which route, tenant role, record, expected result, failure condition, and acceptable bounds? Record all mapping details in the criterion's evidence packet. One root cause can explain several distinct failed outcomes; deduplicate the repair task, not independent outcome failures. Do not count the same control twice under different departments; where a D control and an L control test the same thing, run the test once and cite the same evidence from both.
278
291
  CONTROL STATUS: PASS means all defined tests passed with applicable evidence. FAIL means an observed behavior violates the defined acceptance condition. UNKNOWN means untested, inaccessible, skipped, inconclusive, or stale evidence; attach the exact reason. N/A requires affirmative non-applicability evidence and rationale. Unknown is not a proven defect. Empty data, disabled features, and unavailable credentials do not establish N/A. Mixed or partial outcomes must be recorded in subtests and resolve to FAIL if a required subtest fails, otherwise UNKNOWN if any required subtest is unresolved. Source existence alone cannot pass an execution requirement.
279
292
  EVIDENCE REQUIREMENTS: each control has scope, plane (platform, client, or both), product and workflow mapping, environment, immutable commit, deployment, and schema identifiers where relevant, UTC time, expected and actual result, executed test, query, or command, sanitized artifact reference, and evaluator identity. Screenshots support UX evidence but cannot alone certify backend security. A PASS or FAIL without sufficient evidence becomes UNKNOWN. N/A without sufficient rationale becomes UNKNOWN.
280
- ACTION PLAN: every FAIL and UNKNOWN has a plain-language problem or missing-proof statement, who is affected, business consequence, confirmed cause or explicit hypothesis, immediate containment where needed, exact file, function, table, or config scope, dependencies, atomic implementation or verification tasks, responsible role, effort basis if estimable, and objective retest. Unknowns get investigate and prove actions rather than fabricated fixes. Link actions to PRDs, commits, tests, and affected criteria. Prioritize P0 for confirmed high-impact exposure with containment; then verification of critical unknowns before dependent release; then dependency-unblocking correctness and recovery work; then ordinary capability and optimization gaps. State evidence for priority rather than ranking by score improvement alone.
293
+ ACTION PLAN: every FAIL and UNKNOWN has a plain-language problem or missing-proof statement, who is affected, business consequence, confirmed cause or explicit hypothesis, immediate containment where needed, exact file, function, table, or config scope, dependencies, atomic implementation or verification tasks, responsible role, effort basis if estimable, and objective retest. Unknowns get investigate and prove actions rather than fabricated fixes. Link actions to PRDs, commits, tests, and affected criteria. Prioritize against the North Star first: anything that breaks a critical-few workflow or the money path outranks anything that does not, whatever its weight. Then P0 for confirmed high-impact exposure with containment; then verification of critical unknowns before dependent release; then dependency-unblocking correctness and recovery work; then ordinary capability and optimization gaps. Every item carries its one-line "why this matters here". State evidence for priority rather than ranking by score improvement alone.
281
294
  FAIR REVIEW: use a separate evaluator when available (a sub-agent that inspects evidence without adopting the builder's conclusions). Record disagreements and resolve them by reproducible tests or leave UNKNOWN. For stochastic AI tests record dataset and version, repetitions, and outcome distribution; do not select only passing runs. Record sampling limits. Never invent a reviewer or claim independent verification when none occurred.
282
295
  B12. ASSESS EXISTING UNATTENDED PROOF AND SPECIFY FUTURE TESTS
283
296
  Scope: inspect existing evidence and document the proposed target state only. Do not apply changes.
@@ -323,12 +336,14 @@ Step 6. Write the report card
323
336
  Create .planning/launch-audit/REPORT-CARD.html (single file, inline CSS and JS, opens from disk, no external requests) and REPORT-CARD.md with identical content. Then copy both to the repository root as "The Big Sean - <App Name>.html" and "The Big Sean - <App Name>.md", where <App Name> is the product's real name (the brand a customer would recognize, from the manifest, the README, or the UI; the folder name only if nothing better exists). Those two root files are the deliverable a person shares; the report folder keeps the working artifacts. Build both from a grade.json you write first, so every number is computed. Structure, in this order:
324
337
  • Headline: verified score and band, ceiling, coverage, gate status (LAUNCH BLOCKED, NOT VERIFIED FOR LAUNCH, or GATES CLEAR), the launch-readiness score and the autonomy score side by side, the unattended verdict with the duty counts, the one-sentence verdict, the date, the commit, the sources reached live and not reached.
325
338
  • The thirty-five category cards (L01 to L15, then D01 to D20). Each card shows: the category score and band; the four checks with PASS, FAIL, UNKNOWN, or N/A and their weight; Why you scored this (one plain-language paragraph per check that is not PASS, citing the file and line, the query, or the page); What this costs you (what happens to a customer, to your money, or to your reputation if you launch like this); How to fix it (the exact change, where, and the test that proves it); How to verify it yourself (a command or a click path).
326
- • Your path to 100: every check that is not PASS, across all categories, in the order to do them. Blockers first (weight-5 FAILs), then the weight-5 UNKNOWNs to prove, then everything else in dependency order. Each with the change, the location, the test, and a rough effort (hours, one day, multi-day).
327
- • Your next ten actions: the first ten items from the path, written as instructions a person can start today.
339
+ • Your North Star, first, before any score: the one sentence, the value moment, the money path, the stage, the critical few, and what is out of scope by design, each cited. Every section below is judged against it. If the evidence contradicted itself, that contradiction appears here as a finding.
340
+ • Your path to 100: every check that is not PASS, across all categories, in the order to do them, each with its one-line "why this matters here" naming the workflow, money path, or critical-few item it protects. Blockers first (weight-5 FAILs), then the weight-5 UNKNOWNs to prove, then everything else in dependency order. Each with the change, the location, the test, and a rough effort (hours, one day, multi-day).
341
+ • Your next ten actions: the first ten items from the path, written as instructions a person can start today, ordered by what the North Star says matters, not by the rubric's numbering.
342
+ • Rubric items that do not serve your North Star: every check that is not PASS and could not be connected to the mission, money path, or critical few, with one line each on why it waits. This section proves the audit read the business and not only the checklist; an empty section is a claim that all one hundred forty checks matter to this app, so only write that if it is true.
328
343
  • Human work to remove (from Step 4A): two columns, platform owner and tenant staff; every duty a human touches today with its tag removable, partial, or by design, and the change that removes it. Then what still needs a human for this audit: decisions, credentials, dashboard clicks, legal copy, content, and anything I chose not to test.
329
344
  • Everything I found: every finding with an ID, category, severity, file and line or query, expected versus observed, and impact.
330
345
  • What I tested, what I could not, and what I assumed, with counts.
331
- • Show the math: the validator output from score.mjs, unedited, for LAUNCH, for AUTONOMY, and combined. 8A. The autonomy artifacts from Part B, every one, linked from the report card: RUN-STATE, EVIDENCE-REGISTER, AUTONOMY-AUDIT, WORKFLOW-COVERAGE (machine-readable plus the readable matrix), AUTONOMY-CONTRACT, SAFETY-AND-APPLICABILITY (jurisdiction matrix, hazard register, guardrail registry, consent and evidence design, adversarial test suite, incident-response runbook), DATA-INTELLIGENCE-ARCHITECTURE (all eight areas, current versus proposed), ONBOARDING-ACTIVATION-PLAN, WIRING-MAP, IMPLEMENTATION-SERIES, the PRDs, ACCEPTANCE-AND-FAILURE-TESTS, ROLLOUT-RECOVERY, BLOCKERS-AND-DECISIONS, CURRENT-GRADE-SHEET, RANKED-FINDINGS, ACTION-PLAN, FUTURE-RETEST-PLAN, and the machine-readable audit records with rubricVersion, product, environment, commit, assessedAt, evaluator, and every control's id, department, requirement, weight, critical, status, evidence[], reason, expected, actual, causeStatus, rootCauseId, action, dependencies[], retest, owner, effort, naReason, reviewer. A report card without these is a summary, not the deliverable.
346
+ • Show the math: the validator output from score.mjs, unedited, for LAUNCH, for AUTONOMY, and combined. 8A. The artifacts, every one, linked from the report card: NORTH-STAR.md and WHAT-THIS-APP-DOES.md first, then the autonomy artifacts from Part B: RUN-STATE, EVIDENCE-REGISTER, AUTONOMY-AUDIT, WORKFLOW-COVERAGE (machine-readable plus the readable matrix), AUTONOMY-CONTRACT, SAFETY-AND-APPLICABILITY (jurisdiction matrix, hazard register, guardrail registry, consent and evidence design, adversarial test suite, incident-response runbook), DATA-INTELLIGENCE-ARCHITECTURE (all eight areas, current versus proposed), ONBOARDING-ACTIVATION-PLAN, WIRING-MAP, IMPLEMENTATION-SERIES, the PRDs, ACCEPTANCE-AND-FAILURE-TESTS, ROLLOUT-RECOVERY, BLOCKERS-AND-DECISIONS, CURRENT-GRADE-SHEET, RANKED-FINDINGS, ACTION-PLAN, FUTURE-RETEST-PLAN, and the machine-readable audit records with rubricVersion, product, environment, commit, assessedAt, evaluator, and every control's id, department, requirement, weight, critical, status, evidence[], reason, expected, actual, causeStatus, rootCauseId, action, dependencies[], retest, owner, effort, naReason, reviewer. A report card without these is a summary, not the deliverable.
332
347
  • The evidence register: every evidence entry with its ID, the exact query or command, the exact output excerpt (secrets and personal data redacted), the file and line, the time, and which checks cite it. Every PASS and every FAIL on the page links to at least one entry here.
333
348
  • The full inventories (this is where the depth lives; a highlight reel is not a report card):
334
349
  • Every route and page, with its auth requirement, ownership check, input validation, rate limit, and the check IDs that touched it.
@@ -365,6 +380,7 @@ Do not make me look for anything.
365
380
  • Whether or not that worked, open the report in the default browser automatically: start "" "<full path>" on Windows, open "<full path>" on macOS, xdg-open "<full path>" on Linux. Then print the clickable link on its own line: file:///<full path to REPORT-CARD.html>.
366
381
  • Under the link, print the summary card in chat:
367
382
  LAUNCH REPORT CARD
383
+ North Star: <the one sentence, or UNKNOWN with what would settle it>
368
384
  Score: <verified> / 100 (<band>) <INCOMPLETE if any unknowns> Ceiling: <ceiling> Proven either way: <coverage>%
369
385
  Checks: <n> PASS, <n> FAIL, <n> UNKNOWN, <n> N/A of 140 Findings: <count> Evidence entries: <count>
370
386
  Launch readiness (L01 to L15): <score> Autonomy (D01 to D20): <score> Combined: <score>