@produtype/core 0.2.2 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +32 -132
- package/dist/analyzer/analyzeProject.js +7 -1
- package/dist/analyzer/detectAiSafety.js +21 -1
- package/dist/analyzer/detectAudit.d.ts +3 -0
- package/dist/analyzer/detectAudit.js +62 -0
- package/dist/analyzer/detectBackend.js +56 -0
- package/dist/analyzer/detectBilling.js +58 -3
- package/dist/analyzer/detectDatabase.d.ts +1 -0
- package/dist/analyzer/detectDatabase.js +110 -1
- package/dist/analyzer/detectEnv.js +25 -2
- package/dist/analyzer/detectFrontend.js +6 -0
- package/dist/analyzer/detectJobs.js +120 -13
- package/dist/analyzer/detectStack.d.ts +2 -0
- package/dist/analyzer/detectStack.js +2 -0
- package/dist/analyzer/types.d.ts +7 -0
- package/dist/cli.js +9 -2
- package/dist/expectations/evaluateExpectations.js +18 -2
- package/dist/expectations/inferProductProfile.js +26 -6
- package/dist/rules/rules.js +26 -8
- package/dist/utils/fileScanner.js +16 -0
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -4,7 +4,7 @@ ProdKit is a deterministic local CLI that analyzes an existing web application r
|
|
|
4
4
|
|
|
5
5
|
It is designed for early-stage and AI-generated apps where architecture and security quality can vary significantly.
|
|
6
6
|
|
|
7
|
-
ProdKit
|
|
7
|
+
ProdKit is read-only: it never modifies the target repository.
|
|
8
8
|
|
|
9
9
|
## What ProdKit does
|
|
10
10
|
|
|
@@ -110,11 +110,13 @@ Optional AI layer:
|
|
|
110
110
|
- **A separate package.** The AI features live in `@produtype/ai`, which is commercial
|
|
111
111
|
and is not a dependency of this package. An open source install runs the
|
|
112
112
|
deterministic analysis only, and `--ai` reports that the package is not installed
|
|
113
|
-
rather than failing.
|
|
113
|
+
rather than failing. That package brings its own model access; nothing needs to be
|
|
114
|
+
configured here to run the analysis.
|
|
114
115
|
- ProdKit stays deterministic, offline, and read-only: nothing here makes a network
|
|
115
116
|
call.
|
|
116
|
-
-
|
|
117
|
-
|
|
117
|
+
- When `@produtype/ai` is installed, it redacts secrets from repository content before
|
|
118
|
+
sending anything, and its output stays advisory — it never changes the deterministic
|
|
119
|
+
score.
|
|
118
120
|
|
|
119
121
|
Profile warning:
|
|
120
122
|
|
|
@@ -145,145 +147,43 @@ prodkit plan ../my-app --format json
|
|
|
145
147
|
prodkit plan ../my-app --output prodkit-plan.md
|
|
146
148
|
```
|
|
147
149
|
|
|
148
|
-
## Supported stacks
|
|
150
|
+
## Supported stacks
|
|
149
151
|
|
|
150
|
-
- Backend
|
|
151
|
-
|
|
152
|
+
- **Backend (Node):** Express, Next.js, NestJS, Fastify, Hono, Elysia, Koa, AdonisJS,
|
|
153
|
+
SvelteKit, Remix, Nuxt, Nitro
|
|
154
|
+
- **Backend (Python):** Django, FastAPI, Flask, Litestar, Starlette, Sanic, Tornado,
|
|
155
|
+
aiohttp
|
|
156
|
+
- **Frontend:** React, Vue, Svelte, Angular, Solid, Qwik, Preact, Astro, Nuxt, Remix,
|
|
157
|
+
Vite, Tailwind, htmx
|
|
158
|
+
- **Databases:** Postgres, MySQL, SQLite, MongoDB, Redis
|
|
159
|
+
- **Hosted data platforms:** Supabase, Firebase, Neon, PlanetScale, Vercel Postgres,
|
|
160
|
+
Turso, Upstash, DynamoDB, Convex
|
|
161
|
+
- **ORMs:** Prisma, Drizzle, TypeORM, Sequelize, Knex, MikroORM, Kysely, SQLAlchemy,
|
|
162
|
+
Tortoise, Peewee
|
|
152
163
|
- Generic unknown app fallback
|
|
153
164
|
|
|
154
|
-
|
|
165
|
+
Two of these distinctions are deliberate rather than incidental.
|
|
155
166
|
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
prodkit plan tests/fixtures/express-basic
|
|
165
|
-
prodkit plan tests/fixtures/express-basic --format json
|
|
166
|
-
prodkit plan tests/fixtures/express-basic --output prodkit-plan.md
|
|
167
|
-
```
|
|
167
|
+
A hosted platform is recorded separately from the engine underneath it. Supabase is
|
|
168
|
+
Postgres and Turso is SQLite, so rules written about an engine keep working without
|
|
169
|
+
knowing about the host, while "this data lives on infrastructure someone else
|
|
170
|
+
operates" stays a question the report can ask on its own.
|
|
171
|
+
|
|
172
|
+
Astro counts as a backend only when it is configured to serve requests — `output` set
|
|
173
|
+
to `server` or `hybrid`, or an adapter installed. A static Astro site is a static
|
|
174
|
+
site, and is not marked down for missing the things an application needs.
|
|
168
175
|
|
|
169
176
|
## Current limitations
|
|
170
177
|
|
|
171
|
-
- Deterministic heuristics only
|
|
178
|
+
- Deterministic heuristics only: the AI layer is a separate package (see above)
|
|
172
179
|
- Signal-based stack coverage across common Node/Python/JS frameworks + fallback
|
|
173
180
|
- Signal-based detection can produce false positives/negatives
|
|
174
181
|
- Plan output is deterministic and read-only only
|
|
175
|
-
- No cloud dashboard
|
|
182
|
+
- No cloud dashboard or UI: this package is the CLI and the library
|
|
176
183
|
|
|
177
184
|
## Roadmap
|
|
178
185
|
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
3. AI explanation layer
|
|
182
|
-
4. Apply/generate PRs
|
|
183
|
-
5. Cloud GitHub integration
|
|
184
|
-
6. Framework adapters
|
|
185
|
-
7. Dashboard
|
|
186
|
-
|
|
187
|
-
## Project layout
|
|
188
|
-
|
|
189
|
-
```text
|
|
190
|
-
prodkit/
|
|
191
|
-
package.json
|
|
192
|
-
tsconfig.json
|
|
193
|
-
README.md
|
|
194
|
-
src/
|
|
195
|
-
index.ts
|
|
196
|
-
cli.ts
|
|
197
|
-
analyzer/
|
|
198
|
-
planner/
|
|
199
|
-
rules/
|
|
200
|
-
report/
|
|
201
|
-
utils/
|
|
202
|
-
tests/
|
|
203
|
-
fixtures/
|
|
204
|
-
analyzer.test.ts
|
|
205
|
-
report.test.ts
|
|
206
|
-
planner.test.ts
|
|
207
|
-
```
|
|
208
|
-
|
|
209
|
-
## GitHub Action
|
|
210
|
-
|
|
211
|
-
Gate a pull request on production readiness:
|
|
212
|
-
|
|
213
|
-
```yaml
|
|
214
|
-
- uses: denisbilli/prodkit@main
|
|
215
|
-
with:
|
|
216
|
-
profile: b2b-saas
|
|
217
|
-
fail-under: "65"
|
|
218
|
-
```
|
|
219
|
-
|
|
220
|
-
Inputs: `path`, `profile`, `fail-under`, `min-maturity`, `comment-summary`.
|
|
221
|
-
Outputs: `score`, `maturity`, `launch-ready`.
|
|
222
|
-
|
|
223
|
-
The executive summary is written to the job summary even when the thresholds fail the
|
|
224
|
-
job — a check that fails without saying why is worse than no check at all.
|
|
225
|
-
|
|
226
|
-
Requires the package to be published; until then, point the action at a local
|
|
227
|
-
checkout.
|
|
228
|
-
|
|
229
|
-
## MCP server
|
|
230
|
-
|
|
231
|
-
ProdKit ships an MCP server so an agent can assess a repository without leaving the
|
|
232
|
-
editor. It exposes the deterministic analysis only, makes no network calls, and the
|
|
233
|
-
code being analysed never leaves the machine.
|
|
234
|
-
|
|
235
|
-
Tools: `analyze_project`, `plan_remediation`, `compare_profiles`, `list_profiles`.
|
|
236
|
-
|
|
237
|
-
ProdKit is not on npm yet — the name `prodkit` is taken by an unrelated package — so
|
|
238
|
-
install it from a local checkout:
|
|
239
|
-
|
|
240
|
-
```bash
|
|
241
|
-
npm install && npm run build && npm link
|
|
242
|
-
```
|
|
243
|
-
|
|
244
|
-
Claude Code:
|
|
245
|
-
|
|
246
|
-
```bash
|
|
247
|
-
claude mcp add prodkit -- prodkit-mcp
|
|
248
|
-
```
|
|
249
|
-
|
|
250
|
-
Or, in a client that reads a JSON config:
|
|
251
|
-
|
|
252
|
-
```json
|
|
253
|
-
{
|
|
254
|
-
"mcpServers": {
|
|
255
|
-
"prodkit": {
|
|
256
|
-
"command": "prodkit-mcp"
|
|
257
|
-
}
|
|
258
|
-
}
|
|
259
|
-
}
|
|
260
|
-
```
|
|
261
|
-
|
|
262
|
-
`compare_profiles` is the one to reach for first: it scores the same repository
|
|
263
|
-
against every product profile in a single call, which is the question ProdKit exists
|
|
264
|
-
to answer.
|
|
265
|
-
|
|
266
|
-
## Releases
|
|
267
|
-
|
|
268
|
-
Published from CI on a version tag, with
|
|
269
|
-
[npm provenance](https://docs.npmjs.com/generating-provenance-statements): the tarball
|
|
270
|
-
is cryptographically linked to the commit and the workflow run that produced it, so it
|
|
271
|
-
can be verified rather than merely trusted. No publish token ever sits on a developer
|
|
272
|
-
machine.
|
|
273
|
-
|
|
274
|
-
To cut a release:
|
|
275
|
-
|
|
276
|
-
```bash
|
|
277
|
-
npm version patch # or minor / major — commits and tags
|
|
278
|
-
git push --follow-tags
|
|
279
|
-
```
|
|
280
|
-
|
|
281
|
-
The workflow refuses to publish if the tag and `package.json` disagree, and runs a
|
|
282
|
-
check that the commercial AI layer is absent from `dist` before anything leaves.
|
|
283
|
-
Release notes are generated from the commits since the previous tag.
|
|
284
|
-
|
|
285
|
-
## License
|
|
286
|
-
|
|
287
|
-
|
|
186
|
+
Shipped: the deterministic analyzer, deterministic remediation planning, and the
|
|
187
|
+
optional AI layer as `@produtype/ai`.
|
|
288
188
|
|
|
289
|
-
|
|
189
|
+
Next: applying fixes as generated pull requests, and a hosted GitHub integration.
|
|
@@ -42,6 +42,7 @@ const detectPackageManager_1 = require("./detectPackageManager");
|
|
|
42
42
|
const detectFrontend_1 = require("./detectFrontend");
|
|
43
43
|
const detectBackend_1 = require("./detectBackend");
|
|
44
44
|
const detectDatabase_1 = require("./detectDatabase");
|
|
45
|
+
const detectAudit_1 = require("./detectAudit");
|
|
45
46
|
const detectDocker_1 = require("./detectDocker");
|
|
46
47
|
const detectEnv_1 = require("./detectEnv");
|
|
47
48
|
const detectAuth_1 = require("./detectAuth");
|
|
@@ -301,7 +302,7 @@ async function analyzeProject(projectPath) {
|
|
|
301
302
|
npmDeps,
|
|
302
303
|
workspaces,
|
|
303
304
|
};
|
|
304
|
-
const [pm, frontend, backend, database, docker, env, auth, security, uploads, gdpr, billing, observability, jobs, marketplace, aiSafety, engagement, deployment] = await Promise.all([
|
|
305
|
+
const [pm, frontend, backend, database, docker, env, auth, security, uploads, gdpr, billing, observability, jobs, marketplace, aiSafety, engagement, deployment, audit] = await Promise.all([
|
|
305
306
|
(0, detectPackageManager_1.detectPackageManager)(ctx),
|
|
306
307
|
(0, detectFrontend_1.detectFrontend)(ctx),
|
|
307
308
|
(0, detectBackend_1.detectBackend)(ctx),
|
|
@@ -319,12 +320,15 @@ async function analyzeProject(projectPath) {
|
|
|
319
320
|
(0, detectAiSafety_1.detectAiSafety)(ctx),
|
|
320
321
|
(0, detectNotifications_1.detectEngagement)(ctx),
|
|
321
322
|
(0, detectDeployment_1.detectDeployment)(ctx),
|
|
323
|
+
(0, detectAudit_1.detectAudit)(ctx),
|
|
322
324
|
]);
|
|
323
325
|
const detectors = (0, detectStack_1.mergeDetectors)([
|
|
324
326
|
pm.result,
|
|
325
327
|
frontend.result,
|
|
326
328
|
backend.result,
|
|
327
329
|
database.result,
|
|
330
|
+
...database.extra,
|
|
331
|
+
audit,
|
|
328
332
|
docker,
|
|
329
333
|
...env,
|
|
330
334
|
...auth,
|
|
@@ -346,6 +350,8 @@ async function analyzeProject(projectPath) {
|
|
|
346
350
|
frontend: frontend.frameworks,
|
|
347
351
|
backend: backend.frameworks,
|
|
348
352
|
databases: database.databases,
|
|
353
|
+
dataPlatforms: database.extra.find((d) => d.key === 'stack.dataPlatform')?.details?.platforms ?? [],
|
|
354
|
+
orms: database.extra.find((d) => d.key === 'stack.orm')?.details?.orms ?? [],
|
|
349
355
|
packageManager: pm.manager,
|
|
350
356
|
packageManagerConfidence: pm.confidence,
|
|
351
357
|
warnings: pm.warnings,
|
|
@@ -69,6 +69,26 @@ async function detectPromptSafety(ctx) {
|
|
|
69
69
|
},
|
|
70
70
|
};
|
|
71
71
|
}
|
|
72
|
+
/**
|
|
73
|
+
* Whether this product calls a model at all.
|
|
74
|
+
*
|
|
75
|
+
* Kept separate from the two safety checks because it answers a different question.
|
|
76
|
+
* Those ask whether an AI product is built safely; this one asks whether it is an AI
|
|
77
|
+
* product, which is what decides the profile it gets judged against — and profile
|
|
78
|
+
* inference was working it out by matching words like "model" and "prompt" against
|
|
79
|
+
* the evidence strings of the billing, API-key and background-job detectors. A
|
|
80
|
+
* dependency on a model SDK is the thing itself rather than a trace of it.
|
|
81
|
+
*/
|
|
82
|
+
function detectModelProvider(ctx) {
|
|
83
|
+
const deps = modelDependencies(ctx);
|
|
84
|
+
return {
|
|
85
|
+
key: 'ai.modelProvider',
|
|
86
|
+
present: deps.length > 0,
|
|
87
|
+
evidence: deps.map((dep) => ({ type: 'dependency', value: dep })),
|
|
88
|
+
details: { providers: deps },
|
|
89
|
+
};
|
|
90
|
+
}
|
|
72
91
|
async function detectAiSafety(ctx) {
|
|
73
|
-
|
|
92
|
+
const [costControl, promptSafety] = await Promise.all([detectCostControl(ctx), detectPromptSafety(ctx)]);
|
|
93
|
+
return [costControl, promptSafety, detectModelProvider(ctx)];
|
|
74
94
|
}
|
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
"use strict";
|
|
2
|
+
Object.defineProperty(exports, "__esModule", { value: true });
|
|
3
|
+
exports.detectAudit = detectAudit;
|
|
4
|
+
const detectContext_1 = require("./detectContext");
|
|
5
|
+
const textSearch_1 = require("../utils/textSearch");
|
|
6
|
+
/**
|
|
7
|
+
* A durable record of who did what.
|
|
8
|
+
*
|
|
9
|
+
* This is not observability, and conflating the two was the bug this replaces:
|
|
10
|
+
* `audit.baseline` was evaluated from structured logging and a request id, which
|
|
11
|
+
* answer "can we debug this?" rather than "can we say, months later, who deleted that
|
|
12
|
+
* organisation?". Logs rotate; an audit trail is meant not to.
|
|
13
|
+
*
|
|
14
|
+
* Two kinds of evidence, in order of strength. A place to keep the record — a table or
|
|
15
|
+
* a model whose name says what it is — and calls that write to it. A table with
|
|
16
|
+
* nothing writing to it is a good intention; writes with no durable store behind them
|
|
17
|
+
* are log lines.
|
|
18
|
+
*/
|
|
19
|
+
const AUDIT_DEPS = ['audit-log', '@casl/ability', 'express-winston'];
|
|
20
|
+
const AUDIT_PY_DEPS = ['django-auditlog', 'django-simple-history', 'sqlalchemy-continuum'];
|
|
21
|
+
/** A table or model that exists to hold the record. */
|
|
22
|
+
const AUDIT_STORE = [
|
|
23
|
+
/CREATE\s+TABLE\s+(?:IF\s+NOT\s+EXISTS\s+)?["'`]?\w*audit\w*/i,
|
|
24
|
+
/\bmodel\s+\w*Audit\w*\s*\{/i,
|
|
25
|
+
/class\s+\w*Audit(?:Log|Event|Trail)\w*\b/i,
|
|
26
|
+
/\b(audit_events?|audit_logs?|auditlog|audit_trail)\b/i,
|
|
27
|
+
];
|
|
28
|
+
/** Calls that write an entry. */
|
|
29
|
+
const AUDIT_WRITE = [
|
|
30
|
+
/\b(record|write|log|create|append|emit)[A-Z_]?\w*audit\w*\s*\(/i,
|
|
31
|
+
/\baudit[._]?(log|event|trail)\s*\(/i,
|
|
32
|
+
/\bauditLog\s*\(/,
|
|
33
|
+
/\blog_audit\b/i,
|
|
34
|
+
];
|
|
35
|
+
async function detectAudit(ctx) {
|
|
36
|
+
const evidence = [];
|
|
37
|
+
const deps = [...(0, detectContext_1.hasAnyDep)(ctx, AUDIT_DEPS), ...(0, detectContext_1.hasAnyPyDep)(ctx, AUDIT_PY_DEPS)];
|
|
38
|
+
for (const dep of deps)
|
|
39
|
+
evidence.push({ type: 'dependency', value: dep });
|
|
40
|
+
// Schema and migration files are where a store is declared, and the analyzer
|
|
41
|
+
// classifies them as neither source nor config, so they are taken from `all`.
|
|
42
|
+
const schemaFiles = ctx.files.all.filter((file) => /(^|\/)(schema\.prisma|.*schema\.[cm]?[jt]s|.*\.sql)$/i.test(file));
|
|
43
|
+
const storeHits = await (0, textSearch_1.searchInFiles)(ctx.root, [...ctx.files.source, ...schemaFiles], AUDIT_STORE, 15);
|
|
44
|
+
for (const hit of storeHits) {
|
|
45
|
+
evidence.push({ type: 'snippet', value: hit.snippet, file: hit.file, line: hit.line });
|
|
46
|
+
}
|
|
47
|
+
const writeHits = await (0, textSearch_1.searchInFiles)(ctx.root, ctx.files.source, AUDIT_WRITE, 15);
|
|
48
|
+
for (const hit of writeHits) {
|
|
49
|
+
evidence.push({ type: 'snippet', value: hit.snippet, file: hit.file, line: hit.line });
|
|
50
|
+
}
|
|
51
|
+
const hasStore = storeHits.length > 0 || deps.length > 0;
|
|
52
|
+
const hasWrites = writeHits.length > 0;
|
|
53
|
+
// `complete` is what lets the expectation say `present` rather than `partial`: a
|
|
54
|
+
// store that is written to is a working audit trail, and either half alone is not.
|
|
55
|
+
return {
|
|
56
|
+
key: 'audit.trail',
|
|
57
|
+
present: hasStore || hasWrites,
|
|
58
|
+
complete: hasStore && hasWrites,
|
|
59
|
+
evidence,
|
|
60
|
+
details: { store: hasStore, writes: hasWrites, writeSites: writeHits.length },
|
|
61
|
+
};
|
|
62
|
+
}
|
|
@@ -3,6 +3,7 @@ Object.defineProperty(exports, "__esModule", { value: true });
|
|
|
3
3
|
exports.detectBackend = detectBackend;
|
|
4
4
|
const detectContext_1 = require("./detectContext");
|
|
5
5
|
const textSearch_1 = require("../utils/textSearch");
|
|
6
|
+
const readTextFileSafe_1 = require("../utils/readTextFileSafe");
|
|
6
7
|
async function detectBackend(ctx) {
|
|
7
8
|
const frameworks = [];
|
|
8
9
|
const evidence = [];
|
|
@@ -21,10 +22,25 @@ async function detectBackend(ctx) {
|
|
|
21
22
|
}
|
|
22
23
|
}
|
|
23
24
|
// Node backend frameworks detected purely from dependencies.
|
|
25
|
+
//
|
|
26
|
+
// SvelteKit, Remix and Nuxt appear here as well as in the frontend detector, and
|
|
27
|
+
// that is not a mistake: they serve their own routes, so a repository built on one
|
|
28
|
+
// has a backend to be judged even though it also has a UI. Treating them as
|
|
29
|
+
// frontend-only meant an application with server routes, sessions and database
|
|
30
|
+
// access was scored as though it had none of them.
|
|
24
31
|
const nodeFrameworkDeps = [
|
|
25
32
|
['next', 'next'],
|
|
26
33
|
['nestjs', '@nestjs/core'],
|
|
27
34
|
['fastify', 'fastify'],
|
|
35
|
+
['hono', 'hono'],
|
|
36
|
+
['elysia', 'elysia'],
|
|
37
|
+
['koa', 'koa'],
|
|
38
|
+
['adonis', '@adonisjs/core'],
|
|
39
|
+
['sveltekit', '@sveltejs/kit'],
|
|
40
|
+
['remix', '@remix-run/node'],
|
|
41
|
+
['remix', '@remix-run/server-runtime'],
|
|
42
|
+
['nuxt', 'nuxt'],
|
|
43
|
+
['nitro', 'nitropack'],
|
|
28
44
|
];
|
|
29
45
|
for (const [framework, dep] of nodeFrameworkDeps) {
|
|
30
46
|
if ((0, detectContext_1.hasDep)(ctx, dep)) {
|
|
@@ -32,6 +48,46 @@ async function detectBackend(ctx) {
|
|
|
32
48
|
evidence.push({ type: 'dependency', value: dep });
|
|
33
49
|
}
|
|
34
50
|
}
|
|
51
|
+
/**
|
|
52
|
+
* Astro, which is a backend only when it is configured to be one.
|
|
53
|
+
*
|
|
54
|
+
* Astro builds a static site by default and becomes a server when `output` is set
|
|
55
|
+
* to `server` or `hybrid`, or when an adapter is installed. Listing it as a backend
|
|
56
|
+
* unconditionally would be worse than not listing it at all: this tool scores a
|
|
57
|
+
* project against what its kind of product is expected to have, so a static
|
|
58
|
+
* brochure site would start being marked down for missing sessions, tenant
|
|
59
|
+
* isolation and an audit trail it has no reason to want.
|
|
60
|
+
*/
|
|
61
|
+
if ((0, detectContext_1.hasDep)(ctx, 'astro')) {
|
|
62
|
+
const configFile = ctx.files.all.find((f) => /(^|\/)astro\.config\.[cm]?[jt]s$/.test(f));
|
|
63
|
+
const config = configFile ? ((await (0, readTextFileSafe_1.readTextFileSafe)(ctx.root, configFile)) ?? '') : '';
|
|
64
|
+
const servesRequests = /output\s*:\s*['"](?:server|hybrid)['"]/.test(config)
|
|
65
|
+
|| /adapter\s*:/.test(config)
|
|
66
|
+
|| ctx.files.all.some((f) => /(^|\/)src\/pages\/api\//.test(f));
|
|
67
|
+
if (servesRequests) {
|
|
68
|
+
frameworks.push('astro');
|
|
69
|
+
evidence.push({
|
|
70
|
+
type: configFile ? 'file' : 'dependency',
|
|
71
|
+
value: configFile ? `astro configured to serve requests (${configFile})` : 'astro',
|
|
72
|
+
...(configFile ? { file: configFile } : {}),
|
|
73
|
+
});
|
|
74
|
+
}
|
|
75
|
+
}
|
|
76
|
+
// Python backend frameworks detected purely from dependencies. Django, Flask and
|
|
77
|
+
// FastAPI have their own blocks below because each also has a source-level fallback.
|
|
78
|
+
const pyFrameworkDeps = [
|
|
79
|
+
['litestar', 'litestar'],
|
|
80
|
+
['sanic', 'sanic'],
|
|
81
|
+
['tornado', 'tornado'],
|
|
82
|
+
['aiohttp', 'aiohttp'],
|
|
83
|
+
['starlette', 'starlette'],
|
|
84
|
+
];
|
|
85
|
+
for (const [framework, dep] of pyFrameworkDeps) {
|
|
86
|
+
if ((0, detectContext_1.hasPyDep)(ctx, dep)) {
|
|
87
|
+
frameworks.push(framework);
|
|
88
|
+
evidence.push({ type: 'dependency', value: dep });
|
|
89
|
+
}
|
|
90
|
+
}
|
|
35
91
|
// Flask
|
|
36
92
|
if ((0, detectContext_1.hasPyDep)(ctx, 'flask')) {
|
|
37
93
|
frameworks.push('flask');
|
|
@@ -28,8 +28,57 @@ async function detectBilling(ctx) {
|
|
|
28
28
|
/router\.post\(\s*['"][^'"]*webhook/i,
|
|
29
29
|
], 40)
|
|
30
30
|
: [];
|
|
31
|
+
/**
|
|
32
|
+
* Routes that exist as a file path rather than as a string in the code.
|
|
33
|
+
*
|
|
34
|
+
* The patterns above all look for the URL written out somewhere — `app.post('/webhook')`
|
|
35
|
+
* and its relatives. In a file-system-routed framework the URL is never written
|
|
36
|
+
* anywhere: `src/app/api/stripe/webhook/route.ts` IS the route. That covers the
|
|
37
|
+
* Next.js App Router, Remix, SvelteKit, Nuxt and Astro, so the check reported "no
|
|
38
|
+
* webhook route" for a correctly verified webhook and marked the whole control
|
|
39
|
+
* partial.
|
|
40
|
+
*
|
|
41
|
+
* A `webhook` path segment is required, not a substring, so `webhooks.test.ts`
|
|
42
|
+
* beside a handler does not count as a route on its own.
|
|
43
|
+
*/
|
|
44
|
+
const webhookPathHits = hasStrongStripeSignal
|
|
45
|
+
? ctx.files.source.filter((file) => /(?:^|\/)webhooks?(?:\/|\.[cm]?[jt]sx?$)/i.test(file))
|
|
46
|
+
: [];
|
|
47
|
+
/**
|
|
48
|
+
* Reading the body without letting a JSON parser touch it first.
|
|
49
|
+
*
|
|
50
|
+
* This used to look for `express.raw(` and `req.rawBody` and nothing else, which
|
|
51
|
+
* made the check Express-only. A Stripe signature covers the exact bytes that were
|
|
52
|
+
* sent, so every framework has an idiom for this, and outside Express none of them
|
|
53
|
+
* mention "raw": a Web-standard handler (Next.js route handlers, Remix, SvelteKit,
|
|
54
|
+
* Hono, Bun, Deno, Cloudflare Workers) awaits `request.text()`, Django reads
|
|
55
|
+
* `request.body`, FastAPI awaits `request.body()`, Flask calls
|
|
56
|
+
* `request.get_data()`. Next.js is the most common stack among the repositories
|
|
57
|
+
* this tool is pointed at, so the omission failed the case it meets most often —
|
|
58
|
+
* and it failed it quietly, reporting `partial` on a webhook that was correctly
|
|
59
|
+
* verified.
|
|
60
|
+
*
|
|
61
|
+
* The Web-standard patterns are anchored to a request identifier rather than
|
|
62
|
+
* matching `.text()` anywhere: a `.text()` on a fetch *response* is a different
|
|
63
|
+
* thing entirely and is conventionally held in `res` or `response`.
|
|
64
|
+
*/
|
|
31
65
|
const rawBodyHits = hasStrongStripeSignal
|
|
32
|
-
? await (0, textSearch_1.searchInFiles)(ctx.root, ctx.files.source, [
|
|
66
|
+
? await (0, textSearch_1.searchInFiles)(ctx.root, ctx.files.source, [
|
|
67
|
+
// Express, and the body-parser spelling of the same thing.
|
|
68
|
+
/express\.raw\(/i,
|
|
69
|
+
/bodyParser\.raw\(/i,
|
|
70
|
+
/\braw_?[bB]ody\b/,
|
|
71
|
+
/getRawBody\(/i,
|
|
72
|
+
// Web-standard Request: Next.js route handlers, Remix, SvelteKit, Hono,
|
|
73
|
+
// Bun, Deno, Cloudflare Workers.
|
|
74
|
+
/\b(?:request|req)\.(?:text|arrayBuffer|blob)\(\s*\)/,
|
|
75
|
+
// Python: Django, FastAPI, Flask.
|
|
76
|
+
/\brequest\.body\b/,
|
|
77
|
+
/\brequest\.get_data\(/,
|
|
78
|
+
/\bawait\s+request\.body\(\s*\)/,
|
|
79
|
+
// Go's net/http.
|
|
80
|
+
/io\.ReadAll\(\s*r\.Body\s*\)/,
|
|
81
|
+
], 20)
|
|
33
82
|
: [];
|
|
34
83
|
const secretHits = hasStrongStripeSignal
|
|
35
84
|
? await (0, textSearch_1.searchInFiles)(ctx.root, ctx.files.source, [/STRIPE_WEBHOOK_SECRET/i, /stripeWebhookSecret/i], 20)
|
|
@@ -45,6 +94,9 @@ async function detectBilling(ctx) {
|
|
|
45
94
|
for (const m of [...stripeContextHits, ...webhookRouteHits, ...rawBodyHits, ...secretHits, ...signatureHits]) {
|
|
46
95
|
evidence.push({ type: 'snippet', value: m.snippet, file: m.file, line: m.line });
|
|
47
96
|
}
|
|
97
|
+
for (const file of webhookPathHits) {
|
|
98
|
+
evidence.push({ type: 'file', value: file, file });
|
|
99
|
+
}
|
|
48
100
|
return [
|
|
49
101
|
{
|
|
50
102
|
key: 'billing.stripe',
|
|
@@ -56,8 +108,11 @@ async function detectBilling(ctx) {
|
|
|
56
108
|
},
|
|
57
109
|
{
|
|
58
110
|
key: 'billing.webhook.route',
|
|
59
|
-
present: webhookRouteHits.length > 0,
|
|
60
|
-
evidence:
|
|
111
|
+
present: webhookRouteHits.length > 0 || webhookPathHits.length > 0,
|
|
112
|
+
evidence: [
|
|
113
|
+
...toEvidence(webhookRouteHits),
|
|
114
|
+
...webhookPathHits.map((file) => ({ type: 'file', value: file, file })),
|
|
115
|
+
],
|
|
61
116
|
},
|
|
62
117
|
{
|
|
63
118
|
key: 'billing.webhook.rawBody',
|
|
@@ -3,6 +3,65 @@ Object.defineProperty(exports, "__esModule", { value: true });
|
|
|
3
3
|
exports.detectDatabase = detectDatabase;
|
|
4
4
|
const detectContext_1 = require("./detectContext");
|
|
5
5
|
const readTextFileSafe_1 = require("../utils/readTextFileSafe");
|
|
6
|
+
/**
|
|
7
|
+
* Managed data platforms, and the engine each one actually is.
|
|
8
|
+
*
|
|
9
|
+
* A repository built with an AI tool often has no database driver at all: the whole
|
|
10
|
+
* data layer is a hosted service reached over HTTP. Read only for drivers, the
|
|
11
|
+
* analyzer reported "Detected databases: unknown" for an app whose data layer was
|
|
12
|
+
* perfectly clear — the first line of the report telling the reader the tool had not
|
|
13
|
+
* understood their project. Supabase, the default in Lovable, Bolt and v0, was the
|
|
14
|
+
* most common case of it.
|
|
15
|
+
*
|
|
16
|
+
* The engine matters as much as the platform: Supabase is Postgres and Turso is
|
|
17
|
+
* SQLite, so rules written about an engine keep working without knowing about the
|
|
18
|
+
* host. The platform is recorded separately, because whether the data sits on
|
|
19
|
+
* infrastructure someone else operates is its own question for a readiness report.
|
|
20
|
+
*/
|
|
21
|
+
const MANAGED_PLATFORMS = [
|
|
22
|
+
{
|
|
23
|
+
platform: 'supabase',
|
|
24
|
+
engine: 'postgres',
|
|
25
|
+
npm: ['@supabase/supabase-js', '@supabase/ssr', '@supabase/auth-helpers-nextjs', '@supabase/postgrest-js'],
|
|
26
|
+
py: ['supabase', 'supabase-py'],
|
|
27
|
+
},
|
|
28
|
+
{ platform: 'firebase', engine: 'firestore', npm: ['firebase', 'firebase-admin', '@angular/fire'], py: ['firebase-admin'] },
|
|
29
|
+
{ platform: 'planetscale', engine: 'mysql', npm: ['@planetscale/database'], py: [] },
|
|
30
|
+
{ platform: 'neon', engine: 'postgres', npm: ['@neondatabase/serverless'], py: [] },
|
|
31
|
+
{ platform: 'vercel-postgres', engine: 'postgres', npm: ['@vercel/postgres'], py: [] },
|
|
32
|
+
{ platform: 'turso', engine: 'sqlite', npm: ['@libsql/client', 'libsql'], py: ['libsql-client'] },
|
|
33
|
+
{ platform: 'upstash', engine: 'redis', npm: ['@upstash/redis'], py: ['upstash-redis'] },
|
|
34
|
+
{ platform: 'dynamodb', engine: 'dynamodb', npm: ['@aws-sdk/client-dynamodb', 'dynamoose'], py: [] },
|
|
35
|
+
{ platform: 'convex', engine: 'convex', npm: ['convex'], py: [] },
|
|
36
|
+
];
|
|
37
|
+
/**
|
|
38
|
+
* Object-relational mappers. These prove there is a data layer without naming the
|
|
39
|
+
* engine, so they are recorded on their own key; where the engine is written in a
|
|
40
|
+
* schema file, it is read from there below.
|
|
41
|
+
*/
|
|
42
|
+
const ORMS = [
|
|
43
|
+
{ orm: 'prisma', npm: ['@prisma/client', 'prisma'], py: [] },
|
|
44
|
+
{ orm: 'drizzle', npm: ['drizzle-orm'], py: [] },
|
|
45
|
+
{ orm: 'typeorm', npm: ['typeorm'], py: [] },
|
|
46
|
+
{ orm: 'sequelize', npm: ['sequelize'], py: [] },
|
|
47
|
+
{ orm: 'knex', npm: ['knex'], py: [] },
|
|
48
|
+
{ orm: 'mikro-orm', npm: ['@mikro-orm/core'], py: [] },
|
|
49
|
+
{ orm: 'kysely', npm: ['kysely'], py: [] },
|
|
50
|
+
{ orm: 'sqlalchemy', npm: [], py: ['sqlalchemy', 'alembic'] },
|
|
51
|
+
{ orm: 'tortoise', npm: [], py: ['tortoise-orm'] },
|
|
52
|
+
{ orm: 'peewee', npm: [], py: ['peewee'] },
|
|
53
|
+
];
|
|
54
|
+
/** Prisma and Drizzle both write the engine into a config file; this reads it. */
|
|
55
|
+
const SCHEMA_ENGINES = [
|
|
56
|
+
[/provider\s*=\s*['"]postgresql['"]/i, 'postgres'],
|
|
57
|
+
[/provider\s*=\s*['"]mysql['"]/i, 'mysql'],
|
|
58
|
+
[/provider\s*=\s*['"]sqlite['"]/i, 'sqlite'],
|
|
59
|
+
[/provider\s*=\s*['"]mongodb['"]/i, 'mongodb'],
|
|
60
|
+
[/provider\s*=\s*['"]cockroachdb['"]/i, 'postgres'],
|
|
61
|
+
[/dialect\s*:\s*['"]postgresql['"]/i, 'postgres'],
|
|
62
|
+
[/dialect\s*:\s*['"]mysql['"]/i, 'mysql'],
|
|
63
|
+
[/dialect\s*:\s*['"]sqlite['"]/i, 'sqlite'],
|
|
64
|
+
];
|
|
6
65
|
async function detectDatabase(ctx) {
|
|
7
66
|
const databases = new Set();
|
|
8
67
|
const evidence = [];
|
|
@@ -64,14 +123,64 @@ async function detectDatabase(ctx) {
|
|
|
64
123
|
}
|
|
65
124
|
// DATABASE_URL env reference
|
|
66
125
|
// (kept lightweight; not adding DB type from this alone)
|
|
126
|
+
const platforms = new Set();
|
|
127
|
+
const platformEvidence = [];
|
|
128
|
+
for (const entry of MANAGED_PLATFORMS) {
|
|
129
|
+
const hits = [...(0, detectContext_1.hasAnyDep)(ctx, entry.npm), ...(0, detectContext_1.hasAnyPyDep)(ctx, entry.py)];
|
|
130
|
+
if (!hits.length)
|
|
131
|
+
continue;
|
|
132
|
+
platforms.add(entry.platform);
|
|
133
|
+
if (entry.engine)
|
|
134
|
+
databases.add(entry.engine);
|
|
135
|
+
for (const h of hits)
|
|
136
|
+
platformEvidence.push({ type: 'dependency', value: h });
|
|
137
|
+
}
|
|
138
|
+
const orms = new Set();
|
|
139
|
+
const ormEvidence = [];
|
|
140
|
+
for (const entry of ORMS) {
|
|
141
|
+
const hits = [...(0, detectContext_1.hasAnyDep)(ctx, entry.npm), ...(0, detectContext_1.hasAnyPyDep)(ctx, entry.py)];
|
|
142
|
+
if (!hits.length)
|
|
143
|
+
continue;
|
|
144
|
+
orms.add(entry.orm);
|
|
145
|
+
for (const h of hits)
|
|
146
|
+
ormEvidence.push({ type: 'dependency', value: h });
|
|
147
|
+
}
|
|
148
|
+
// The ORM names the engine in its own schema, which is more reliable than guessing
|
|
149
|
+
// from a driver that a project using an ORM often does not depend on directly.
|
|
150
|
+
const schemaFiles = ctx.files.all.filter((f) => /(^|\/)schema\.prisma$/.test(f) || /(^|\/)drizzle\.config\.[cm]?[jt]s$/.test(f));
|
|
151
|
+
for (const file of schemaFiles) {
|
|
152
|
+
const text = (await (0, readTextFileSafe_1.readTextFileSafe)(ctx.root, file)) ?? '';
|
|
153
|
+
for (const [pattern, engine] of SCHEMA_ENGINES) {
|
|
154
|
+
if (!pattern.test(text))
|
|
155
|
+
continue;
|
|
156
|
+
databases.add(engine);
|
|
157
|
+
ormEvidence.push({ type: 'file', value: `${engine} declared in ${file}`, file });
|
|
158
|
+
}
|
|
159
|
+
}
|
|
67
160
|
const dbs = Array.from(databases);
|
|
161
|
+
const platformList = Array.from(platforms);
|
|
162
|
+
const ormList = Array.from(orms);
|
|
68
163
|
return {
|
|
69
164
|
databases: dbs,
|
|
70
165
|
result: {
|
|
71
166
|
key: 'stack.database',
|
|
72
167
|
present: dbs.length > 0,
|
|
73
|
-
evidence,
|
|
168
|
+
evidence: [...evidence, ...platformEvidence, ...ormEvidence],
|
|
74
169
|
details: { databases: dbs },
|
|
75
170
|
},
|
|
171
|
+
extra: [
|
|
172
|
+
{
|
|
173
|
+
key: 'stack.dataPlatform',
|
|
174
|
+
present: platformList.length > 0,
|
|
175
|
+
evidence: platformEvidence,
|
|
176
|
+
details: { platforms: platformList },
|
|
177
|
+
},
|
|
178
|
+
{
|
|
179
|
+
key: 'stack.orm',
|
|
180
|
+
present: ormList.length > 0,
|
|
181
|
+
evidence: ormEvidence,
|
|
182
|
+
details: { orms: ormList },
|
|
183
|
+
},
|
|
184
|
+
],
|
|
76
185
|
};
|
|
77
186
|
}
|
|
@@ -5,6 +5,12 @@ const textSearch_1 = require("../utils/textSearch");
|
|
|
5
5
|
const WEAK_SECRET_VALUE_RE = /(changeme|your[_-]?secret|fallback-secret(?:-change-in-production)?|change[_-]in[_-]production|your_jwt_secret_key_change_in_production|local[-_]?secret|development[-_]?secret|dev[-_]?secret|not[_-]?for[_-]?production|test123|secret)/i;
|
|
6
6
|
const FALLBACK_SECRET_RE = /(jwt_secret|secret_key|session_secret)\s*[:=]\s*['"][^'"]*(changeme|your[_-]?secret|fallback-secret(?:-change-in-production)?|change[_-]in[_-]production|your_jwt_secret_key_change_in_production|local[-_]?secret|development[-_]?secret|dev[-_]?secret|not[_-]?for[_-]?production|test123|secret)[^'"]*['"]/i;
|
|
7
7
|
const ENV_FALLBACK_RE = /(process\.env\.(JWT_SECRET|SECRET_KEY|SESSION_SECRET)\s*(\|\||\?\?)\s*['"][^'"]*(changeme|your[_-]?secret|fallback-secret(?:-change-in-production)?|change[_-]in[_-]production|your_jwt_secret_key_change_in_production|local[-_]?secret|development[-_]?secret|dev[-_]?secret|not[_-]?for[_-]?production|test123|secret)[^'"]*['"])/i;
|
|
8
|
+
/**
|
|
9
|
+
* The identifier being assigned has to be a secret. Belt and braces with the anchored
|
|
10
|
+
* patterns above: a future pattern added to that list cannot reintroduce the problem
|
|
11
|
+
* of a snippet qualifying on the word "secret" appearing anywhere in it.
|
|
12
|
+
*/
|
|
13
|
+
const SECRET_ASSIGNMENT_CONTEXT_RE = /(JWT_SECRET|SECRET_KEY|SESSION_SECRET|API_KEY|jwt_secret|secret_key|session_secret|api_key)\s*[:=]|process\.env\.[A-Z0-9_]*(SECRET|API_KEY)/;
|
|
8
14
|
const GENERIC_SECRET_ASSIGNMENT_RE = /(JWT_SECRET|SECRET_KEY|SESSION_SECRET|API_KEY)\s*[:=]\s*['"][^'"]+['"]/i;
|
|
9
15
|
function classifySecretFallback(snippet) {
|
|
10
16
|
if (/JWT_SECRET/i.test(snippet))
|
|
@@ -38,8 +44,25 @@ async function detectEnv(ctx) {
|
|
|
38
44
|
for (const m of envReads) {
|
|
39
45
|
evidence.push({ type: 'snippet', value: m.snippet, file: m.file, line: m.line });
|
|
40
46
|
}
|
|
41
|
-
|
|
42
|
-
|
|
47
|
+
/**
|
|
48
|
+
* Only patterns that name a secret. The list used to include a bare
|
|
49
|
+
* `|| 'anything'` and `?? 'anything'`, which match every nullish default in every
|
|
50
|
+
* codebase and carry no signal about secrets at all. What qualified such a line as
|
|
51
|
+
* weak was WEAK_SECRET_VALUE_RE, and that matches the bare word `secret` — so any
|
|
52
|
+
* line holding a default and the word "secret" anywhere in it was reported.
|
|
53
|
+
*
|
|
54
|
+
* That produced `critical`, the loudest severity there is, and one it caps maturity
|
|
55
|
+
* with. It fired on this very repository, on the string that describes the check;
|
|
56
|
+
* it would fire the same way on an interface label, a translation file, or any
|
|
57
|
+
* application with a "secret question" feature. For a product whose pitch is that
|
|
58
|
+
* every point traces back to a rule and a file you can open, a critical pointing at
|
|
59
|
+
* a sentence about secrets is the worst possible first impression.
|
|
60
|
+
*
|
|
61
|
+
* The three remaining patterns each anchor on a secret-named identifier, which is
|
|
62
|
+
* the thing actually being claimed.
|
|
63
|
+
*/
|
|
64
|
+
const fallbackHits = await (0, textSearch_1.searchInFiles)(ctx.root, sourceFiles, [FALLBACK_SECRET_RE, ENV_FALLBACK_RE, GENERIC_SECRET_ASSIGNMENT_RE], 30);
|
|
65
|
+
const weakHits = fallbackHits.filter((m) => WEAK_SECRET_VALUE_RE.test(m.snippet) && SECRET_ASSIGNMENT_CONTEXT_RE.test(m.snippet));
|
|
43
66
|
for (const m of weakHits) {
|
|
44
67
|
const hitEvidence = { type: 'snippet', value: m.snippet, file: m.file, line: m.line };
|
|
45
68
|
evidence.push(hitEvidence);
|
|
@@ -23,6 +23,12 @@ async function detectFrontend(ctx) {
|
|
|
23
23
|
['nuxt', ['nuxt']],
|
|
24
24
|
['svelte', ['svelte', '@sveltejs/kit']],
|
|
25
25
|
['angular', ['@angular/core']],
|
|
26
|
+
['astro', ['astro']],
|
|
27
|
+
['solid', ['solid-js']],
|
|
28
|
+
['qwik', ['@builder.io/qwik']],
|
|
29
|
+
['preact', ['preact']],
|
|
30
|
+
['remix', ['@remix-run/react']],
|
|
31
|
+
['htmx', ['htmx.org']],
|
|
26
32
|
];
|
|
27
33
|
for (const [framework, names] of frameworkDeps) {
|
|
28
34
|
const hits = (0, detectContext_1.hasAnyDep)(ctx, names);
|
|
@@ -3,28 +3,135 @@ Object.defineProperty(exports, "__esModule", { value: true });
|
|
|
3
3
|
exports.detectJobs = detectJobs;
|
|
4
4
|
const detectContext_1 = require("./detectContext");
|
|
5
5
|
const textSearch_1 = require("../utils/textSearch");
|
|
6
|
+
const readTextFileSafe_1 = require("../utils/readTextFileSafe");
|
|
7
|
+
/**
|
|
8
|
+
* Work that happens outside a request.
|
|
9
|
+
*
|
|
10
|
+
* `present` used to be satisfied by a bare /queue/i, /worker/i or /cron/i anywhere in
|
|
11
|
+
* any source file. That matches a breadth-first search holding a `queue`, a browser
|
|
12
|
+
* `Worker` in frontend code, and the word "cron" in a comment. Measured: a pirate game
|
|
13
|
+
* was reported as having background jobs because of
|
|
14
|
+
* `queue: deque[tuple[int, int]] = deque()` in a flood fill, and that single signal was
|
|
15
|
+
* then enough to have the whole project judged as an AI SaaS.
|
|
16
|
+
*
|
|
17
|
+
* Presence now needs something declared rather than mentioned: a dependency on a queue
|
|
18
|
+
* or scheduler, or a job actually being defined. The loose terms are still collected,
|
|
19
|
+
* because they are useful context in a report — but context is all they are, and they
|
|
20
|
+
* no longer decide the answer.
|
|
21
|
+
*/
|
|
22
|
+
const QUEUE_DEPS = [
|
|
23
|
+
'bull',
|
|
24
|
+
'bullmq',
|
|
25
|
+
'node-cron',
|
|
26
|
+
'agenda',
|
|
27
|
+
'bee-queue',
|
|
28
|
+
'kue',
|
|
29
|
+
'pg-boss',
|
|
30
|
+
'graphile-worker',
|
|
31
|
+
'inngest',
|
|
32
|
+
'@trigger.dev/sdk',
|
|
33
|
+
'@temporalio/client',
|
|
34
|
+
'@temporalio/worker',
|
|
35
|
+
'quirrel',
|
|
36
|
+
'croner',
|
|
37
|
+
];
|
|
38
|
+
const QUEUE_PY_DEPS = [
|
|
39
|
+
'celery',
|
|
40
|
+
'apscheduler',
|
|
41
|
+
'rq',
|
|
42
|
+
'dramatiq',
|
|
43
|
+
'huey',
|
|
44
|
+
'arq',
|
|
45
|
+
'prefect',
|
|
46
|
+
'apache-airflow',
|
|
47
|
+
'django-q',
|
|
48
|
+
];
|
|
49
|
+
/**
|
|
50
|
+
* A job being defined. Each of these is a call or a decorator, so none of them match
|
|
51
|
+
* the word appearing in prose or in an unrelated identifier.
|
|
52
|
+
*
|
|
53
|
+
* `new Worker(` is deliberately absent: the browser spells a Web Worker exactly that
|
|
54
|
+
* way, and a Web Worker is not a background job. Where a queue library is in use, its
|
|
55
|
+
* dependency has already answered the question.
|
|
56
|
+
*/
|
|
57
|
+
const JOB_DECLARATIONS = [
|
|
58
|
+
/new\s+Queue\s*\(/,
|
|
59
|
+
/createQueue\s*\(/,
|
|
60
|
+
/cron\.schedule\s*\(/,
|
|
61
|
+
/CronJob\s*\(/,
|
|
62
|
+
/@shared_task\b/,
|
|
63
|
+
/@(?:celery|app)\.task\b/,
|
|
64
|
+
/@periodic_task\b/,
|
|
65
|
+
/defineJob\s*\(/,
|
|
66
|
+
/\bschedule\.every\s*\(/,
|
|
67
|
+
/add_periodic_task\s*\(/,
|
|
68
|
+
];
|
|
69
|
+
/**
|
|
70
|
+
* A long-running process declared outside the code.
|
|
71
|
+
*
|
|
72
|
+
* Not every background worker uses a library. Ours does not: it polls Postgres with
|
|
73
|
+
* FOR UPDATE SKIP LOCKED and has no queue dependency and no decorator to find. What it
|
|
74
|
+
* does have — what any shipped worker has — is somewhere that says a second process
|
|
75
|
+
* exists: a script that starts it, and a service that runs it. Those are declarations
|
|
76
|
+
* too, and structural ones.
|
|
77
|
+
*
|
|
78
|
+
* Requiring the library was a false negative on this product's own repository, which
|
|
79
|
+
* is a good sign a rule is wrong.
|
|
80
|
+
*/
|
|
81
|
+
const PROCESS_NAME = /^(worker|jobs?|queue|scheduler|cron|consumer)(:|-|_|$)/i;
|
|
82
|
+
/** Files whose name says they hold out-of-request work. */
|
|
83
|
+
const WORKER_FILE = /(^|\/)(workers?|jobs|queues?|tasks)(\/|\.[cm]?[jt]sx?$|\.py$)/i;
|
|
6
84
|
async function detectJobs(ctx) {
|
|
7
85
|
const evidence = [];
|
|
8
|
-
const nodeDeps = (0, detectContext_1.hasAnyDep)(ctx,
|
|
9
|
-
const pyDeps = (0, detectContext_1.hasAnyPyDep)(ctx,
|
|
10
|
-
for (const
|
|
11
|
-
evidence.push({ type: 'dependency', value:
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
const
|
|
18
|
-
for (const
|
|
19
|
-
evidence.push({ type: '
|
|
86
|
+
const nodeDeps = (0, detectContext_1.hasAnyDep)(ctx, QUEUE_DEPS);
|
|
87
|
+
const pyDeps = (0, detectContext_1.hasAnyPyDep)(ctx, QUEUE_PY_DEPS);
|
|
88
|
+
for (const dep of [...nodeDeps, ...pyDeps])
|
|
89
|
+
evidence.push({ type: 'dependency', value: dep });
|
|
90
|
+
const workerFiles = ctx.files.all.filter((file) => WORKER_FILE.test(file)).slice(0, 20);
|
|
91
|
+
for (const file of workerFiles)
|
|
92
|
+
evidence.push({ type: 'file', value: file, file });
|
|
93
|
+
// A script that starts a separate process.
|
|
94
|
+
const scripts = ctx.packageJson?.scripts ?? {};
|
|
95
|
+
const processScripts = Object.keys(scripts).filter((name) => PROCESS_NAME.test(name) || /(^|\/)workers?\//i.test(String(scripts[name] ?? '')));
|
|
96
|
+
for (const name of processScripts) {
|
|
97
|
+
evidence.push({ type: 'note', value: `package.json script "${name}" starts a separate process` });
|
|
98
|
+
}
|
|
99
|
+
// A compose service that runs one.
|
|
100
|
+
const composeFile = ctx.files.all.find((file) => /(^|\/)(docker-compose\.ya?ml|compose\.ya?ml)$/.test(file));
|
|
101
|
+
const composeText = composeFile ? ((await (0, readTextFileSafe_1.readTextFileSafe)(ctx.root, composeFile)) ?? '') : '';
|
|
102
|
+
const composeWorker = /^\s{2}(worker|jobs?|scheduler|cron|consumer)[a-z0-9_-]*:/im.test(composeText);
|
|
103
|
+
if (composeWorker && composeFile) {
|
|
104
|
+
evidence.push({ type: 'file', value: 'a worker service in compose', file: composeFile });
|
|
105
|
+
}
|
|
106
|
+
const declarations = await (0, textSearch_1.searchInFiles)(ctx.root, ctx.files.source, JOB_DECLARATIONS, 20);
|
|
107
|
+
for (const hit of declarations) {
|
|
108
|
+
evidence.push({ type: 'snippet', value: hit.snippet, file: hit.file, line: hit.line });
|
|
109
|
+
}
|
|
110
|
+
// Collected for the reader, never to decide. A report that says "background jobs
|
|
111
|
+
// detected" and shows only a variable called `queue` has told the reader nothing
|
|
112
|
+
// they can act on.
|
|
113
|
+
const mentions = await (0, textSearch_1.searchInFiles)(ctx.root, ctx.files.source, [/\bqueue\b/i, /\bworker\b/i, /\bcron\b/i], 10);
|
|
114
|
+
for (const hit of mentions) {
|
|
115
|
+
evidence.push({ type: 'snippet', value: hit.snippet, file: hit.file, line: hit.line });
|
|
116
|
+
}
|
|
117
|
+
const declared = nodeDeps.length > 0
|
|
118
|
+
|| pyDeps.length > 0
|
|
119
|
+
|| declarations.length > 0
|
|
120
|
+
|| processScripts.length > 0
|
|
121
|
+
|| composeWorker;
|
|
20
122
|
return {
|
|
21
123
|
key: 'jobs.background',
|
|
22
|
-
|
|
124
|
+
// A file named `workers/` is only taken as proof alongside something declared in
|
|
125
|
+
// the code: the name is a convention, and conventions are borrowed.
|
|
126
|
+
present: declared,
|
|
23
127
|
evidence,
|
|
24
128
|
details: {
|
|
25
129
|
nodeQueue: nodeDeps.length > 0,
|
|
26
130
|
pythonQueue: pyDeps.length > 0,
|
|
27
131
|
workerFiles: workerFiles.length,
|
|
132
|
+
declarations: declarations.length,
|
|
133
|
+
processScripts: processScripts.length,
|
|
134
|
+
composeWorker,
|
|
28
135
|
},
|
|
29
136
|
};
|
|
30
137
|
}
|
|
@@ -4,6 +4,8 @@ export declare function buildStackInfo(input: {
|
|
|
4
4
|
frontend: string[];
|
|
5
5
|
backend: string[];
|
|
6
6
|
databases: string[];
|
|
7
|
+
dataPlatforms?: string[];
|
|
8
|
+
orms?: string[];
|
|
7
9
|
packageManager: StackInfo['packageManager'];
|
|
8
10
|
packageManagerConfidence: StackInfo['packageManagerConfidence'];
|
|
9
11
|
warnings: string[];
|
|
@@ -21,6 +21,8 @@ function buildStackInfo(input) {
|
|
|
21
21
|
frontend: unique(input.frontend),
|
|
22
22
|
backend: unique(input.backend),
|
|
23
23
|
databases: unique(input.databases),
|
|
24
|
+
dataPlatforms: unique(input.dataPlatforms ?? []),
|
|
25
|
+
orms: unique(input.orms ?? []),
|
|
24
26
|
packageManager: input.packageManager,
|
|
25
27
|
packageManagerConfidence: input.packageManagerConfidence,
|
|
26
28
|
warnings: unique(input.warnings),
|
package/dist/analyzer/types.d.ts
CHANGED
|
@@ -13,6 +13,13 @@ export interface StackInfo {
|
|
|
13
13
|
frontend: string[];
|
|
14
14
|
backend: string[];
|
|
15
15
|
databases: string[];
|
|
16
|
+
/**
|
|
17
|
+
* Hosted data services: ['supabase'], ['firebase']. Separate from `databases`,
|
|
18
|
+
* which holds the engine — Supabase appears in both, as 'supabase' here and
|
|
19
|
+
* 'postgres' there, because it is both.
|
|
20
|
+
*/
|
|
21
|
+
dataPlatforms: string[];
|
|
22
|
+
orms: string[];
|
|
16
23
|
languages: string[];
|
|
17
24
|
packageManager: PackageManager;
|
|
18
25
|
packageManagerConfidence: PackageManagerConfidence;
|
package/dist/cli.js
CHANGED
|
@@ -117,6 +117,13 @@ function summarize(report) {
|
|
|
117
117
|
`Detected frontend: ${report.detectedStack.frontend.join(', ') || 'unknown'}`,
|
|
118
118
|
`Detected backend: ${report.detectedStack.backend.join(', ') || 'unknown'}`,
|
|
119
119
|
`Detected databases: ${report.detectedStack.databases.join(', ') || 'unknown'}`,
|
|
120
|
+
// Named on their own line rather than folded into the databases list. A reader
|
|
121
|
+
// whose data layer is entirely Supabase needs to see that the tool recognised
|
|
122
|
+
// Supabase, not only that it worked out the engine underneath is Postgres.
|
|
123
|
+
...(report.detectedStack.dataPlatforms.length > 0
|
|
124
|
+
? [`Data platform: ${report.detectedStack.dataPlatforms.join(', ')}`]
|
|
125
|
+
: []),
|
|
126
|
+
...(report.detectedStack.orms.length > 0 ? [`ORM: ${report.detectedStack.orms.join(', ')}`] : []),
|
|
120
127
|
`Package manager: ${report.detectedStack.packageManager} (${report.detectedStack.packageManagerConfidence})`,
|
|
121
128
|
report.detectedStack.warnings.length > 0 ? `Warnings: ${report.detectedStack.warnings.join('; ')}` : 'Warnings: none',
|
|
122
129
|
`Workspaces: ${workspaceSummary}`,
|
|
@@ -206,8 +213,8 @@ async function runCli(argv = process.argv) {
|
|
|
206
213
|
.option('--output <path>', 'Output file path (optional)')
|
|
207
214
|
.option('--fail-under <score>', 'Exit with an error if overall score is below this threshold (0-100)')
|
|
208
215
|
.option('--min-maturity <level>', `Exit with an error if maturity is below this level: ${maturityOrder.join('|')}`)
|
|
209
|
-
.option('--ai', 'Add an AI stack/architecture insight (opt-in; requires
|
|
210
|
-
.option('--ai-review', 'Add an AI semantic code review of fine-grained issues (opt-in; requires
|
|
216
|
+
.option('--ai', 'Add an AI stack/architecture insight (opt-in; requires @produtype/ai; advisory only)')
|
|
217
|
+
.option('--ai-review', 'Add an AI semantic code review of fine-grained issues (opt-in; requires @produtype/ai; advisory only)')
|
|
211
218
|
.action(async (targetPath, options) => {
|
|
212
219
|
const format = options.format === 'json' ? 'json' : 'markdown';
|
|
213
220
|
const formatSpecified = typeof options.format === 'string';
|
|
@@ -140,12 +140,28 @@ function deriveStatus(analysis, capability) {
|
|
|
140
140
|
return 'missing';
|
|
141
141
|
}
|
|
142
142
|
case 'audit.baseline': {
|
|
143
|
+
/**
|
|
144
|
+
* Read from the audit trail, not from the logs.
|
|
145
|
+
*
|
|
146
|
+
* This used to be decided by structured logging and a request id, which answer
|
|
147
|
+
* "can we debug this?" — a different question from "can we say, months later,
|
|
148
|
+
* who deleted that organisation?". Logs rotate; an audit trail is meant not to.
|
|
149
|
+
*
|
|
150
|
+
* It also could not return `present`. Both branches returned `partial`, so a
|
|
151
|
+
* project with a complete audit trail was marked down for it permanently, on a
|
|
152
|
+
* capability the b2b-saas profile asks for.
|
|
153
|
+
*/
|
|
154
|
+
const trail = detector(analysis, 'audit.trail');
|
|
155
|
+
if (trail?.present && trail.complete)
|
|
156
|
+
return 'present';
|
|
157
|
+
if (trail?.present)
|
|
158
|
+
return 'partial';
|
|
159
|
+
// Structured logging with a request id is not an audit trail, but it is not
|
|
160
|
+
// nothing either: the events can often be reconstructed from it.
|
|
143
161
|
const structured = boolDetail(obs, 'structuredLogging');
|
|
144
162
|
const requestId = boolDetail(obs, 'requestId');
|
|
145
163
|
if (structured && requestId)
|
|
146
164
|
return 'partial';
|
|
147
|
-
if (structured || requestId)
|
|
148
|
-
return 'partial';
|
|
149
165
|
return 'missing';
|
|
150
166
|
}
|
|
151
167
|
case 'jobs.background': {
|
|
@@ -10,18 +10,38 @@ function inferProductProfile(analysis) {
|
|
|
10
10
|
const tenancy = analysis.detectors['tenancy.organization']?.present === true || analysis.detectors['tenancy.membership']?.present === true;
|
|
11
11
|
const jobs = analysis.detectors['jobs.background']?.present === true;
|
|
12
12
|
const uploads = analysis.detectors['uploads.exposure']?.present === true;
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
13
|
+
/**
|
|
14
|
+
* A dependency on a model SDK, not a word that suggests one.
|
|
15
|
+
*
|
|
16
|
+
* This used to match /openai|anthropic|claude|gemini|llm|transcrib|generation|prompt|model/
|
|
17
|
+
* against the evidence strings of the billing, API-key and background-job detectors.
|
|
18
|
+
* "model" is in every ORM, "generation" and "prompt" are ordinary English, and the
|
|
19
|
+
* strings being searched were written to describe something else entirely.
|
|
20
|
+
*/
|
|
21
|
+
const callsAModel = analysis.detectors['ai.modelProvider']?.present === true;
|
|
17
22
|
const sourceHints = analysis.files.source.join('\n');
|
|
18
23
|
const marketplaceHints = /seller|buyer|vendor|listing|order/i.test(sourceHints);
|
|
19
24
|
const adminHints = /admin|\/users|subscription|billing/i.test(sourceHints);
|
|
20
25
|
if (frontendPresent && !backendPresent && !dbPresent && !auth) {
|
|
21
26
|
return { inferredProfile: 'static-site', confidence: 'high', reason: 'Frontend-only structure with no backend/db/auth signals.' };
|
|
22
27
|
}
|
|
23
|
-
|
|
24
|
-
|
|
28
|
+
/**
|
|
29
|
+
* Background jobs used to be sufficient here, in `(aiEvidence || jobs)`. They are
|
|
30
|
+
* orthogonal: every serious application has a queue, and this branch sits second,
|
|
31
|
+
* ahead of b2b-saas, marketplace and b2c — so an ordinary application with a worker
|
|
32
|
+
* was judged against the most demanding profile in the catalogue, which accumulates
|
|
33
|
+
* roughly 185 points of expectation. The score came out wrong for a reason the
|
|
34
|
+
* reader had no way to see.
|
|
35
|
+
*
|
|
36
|
+
* Measured: of ten unrelated local repositories, five were called ai-saas. One was a
|
|
37
|
+
* pirate game, promoted on a local variable named `queue` in a flood fill.
|
|
38
|
+
*/
|
|
39
|
+
if (callsAModel && (uploads || jobs || backendPresent)) {
|
|
40
|
+
return {
|
|
41
|
+
inferredProfile: 'ai-saas',
|
|
42
|
+
confidence: 'medium',
|
|
43
|
+
reason: 'A model SDK dependency with a backend or a processing pipeline.',
|
|
44
|
+
};
|
|
25
45
|
}
|
|
26
46
|
if (marketplaceHints && billing) {
|
|
27
47
|
return { inferredProfile: 'marketplace', confidence: 'medium', reason: 'Marketplace vocabulary and billing signals detected.' };
|
package/dist/rules/rules.js
CHANGED
|
@@ -10,8 +10,26 @@ function statusFromFlags(present, complete) {
|
|
|
10
10
|
return 'passed';
|
|
11
11
|
return 'unknown';
|
|
12
12
|
}
|
|
13
|
+
/**
|
|
14
|
+
* `unknown` is not a severity, it is the absence of one.
|
|
15
|
+
*
|
|
16
|
+
* A rule that does not apply — the Django DEBUG check on an Express project, say —
|
|
17
|
+
* returned `unknown` while keeping the severity it would have had if it had applied.
|
|
18
|
+
* The JSON report therefore carried `severity: 'critical'` on a finding whose own
|
|
19
|
+
* evidence read "Django stack not detected", and a scan of a repository with no
|
|
20
|
+
* Python in it published three of them.
|
|
21
|
+
*
|
|
22
|
+
* Nothing shipped was misled: every consumer, in the CLI summary, the markdown report
|
|
23
|
+
* and the cloud, filters `status !== 'unknown'` alongside the severity. But that pair
|
|
24
|
+
* is written out in four separate places, and each one is a chance to write only half
|
|
25
|
+
* of it — which anyone reading the JSON from outside this repository, through the MCP
|
|
26
|
+
* server or the Action, would have no reason to know they had to do at all.
|
|
27
|
+
*
|
|
28
|
+
* Carrying `info` makes `severity === 'critical'` a safe question on its own. Scoring
|
|
29
|
+
* is unaffected: score.ts already skips `unknown` before it looks at severity.
|
|
30
|
+
*/
|
|
13
31
|
function sevForStatus(status, missing) {
|
|
14
|
-
return status === 'passed' ? 'info' : missing;
|
|
32
|
+
return status === 'passed' || status === 'unknown' ? 'info' : missing;
|
|
15
33
|
}
|
|
16
34
|
function detectorEvidence(evidence) {
|
|
17
35
|
return evidence && evidence.length > 0 ? evidence : [{ type: 'note', value: 'no direct evidence captured' }];
|
|
@@ -69,8 +87,8 @@ exports.rules = [
|
|
|
69
87
|
title: 'Application stack fingerprint',
|
|
70
88
|
category: 'stack',
|
|
71
89
|
status,
|
|
72
|
-
severity: status
|
|
73
|
-
description: hasStack ? 'ProdKit identified known stack signals.' : 'Stack is generic
|
|
90
|
+
severity: sevForStatus(status, 'low'),
|
|
91
|
+
description: hasStack ? 'ProdKit identified known stack signals.' : 'Stack is generic or unknown to ProdKit.',
|
|
74
92
|
recommendation: hasStack
|
|
75
93
|
? 'No action required.'
|
|
76
94
|
: 'Add explicit framework manifests or keep this as generic app baseline.',
|
|
@@ -96,7 +114,7 @@ exports.rules = [
|
|
|
96
114
|
title: 'Environment template file',
|
|
97
115
|
category: 'env',
|
|
98
116
|
status,
|
|
99
|
-
severity: status
|
|
117
|
+
severity: sevForStatus(status, 'medium'),
|
|
100
118
|
description: missing
|
|
101
119
|
? 'The project reads environment variables but no .env.example was detected.'
|
|
102
120
|
: 'Environment template looks available or env usage was not detected.',
|
|
@@ -206,7 +224,7 @@ exports.rules = [
|
|
|
206
224
|
title: 'CORS origin restrictions',
|
|
207
225
|
category: 'security',
|
|
208
226
|
status,
|
|
209
|
-
severity: status === '
|
|
227
|
+
severity: status === 'partial' ? sevForStatus(status, 'high') : sevForStatus(status, 'medium'),
|
|
210
228
|
description: status === 'passed'
|
|
211
229
|
? 'CORS appears configured with explicit origins.'
|
|
212
230
|
: status === 'partial'
|
|
@@ -232,7 +250,7 @@ exports.rules = [
|
|
|
232
250
|
title: 'Django DEBUG hardening',
|
|
233
251
|
category: 'security',
|
|
234
252
|
status,
|
|
235
|
-
severity: status
|
|
253
|
+
severity: sevForStatus(status, 'critical'),
|
|
236
254
|
description: !isDjango
|
|
237
255
|
? 'Django stack not detected.'
|
|
238
256
|
: debugTrue
|
|
@@ -341,7 +359,7 @@ exports.rules = [
|
|
|
341
359
|
title: 'Authorization depth',
|
|
342
360
|
category: 'authz',
|
|
343
361
|
status,
|
|
344
|
-
severity: status
|
|
362
|
+
severity: sevForStatus(status, 'medium'),
|
|
345
363
|
description: status === 'passed'
|
|
346
364
|
? 'Permission-level authorization signals detected.'
|
|
347
365
|
: status === 'partial'
|
|
@@ -372,7 +390,7 @@ exports.rules = [
|
|
|
372
390
|
title: 'Tenant and organization boundaries',
|
|
373
391
|
category: 'tenancy',
|
|
374
392
|
status,
|
|
375
|
-
severity: status === '
|
|
393
|
+
severity: status === 'partial' ? sevForStatus(status, 'medium') : sevForStatus(status, 'high'),
|
|
376
394
|
description: status === 'unknown'
|
|
377
395
|
? (backendDetected
|
|
378
396
|
? 'No B2B/SaaS signals detected.'
|
|
@@ -54,6 +54,22 @@ exports.DEFAULT_IGNORE = [
|
|
|
54
54
|
'**/.turbo/**',
|
|
55
55
|
'**/.cache/**',
|
|
56
56
|
'**/.parcel-cache/**',
|
|
57
|
+
// Build output that is not literally called `dist`. A bundle is a copy of the
|
|
58
|
+
// source, so scanning it counts every signal twice and — worse — produces evidence
|
|
59
|
+
// pointing at a generated file. This product's promise is that a finding names a
|
|
60
|
+
// file you can open and argue with; `dist-worker/index.js:254` is not that. Found by
|
|
61
|
+
// analysing prodkit-cloud, whose worker bundle is emitted to `dist-worker/`.
|
|
62
|
+
'**/dist-*/**',
|
|
63
|
+
'**/out/**',
|
|
64
|
+
'**/.output/**',
|
|
65
|
+
'**/.svelte-kit/**',
|
|
66
|
+
'**/.astro/**',
|
|
67
|
+
'**/.nuxt/**',
|
|
68
|
+
'**/.vercel/**',
|
|
69
|
+
'**/.netlify/**',
|
|
70
|
+
'**/target/**',
|
|
71
|
+
'**/*.min.js',
|
|
72
|
+
'**/*.bundle.js',
|
|
57
73
|
// Test fixtures are sample applications, often deliberately insecure, and they are
|
|
58
74
|
// not the product. Scanning them makes a repository inherit the stack and the
|
|
59
75
|
// defects of its own test data: prodkit analysing itself reported express, next,
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@produtype/core",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.3.0",
|
|
4
4
|
"description": "Deterministic CLI and library that analyzes a web application repository and reports how far it is from production-ready for the kind of product it is meant to be.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"bin": {
|