@titan-design/style-analyzer 0.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Henry Jewkes
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,124 @@
1
+ # @titan-design/style-analyzer
2
+
3
+ Tree-sitter style extractors, and the aggregator that turns their output into a style
4
+ profile. Each extractor reads a parsed TypeScript or Python file and emits `Observation`
5
+ records: one per naming choice, import, branch, comment, function metric, and so on. The
6
+ aggregator groups observations by type into features with a dominant convention, a
7
+ confidence, a stability and a severity. An optional enricher asks an injected LLM to
8
+ describe the features that numbers alone cannot capture.
9
+
10
+ Tier 2 of the titan-platform DAG (TP-134). Depends on `@titan-design/code-parser` (tier 0)
11
+ and `@titan-design/style-profile` (tier 2). `zod` v4 is a peer because style-profile
12
+ needs it at runtime. Ported unchanged from codewatch's
13
+ `@codewatch/analyzer`.
14
+
15
+ ```sh
16
+ npm install @titan-design/style-analyzer web-tree-sitter@^0.26.6 zod
17
+ ```
18
+
19
+ ```ts
20
+ import { Aggregator, createStyleExtractors, parseFile } from "@titan-design/style-analyzer";
21
+
22
+ const file = await parseFile(source, "src/users.ts", "typescript");
23
+ const observations = createStyleExtractors().flatMap((extractor) => extractor.extract(file));
24
+ // { type: "naming.function", category: "naming", value: "camelCase", file: "src/users.ts", line: 3 }
25
+
26
+ const { features, reviewQueue, summary } = new Aggregator().aggregate(observations);
27
+ features.get("formatting.quoteStyle"); // { convention: "double", confidence: 1, severity: "error", ... }
28
+ ```
29
+
30
+ ## API
31
+
32
+ - `Observation` is `{ type, category, value, file, line, metadata? }`. `type` is the feature
33
+ (`naming.variable`, `complexity.cyclomatic`), `value` is a string, number or boolean, and
34
+ `line` is 1-based.
35
+ - `StyleExtractor` is `Extractor<Observation>` from code-parser. `Extractor` is its
36
+ deprecated alias, kept for compatibility.
37
+ - `createStyleExtractors()` returns the canonical nine, in a fixed order: `NamingExtractor`,
38
+ `StructureExtractor`, `ControlFlowExtractor`, `DocumentationExtractor`,
39
+ `ErrorHandlingExtractor`, `FormattingExtractor`, `ComplexityExtractor`, `IdiomsExtractor`,
40
+ `ReviewVoiceExtractor`.
41
+ - `FormattingExtractor.extractFromConfig(path)` reads a `.prettierrc` or `.editorconfig`
42
+ from disk. `extractFromSource(text, path)` works on raw text.
43
+ - `IdiomsExtractor.extractFromSources([{ path, content, language }])` finds repeated code
44
+ across files with jscpd.
45
+ - `ReviewVoiceExtractor.extractFromComments([{ body }])` classifies review comments by
46
+ topic and keyword.
47
+ - `new Aggregator(config?).aggregate(observations)` returns `{ features, reviewQueue,
48
+ summary }`. `features` maps each observation type to an `AggregatedFeature`.
49
+ - Confidence is `min(1, consistency * weight)`. Consistency is the dominant value's
50
+ share, and the weight comes from the type's stability (high 1.0, medium 0.85, low 0.7).
51
+ - Severity follows style-profile's thresholds.
52
+ - A feature below `reviewThreshold` (default 0.6) joins `reviewQueue`, lowest first.
53
+ - `computeConfidence`, `mapSeverity` and `lookupStability` are exported on their own.
54
+ - `new Enricher({ provider, enabled?, totalTokenBudget? }).enrich(features)` sends one
55
+ prompt per AI-enriched feature (`AI_ENRICHED_FEATURES`, `needsAiEnrichment`).
56
+ - Prompts run sequentially and stop once the token budget is used up (default 20,000).
57
+ - A failed call becomes an entry in `errors` instead of a throw.
58
+ - `enabled: false` returns `skipped: true` without calling the provider.
59
+ - `IngestConfig`, `CodeCorpus`, `CodeFile`, `ReviewComment`, `PullRequest`,
60
+ `PullRequestFile` and `IngestMetadata` are the corpus types codewatch's ingestion
61
+ produced. They are types only; the GitHub ingestion itself was not ported.
62
+ - `parseFile`, `getSupportedLanguages`, `shouldIncludeFile` and `getLanguageFromPath` are
63
+ re-exported from code-parser, as the original re-exported them from `@codewatch/core`.
64
+
65
+ ## Plugging in an LLM
66
+
67
+ The enricher never picks a model. It takes an `LlmProvider`:
68
+
69
+ ```ts
70
+ interface LlmProvider {
71
+ name: string;
72
+ generate(messages: LlmMessage[], options: { maxTokens: number }): Promise<LlmResponse>;
73
+ }
74
+ // LlmMessage is { role: "system" | "user" | "assistant", content }
75
+ // LlmResponse is { content, tokensUsed }
76
+ ```
77
+
78
+ A product that uses `@titan-design/agent` adapts it in a few lines. This package does not
79
+ depend on the agent package:
80
+
81
+ ```ts
82
+ import { runAgent } from "@titan-design/agent";
83
+ import type { LlmProvider } from "@titan-design/style-analyzer";
84
+
85
+ const agentProvider: LlmProvider = {
86
+ name: "claude-agent",
87
+ async generate(messages) {
88
+ const prompt = messages.map((m) => m.content).join("\n\n");
89
+ const result = await runAgent({ prompt, cwd: process.cwd(), maxTurns: 1, maxBudgetUsd: 0.05 });
90
+ if (!result.ok) throw new Error(result.failure.kind);
91
+ const tokensUsed = Object.values(result.usage.modelUsage)
92
+ .reduce((sum, u) => sum + u.inputTokens + u.outputTokens, 0);
93
+ return { content: result.output, tokensUsed };
94
+ },
95
+ };
96
+ ```
97
+
98
+ `runAgent` has no system prompt field and budgets in turns and dollars, not tokens. The
99
+ adapter therefore folds the system message into the prompt and ignores `maxTokens`.
100
+ Throwing on a failed run is what lets the enricher record the failure and carry on.
101
+
102
+ ## Dependencies
103
+
104
+ `web-tree-sitter` is a peer dependency, for the same reason code-parser makes it one:
105
+ `ParsedFile.tree` is a web-tree-sitter `Tree`, and the extractors walk it with that
106
+ package's `Node` type, so the parser and the extractors must see one copy. This package
107
+ imports it for types only. `@jscpd/core` 3.5.10 and `@jscpd/tokenizer` 3.5.4 stay pinned
108
+ at the original's versions.
109
+
110
+ ## Gotchas
111
+
112
+ - `IdiomsExtractor.extract` and `ReviewVoiceExtractor.extract` return `[]`. Their real
113
+ inputs are a set of sources and a list of review comments, not one parsed file.
114
+ - `FormattingExtractor` works on text with regular expressions, so it also reports
115
+ semicolons and quote style for Python files.
116
+ - `extractFromConfig` swallows every error, including a missing file or malformed JSON,
117
+ and returns `[]`.
118
+ - Review-voice observations use `file: "_reviews"` and `line: 0`.
119
+ - `STABILITY_MAP` uses spellings the extractors do not emit. Examples are
120
+ `naming.variables` against the emitted `naming.variable`, and `controlFlow.*` against
121
+ `control-flow.*`. Those types fall back to `medium`. This is the original's behaviour.
122
+
123
+ The worked example and the full observation list are in the site reference page,
124
+ `site/reference/style-analyzer.md`.
@@ -0,0 +1,351 @@
1
+ import { Extractor as Extractor$1, ParsedFile } from '@titan-design/code-parser';
2
+ export { ParsedFile, getLanguageFromPath, getSupportedLanguages, parseFile, shouldIncludeFile } from '@titan-design/code-parser';
3
+ import { ProfileCategory, SeverityThresholds, Severity } from '@titan-design/style-profile';
4
+ export { Severity, SeverityThresholds } from '@titan-design/style-profile';
5
+
6
+ interface IngestConfig {
7
+ repos: string[];
8
+ since?: string;
9
+ until?: string;
10
+ languages: string[];
11
+ githubToken: string;
12
+ cacheDir?: string;
13
+ }
14
+ interface CodeFile {
15
+ path: string;
16
+ content: string;
17
+ language: string;
18
+ repo: string;
19
+ sha: string;
20
+ }
21
+ interface ReviewComment$1 {
22
+ body: string;
23
+ path: string;
24
+ line: number | null;
25
+ author: string;
26
+ prNumber: number;
27
+ repo: string;
28
+ createdAt: string;
29
+ }
30
+ interface PullRequest {
31
+ number: number;
32
+ title: string;
33
+ repo: string;
34
+ author: string;
35
+ files: PullRequestFile[];
36
+ comments: ReviewComment$1[];
37
+ }
38
+ interface PullRequestFile {
39
+ filename: string;
40
+ status: "added" | "modified" | "removed" | "renamed";
41
+ patch?: string;
42
+ additions: number;
43
+ deletions: number;
44
+ }
45
+ interface CodeCorpus {
46
+ files: CodeFile[];
47
+ pullRequests: PullRequest[];
48
+ reviewComments: ReviewComment$1[];
49
+ metadata: IngestMetadata;
50
+ }
51
+ interface IngestMetadata {
52
+ repos: string[];
53
+ author: string;
54
+ since?: string;
55
+ until?: string;
56
+ fetchedAt: string;
57
+ totalCommits: number;
58
+ totalFiles: number;
59
+ totalReviewComments: number;
60
+ }
61
+
62
+ /**
63
+ * Categories that extractors emit. Includes all ProfileCategory values
64
+ * plus extractor-specific categories that don't map directly to profile sections.
65
+ */
66
+ type ObservationCategory = ProfileCategory | "control-flow" | "error-handling" | "reviewVoice" | "idioms" | "complexity";
67
+ interface Observation {
68
+ /** Feature type, e.g. "naming.variable", "naming.function" */
69
+ type: string;
70
+ /** Top-level category, e.g. "naming", "structure" */
71
+ category: ObservationCategory;
72
+ /** Detected value, e.g. "camelCase", true, 28 */
73
+ value: string | number | boolean;
74
+ /** Source file path */
75
+ file: string;
76
+ /** Line number (1-based) */
77
+ line: number;
78
+ /** Additional context for aggregation */
79
+ metadata?: Record<string, unknown>;
80
+ }
81
+ /** Style extractor — produces per-file style Observations. */
82
+ type StyleExtractor = Extractor$1<Observation>;
83
+ /** @deprecated Use StyleExtractor. Retained for backward compatibility. */
84
+ type Extractor = StyleExtractor;
85
+
86
+ declare class NamingExtractor implements StyleExtractor {
87
+ readonly name = "naming";
88
+ extract(file: ParsedFile): Observation[];
89
+ private processNode;
90
+ private processTypeScriptNode;
91
+ private processTypeScriptVariable;
92
+ private processTypeScriptParameter;
93
+ private processPythonNode;
94
+ private processPythonAssignment;
95
+ private processPythonParameters;
96
+ private observeVariable;
97
+ private observeDeclarationName;
98
+ private detectPrivateMembers;
99
+ private addObservation;
100
+ }
101
+
102
+ declare class StructureExtractor implements StyleExtractor {
103
+ readonly name = "structure";
104
+ extract(file: ParsedFile): Observation[];
105
+ private extractImports;
106
+ private addImportOrder;
107
+ private importSource;
108
+ private extractExports;
109
+ private processExport;
110
+ private exportProximity;
111
+ private detectBarrelFile;
112
+ }
113
+
114
+ declare class ControlFlowExtractor implements StyleExtractor {
115
+ readonly name = "control-flow";
116
+ extract(file: ParsedFile): Observation[];
117
+ private walk;
118
+ private processNode;
119
+ private processBranch;
120
+ private processLoop;
121
+ private processCall;
122
+ private processMemberCall;
123
+ private detectGuardClause;
124
+ private isFunctionBody;
125
+ private detectElseAfterReturn;
126
+ private containsReturn;
127
+ private emit;
128
+ }
129
+
130
+ declare class DocumentationExtractor implements StyleExtractor {
131
+ readonly name = "documentation";
132
+ extract(file: ParsedFile): Observation[];
133
+ private walkDeclarations;
134
+ private isDeclaration;
135
+ private processDeclaration;
136
+ private hasLeadingDoc;
137
+ private hasJSDoc;
138
+ private hasPythonDocstring;
139
+ private extractTags;
140
+ private extractPythonDocTags;
141
+ private walkForComments;
142
+ private getCommentPlacement;
143
+ private isExported;
144
+ private getLeadingComment;
145
+ private emit;
146
+ }
147
+
148
+ declare class ErrorHandlingExtractor implements StyleExtractor {
149
+ readonly name = "error-handling";
150
+ extract(file: ParsedFile): Observation[];
151
+ private walk;
152
+ private processNode;
153
+ private processDeclaration;
154
+ private processFunctionCheck;
155
+ private analyzeCatchClauses;
156
+ private hasInstanceofCheck;
157
+ private detectCustomErrorClass;
158
+ private detectResultType;
159
+ private detectResultReturnType;
160
+ private detectAssertNever;
161
+ private detectExhaustiveSwitch;
162
+ private emit;
163
+ }
164
+
165
+ declare class FormattingExtractor implements StyleExtractor {
166
+ readonly name = "formatting";
167
+ extract(file: ParsedFile): Observation[];
168
+ extractFromConfig(configPath: string): Promise<Observation[]>;
169
+ extractFromSource(source: string, filePath: string): Observation[];
170
+ private detectSemicolons;
171
+ private detectQuoteStyle;
172
+ private detectTrailingCommas;
173
+ private detectBraceStyle;
174
+ private detectIndentation;
175
+ private countIndentation;
176
+ private isComment;
177
+ private isStructuralLine;
178
+ private isStatementStart;
179
+ private findGcdOfArray;
180
+ private gcd;
181
+ }
182
+
183
+ declare class ComplexityExtractor implements StyleExtractor {
184
+ readonly name = "complexity";
185
+ extract(file: ParsedFile): Observation[];
186
+ private functionObservations;
187
+ private getFunctionTypes;
188
+ private findFunctions;
189
+ private analyzeFunctionNode;
190
+ private countStatements;
191
+ private isStatement;
192
+ private isTypeScriptStatement;
193
+ private isPythonStatement;
194
+ private measureNestingDepth;
195
+ private measureCyclomaticComplexity;
196
+ private branchIncrement;
197
+ }
198
+
199
+ interface SourceFile {
200
+ content: string;
201
+ path: string;
202
+ language: string;
203
+ }
204
+ declare class IdiomsExtractor implements StyleExtractor {
205
+ readonly name = "idioms";
206
+ private minLines;
207
+ private minTokens;
208
+ constructor(options?: {
209
+ minLines?: number;
210
+ minTokens?: number;
211
+ });
212
+ extract(_file: ParsedFile): Observation[];
213
+ extractFromSources(sources: SourceFile[]): Promise<Observation[]>;
214
+ private cloneObservation;
215
+ private detectClones;
216
+ private runDetector;
217
+ private toCloneGroup;
218
+ private toInstance;
219
+ private extractFragment;
220
+ private languageToFormat;
221
+ private groupClones;
222
+ private mergeInstances;
223
+ private normalizeFragment;
224
+ private summarizeClone;
225
+ }
226
+
227
+ interface ReviewComment {
228
+ body: string;
229
+ file?: string;
230
+ }
231
+ declare class ReviewVoiceExtractor implements StyleExtractor {
232
+ readonly name = "reviewVoice";
233
+ extract(_file: ParsedFile): Observation[];
234
+ extractFromComments(comments: ReviewComment[]): Observation[];
235
+ private categorizeTopics;
236
+ private countTopics;
237
+ private matchTopics;
238
+ private extractKeywords;
239
+ }
240
+
241
+ /**
242
+ * Canonical set of style extractors. Single source of truth — CLI commands,
243
+ * scripts, and tests should import this rather than reconstructing the list.
244
+ */
245
+ declare function createStyleExtractors(): StyleExtractor[];
246
+
247
+ type Stability = "high" | "medium" | "low";
248
+ declare function lookupStability(type: string): Stability;
249
+
250
+ interface StabilityWeights {
251
+ high: number;
252
+ medium: number;
253
+ low: number;
254
+ }
255
+ declare function computeConfidence(consistency: number, stability: Stability, weights?: StabilityWeights): number;
256
+ declare function mapSeverity(confidence: number, thresholds?: SeverityThresholds): Severity;
257
+
258
+ interface FrequencyDistribution {
259
+ values: Map<string | number | boolean, number>;
260
+ total: number;
261
+ dominant: string | number | boolean;
262
+ consistency: number;
263
+ }
264
+
265
+ interface AggregatedFeature {
266
+ type: string;
267
+ category: ObservationCategory;
268
+ convention: string | number | boolean | string[];
269
+ distribution: FrequencyDistribution;
270
+ confidence: number;
271
+ stability: Stability;
272
+ severity: Severity;
273
+ needsReview: boolean;
274
+ examples: Observation[];
275
+ }
276
+ interface AggregatorConfig {
277
+ stabilityWeights?: StabilityWeights;
278
+ severityThresholds?: SeverityThresholds;
279
+ reviewThreshold?: number;
280
+ maxExamples?: number;
281
+ }
282
+ interface AggregatorResult {
283
+ features: Map<string, AggregatedFeature>;
284
+ reviewQueue: AggregatedFeature[];
285
+ summary: {
286
+ totalObservations: number;
287
+ totalFeatures: number;
288
+ avgConfidence: number;
289
+ featuresNeedingReview: number;
290
+ };
291
+ }
292
+ declare class Aggregator {
293
+ private stabilityWeights;
294
+ private severityThresholds;
295
+ private reviewThreshold;
296
+ private maxExamples;
297
+ constructor(config?: AggregatorConfig);
298
+ aggregate(observations: Observation[]): AggregatorResult;
299
+ private buildFeature;
300
+ private extractCategory;
301
+ }
302
+
303
+ interface LlmMessage {
304
+ role: "system" | "user" | "assistant";
305
+ content: string;
306
+ }
307
+ interface LlmResponse {
308
+ content: string;
309
+ tokensUsed: number;
310
+ }
311
+ interface LlmProvider {
312
+ name: string;
313
+ generate(messages: LlmMessage[], options: {
314
+ maxTokens: number;
315
+ }): Promise<LlmResponse>;
316
+ }
317
+
318
+ declare const AI_ENRICHED_FEATURES: string[];
319
+ declare function needsAiEnrichment(featureType: string): boolean;
320
+
321
+ interface EnrichmentEntry {
322
+ featureType: string;
323
+ description: string;
324
+ tokensUsed: number;
325
+ }
326
+ interface EnrichmentError {
327
+ featureType: string;
328
+ error: string;
329
+ }
330
+ interface EnrichmentResult {
331
+ enriched: Map<string, EnrichmentEntry>;
332
+ errors: EnrichmentError[];
333
+ totalTokensUsed: number;
334
+ budgetExceeded: boolean;
335
+ skipped: boolean;
336
+ }
337
+ interface EnricherConfig {
338
+ provider: LlmProvider;
339
+ enabled?: boolean;
340
+ totalTokenBudget?: number;
341
+ }
342
+ declare class Enricher {
343
+ private runner;
344
+ private enabled;
345
+ constructor(config: EnricherConfig);
346
+ enrich(features: Map<string, AggregatedFeature>): Promise<EnrichmentResult>;
347
+ private buildJobs;
348
+ private buildPromptInput;
349
+ }
350
+
351
+ export { AI_ENRICHED_FEATURES, type AggregatedFeature, Aggregator, type AggregatorConfig, type AggregatorResult, type CodeCorpus, type CodeFile, ComplexityExtractor, ControlFlowExtractor, DocumentationExtractor, Enricher, type EnricherConfig, type EnrichmentEntry, type EnrichmentError, type EnrichmentResult, ErrorHandlingExtractor, type Extractor, FormattingExtractor, type FrequencyDistribution, IdiomsExtractor, type IngestConfig, type IngestMetadata, type LlmMessage, type LlmProvider, type LlmResponse, NamingExtractor, type Observation, type ObservationCategory, type PullRequest, type PullRequestFile, type ReviewComment$1 as ReviewComment, ReviewVoiceExtractor, type Stability, type StabilityWeights, StructureExtractor, type StyleExtractor, computeConfidence, createStyleExtractors, lookupStability, mapSeverity, needsAiEnrichment };