tribunal-kit 4.5.0 → 4.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (217) hide show
  1. package/.agent/.shared/ui-ux-pro-max/README.md +4 -4
  2. package/.agent/ARCHITECTURE.md +279 -277
  3. package/.agent/GEMINI.md +127 -121
  4. package/.agent/agents/accessibility-reviewer.md +187 -187
  5. package/.agent/agents/ai-code-reviewer.md +199 -199
  6. package/.agent/agents/api-architect.md +71 -66
  7. package/.agent/agents/backend-specialist.md +219 -215
  8. package/.agent/agents/cloud-engineer.md +98 -0
  9. package/.agent/agents/code-archaeologist.md +168 -161
  10. package/.agent/agents/database-architect.md +184 -184
  11. package/.agent/agents/db-latency-auditor.md +213 -216
  12. package/.agent/agents/debugger.md +198 -191
  13. package/.agent/agents/dependency-reviewer.md +106 -103
  14. package/.agent/agents/devops-engineer.md +218 -218
  15. package/.agent/agents/documentation-writer.md +209 -201
  16. package/.agent/agents/explorer-agent.md +167 -160
  17. package/.agent/agents/frontend-reviewer.md +162 -160
  18. package/.agent/agents/frontend-specialist.md +257 -248
  19. package/.agent/agents/game-developer.md +48 -48
  20. package/.agent/agents/logic-reviewer.md +118 -116
  21. package/.agent/agents/mobile-developer.md +197 -200
  22. package/.agent/agents/mobile-reviewer.md +159 -162
  23. package/.agent/agents/orchestrator.md +187 -181
  24. package/.agent/agents/penetration-tester.md +160 -157
  25. package/.agent/agents/performance-optimizer.md +183 -183
  26. package/.agent/agents/performance-reviewer.md +178 -178
  27. package/.agent/agents/precedence-reviewer.md +251 -250
  28. package/.agent/agents/product-manager.md +149 -142
  29. package/.agent/agents/product-owner.md +81 -80
  30. package/.agent/agents/project-planner.md +152 -142
  31. package/.agent/agents/qa-automation-engineer.md +216 -225
  32. package/.agent/agents/resilience-reviewer.md +88 -88
  33. package/.agent/agents/schema-reviewer.md +67 -67
  34. package/.agent/agents/security-auditor.md +180 -174
  35. package/.agent/agents/seo-specialist.md +188 -193
  36. package/.agent/agents/sql-reviewer.md +159 -161
  37. package/.agent/agents/supervisor-agent.md +173 -184
  38. package/.agent/agents/swarm-worker-contracts.md +170 -166
  39. package/.agent/agents/swarm-worker-registry.md +92 -92
  40. package/.agent/agents/system-architect.md +85 -0
  41. package/.agent/agents/test-coverage-reviewer.md +158 -160
  42. package/.agent/agents/test-engineer.md +118 -118
  43. package/.agent/agents/throughput-optimizer.md +291 -299
  44. package/.agent/agents/type-safety-reviewer.md +182 -175
  45. package/.agent/agents/ui-ux-auditor.md +300 -292
  46. package/.agent/agents/vitals-reviewer.md +223 -223
  47. package/.agent/mcp_config.json +37 -40
  48. package/.agent/patterns/generator.md +11 -9
  49. package/.agent/patterns/inversion.md +14 -12
  50. package/.agent/patterns/pipeline.md +11 -9
  51. package/.agent/patterns/reviewer.md +15 -13
  52. package/.agent/patterns/tool-wrapper.md +11 -9
  53. package/.agent/routing_index.json +654 -0
  54. package/.agent/rules/GEMINI.md +358 -352
  55. package/.agent/scripts/compile_router.py +112 -0
  56. package/.agent/scripts/migrate_skills_frontmatter.py +64 -0
  57. package/.agent/scripts/strengthen_skills.js +1 -1
  58. package/.agent/skills/advanced-rag-pipelines/SKILL.md +56 -0
  59. package/.agent/skills/agent-organizer/SKILL.md +156 -150
  60. package/.agent/skills/agentic-patterns/SKILL.md +313 -315
  61. package/.agent/skills/ai-prompt-injection-defense/SKILL.md +190 -184
  62. package/.agent/skills/api-patterns/SKILL.md +253 -247
  63. package/.agent/skills/api-security-auditor/SKILL.md +195 -193
  64. package/.agent/skills/app-builder/SKILL.md +573 -572
  65. package/.agent/skills/app-builder/templates/SKILL.md +108 -115
  66. package/.agent/skills/app-builder/templates/astro-static/TEMPLATE.md +76 -76
  67. package/.agent/skills/app-builder/templates/chrome-extension/TEMPLATE.md +92 -92
  68. package/.agent/skills/app-builder/templates/cli-tool/TEMPLATE.md +88 -88
  69. package/.agent/skills/app-builder/templates/electron-desktop/TEMPLATE.md +88 -88
  70. package/.agent/skills/app-builder/templates/express-api/TEMPLATE.md +83 -83
  71. package/.agent/skills/app-builder/templates/flutter-app/TEMPLATE.md +90 -90
  72. package/.agent/skills/app-builder/templates/monorepo-turborepo/TEMPLATE.md +90 -90
  73. package/.agent/skills/app-builder/templates/nextjs-fullstack/TEMPLATE.md +126 -122
  74. package/.agent/skills/app-builder/templates/nextjs-saas/TEMPLATE.md +127 -122
  75. package/.agent/skills/app-builder/templates/nextjs-static/TEMPLATE.md +172 -169
  76. package/.agent/skills/app-builder/templates/nuxt-app/TEMPLATE.md +139 -134
  77. package/.agent/skills/app-builder/templates/python-fastapi/TEMPLATE.md +83 -83
  78. package/.agent/skills/app-builder/templates/react-native-app/TEMPLATE.md +122 -119
  79. package/.agent/skills/appflow-wireframe/SKILL.md +146 -145
  80. package/.agent/skills/architecture/SKILL.md +226 -219
  81. package/.agent/skills/authentication-best-practices/SKILL.md +197 -189
  82. package/.agent/skills/backend-security-expert/SKILL.md +16 -2
  83. package/.agent/skills/bash-linux/SKILL.md +179 -179
  84. package/.agent/skills/behavioral-modes/SKILL.md +239 -223
  85. package/.agent/skills/brainstorming/SKILL.md +498 -486
  86. package/.agent/skills/browser-native-ai/SKILL.md +57 -4
  87. package/.agent/skills/building-native-ui/SKILL.md +202 -202
  88. package/.agent/skills/cicd-pro/SKILL.md +442 -0
  89. package/.agent/skills/clean-code/SKILL.md +400 -381
  90. package/.agent/skills/cloud-architect/SKILL.md +439 -0
  91. package/.agent/skills/code-review-checklist/SKILL.md +203 -194
  92. package/.agent/skills/config-validator/SKILL.md +165 -165
  93. package/.agent/skills/containerization-pro/SKILL.md +452 -0
  94. package/.agent/skills/csharp-developer/SKILL.md +518 -518
  95. package/.agent/skills/data-validation-schemas/SKILL.md +333 -328
  96. package/.agent/skills/database-design/SKILL.md +247 -240
  97. package/.agent/skills/deployment-procedures/SKILL.md +172 -169
  98. package/.agent/skills/devops-engineer/SKILL.md +345 -345
  99. package/.agent/skills/devops-incident-responder/SKILL.md +143 -137
  100. package/.agent/skills/doc.md +209 -177
  101. package/.agent/skills/documentation-templates/SKILL.md +291 -279
  102. package/.agent/skills/edge-computing/SKILL.md +183 -181
  103. package/.agent/skills/error-resilience/SKILL.md +411 -428
  104. package/.agent/skills/extract-design-system/SKILL.md +160 -158
  105. package/.agent/skills/framer-motion-expert/SKILL.md +253 -244
  106. package/.agent/skills/frontend-design/SKILL.md +208 -201
  107. package/.agent/skills/frontend-security-expert/SKILL.md +16 -3
  108. package/.agent/skills/game-design-expert/SKILL.md +132 -129
  109. package/.agent/skills/game-engineering-expert/SKILL.md +148 -146
  110. package/.agent/skills/generative-ui-expert/SKILL.md +57 -1
  111. package/.agent/skills/geo-fundamentals/SKILL.md +148 -147
  112. package/.agent/skills/git-pro/SKILL.md +435 -0
  113. package/.agent/skills/github-operations/SKILL.md +335 -329
  114. package/.agent/skills/gsap-core/SKILL.md +319 -308
  115. package/.agent/skills/gsap-frameworks/SKILL.md +213 -207
  116. package/.agent/skills/gsap-performance/SKILL.md +139 -133
  117. package/.agent/skills/gsap-plugins/SKILL.md +486 -480
  118. package/.agent/skills/gsap-react/SKILL.md +202 -189
  119. package/.agent/skills/gsap-scrolltrigger/SKILL.md +357 -350
  120. package/.agent/skills/gsap-timeline/SKILL.md +165 -161
  121. package/.agent/skills/gsap-utils/SKILL.md +344 -338
  122. package/.agent/skills/harness-protocol/SKILL.md +48 -0
  123. package/.agent/skills/i18n-localization/SKILL.md +174 -163
  124. package/.agent/skills/intelligent-routing/SKILL.md +202 -246
  125. package/.agent/skills/knowledge-graph/SKILL.md +60 -52
  126. package/.agent/skills/lint-and-validate/SKILL.md +261 -261
  127. package/.agent/skills/llm-engineering/SKILL.md +400 -394
  128. package/.agent/skills/local-first/SKILL.md +178 -178
  129. package/.agent/skills/mcp-builder/SKILL.md +143 -142
  130. package/.agent/skills/mobile-design/SKILL.md +272 -263
  131. package/.agent/skills/monorepo-management/SKILL.md +335 -334
  132. package/.agent/skills/motion-engineering/SKILL.md +266 -234
  133. package/.agent/skills/nextjs-react-expert/SKILL.md +236 -234
  134. package/.agent/skills/nodejs-best-practices/SKILL.md +547 -548
  135. package/.agent/skills/observability/SKILL.md +343 -343
  136. package/.agent/skills/parallel-agents/SKILL.md +143 -146
  137. package/.agent/skills/performance-profiling/SKILL.md +259 -267
  138. package/.agent/skills/plan-writing/SKILL.md +150 -142
  139. package/.agent/skills/platform-engineer/SKILL.md +148 -147
  140. package/.agent/skills/playwright-best-practices/SKILL.md +188 -187
  141. package/.agent/skills/powershell-windows/SKILL.md +162 -162
  142. package/.agent/skills/project-idioms/SKILL.md +137 -137
  143. package/.agent/skills/python-patterns/SKILL.md +260 -259
  144. package/.agent/skills/python-pro/SKILL.md +324 -323
  145. package/.agent/skills/react-specialist/SKILL.md +305 -277
  146. package/.agent/skills/readme-builder/SKILL.md +310 -300
  147. package/.agent/skills/realtime-patterns/SKILL.md +323 -319
  148. package/.agent/skills/red-team-tactics/SKILL.md +231 -218
  149. package/.agent/skills/rust-pro/SKILL.md +671 -673
  150. package/.agent/skills/seo-fundamentals/SKILL.md +179 -179
  151. package/.agent/skills/server-management/SKILL.md +218 -214
  152. package/.agent/skills/shadcn-ui-expert/SKILL.md +231 -231
  153. package/.agent/skills/skill-creator/SKILL.md +87 -86
  154. package/.agent/skills/sql-pro/SKILL.md +629 -629
  155. package/.agent/skills/supabase-postgres-best-practices/SKILL.md +97 -97
  156. package/.agent/skills/swiftui-expert/SKILL.md +204 -201
  157. package/.agent/skills/system-design-pro/SKILL.md +345 -0
  158. package/.agent/skills/systematic-debugging/SKILL.md +153 -142
  159. package/.agent/skills/tailwind-patterns/SKILL.md +610 -566
  160. package/.agent/skills/tdd-workflow/SKILL.md +169 -161
  161. package/.agent/skills/test-result-analyzer/SKILL.md +313 -309
  162. package/.agent/skills/testing-patterns/SKILL.md +566 -579
  163. package/.agent/skills/trend-researcher/SKILL.md +243 -237
  164. package/.agent/skills/typescript-advanced/SKILL.md +336 -335
  165. package/.agent/skills/ui-ux-pro-max/SKILL.md +590 -562
  166. package/.agent/skills/ui-ux-researcher/SKILL.md +244 -244
  167. package/.agent/skills/vue-expert/SKILL.md +294 -275
  168. package/.agent/skills/vulnerability-scanner/SKILL.md +416 -404
  169. package/.agent/skills/web-accessibility-auditor/SKILL.md +219 -218
  170. package/.agent/skills/web-design-guidelines/SKILL.md +192 -186
  171. package/.agent/skills/webapp-testing/SKILL.md +167 -169
  172. package/.agent/skills/webgpu-performance/SKILL.md +56 -2
  173. package/.agent/skills/whimsy-injector/SKILL.md +346 -325
  174. package/.agent/skills/workflow-optimizer/SKILL.md +231 -229
  175. package/.agent/workflows/acf.md +141 -0
  176. package/.agent/workflows/api-tester.md +176 -151
  177. package/.agent/workflows/audit.md +150 -127
  178. package/.agent/workflows/brainstorm.md +134 -110
  179. package/.agent/workflows/changelog.md +140 -112
  180. package/.agent/workflows/create.md +168 -124
  181. package/.agent/workflows/debug.md +190 -165
  182. package/.agent/workflows/deploy.md +201 -180
  183. package/.agent/workflows/enhance.md +154 -128
  184. package/.agent/workflows/fix.md +136 -114
  185. package/.agent/workflows/generate.md +198 -183
  186. package/.agent/workflows/marathon.md +37 -11
  187. package/.agent/workflows/migrate.md +184 -160
  188. package/.agent/workflows/orchestrate.md +192 -168
  189. package/.agent/workflows/performance-benchmarker.md +135 -114
  190. package/.agent/workflows/plan.md +196 -173
  191. package/.agent/workflows/preview.md +103 -80
  192. package/.agent/workflows/refactor.md +192 -161
  193. package/.agent/workflows/review-ai.md +125 -101
  194. package/.agent/workflows/review.md +141 -116
  195. package/.agent/workflows/session.md +122 -94
  196. package/.agent/workflows/status.md +101 -79
  197. package/.agent/workflows/strengthen-skills.md +164 -138
  198. package/.agent/workflows/super-prompt.md +24 -0
  199. package/.agent/workflows/swarm.md +193 -179
  200. package/.agent/workflows/test.md +211 -189
  201. package/.agent/workflows/tribunal-backend.md +136 -105
  202. package/.agent/workflows/tribunal-database.md +122 -95
  203. package/.agent/workflows/tribunal-frontend.md +221 -96
  204. package/.agent/workflows/tribunal-full.md +129 -100
  205. package/.agent/workflows/tribunal-mobile.md +122 -95
  206. package/.agent/workflows/tribunal-performance.md +136 -110
  207. package/.agent/workflows/tribunal-speed.md +209 -183
  208. package/.agent/workflows/ui-ux-pro-max.md +145 -122
  209. package/README.md +107 -55
  210. package/bin/mcp-server.js +159 -0
  211. package/bin/tribunal-kit.js +105 -29
  212. package/bin/wrapper.js +16 -7
  213. package/mcp_config.json +9 -0
  214. package/package.json +94 -86
  215. package/scripts/changelog.js +4 -3
  216. package/scripts/validate-payload.js +6 -1
  217. package/scripts/postinstall.js +0 -127
@@ -1,398 +1,402 @@
1
- ---
2
- name: llm-engineering
3
- description: LLM engineering mastery for production AI systems. Prompt engineering, RAG pipeline design, vector store selection, embedding strategies, chunking, reranking, structured output, function calling, streaming, evals, guard-rails, cost optimization, and LLMOps. Use when building AI features, chat interfaces, semantic search, or any system calling an LLM API.
4
- allowed-tools: Read, Write, Edit, Glob, Grep
5
- version: 3.2.0
6
- last-updated: 2026-04-07
7
- applies-to-model: gemini-3-1-pro, claude-3-7-sonnet
8
- ---
9
-
10
- # LLM Engineering — Production AI Systems Mastery
11
-
12
- ---
13
-
14
- ## Model Selection
15
-
16
- ```
17
- Model │ Use Case │ Cost Tier
18
- ─────────────────────────┼───────────────────────────────────────┼──────────
19
- GPT-4o │ Complex reasoning, vision, code │ $$$
20
- GPT-4o-mini │ Classification, summaries, chat │ $
21
- o3-mini │ Deep reasoning, math, code review │ $$
22
- Claude 3.7 Sonnet │ Long documents, analysis, code │ $$$
23
- Claude 3.5 Haiku │ Fast responses, simple tasks │ $
24
- Gemini 3.1 Pro (High) │ Large context, multimodal, code │ $$$
25
- Gemini 3.0 Flash │ High throughput, cost-efficient │ $
26
- Llama 3.3 70B (open) │ Self-hosted, data privacy │ Free*
27
- Mistral Large 2 │ European data residency, code │ $$
28
-
29
- * = compute costs only
30
-
31
- Selection rules:
32
- 1. Start with the cheapest model that passes your evals
33
- 2. Upgrade only when eval scores require it
34
- 3. Use large models for complex reasoning, small for classification/routing
35
- 4. Fine-tune ONLY after prompt engineering and RAG are exhausted
36
- 5. ❌ HALLUCINATION TRAP: Model names change frequently — always verify current names
37
- from provider docs before hardcoding (e.g. "gpt-4o" vs "gpt-4o-2024-11-20")
38
- ```
39
-
40
- ---
41
-
42
- ## Prompt Engineering
43
-
44
- ### System Prompt Design
45
-
46
- ```typescript
47
- const SYSTEM_PROMPT = `You are a customer support agent for Acme Corp.
48
-
49
- ## Rules
50
- 1. Answer ONLY questions about Acme products and services.
51
- 2. If you don't know the answer, say "I'll connect you with a specialist."
52
- 3. Never discuss competitors.
53
- 4. Never make up product features or pricing.
54
- 5. Keep responses under 200 words.
55
-
56
- ## Response Format
57
- - Use bullet points for lists
58
- - Include product links when relevant
59
- - End with a follow-up question
60
-
61
- ## Context
62
- Current date: ${new Date().toISOString().split("T")[0]}
63
- User plan: {{user_plan}}
64
- `;
65
-
66
- // ❌ HALLUCINATION TRAP: System prompts are NOT secrets
67
- // Users can extract system prompts with jailbreak techniques
68
- // Never put API keys, internal URLs, or secrets in system prompts
69
- ```
70
-
71
- ### Structured Output (JSON Mode)
72
-
73
- ```typescript
74
- import { z } from "zod";
75
- import OpenAI from "openai";
76
-
77
- const SentimentSchema = z.object({
78
- sentiment: z.enum(["positive", "negative", "neutral"]),
79
- confidence: z.number().min(0).max(1),
80
- reasoning: z.string(),
81
- topics: z.array(z.string()),
82
- });
83
-
84
- // OpenAI — json_schema mode (strict = true enforces schema exactly)
85
- async function analyzeSentiment(text: string) {
86
- const response = await openai.chat.completions.create({
87
- model: "gpt-4o-mini",
88
- response_format: {
89
- type: "json_schema",
90
- json_schema: {
91
- name: "sentiment_analysis",
92
- strict: true,
93
- schema: {
94
- type: "object",
95
- properties: {
96
- sentiment: { type: "string", enum: ["positive", "negative", "neutral"] },
97
- confidence: { type: "number" },
98
- reasoning: { type: "string" },
99
- topics: { type: "array", items: { type: "string" } },
100
- },
101
- required: ["sentiment", "confidence", "reasoning", "topics"],
102
- additionalProperties: false, // required for strict mode
103
- },
104
- },
105
- },
106
- messages: [{ role: "system", content: "Analyze sentiment." }, { role: "user", content: text }],
107
- });
108
- const raw = JSON.parse(response.choices[0].message.content ?? "{}");
109
- return SentimentSchema.parse(raw); // always validate with Zod even in strict mode
110
- }
111
-
112
- // Gemini — response_mime_type + response_schema
113
- import { GoogleGenerativeAI, SchemaType } from "@google/generative-ai";
114
- const genAI = new GoogleGenerativeAI(process.env.GEMINI_API_KEY!);
115
- const model = genAI.getGenerativeModel({
116
- model: "gemini-2.0-flash",
117
- generationConfig: {
118
- responseMimeType: "application/json",
119
- responseSchema: {
120
- type: SchemaType.OBJECT,
121
- properties: {
122
- sentiment: { type: SchemaType.STRING, enum: ["positive", "negative", "neutral"] },
123
- confidence: { type: SchemaType.NUMBER },
124
- topics: { type: SchemaType.ARRAY, items: { type: SchemaType.STRING } },
125
- },
126
- required: ["sentiment", "confidence", "topics"],
127
- },
128
- },
129
- });
130
-
131
- // ❌ HALLUCINATION TRAP: Always validate LLM JSON output with Zod/schema
132
- // LLMs produce malformed JSON, wrong types, missing fields even with strict mode
133
- // ❌ const result = JSON.parse(response); // trust blindly
134
- // ✅ const result = Schema.parse(JSON.parse(response)); // validate always
135
- ```
136
-
137
- ### Function Calling / Tool Use
138
-
139
- ```typescript
140
- const tools: OpenAI.ChatCompletionTool[] = [
141
- {
142
- type: "function",
143
- function: {
144
- name: "search_products",
145
- description: "Search products by name, category, or price range",
146
- parameters: {
147
- type: "object",
148
- properties: {
149
- query: { type: "string", description: "Search query" },
150
- category: { type: "string", enum: ["electronics", "clothing", "home"] },
151
- max_price: { type: "number", description: "Maximum price in USD" },
152
- },
153
- required: ["query"],
154
- },
155
- },
156
- },
157
- {
158
- type: "function",
159
- function: {
160
- name: "get_order_status",
161
- description: "Get the status of an order by order ID",
162
- parameters: {
163
- type: "object",
164
- properties: {
165
- order_id: { type: "string", description: "The order ID (e.g., ORD-12345)" },
166
- },
167
- required: ["order_id"],
168
- },
169
- },
170
- },
171
- ];
172
-
173
- // Tool execution loop
174
- async function chatWithTools(userMessage: string) {
175
- const messages: OpenAI.ChatCompletionMessageParam[] = [
176
- { role: "system", content: SYSTEM_PROMPT },
177
- { role: "user", content: userMessage },
178
- ];
179
-
180
- let response = await openai.chat.completions.create({
181
- model: "gpt-4o-mini",
182
- messages,
183
- tools,
184
- });
185
-
186
- // Process tool calls
187
- while (response.choices[0].finish_reason === "tool_calls") {
188
- const toolCalls = response.choices[0].message.tool_calls ?? [];
189
- messages.push(response.choices[0].message);
190
-
191
- for (const call of toolCalls) {
192
- const args = JSON.parse(call.function.arguments);
193
- const result = await executeFunction(call.function.name, args);
194
- messages.push({
195
- role: "tool",
196
- tool_call_id: call.id,
197
- content: JSON.stringify(result),
198
- });
199
- }
200
-
201
- response = await openai.chat.completions.create({
202
- model: "gpt-4o-mini",
203
- messages,
204
- tools,
205
- });
206
- }
207
-
208
- return response.choices[0].message.content;
209
- }
210
- ```
211
-
212
- ---
213
-
214
- ## RAG (Retrieval-Augmented Generation)
215
-
216
- ### Pipeline
217
-
218
- ```
219
- User Query
220
-
221
- [1] Embed query → vector
222
-
223
- [2] Search vector DB → top K chunks
224
-
225
- [3] (Optional) Rerank results → top N
226
-
227
- [4] Build prompt: system + context chunks + query
228
-
229
- [5] LLM generates answer with citations
230
-
231
- [6] Validate response (hallucination check)
232
- ```
233
-
234
- ### Chunking Strategy
235
-
236
- ```typescript
237
- // ❌ BAD: Arbitrary character splitting
238
- const chunks = text.match(/.{1,1000}/g); // breaks mid-sentence, mid-word
239
-
240
- // ✅ GOOD: Semantic chunking with overlap
241
- function chunkDocument(text: string, options: ChunkOptions = {}): Chunk[] {
242
- const {
243
- maxTokens = 512, // chunk size
244
- overlapTokens = 50, // overlap between chunks
245
- separator = "\n\n", // split on paragraph boundaries first
246
- } = options;
247
-
248
- const paragraphs = text.split(separator);
249
- const chunks: Chunk[] = [];
250
- let current = "";
251
-
252
- for (const para of paragraphs) {
253
- if (tokenCount(current + para) > maxTokens && current) {
254
- chunks.push({ text: current.trim(), tokens: tokenCount(current) });
255
- // Keep overlap from previous chunk
256
- const words = current.split(" ");
257
- current = words.slice(-overlapTokens).join(" ") + separator + para;
258
- } else {
259
- current += separator + para;
260
- }
261
- }
262
- if (current.trim()) chunks.push({ text: current.trim(), tokens: tokenCount(current) });
263
-
264
- return chunks;
265
- }
266
-
267
- // Chunk size guidelines:
268
- // 256-512 tokens → precise retrieval (Q&A, support)
269
- // 512-1024 tokens → balanced (general RAG)
270
- // 1024-2048 tokens → broad context (summarization)
271
- ```
272
-
273
- ### Vector Store Selection
274
-
275
- ```
276
- pgvector (PostgreSQL) → Already using Postgres, <10M vectors, simple
277
- Pinecone → Managed, serverless, easy scaling
278
- Weaviate → Hybrid search (vector + keyword), multi-model
279
- Qdrant → High performance, Rust-based, self-hostable
280
- Chroma → Local development, prototyping
281
- Milvus → Enterprise scale, GPU acceleration
282
-
283
- // ❌ HALLUCINATION TRAP: Vector search is NOT keyword search
284
- // "Apple CEO" might not find "Tim Cook runs Apple Inc."
285
- // Use HYBRID search (vector + BM25 keyword) for production
286
- ```
287
-
288
- ---
289
-
290
- ## Streaming
291
-
292
- ```typescript
293
- // Server-Sent Events for AI token streaming
294
- app.get("/api/chat", async (req, res) => {
295
- res.setHeader("Content-Type", "text/event-stream");
296
- res.setHeader("Cache-Control", "no-cache");
297
- res.setHeader("Connection", "keep-alive");
298
-
299
- const stream = await openai.chat.completions.create({
300
- model: "gpt-4o-mini",
301
- messages: [{ role: "user", content: req.query.message as string }],
302
- stream: true,
303
- });
304
-
305
- for await (const chunk of stream) {
306
- const content = chunk.choices[0]?.delta?.content;
307
- if (content) {
308
- res.write(`data: ${JSON.stringify({ content })}\n\n`);
309
- }
310
- }
311
-
312
- res.write("data: [DONE]\n\n");
313
- res.end();
314
- });
315
-
316
- // Client-side consumption
317
- const eventSource = new EventSource(`/api/chat?message=${encodeURIComponent(msg)}`);
318
- eventSource.onmessage = (event) => {
319
- if (event.data === "[DONE]") { eventSource.close(); return; }
320
- const { content } = JSON.parse(event.data);
321
- appendToChat(content);
322
- };
323
- ```
324
-
325
- ---
326
-
327
- ## Cost Optimization
328
-
329
- ```
330
- 1. Prompt caching → Cache system prompts (OpenAI, Anthropic support this)
331
- 2. Output token limiting → Set max_tokens to prevent runaway responses
332
- 3. Tiered models → Use cheap models for classification, expensive for reasoning
333
- 4. Batch processing → Use batch APIs for offline processing (50% discount)
334
- 5. Chunked context → Send only relevant chunks, not entire documents
335
- 6. Response streaming → Stream to reduce TTFT (time to first token)
336
- 7. Structured output → Shorter JSON responses vs verbose prose
337
-
338
- // Cost estimation:
339
- // GPT-4o: ~$2.50/1M input, ~$10/1M output
340
- // GPT-4o-mini: ~$0.15/1M input, ~$0.60/1M output
341
- // 1M tokens ≈ 750,000 words ≈ 3,000 pages
342
- ```
343
-
344
- ---
345
-
346
-
347
- ---
348
-
349
-
350
-
351
- AI coding assistants often fall into specific bad habits when dealing with this domain. These are strictly forbidden:
352
-
353
- 1. **Over-engineering:** Proposing complex abstractions or distributed systems when a simpler approach suffices.
354
- 2. **Hallucinated Libraries/Methods:** Using non-existent methods or packages. Always `// VERIFY` or check `package.json` / `requirements.txt`.
355
- 3. **Skipping Edge Cases:** Writing the "happy path" and ignoring error handling, timeouts, or data validation.
356
- 4. **Context Amnesia:** Forgetting the user's constraints and offering generic advice instead of tailored solutions.
357
- 5. **Silent Degradation:** Catching and suppressing errors without logging or re-raising.
358
-
359
- ---
360
-
361
-
362
-
363
- **Slash command: `/review` or `/tribunal-full`**
364
- **Active reviewers: `logic-reviewer` · `security-auditor`**
365
-
366
- ### ❌ Forbidden AI Tropes
367
-
368
- 1. **Blind Assumptions:** Never make an assumption without documenting it clearly with `// VERIFY: [reason]`.
369
- 2. **Silent Degradation:** Catching and suppressing errors without logging or handling.
370
- 3. **Context Amnesia:** Forgetting the user's constraints and offering generic advice instead of tailored solutions.
371
-
372
-
373
-
374
- Review these questions before confirming output:
375
- ```
376
- ✅ Did I rely ONLY on real, verified tools and methods?
377
- ✅ Is this solution appropriately scoped to the user's constraints?
378
- ✅ Did I handle potential failure modes and edge cases?
379
- ✅ Have I avoided generic boilerplate that doesn't add value?
380
- ```
381
-
382
- ### 🛑 Verification-Before-Completion (VBC) Protocol
383
-
384
- **CRITICAL:** You must follow a strict "evidence-based closeout" state machine.
385
- - ❌ **Forbidden:** Declaring a task complete because the output "looks correct."
386
- - ✅ **Required:** You are explicitly forbidden from finalizing any task without providing **concrete evidence** (terminal output, passing tests, compile success, or equivalent proof) that your output works as intended.
387
-
388
-
389
- ## Pre-Flight Checklist
390
- - [ ] Have I reviewed the user's specific constraints and requests?
391
- - [ ] Have I checked the environment for relevant existing implementations?
392
-
393
- ## VBC Protocol (Verification-Before-Completion)
394
- You MUST verify existing code signatures and variables before attempting to modify or call them. No hallucination is permitted.
1
+ ---
2
+ name: llm-engineering
3
+ description: LLM engineering mastery for production AI systems. Prompt engineering, RAG pipeline design, vector store selection, embedding strategies, chunking, reranking, structured output, function calling, streaming, evals, guard-rails, cost optimization, and LLMOps. Use when building AI features, chat interfaces, semantic search, or any system calling an LLM API.
4
+ allowed-tools: Read, Write, Edit, Glob, Grep
5
+ version: 3.2.0
6
+ last-updated: 2026-04-07
7
+ applies-to-model: gemini-3-1-pro, claude-3-7-sonnet
8
+ routing:
9
+ domain: general
10
+ tier: basic
11
+ ---
395
12
 
13
+ # LLM Engineering — Production AI Systems Mastery
14
+
15
+ ---
16
+
17
+ ## Model Selection
18
+
19
+ ```
20
+ Model │ Use Case │ Cost Tier
21
+ ─────────────────────────┼───────────────────────────────────────┼──────────
22
+ GPT-4o │ Complex reasoning, vision, code │ $$$
23
+ GPT-4o-mini │ Classification, summaries, chat │ $
24
+ o3-mini │ Deep reasoning, math, code review │ $$
25
+ Claude 3.7 Sonnet │ Long documents, analysis, code │ $$$
26
+ Claude 3.5 Haiku │ Fast responses, simple tasks │ $
27
+ Gemini 3.1 Pro (High) │ Large context, multimodal, code │ $$$
28
+ Gemini 3.0 Flash │ High throughput, cost-efficient │ $
29
+ Llama 3.3 70B (open) │ Self-hosted, data privacy │ Free*
30
+ Mistral Large 2 │ European data residency, code │ $$
31
+
32
+ * = compute costs only
33
+
34
+ Selection rules:
35
+ 1. Start with the cheapest model that passes your evals
36
+ 2. Upgrade only when eval scores require it
37
+ 3. Use large models for complex reasoning, small for classification/routing
38
+ 4. Fine-tune ONLY after prompt engineering and RAG are exhausted
39
+ 5. ❌ HALLUCINATION TRAP: Model names change frequently — always verify current names
40
+ from provider docs before hardcoding (e.g. "gpt-4o" vs "gpt-4o-2024-11-20")
41
+ ```
42
+
43
+ ---
44
+
45
+ ## Prompt Engineering
46
+
47
+ ### System Prompt Design
48
+
49
+ ```typescript
50
+ const SYSTEM_PROMPT = `You are a customer support agent for Acme Corp.
51
+
52
+ ## Rules
53
+ 1. Answer ONLY questions about Acme products and services.
54
+ 2. If you don't know the answer, say "I'll connect you with a specialist."
55
+ 3. Never discuss competitors.
56
+ 4. Never make up product features or pricing.
57
+ 5. Keep responses under 200 words.
58
+
59
+ ## Response Format
60
+ - Use bullet points for lists
61
+ - Include product links when relevant
62
+ - End with a follow-up question
63
+
64
+ ## Context
65
+ Current date: ${new Date().toISOString().split("T")[0]}
66
+ User plan: {{user_plan}}
67
+ `;
68
+
69
+ // ❌ HALLUCINATION TRAP: System prompts are NOT secrets
70
+ // Users can extract system prompts with jailbreak techniques
71
+ // Never put API keys, internal URLs, or secrets in system prompts
72
+ ```
73
+
74
+ ### Structured Output (JSON Mode)
75
+
76
+ ```typescript
77
+ import { z } from "zod";
78
+ import OpenAI from "openai";
79
+
80
+ const SentimentSchema = z.object({
81
+ sentiment: z.enum(["positive", "negative", "neutral"]),
82
+ confidence: z.number().min(0).max(1),
83
+ reasoning: z.string(),
84
+ topics: z.array(z.string()),
85
+ });
86
+
87
+ // OpenAI — json_schema mode (strict = true enforces schema exactly)
88
+ async function analyzeSentiment(text: string) {
89
+ const response = await openai.chat.completions.create({
90
+ model: "gpt-4o-mini",
91
+ response_format: {
92
+ type: "json_schema",
93
+ json_schema: {
94
+ name: "sentiment_analysis",
95
+ strict: true,
96
+ schema: {
97
+ type: "object",
98
+ properties: {
99
+ sentiment: { type: "string", enum: ["positive", "negative", "neutral"] },
100
+ confidence: { type: "number" },
101
+ reasoning: { type: "string" },
102
+ topics: { type: "array", items: { type: "string" } },
103
+ },
104
+ required: ["sentiment", "confidence", "reasoning", "topics"],
105
+ additionalProperties: false, // required for strict mode
106
+ },
107
+ },
108
+ },
109
+ messages: [
110
+ { role: "system", content: "Analyze sentiment." },
111
+ { role: "user", content: text },
112
+ ],
113
+ });
114
+ const raw = JSON.parse(response.choices[0].message.content ?? "{}");
115
+ return SentimentSchema.parse(raw); // always validate with Zod even in strict mode
116
+ }
117
+
118
+ // Gemini — response_mime_type + response_schema
119
+ import { GoogleGenerativeAI, SchemaType } from "@google/generative-ai";
120
+ const genAI = new GoogleGenerativeAI(process.env.GEMINI_API_KEY!);
121
+ const model = genAI.getGenerativeModel({
122
+ model: "gemini-2.0-flash",
123
+ generationConfig: {
124
+ responseMimeType: "application/json",
125
+ responseSchema: {
126
+ type: SchemaType.OBJECT,
127
+ properties: {
128
+ sentiment: { type: SchemaType.STRING, enum: ["positive", "negative", "neutral"] },
129
+ confidence: { type: SchemaType.NUMBER },
130
+ topics: { type: SchemaType.ARRAY, items: { type: SchemaType.STRING } },
131
+ },
132
+ required: ["sentiment", "confidence", "topics"],
133
+ },
134
+ },
135
+ });
136
+
137
+ // ❌ HALLUCINATION TRAP: Always validate LLM JSON output with Zod/schema
138
+ // LLMs produce malformed JSON, wrong types, missing fields even with strict mode
139
+ // ❌ const result = JSON.parse(response); // trust blindly
140
+ // ✅ const result = Schema.parse(JSON.parse(response)); // validate always
141
+ ```
142
+
143
+ ### Function Calling / Tool Use
144
+
145
+ ```typescript
146
+ const tools: OpenAI.ChatCompletionTool[] = [
147
+ {
148
+ type: "function",
149
+ function: {
150
+ name: "search_products",
151
+ description: "Search products by name, category, or price range",
152
+ parameters: {
153
+ type: "object",
154
+ properties: {
155
+ query: { type: "string", description: "Search query" },
156
+ category: { type: "string", enum: ["electronics", "clothing", "home"] },
157
+ max_price: { type: "number", description: "Maximum price in USD" },
158
+ },
159
+ required: ["query"],
160
+ },
161
+ },
162
+ },
163
+ {
164
+ type: "function",
165
+ function: {
166
+ name: "get_order_status",
167
+ description: "Get the status of an order by order ID",
168
+ parameters: {
169
+ type: "object",
170
+ properties: {
171
+ order_id: { type: "string", description: "The order ID (e.g., ORD-12345)" },
172
+ },
173
+ required: ["order_id"],
174
+ },
175
+ },
176
+ },
177
+ ];
178
+
179
+ // Tool execution loop
180
+ async function chatWithTools(userMessage: string) {
181
+ const messages: OpenAI.ChatCompletionMessageParam[] = [
182
+ { role: "system", content: SYSTEM_PROMPT },
183
+ { role: "user", content: userMessage },
184
+ ];
185
+
186
+ let response = await openai.chat.completions.create({
187
+ model: "gpt-4o-mini",
188
+ messages,
189
+ tools,
190
+ });
191
+
192
+ // Process tool calls
193
+ while (response.choices[0].finish_reason === "tool_calls") {
194
+ const toolCalls = response.choices[0].message.tool_calls ?? [];
195
+ messages.push(response.choices[0].message);
196
+
197
+ for (const call of toolCalls) {
198
+ const args = JSON.parse(call.function.arguments);
199
+ const result = await executeFunction(call.function.name, args);
200
+ messages.push({
201
+ role: "tool",
202
+ tool_call_id: call.id,
203
+ content: JSON.stringify(result),
204
+ });
205
+ }
206
+
207
+ response = await openai.chat.completions.create({
208
+ model: "gpt-4o-mini",
209
+ messages,
210
+ tools,
211
+ });
212
+ }
213
+
214
+ return response.choices[0].message.content;
215
+ }
216
+ ```
217
+
218
+ ---
219
+
220
+ ## RAG (Retrieval-Augmented Generation)
221
+
222
+ ### Pipeline
223
+
224
+ ```
225
+ User Query
226
+
227
+ [1] Embed query → vector
228
+
229
+ [2] Search vector DB → top K chunks
230
+
231
+ [3] (Optional) Rerank results → top N
232
+
233
+ [4] Build prompt: system + context chunks + query
234
+
235
+ [5] LLM generates answer with citations
236
+
237
+ [6] Validate response (hallucination check)
238
+ ```
239
+
240
+ ### Chunking Strategy
241
+
242
+ ```typescript
243
+ // ❌ BAD: Arbitrary character splitting
244
+ const chunks = text.match(/.{1,1000}/g); // breaks mid-sentence, mid-word
245
+
246
+ // ✅ GOOD: Semantic chunking with overlap
247
+ function chunkDocument(text: string, options: ChunkOptions = {}): Chunk[] {
248
+ const {
249
+ maxTokens = 512, // chunk size
250
+ overlapTokens = 50, // overlap between chunks
251
+ separator = "\n\n", // split on paragraph boundaries first
252
+ } = options;
253
+
254
+ const paragraphs = text.split(separator);
255
+ const chunks: Chunk[] = [];
256
+ let current = "";
257
+
258
+ for (const para of paragraphs) {
259
+ if (tokenCount(current + para) > maxTokens && current) {
260
+ chunks.push({ text: current.trim(), tokens: tokenCount(current) });
261
+ // Keep overlap from previous chunk
262
+ const words = current.split(" ");
263
+ current = words.slice(-overlapTokens).join(" ") + separator + para;
264
+ } else {
265
+ current += separator + para;
266
+ }
267
+ }
268
+ if (current.trim()) chunks.push({ text: current.trim(), tokens: tokenCount(current) });
269
+
270
+ return chunks;
271
+ }
272
+
273
+ // Chunk size guidelines:
274
+ // 256-512 tokens → precise retrieval (Q&A, support)
275
+ // 512-1024 tokens → balanced (general RAG)
276
+ // 1024-2048 tokens → broad context (summarization)
277
+ ```
278
+
279
+ ### Vector Store Selection
280
+
281
+ ```
282
+ pgvector (PostgreSQL) → Already using Postgres, <10M vectors, simple
283
+ Pinecone → Managed, serverless, easy scaling
284
+ Weaviate → Hybrid search (vector + keyword), multi-model
285
+ Qdrant → High performance, Rust-based, self-hostable
286
+ Chroma → Local development, prototyping
287
+ Milvus → Enterprise scale, GPU acceleration
288
+
289
+ // ❌ HALLUCINATION TRAP: Vector search is NOT keyword search
290
+ // "Apple CEO" might not find "Tim Cook runs Apple Inc."
291
+ // Use HYBRID search (vector + BM25 keyword) for production
292
+ ```
293
+
294
+ ---
295
+
296
+ ## Streaming
297
+
298
+ ```typescript
299
+ // Server-Sent Events for AI token streaming
300
+ app.get("/api/chat", async (req, res) => {
301
+ res.setHeader("Content-Type", "text/event-stream");
302
+ res.setHeader("Cache-Control", "no-cache");
303
+ res.setHeader("Connection", "keep-alive");
304
+
305
+ const stream = await openai.chat.completions.create({
306
+ model: "gpt-4o-mini",
307
+ messages: [{ role: "user", content: req.query.message as string }],
308
+ stream: true,
309
+ });
310
+
311
+ for await (const chunk of stream) {
312
+ const content = chunk.choices[0]?.delta?.content;
313
+ if (content) {
314
+ res.write(`data: ${JSON.stringify({ content })}\n\n`);
315
+ }
316
+ }
317
+
318
+ res.write("data: [DONE]\n\n");
319
+ res.end();
320
+ });
321
+
322
+ // Client-side consumption
323
+ const eventSource = new EventSource(`/api/chat?message=${encodeURIComponent(msg)}`);
324
+ eventSource.onmessage = (event) => {
325
+ if (event.data === "[DONE]") {
326
+ eventSource.close();
327
+ return;
328
+ }
329
+ const { content } = JSON.parse(event.data);
330
+ appendToChat(content);
331
+ };
332
+ ```
333
+
334
+ ---
335
+
336
+ ## Cost Optimization
337
+
338
+ ```
339
+ 1. Prompt caching → Cache system prompts (OpenAI, Anthropic support this)
340
+ 2. Output token limiting → Set max_tokens to prevent runaway responses
341
+ 3. Tiered models → Use cheap models for classification, expensive for reasoning
342
+ 4. Batch processing → Use batch APIs for offline processing (50% discount)
343
+ 5. Chunked context → Send only relevant chunks, not entire documents
344
+ 6. Response streaming → Stream to reduce TTFT (time to first token)
345
+ 7. Structured output → Shorter JSON responses vs verbose prose
346
+
347
+ // Cost estimation:
348
+ // GPT-4o: ~$2.50/1M input, ~$10/1M output
349
+ // GPT-4o-mini: ~$0.15/1M input, ~$0.60/1M output
350
+ // 1M tokens ≈ 750,000 words ≈ 3,000 pages
351
+ ```
352
+
353
+ ---
354
+
355
+ ---
356
+
357
+ AI coding assistants often fall into specific bad habits when dealing with this domain. These are strictly forbidden:
358
+
359
+ 1. **Over-engineering:** Proposing complex abstractions or distributed systems when a simpler approach suffices.
360
+ 2. **Hallucinated Libraries/Methods:** Using non-existent methods or packages. Always `// VERIFY` or check `package.json` / `requirements.txt`.
361
+ 3. **Skipping Edge Cases:** Writing the "happy path" and ignoring error handling, timeouts, or data validation.
362
+ 4. **Context Amnesia:** Forgetting the user's constraints and offering generic advice instead of tailored solutions.
363
+ 5. **Silent Degradation:** Catching and suppressing errors without logging or re-raising.
364
+
365
+ ---
366
+
367
+ **Slash command: `/review` or `/tribunal-full`**
368
+ **Active reviewers: `logic-reviewer` · `security-auditor`**
369
+
370
+ ### ❌ Forbidden AI Tropes
371
+
372
+ 1. **Blind Assumptions:** Never make an assumption without documenting it clearly with `// VERIFY: [reason]`.
373
+ 2. **Silent Degradation:** Catching and suppressing errors without logging or handling.
374
+ 3. **Context Amnesia:** Forgetting the user's constraints and offering generic advice instead of tailored solutions.
375
+
376
+ Review these questions before confirming output:
377
+
378
+ ```
379
+ ✅ Did I rely ONLY on real, verified tools and methods?
380
+ ✅ Is this solution appropriately scoped to the user's constraints?
381
+ ✅ Did I handle potential failure modes and edge cases?
382
+ ✅ Have I avoided generic boilerplate that doesn't add value?
383
+ ```
384
+
385
+ ### 🛑 Verification-Before-Completion (VBC) Protocol
386
+
387
+ **CRITICAL:** You must follow a strict "evidence-based closeout" state machine.
388
+
389
+ - ❌ **Forbidden:** Declaring a task complete because the output "looks correct."
390
+ - ✅ **Required:** You are explicitly forbidden from finalizing any task without providing **concrete evidence** (terminal output, passing tests, compile success, or equivalent proof) that your output works as intended.
391
+
392
+ ## Pre-Flight Checklist
393
+
394
+ - [ ] Have I reviewed the user's specific constraints and requests?
395
+ - [ ] Have I checked the environment for relevant existing implementations?
396
+
397
+ ## VBC Protocol (Verification-Before-Completion)
398
+
399
+ You MUST verify existing code signatures and variables before attempting to modify or call them. No hallucination is permitted.
396
400
 
397
401
  ---
398
402
 
@@ -422,6 +426,7 @@ AI coding assistants often fall into specific bad habits when dealing with this
422
426
  ### ✅ Pre-Flight Self-Audit
423
427
 
424
428
  Review these questions before confirming output:
429
+
425
430
  ```
426
431
  ✅ Did I rely ONLY on real, verified tools and methods?
427
432
  ✅ Is this solution appropriately scoped to the user's constraints?
@@ -432,5 +437,6 @@ Review these questions before confirming output:
432
437
  ### 🛑 Verification-Before-Completion (VBC) Protocol
433
438
 
434
439
  **CRITICAL:** You must follow a strict "evidence-based closeout" state machine.
440
+
435
441
  - ❌ **Forbidden:** Declaring a task complete because the output "looks correct."
436
442
  - ✅ **Required:** You are explicitly forbidden from finalizing any task without providing **concrete evidence** (terminal output, passing tests, compile success, or equivalent proof) that your output works as intended.