tribunal-kit 4.5.1 → 4.6.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (214) hide show
  1. package/.agent/.shared/ui-ux-pro-max/README.md +4 -4
  2. package/.agent/ARCHITECTURE.md +282 -277
  3. package/.agent/agents/accessibility-reviewer.md +187 -187
  4. package/.agent/agents/ai-code-reviewer.md +199 -199
  5. package/.agent/agents/api-architect.md +71 -66
  6. package/.agent/agents/backend-specialist.md +219 -215
  7. package/.agent/agents/cloud-engineer.md +98 -0
  8. package/.agent/agents/code-archaeologist.md +168 -161
  9. package/.agent/agents/database-architect.md +184 -184
  10. package/.agent/agents/db-latency-auditor.md +213 -216
  11. package/.agent/agents/debugger.md +198 -191
  12. package/.agent/agents/dependency-reviewer.md +106 -103
  13. package/.agent/agents/devops-engineer.md +218 -218
  14. package/.agent/agents/documentation-writer.md +209 -201
  15. package/.agent/agents/explorer-agent.md +167 -160
  16. package/.agent/agents/frontend-reviewer.md +162 -160
  17. package/.agent/agents/frontend-specialist.md +257 -248
  18. package/.agent/agents/game-developer.md +48 -48
  19. package/.agent/agents/logic-reviewer.md +118 -116
  20. package/.agent/agents/mobile-developer.md +197 -200
  21. package/.agent/agents/mobile-reviewer.md +159 -162
  22. package/.agent/agents/orchestrator.md +187 -181
  23. package/.agent/agents/penetration-tester.md +160 -157
  24. package/.agent/agents/performance-optimizer.md +183 -183
  25. package/.agent/agents/performance-reviewer.md +178 -178
  26. package/.agent/agents/precedence-reviewer.md +251 -250
  27. package/.agent/agents/product-manager.md +149 -142
  28. package/.agent/agents/product-owner.md +81 -80
  29. package/.agent/agents/project-planner.md +152 -142
  30. package/.agent/agents/qa-automation-engineer.md +216 -225
  31. package/.agent/agents/resilience-reviewer.md +88 -88
  32. package/.agent/agents/schema-reviewer.md +67 -67
  33. package/.agent/agents/security-auditor.md +180 -174
  34. package/.agent/agents/seo-specialist.md +188 -193
  35. package/.agent/agents/sql-reviewer.md +159 -161
  36. package/.agent/agents/supervisor-agent.md +173 -184
  37. package/.agent/agents/swarm-worker-contracts.md +170 -166
  38. package/.agent/agents/swarm-worker-registry.md +92 -92
  39. package/.agent/agents/system-architect.md +85 -0
  40. package/.agent/agents/test-coverage-reviewer.md +158 -160
  41. package/.agent/agents/test-engineer.md +118 -118
  42. package/.agent/agents/throughput-optimizer.md +291 -299
  43. package/.agent/agents/type-safety-reviewer.md +182 -175
  44. package/.agent/agents/ui-ux-auditor.md +300 -292
  45. package/.agent/agents/vitals-reviewer.md +223 -223
  46. package/.agent/mcp_config.json +37 -40
  47. package/.agent/patterns/generator.md +11 -9
  48. package/.agent/patterns/inversion.md +14 -12
  49. package/.agent/patterns/pipeline.md +11 -9
  50. package/.agent/patterns/reviewer.md +15 -13
  51. package/.agent/patterns/tool-wrapper.md +11 -9
  52. package/.agent/routing_index.json +714 -0
  53. package/.agent/rules/GEMINI.md +359 -352
  54. package/.agent/scripts/compile_router.py +112 -0
  55. package/.agent/scripts/migrate_skills_frontmatter.py +64 -0
  56. package/.agent/scripts/strengthen_skills.js +1 -1
  57. package/.agent/skills/advanced-rag-pipelines/SKILL.md +56 -0
  58. package/.agent/skills/agent-organizer/SKILL.md +156 -150
  59. package/.agent/skills/agentic-patterns/SKILL.md +313 -315
  60. package/.agent/skills/ai-prompt-injection-defense/SKILL.md +190 -184
  61. package/.agent/skills/api-patterns/SKILL.md +253 -247
  62. package/.agent/skills/api-security-auditor/SKILL.md +195 -193
  63. package/.agent/skills/app-builder/SKILL.md +573 -572
  64. package/.agent/skills/app-builder/templates/SKILL.md +108 -115
  65. package/.agent/skills/app-builder/templates/astro-static/TEMPLATE.md +76 -76
  66. package/.agent/skills/app-builder/templates/chrome-extension/TEMPLATE.md +92 -92
  67. package/.agent/skills/app-builder/templates/cli-tool/TEMPLATE.md +88 -88
  68. package/.agent/skills/app-builder/templates/electron-desktop/TEMPLATE.md +88 -88
  69. package/.agent/skills/app-builder/templates/express-api/TEMPLATE.md +83 -83
  70. package/.agent/skills/app-builder/templates/flutter-app/TEMPLATE.md +90 -90
  71. package/.agent/skills/app-builder/templates/monorepo-turborepo/TEMPLATE.md +90 -90
  72. package/.agent/skills/app-builder/templates/nextjs-fullstack/TEMPLATE.md +126 -122
  73. package/.agent/skills/app-builder/templates/nextjs-saas/TEMPLATE.md +127 -122
  74. package/.agent/skills/app-builder/templates/nextjs-static/TEMPLATE.md +172 -169
  75. package/.agent/skills/app-builder/templates/nuxt-app/TEMPLATE.md +139 -134
  76. package/.agent/skills/app-builder/templates/python-fastapi/TEMPLATE.md +83 -83
  77. package/.agent/skills/app-builder/templates/react-native-app/TEMPLATE.md +122 -119
  78. package/.agent/skills/appflow-wireframe/SKILL.md +146 -145
  79. package/.agent/skills/architecture/SKILL.md +226 -219
  80. package/.agent/skills/authentication-best-practices/SKILL.md +197 -189
  81. package/.agent/skills/backend-security-expert/SKILL.md +16 -2
  82. package/.agent/skills/bash-linux/SKILL.md +179 -179
  83. package/.agent/skills/behavioral-modes/SKILL.md +239 -223
  84. package/.agent/skills/brainstorming/SKILL.md +498 -486
  85. package/.agent/skills/browser-native-ai/SKILL.md +57 -4
  86. package/.agent/skills/building-native-ui/SKILL.md +202 -202
  87. package/.agent/skills/cicd-pro/SKILL.md +442 -0
  88. package/.agent/skills/clean-code/SKILL.md +400 -381
  89. package/.agent/skills/cloud-architect/SKILL.md +439 -0
  90. package/.agent/skills/code-review-checklist/SKILL.md +203 -194
  91. package/.agent/skills/config-validator/SKILL.md +165 -165
  92. package/.agent/skills/containerization-pro/SKILL.md +452 -0
  93. package/.agent/skills/csharp-developer/SKILL.md +518 -518
  94. package/.agent/skills/data-validation-schemas/SKILL.md +333 -328
  95. package/.agent/skills/database-design/SKILL.md +247 -240
  96. package/.agent/skills/deployment-procedures/SKILL.md +172 -169
  97. package/.agent/skills/devops-engineer/SKILL.md +345 -345
  98. package/.agent/skills/devops-incident-responder/SKILL.md +143 -137
  99. package/.agent/skills/documentation-templates/SKILL.md +291 -279
  100. package/.agent/skills/edge-computing/SKILL.md +183 -181
  101. package/.agent/skills/emil-design-eng/SKILL.md +147 -0
  102. package/.agent/skills/error-resilience/SKILL.md +411 -428
  103. package/.agent/skills/extract-design-system/SKILL.md +160 -158
  104. package/.agent/skills/framer-motion-expert/SKILL.md +253 -244
  105. package/.agent/skills/frontend-design/SKILL.md +208 -201
  106. package/.agent/skills/frontend-security-expert/SKILL.md +16 -3
  107. package/.agent/skills/game-design-expert/SKILL.md +132 -129
  108. package/.agent/skills/game-engineering-expert/SKILL.md +148 -146
  109. package/.agent/skills/generative-ui-expert/SKILL.md +57 -1
  110. package/.agent/skills/geo-fundamentals/SKILL.md +148 -147
  111. package/.agent/skills/git-pro/SKILL.md +435 -0
  112. package/.agent/skills/github-operations/SKILL.md +335 -329
  113. package/.agent/skills/gsap-core/SKILL.md +319 -308
  114. package/.agent/skills/gsap-frameworks/SKILL.md +213 -207
  115. package/.agent/skills/gsap-performance/SKILL.md +139 -133
  116. package/.agent/skills/gsap-plugins/SKILL.md +486 -480
  117. package/.agent/skills/gsap-react/SKILL.md +202 -189
  118. package/.agent/skills/gsap-scrolltrigger/SKILL.md +357 -350
  119. package/.agent/skills/gsap-timeline/SKILL.md +165 -161
  120. package/.agent/skills/gsap-utils/SKILL.md +344 -338
  121. package/.agent/skills/harness-protocol/SKILL.md +48 -0
  122. package/.agent/skills/i18n-localization/SKILL.md +174 -163
  123. package/.agent/skills/intelligent-routing/SKILL.md +202 -246
  124. package/.agent/skills/knowledge-graph/SKILL.md +60 -52
  125. package/.agent/skills/lint-and-validate/SKILL.md +261 -261
  126. package/.agent/skills/llm-engineering/SKILL.md +400 -394
  127. package/.agent/skills/local-first/SKILL.md +178 -178
  128. package/.agent/skills/mcp-builder/SKILL.md +143 -142
  129. package/.agent/skills/mobile-design/SKILL.md +272 -263
  130. package/.agent/skills/monorepo-management/SKILL.md +335 -334
  131. package/.agent/skills/motion-engineering/SKILL.md +266 -234
  132. package/.agent/skills/nextjs-react-expert/SKILL.md +236 -234
  133. package/.agent/skills/nodejs-best-practices/SKILL.md +547 -548
  134. package/.agent/skills/observability/SKILL.md +343 -343
  135. package/.agent/skills/parallel-agents/SKILL.md +143 -146
  136. package/.agent/skills/performance-profiling/SKILL.md +259 -267
  137. package/.agent/skills/plan-writing/SKILL.md +150 -142
  138. package/.agent/skills/platform-engineer/SKILL.md +148 -147
  139. package/.agent/skills/playwright-best-practices/SKILL.md +188 -187
  140. package/.agent/skills/powershell-windows/SKILL.md +162 -162
  141. package/.agent/skills/project-idioms/SKILL.md +137 -137
  142. package/.agent/skills/python-patterns/SKILL.md +260 -259
  143. package/.agent/skills/python-pro/SKILL.md +324 -323
  144. package/.agent/skills/react-specialist/SKILL.md +305 -277
  145. package/.agent/skills/readme-builder/SKILL.md +310 -300
  146. package/.agent/skills/realtime-patterns/SKILL.md +323 -319
  147. package/.agent/skills/red-team-tactics/SKILL.md +231 -218
  148. package/.agent/skills/review-animations/SKILL.md +72 -0
  149. package/.agent/skills/review-animations/STANDARDS.md +73 -0
  150. package/.agent/skills/rust-pro/SKILL.md +671 -673
  151. package/.agent/skills/seo-fundamentals/SKILL.md +179 -179
  152. package/.agent/skills/server-management/SKILL.md +218 -214
  153. package/.agent/skills/shadcn-ui-expert/SKILL.md +231 -231
  154. package/.agent/skills/skill-creator/SKILL.md +87 -86
  155. package/.agent/skills/sql-pro/SKILL.md +629 -629
  156. package/.agent/skills/supabase-postgres-best-practices/SKILL.md +97 -97
  157. package/.agent/skills/swiftui-expert/SKILL.md +204 -201
  158. package/.agent/skills/system-design-pro/SKILL.md +345 -0
  159. package/.agent/skills/systematic-debugging/SKILL.md +153 -142
  160. package/.agent/skills/tailwind-patterns/SKILL.md +610 -566
  161. package/.agent/skills/tdd-workflow/SKILL.md +169 -161
  162. package/.agent/skills/test-result-analyzer/SKILL.md +313 -309
  163. package/.agent/skills/testing-patterns/SKILL.md +566 -579
  164. package/.agent/skills/trend-researcher/SKILL.md +243 -237
  165. package/.agent/skills/typescript-advanced/SKILL.md +336 -335
  166. package/.agent/skills/ui-ux-pro-max/SKILL.md +590 -562
  167. package/.agent/skills/ui-ux-researcher/SKILL.md +244 -244
  168. package/.agent/skills/vue-expert/SKILL.md +294 -275
  169. package/.agent/skills/vulnerability-scanner/SKILL.md +416 -404
  170. package/.agent/skills/web-accessibility-auditor/SKILL.md +219 -218
  171. package/.agent/skills/web-design-guidelines/SKILL.md +192 -186
  172. package/.agent/skills/webapp-testing/SKILL.md +167 -169
  173. package/.agent/skills/webgpu-performance/SKILL.md +56 -2
  174. package/.agent/skills/whimsy-injector/SKILL.md +346 -325
  175. package/.agent/skills/workflow-optimizer/SKILL.md +231 -229
  176. package/.agent/workflows/acf.md +141 -0
  177. package/.agent/workflows/api-tester.md +176 -151
  178. package/.agent/workflows/audit.md +150 -127
  179. package/.agent/workflows/brainstorm.md +134 -110
  180. package/.agent/workflows/changelog.md +140 -112
  181. package/.agent/workflows/create.md +168 -124
  182. package/.agent/workflows/debug.md +190 -165
  183. package/.agent/workflows/deploy.md +201 -180
  184. package/.agent/workflows/enhance.md +154 -128
  185. package/.agent/workflows/fix.md +136 -114
  186. package/.agent/workflows/generate.md +198 -183
  187. package/.agent/workflows/marathon.md +37 -11
  188. package/.agent/workflows/migrate.md +184 -160
  189. package/.agent/workflows/orchestrate.md +192 -168
  190. package/.agent/workflows/performance-benchmarker.md +135 -114
  191. package/.agent/workflows/plan.md +196 -173
  192. package/.agent/workflows/preview.md +103 -80
  193. package/.agent/workflows/refactor.md +192 -161
  194. package/.agent/workflows/review-ai.md +125 -101
  195. package/.agent/workflows/review.md +141 -116
  196. package/.agent/workflows/session.md +122 -94
  197. package/.agent/workflows/status.md +101 -79
  198. package/.agent/workflows/strengthen-skills.md +164 -138
  199. package/.agent/workflows/super-prompt.md +24 -0
  200. package/.agent/workflows/swarm.md +193 -179
  201. package/.agent/workflows/test.md +211 -189
  202. package/.agent/workflows/tribunal-backend.md +136 -105
  203. package/.agent/workflows/tribunal-database.md +129 -95
  204. package/.agent/workflows/tribunal-frontend.md +140 -96
  205. package/.agent/workflows/tribunal-full.md +131 -100
  206. package/.agent/workflows/tribunal-mobile.md +129 -95
  207. package/.agent/workflows/tribunal-performance.md +136 -110
  208. package/.agent/workflows/tribunal-speed.md +209 -183
  209. package/.agent/workflows/ui-ux-pro-max.md +155 -122
  210. package/README.md +107 -55
  211. package/mcp_config.json +1 -3
  212. package/package.json +94 -94
  213. package/.agent/GEMINI.md +0 -121
  214. package/.agent/skills/doc.md +0 -177
@@ -1,347 +1,345 @@
1
- ---
2
- name: observability
3
- description: Production observability mastery. Structured logging (Pino/Winston), OpenTelemetry tracing, metrics (Prometheus/Grafana), SLIs/SLOs/error budgets, distributed tracing, alerting design, health checks, and AI observability. Use when setting up monitoring, debugging production issues, or designing observable distributed systems.
4
- allowed-tools: Read, Write, Edit, Glob, Grep
5
- version: 2.0.0
6
- last-updated: 2026-04-01
7
- applies-to-model: gemini-2.5-pro, claude-3-7-sonnet
8
- ---
9
-
10
- # Observability — Production Monitoring Mastery
11
-
12
- ---
13
-
14
- ## The Three Pillars
15
-
16
- ```
17
- Logs → WHAT happened (structured events)
18
- Traces → WHERE it happened (request flow across services)
19
- Metrics → HOW MUCH is happening (counters, histograms, gauges)
20
-
21
- All three are needed. Logs alone are not observability.
22
- ```
23
-
24
- ---
25
-
26
- ## Structured Logging
27
-
28
- ```typescript
29
- import pino from "pino";
30
-
31
- // ✅ Structured JSON logging
32
- const logger = pino({
33
- level: process.env.LOG_LEVEL ?? "info",
34
- timestamp: pino.stdTimeFunctions.isoTime,
35
- ...(process.env.NODE_ENV === "development" && {
36
- transport: { target: "pino-pretty" },
37
- }),
38
- });
39
-
40
- // ✅ GOOD: Structured with context
41
- logger.info({ userId: user.id, action: "login", ip: req.ip }, "User logged in");
42
- logger.error({ err, orderId: order.id, paymentGateway: "stripe" }, "Payment failed");
43
- logger.warn({ queueDepth: 1500, threshold: 1000 }, "Queue depth exceeding threshold");
44
-
45
- // BAD: Unstructured string logging
46
- console.log("User " + user.id + " logged in from " + req.ip);
47
- console.log("Error: " + error.message);
48
-
49
- // ❌ HALLUCINATION TRAP: console.log is NOT production logging
50
- // - No severity levels (info/warn/error)
51
- // - No structured fields (can't search/filter)
52
- // - No timestamps in ISO format
53
- // - Can't be collected by log aggregators
54
- // ✅ Use Pino (Node.js) or structlog (Python) for production
55
- ```
56
-
57
- ### Log Levels
58
-
59
- ```
60
- fatal → App is crashing, immediate attention required
61
- error → Operation failed, needs investigation
62
- warn → Something unexpected, but app continues
63
- info → Business events (user login, order placed, deploy)
64
- debug → Technical details (query timing, cache hit/miss)
65
- trace → Verbose debugging (only in development)
66
-
67
- Rules:
68
- - Production default: info
69
- - Never log PII (names, emails, SSNs) at any level
70
- - Never log secrets (tokens, passwords, API keys)
71
- - Log request IDs for correlation
72
- - Log durations for performance tracking
73
- ```
74
-
75
- ### Request Context / Correlation
76
-
77
- ```typescript
78
- import { AsyncLocalStorage } from "node:async_hooks";
79
-
80
- const requestContext = new AsyncLocalStorage<{ requestId: string; userId?: string }>();
81
-
82
- // Middleware: set context per request
83
- app.use((req, res, next) => {
84
- const requestId = req.headers["x-request-id"]?.toString() ?? crypto.randomUUID();
85
- res.setHeader("x-request-id", requestId);
86
- requestContext.run({ requestId, userId: req.user?.id }, next);
87
- });
88
-
89
- // Child logger with context
90
- function getLogger() {
91
- const ctx = requestContext.getStore();
92
- return logger.child({
93
- requestId: ctx?.requestId,
94
- userId: ctx?.userId,
95
- });
96
- }
97
-
98
- // Every log from this request includes requestId and userId
99
- const log = getLogger();
100
- log.info("Processing order"); // { requestId: "abc-123", userId: "42", msg: "Processing order" }
101
- ```
102
-
103
- ---
104
-
105
- ## Distributed Tracing (OpenTelemetry)
106
-
107
- ```typescript
108
- import { NodeSDK } from "@opentelemetry/sdk-node";
109
- import { getNodeAutoInstrumentations } from "@opentelemetry/auto-instrumentations-node";
110
- import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http";
111
-
112
- // Initialize OpenTelemetry
113
- const sdk = new NodeSDK({
114
- traceExporter: new OTLPTraceExporter({
115
- url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT ?? "http://localhost:4318/v1/traces",
116
- }),
117
- instrumentations: [
118
- getNodeAutoInstrumentations({
119
- "@opentelemetry/instrumentation-http": { enabled: true },
120
- "@opentelemetry/instrumentation-express": { enabled: true },
121
- "@opentelemetry/instrumentation-pg": { enabled: true },
122
- "@opentelemetry/instrumentation-redis": { enabled: true },
123
- }),
124
- ],
125
- });
126
-
127
- sdk.start();
128
-
129
- // Manual span for custom business logic
130
- import { trace } from "@opentelemetry/api";
131
-
132
- const tracer = trace.getTracer("order-service");
133
-
134
- async function processOrder(order: Order) {
135
- return tracer.startActiveSpan("processOrder", async (span) => {
136
- try {
137
- span.setAttribute("order.id", order.id);
138
- span.setAttribute("order.total", order.total);
139
- span.setAttribute("order.items.count", order.items.length);
140
-
141
- const result = await executeOrder(order);
142
- span.setStatus({ code: SpanStatusCode.OK });
143
- return result;
144
- } catch (error) {
145
- span.setStatus({ code: SpanStatusCode.ERROR, message: error.message });
146
- span.recordException(error);
147
- throw error;
148
- } finally {
149
- span.end();
150
- }
151
- });
152
- }
153
- ```
154
-
155
- ---
156
-
157
- ## Metrics
158
-
159
- ```typescript
160
- import { metrics } from "@opentelemetry/api";
161
-
162
- const meter = metrics.getMeter("api-server");
163
-
164
- // Counter — things that only go up
165
- const requestCounter = meter.createCounter("http.requests.total", {
166
- description: "Total HTTP requests",
167
- });
168
-
169
- // Histogram — request durations
170
- const requestDuration = meter.createHistogram("http.request.duration_ms", {
171
- description: "HTTP request duration in milliseconds",
172
- unit: "ms",
173
- });
174
-
175
- // Gauge — current values
176
- const activeConnections = meter.createUpDownCounter("db.connections.active", {
177
- description: "Active database connections",
178
- });
179
-
180
- // Middleware to record metrics
181
- app.use((req, res, next) => {
182
- const start = performance.now();
183
- res.on("finish", () => {
184
- const duration = performance.now() - start;
185
- requestCounter.add(1, {
186
- method: req.method,
187
- path: req.route?.path ?? req.path,
188
- status: res.statusCode.toString(),
189
- });
190
- requestDuration.record(duration, {
191
- method: req.method,
192
- status: res.statusCode.toString(),
193
- });
194
- });
195
- next();
196
- });
197
- ```
198
-
199
- ### Key Metrics to Track
200
-
201
- ```
202
- RED method (for services):
203
- Rate → requests per second
204
- Errors → error rate (4xx, 5xx)
205
- Duration → latency percentiles (P50, P95, P99)
206
-
207
- USE method (for resources):
208
- Utilization → CPU %, memory %, disk %
209
- Saturation → queue depth, thread pool saturation
210
- Errors → disk failures, OOM kills
211
-
212
- Business metrics:
213
- - Sign-ups per hour
214
- - Orders processed per minute
215
- - Revenue per day
216
- - API calls per customer
217
- ```
218
-
219
- ---
220
-
221
- ## SLIs, SLOs & Error Budgets
222
-
223
- ```
224
- SLI (Service Level Indicator) → What you measure
225
- "99.2% of requests complete in <500ms"
226
-
227
- SLO (Service Level Objective) → Your target
228
- "99.9% of requests should complete in <500ms"
229
-
230
- SLA (Service Level Agreement) → Your contract (with penalties)
231
- "99.95% uptime or we refund 10%"
232
-
233
- Error Budget = 100% - SLO
234
- SLO: 99.9% → Error budget: 0.1% → 43 min downtime/month
235
- SLO: 99.5% → Error budget: 0.5% → 3.6 hours downtime/month
236
-
237
- Rules:
238
- - Burn error budget too fast → freeze deployments
239
- - Error budget remaining → ship features faster
240
- - Don't set SLOs you can't measure
241
- - SLOs should be slightly below actual performance
242
- ```
243
-
244
- ---
245
-
246
- ## Health Checks
247
-
248
- ```typescript
249
- // Liveness: Is the process running?
250
- app.get("/health/live", (req, res) => {
251
- res.status(200).json({ status: "ok" });
252
- });
253
-
254
- // Readiness: Can it accept traffic?
255
- app.get("/health/ready", async (req, res) => {
256
- try {
257
- await db.raw("SELECT 1"); // database check
258
- await redis.ping(); // cache check
259
- res.status(200).json({
260
- status: "ready",
261
- checks: { database: "ok", cache: "ok" },
262
- });
263
- } catch (error) {
264
- res.status(503).json({
265
- status: "not ready",
266
- checks: { database: error.message },
267
- });
268
- }
269
- });
270
-
271
- // ❌ HALLUCINATION TRAP: Liveness ≠ Readiness
272
- // Liveness fails → container restarts (only for unrecoverable states)
273
- // Readiness fails → stop sending traffic (temporary — DB down, etc.)
274
- // Making liveness check the DB → DB outage restarts all containers → cascade failure
275
- ```
276
-
277
- ---
278
-
279
- ## Alerting
280
-
281
- ```
282
- Alert design rules:
283
- 1. Alert on SYMPTOMS, not causes (high latency, not "CPU is 80%")
284
- 2. Every alert must have a runbook link
285
- 3. Every alert must be ACTIONABLE — if you can't do anything, it's a notification
286
- 4. Use severity levels:
287
- - Critical → page on-call (customer-facing outage)
288
- - Warning → Slack notification (degraded, not broken)
289
- - Info → dashboard only (awareness)
290
- 5. Avoid alert fatigue — fewer, meaningful alerts beat many noisy ones
291
- ```
292
-
293
- ---
294
-
295
-
296
- ---
297
-
298
-
299
-
300
- AI coding assistants often fall into specific bad habits when dealing with this domain. These are strictly forbidden:
301
-
302
- 1. **Over-engineering:** Proposing complex abstractions or distributed systems when a simpler approach suffices.
303
- 2. **Hallucinated Libraries/Methods:** Using non-existent methods or packages. Always `// VERIFY` or check `package.json` / `requirements.txt`.
304
- 3. **Skipping Edge Cases:** Writing the "happy path" and ignoring error handling, timeouts, or data validation.
305
- 4. **Context Amnesia:** Forgetting the user's constraints and offering generic advice instead of tailored solutions.
306
- 5. **Silent Degradation:** Catching and suppressing errors without logging or re-raising.
307
-
308
- ---
309
-
310
-
311
-
312
- **Slash command: `/review` or `/tribunal-full`**
313
- **Active reviewers: `logic-reviewer` · `security-auditor`**
314
-
315
- ### ❌ Forbidden AI Tropes
316
-
317
- 1. **Blind Assumptions:** Never make an assumption without documenting it clearly with `// VERIFY: [reason]`.
318
- 2. **Silent Degradation:** Catching and suppressing errors without logging or handling.
319
- 3. **Context Amnesia:** Forgetting the user's constraints and offering generic advice instead of tailored solutions.
320
-
321
-
322
-
323
- Review these questions before confirming output:
324
- ```
325
- ✅ Did I rely ONLY on real, verified tools and methods?
326
- ✅ Is this solution appropriately scoped to the user's constraints?
327
- ✅ Did I handle potential failure modes and edge cases?
328
- ✅ Have I avoided generic boilerplate that doesn't add value?
329
- ```
330
-
331
- ### 🛑 Verification-Before-Completion (VBC) Protocol
332
-
333
- **CRITICAL:** You must follow a strict "evidence-based closeout" state machine.
334
- - ❌ **Forbidden:** Declaring a task complete because the output "looks correct."
335
- - ✅ **Required:** You are explicitly forbidden from finalizing any task without providing **concrete evidence** (terminal output, passing tests, compile success, or equivalent proof) that your output works as intended.
336
-
337
-
338
- ## Pre-Flight Checklist
339
- - [ ] Have I reviewed the user's specific constraints and requests?
340
- - [ ] Have I checked the environment for relevant existing implementations?
341
-
342
- ## VBC Protocol (Verification-Before-Completion)
343
- You MUST verify existing code signatures and variables before attempting to modify or call them. No hallucination is permitted.
1
+ ---
2
+ name: observability
3
+ description: Production observability mastery. Structured logging (Pino/Winston), OpenTelemetry tracing, metrics (Prometheus/Grafana), SLIs/SLOs/error budgets, distributed tracing, alerting design, health checks, and AI observability. Use when setting up monitoring, debugging production issues, or designing observable distributed systems.
4
+ allowed-tools: Read, Write, Edit, Glob, Grep
5
+ version: 2.0.0
6
+ last-updated: 2026-04-01
7
+ applies-to-model: gemini-2.5-pro, claude-3-7-sonnet
8
+ routing:
9
+ domain: general
10
+ tier: basic
11
+ ---
12
+
13
+ # Observability — Production Monitoring Mastery
14
+
15
+ ---
16
+
17
+ ## The Three Pillars
18
+
19
+ ```
20
+ Logs → WHAT happened (structured events)
21
+ Traces → WHERE it happened (request flow across services)
22
+ Metrics → HOW MUCH is happening (counters, histograms, gauges)
23
+
24
+ All three are needed. Logs alone are not observability.
25
+ ```
26
+
27
+ ---
28
+
29
+ ## Structured Logging
30
+
31
+ ```typescript
32
+ import pino from "pino";
33
+
34
+ // ✅ Structured JSON logging
35
+ const logger = pino({
36
+ level: process.env.LOG_LEVEL ?? "info",
37
+ timestamp: pino.stdTimeFunctions.isoTime,
38
+ ...(process.env.NODE_ENV === "development" && {
39
+ transport: { target: "pino-pretty" },
40
+ }),
41
+ });
42
+
43
+ // GOOD: Structured with context
44
+ logger.info({ userId: user.id, action: "login", ip: req.ip }, "User logged in");
45
+ logger.error({ err, orderId: order.id, paymentGateway: "stripe" }, "Payment failed");
46
+ logger.warn({ queueDepth: 1500, threshold: 1000 }, "Queue depth exceeding threshold");
344
47
 
48
+ // ❌ BAD: Unstructured string logging
49
+ console.log("User " + user.id + " logged in from " + req.ip);
50
+ console.log("Error: " + error.message);
51
+
52
+ // ❌ HALLUCINATION TRAP: console.log is NOT production logging
53
+ // - No severity levels (info/warn/error)
54
+ // - No structured fields (can't search/filter)
55
+ // - No timestamps in ISO format
56
+ // - Can't be collected by log aggregators
57
+ // ✅ Use Pino (Node.js) or structlog (Python) for production
58
+ ```
59
+
60
+ ### Log Levels
61
+
62
+ ```
63
+ fatal → App is crashing, immediate attention required
64
+ error → Operation failed, needs investigation
65
+ warn → Something unexpected, but app continues
66
+ info → Business events (user login, order placed, deploy)
67
+ debug → Technical details (query timing, cache hit/miss)
68
+ trace → Verbose debugging (only in development)
69
+
70
+ Rules:
71
+ - Production default: info
72
+ - Never log PII (names, emails, SSNs) at any level
73
+ - Never log secrets (tokens, passwords, API keys)
74
+ - Log request IDs for correlation
75
+ - Log durations for performance tracking
76
+ ```
77
+
78
+ ### Request Context / Correlation
79
+
80
+ ```typescript
81
+ import { AsyncLocalStorage } from "node:async_hooks";
82
+
83
+ const requestContext = new AsyncLocalStorage<{ requestId: string; userId?: string }>();
84
+
85
+ // Middleware: set context per request
86
+ app.use((req, res, next) => {
87
+ const requestId = req.headers["x-request-id"]?.toString() ?? crypto.randomUUID();
88
+ res.setHeader("x-request-id", requestId);
89
+ requestContext.run({ requestId, userId: req.user?.id }, next);
90
+ });
91
+
92
+ // Child logger with context
93
+ function getLogger() {
94
+ const ctx = requestContext.getStore();
95
+ return logger.child({
96
+ requestId: ctx?.requestId,
97
+ userId: ctx?.userId,
98
+ });
99
+ }
100
+
101
+ // Every log from this request includes requestId and userId
102
+ const log = getLogger();
103
+ log.info("Processing order"); // { requestId: "abc-123", userId: "42", msg: "Processing order" }
104
+ ```
105
+
106
+ ---
107
+
108
+ ## Distributed Tracing (OpenTelemetry)
109
+
110
+ ```typescript
111
+ import { NodeSDK } from "@opentelemetry/sdk-node";
112
+ import { getNodeAutoInstrumentations } from "@opentelemetry/auto-instrumentations-node";
113
+ import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http";
114
+
115
+ // Initialize OpenTelemetry
116
+ const sdk = new NodeSDK({
117
+ traceExporter: new OTLPTraceExporter({
118
+ url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT ?? "http://localhost:4318/v1/traces",
119
+ }),
120
+ instrumentations: [
121
+ getNodeAutoInstrumentations({
122
+ "@opentelemetry/instrumentation-http": { enabled: true },
123
+ "@opentelemetry/instrumentation-express": { enabled: true },
124
+ "@opentelemetry/instrumentation-pg": { enabled: true },
125
+ "@opentelemetry/instrumentation-redis": { enabled: true },
126
+ }),
127
+ ],
128
+ });
129
+
130
+ sdk.start();
131
+
132
+ // Manual span for custom business logic
133
+ import { trace } from "@opentelemetry/api";
134
+
135
+ const tracer = trace.getTracer("order-service");
136
+
137
+ async function processOrder(order: Order) {
138
+ return tracer.startActiveSpan("processOrder", async (span) => {
139
+ try {
140
+ span.setAttribute("order.id", order.id);
141
+ span.setAttribute("order.total", order.total);
142
+ span.setAttribute("order.items.count", order.items.length);
143
+
144
+ const result = await executeOrder(order);
145
+ span.setStatus({ code: SpanStatusCode.OK });
146
+ return result;
147
+ } catch (error) {
148
+ span.setStatus({ code: SpanStatusCode.ERROR, message: error.message });
149
+ span.recordException(error);
150
+ throw error;
151
+ } finally {
152
+ span.end();
153
+ }
154
+ });
155
+ }
156
+ ```
157
+
158
+ ---
159
+
160
+ ## Metrics
161
+
162
+ ```typescript
163
+ import { metrics } from "@opentelemetry/api";
164
+
165
+ const meter = metrics.getMeter("api-server");
166
+
167
+ // Counter — things that only go up
168
+ const requestCounter = meter.createCounter("http.requests.total", {
169
+ description: "Total HTTP requests",
170
+ });
171
+
172
+ // Histogram — request durations
173
+ const requestDuration = meter.createHistogram("http.request.duration_ms", {
174
+ description: "HTTP request duration in milliseconds",
175
+ unit: "ms",
176
+ });
177
+
178
+ // Gauge — current values
179
+ const activeConnections = meter.createUpDownCounter("db.connections.active", {
180
+ description: "Active database connections",
181
+ });
182
+
183
+ // Middleware to record metrics
184
+ app.use((req, res, next) => {
185
+ const start = performance.now();
186
+ res.on("finish", () => {
187
+ const duration = performance.now() - start;
188
+ requestCounter.add(1, {
189
+ method: req.method,
190
+ path: req.route?.path ?? req.path,
191
+ status: res.statusCode.toString(),
192
+ });
193
+ requestDuration.record(duration, {
194
+ method: req.method,
195
+ status: res.statusCode.toString(),
196
+ });
197
+ });
198
+ next();
199
+ });
200
+ ```
201
+
202
+ ### Key Metrics to Track
203
+
204
+ ```
205
+ RED method (for services):
206
+ Rate → requests per second
207
+ Errors → error rate (4xx, 5xx)
208
+ Duration → latency percentiles (P50, P95, P99)
209
+
210
+ USE method (for resources):
211
+ Utilization → CPU %, memory %, disk %
212
+ Saturation → queue depth, thread pool saturation
213
+ Errors → disk failures, OOM kills
214
+
215
+ Business metrics:
216
+ - Sign-ups per hour
217
+ - Orders processed per minute
218
+ - Revenue per day
219
+ - API calls per customer
220
+ ```
221
+
222
+ ---
223
+
224
+ ## SLIs, SLOs & Error Budgets
225
+
226
+ ```
227
+ SLI (Service Level Indicator) → What you measure
228
+ "99.2% of requests complete in <500ms"
229
+
230
+ SLO (Service Level Objective) → Your target
231
+ "99.9% of requests should complete in <500ms"
232
+
233
+ SLA (Service Level Agreement) → Your contract (with penalties)
234
+ "99.95% uptime or we refund 10%"
235
+
236
+ Error Budget = 100% - SLO
237
+ SLO: 99.9% → Error budget: 0.1% → 43 min downtime/month
238
+ SLO: 99.5% → Error budget: 0.5% → 3.6 hours downtime/month
239
+
240
+ Rules:
241
+ - Burn error budget too fast → freeze deployments
242
+ - Error budget remaining → ship features faster
243
+ - Don't set SLOs you can't measure
244
+ - SLOs should be slightly below actual performance
245
+ ```
246
+
247
+ ---
248
+
249
+ ## Health Checks
250
+
251
+ ```typescript
252
+ // Liveness: Is the process running?
253
+ app.get("/health/live", (req, res) => {
254
+ res.status(200).json({ status: "ok" });
255
+ });
256
+
257
+ // Readiness: Can it accept traffic?
258
+ app.get("/health/ready", async (req, res) => {
259
+ try {
260
+ await db.raw("SELECT 1"); // database check
261
+ await redis.ping(); // cache check
262
+ res.status(200).json({
263
+ status: "ready",
264
+ checks: { database: "ok", cache: "ok" },
265
+ });
266
+ } catch (error) {
267
+ res.status(503).json({
268
+ status: "not ready",
269
+ checks: { database: error.message },
270
+ });
271
+ }
272
+ });
273
+
274
+ // ❌ HALLUCINATION TRAP: Liveness ≠ Readiness
275
+ // Liveness fails → container restarts (only for unrecoverable states)
276
+ // Readiness fails → stop sending traffic (temporary — DB down, etc.)
277
+ // Making liveness check the DB → DB outage restarts all containers → cascade failure
278
+ ```
279
+
280
+ ---
281
+
282
+ ## Alerting
283
+
284
+ ```
285
+ Alert design rules:
286
+ 1. Alert on SYMPTOMS, not causes (high latency, not "CPU is 80%")
287
+ 2. Every alert must have a runbook link
288
+ 3. Every alert must be ACTIONABLE — if you can't do anything, it's a notification
289
+ 4. Use severity levels:
290
+ - Critical → page on-call (customer-facing outage)
291
+ - Warning → Slack notification (degraded, not broken)
292
+ - Info → dashboard only (awareness)
293
+ 5. Avoid alert fatigue — fewer, meaningful alerts beat many noisy ones
294
+ ```
295
+
296
+ ---
297
+
298
+ ---
299
+
300
+ AI coding assistants often fall into specific bad habits when dealing with this domain. These are strictly forbidden:
301
+
302
+ 1. **Over-engineering:** Proposing complex abstractions or distributed systems when a simpler approach suffices.
303
+ 2. **Hallucinated Libraries/Methods:** Using non-existent methods or packages. Always `// VERIFY` or check `package.json` / `requirements.txt`.
304
+ 3. **Skipping Edge Cases:** Writing the "happy path" and ignoring error handling, timeouts, or data validation.
305
+ 4. **Context Amnesia:** Forgetting the user's constraints and offering generic advice instead of tailored solutions.
306
+ 5. **Silent Degradation:** Catching and suppressing errors without logging or re-raising.
307
+
308
+ ---
309
+
310
+ **Slash command: `/review` or `/tribunal-full`**
311
+ **Active reviewers: `logic-reviewer` · `security-auditor`**
312
+
313
+ ### ❌ Forbidden AI Tropes
314
+
315
+ 1. **Blind Assumptions:** Never make an assumption without documenting it clearly with `// VERIFY: [reason]`.
316
+ 2. **Silent Degradation:** Catching and suppressing errors without logging or handling.
317
+ 3. **Context Amnesia:** Forgetting the user's constraints and offering generic advice instead of tailored solutions.
318
+
319
+ Review these questions before confirming output:
320
+
321
+ ```
322
+ ✅ Did I rely ONLY on real, verified tools and methods?
323
+ ✅ Is this solution appropriately scoped to the user's constraints?
324
+ ✅ Did I handle potential failure modes and edge cases?
325
+ ✅ Have I avoided generic boilerplate that doesn't add value?
326
+ ```
327
+
328
+ ### 🛑 Verification-Before-Completion (VBC) Protocol
329
+
330
+ **CRITICAL:** You must follow a strict "evidence-based closeout" state machine.
331
+
332
+ - ❌ **Forbidden:** Declaring a task complete because the output "looks correct."
333
+ - ✅ **Required:** You are explicitly forbidden from finalizing any task without providing **concrete evidence** (terminal output, passing tests, compile success, or equivalent proof) that your output works as intended.
334
+
335
+ ## Pre-Flight Checklist
336
+
337
+ - [ ] Have I reviewed the user's specific constraints and requests?
338
+ - [ ] Have I checked the environment for relevant existing implementations?
339
+
340
+ ## VBC Protocol (Verification-Before-Completion)
341
+
342
+ You MUST verify existing code signatures and variables before attempting to modify or call them. No hallucination is permitted.
345
343
 
346
344
  ---
347
345
 
@@ -371,6 +369,7 @@ AI coding assistants often fall into specific bad habits when dealing with this
371
369
  ### ✅ Pre-Flight Self-Audit
372
370
 
373
371
  Review these questions before confirming output:
372
+
374
373
  ```
375
374
  ✅ Did I rely ONLY on real, verified tools and methods?
376
375
  ✅ Is this solution appropriately scoped to the user's constraints?
@@ -381,5 +380,6 @@ Review these questions before confirming output:
381
380
  ### 🛑 Verification-Before-Completion (VBC) Protocol
382
381
 
383
382
  **CRITICAL:** You must follow a strict "evidence-based closeout" state machine.
383
+
384
384
  - ❌ **Forbidden:** Declaring a task complete because the output "looks correct."
385
385
  - ✅ **Required:** You are explicitly forbidden from finalizing any task without providing **concrete evidence** (terminal output, passing tests, compile success, or equivalent proof) that your output works as intended.