tribunal-kit 4.5.1 → 4.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (211) hide show
  1. package/.agent/.shared/ui-ux-pro-max/README.md +4 -4
  2. package/.agent/ARCHITECTURE.md +279 -277
  3. package/.agent/GEMINI.md +127 -121
  4. package/.agent/agents/accessibility-reviewer.md +187 -187
  5. package/.agent/agents/ai-code-reviewer.md +199 -199
  6. package/.agent/agents/api-architect.md +71 -66
  7. package/.agent/agents/backend-specialist.md +219 -215
  8. package/.agent/agents/cloud-engineer.md +98 -0
  9. package/.agent/agents/code-archaeologist.md +168 -161
  10. package/.agent/agents/database-architect.md +184 -184
  11. package/.agent/agents/db-latency-auditor.md +213 -216
  12. package/.agent/agents/debugger.md +198 -191
  13. package/.agent/agents/dependency-reviewer.md +106 -103
  14. package/.agent/agents/devops-engineer.md +218 -218
  15. package/.agent/agents/documentation-writer.md +209 -201
  16. package/.agent/agents/explorer-agent.md +167 -160
  17. package/.agent/agents/frontend-reviewer.md +162 -160
  18. package/.agent/agents/frontend-specialist.md +257 -248
  19. package/.agent/agents/game-developer.md +48 -48
  20. package/.agent/agents/logic-reviewer.md +118 -116
  21. package/.agent/agents/mobile-developer.md +197 -200
  22. package/.agent/agents/mobile-reviewer.md +159 -162
  23. package/.agent/agents/orchestrator.md +187 -181
  24. package/.agent/agents/penetration-tester.md +160 -157
  25. package/.agent/agents/performance-optimizer.md +183 -183
  26. package/.agent/agents/performance-reviewer.md +178 -178
  27. package/.agent/agents/precedence-reviewer.md +251 -250
  28. package/.agent/agents/product-manager.md +149 -142
  29. package/.agent/agents/product-owner.md +81 -80
  30. package/.agent/agents/project-planner.md +152 -142
  31. package/.agent/agents/qa-automation-engineer.md +216 -225
  32. package/.agent/agents/resilience-reviewer.md +88 -88
  33. package/.agent/agents/schema-reviewer.md +67 -67
  34. package/.agent/agents/security-auditor.md +180 -174
  35. package/.agent/agents/seo-specialist.md +188 -193
  36. package/.agent/agents/sql-reviewer.md +159 -161
  37. package/.agent/agents/supervisor-agent.md +173 -184
  38. package/.agent/agents/swarm-worker-contracts.md +170 -166
  39. package/.agent/agents/swarm-worker-registry.md +92 -92
  40. package/.agent/agents/system-architect.md +85 -0
  41. package/.agent/agents/test-coverage-reviewer.md +158 -160
  42. package/.agent/agents/test-engineer.md +118 -118
  43. package/.agent/agents/throughput-optimizer.md +291 -299
  44. package/.agent/agents/type-safety-reviewer.md +182 -175
  45. package/.agent/agents/ui-ux-auditor.md +300 -292
  46. package/.agent/agents/vitals-reviewer.md +223 -223
  47. package/.agent/mcp_config.json +37 -40
  48. package/.agent/patterns/generator.md +11 -9
  49. package/.agent/patterns/inversion.md +14 -12
  50. package/.agent/patterns/pipeline.md +11 -9
  51. package/.agent/patterns/reviewer.md +15 -13
  52. package/.agent/patterns/tool-wrapper.md +11 -9
  53. package/.agent/routing_index.json +654 -0
  54. package/.agent/rules/GEMINI.md +358 -352
  55. package/.agent/scripts/compile_router.py +112 -0
  56. package/.agent/scripts/migrate_skills_frontmatter.py +64 -0
  57. package/.agent/scripts/strengthen_skills.js +1 -1
  58. package/.agent/skills/advanced-rag-pipelines/SKILL.md +56 -0
  59. package/.agent/skills/agent-organizer/SKILL.md +156 -150
  60. package/.agent/skills/agentic-patterns/SKILL.md +313 -315
  61. package/.agent/skills/ai-prompt-injection-defense/SKILL.md +190 -184
  62. package/.agent/skills/api-patterns/SKILL.md +253 -247
  63. package/.agent/skills/api-security-auditor/SKILL.md +195 -193
  64. package/.agent/skills/app-builder/SKILL.md +573 -572
  65. package/.agent/skills/app-builder/templates/SKILL.md +108 -115
  66. package/.agent/skills/app-builder/templates/astro-static/TEMPLATE.md +76 -76
  67. package/.agent/skills/app-builder/templates/chrome-extension/TEMPLATE.md +92 -92
  68. package/.agent/skills/app-builder/templates/cli-tool/TEMPLATE.md +88 -88
  69. package/.agent/skills/app-builder/templates/electron-desktop/TEMPLATE.md +88 -88
  70. package/.agent/skills/app-builder/templates/express-api/TEMPLATE.md +83 -83
  71. package/.agent/skills/app-builder/templates/flutter-app/TEMPLATE.md +90 -90
  72. package/.agent/skills/app-builder/templates/monorepo-turborepo/TEMPLATE.md +90 -90
  73. package/.agent/skills/app-builder/templates/nextjs-fullstack/TEMPLATE.md +126 -122
  74. package/.agent/skills/app-builder/templates/nextjs-saas/TEMPLATE.md +127 -122
  75. package/.agent/skills/app-builder/templates/nextjs-static/TEMPLATE.md +172 -169
  76. package/.agent/skills/app-builder/templates/nuxt-app/TEMPLATE.md +139 -134
  77. package/.agent/skills/app-builder/templates/python-fastapi/TEMPLATE.md +83 -83
  78. package/.agent/skills/app-builder/templates/react-native-app/TEMPLATE.md +122 -119
  79. package/.agent/skills/appflow-wireframe/SKILL.md +146 -145
  80. package/.agent/skills/architecture/SKILL.md +226 -219
  81. package/.agent/skills/authentication-best-practices/SKILL.md +197 -189
  82. package/.agent/skills/backend-security-expert/SKILL.md +16 -2
  83. package/.agent/skills/bash-linux/SKILL.md +179 -179
  84. package/.agent/skills/behavioral-modes/SKILL.md +239 -223
  85. package/.agent/skills/brainstorming/SKILL.md +498 -486
  86. package/.agent/skills/browser-native-ai/SKILL.md +57 -4
  87. package/.agent/skills/building-native-ui/SKILL.md +202 -202
  88. package/.agent/skills/cicd-pro/SKILL.md +442 -0
  89. package/.agent/skills/clean-code/SKILL.md +400 -381
  90. package/.agent/skills/cloud-architect/SKILL.md +439 -0
  91. package/.agent/skills/code-review-checklist/SKILL.md +203 -194
  92. package/.agent/skills/config-validator/SKILL.md +165 -165
  93. package/.agent/skills/containerization-pro/SKILL.md +452 -0
  94. package/.agent/skills/csharp-developer/SKILL.md +518 -518
  95. package/.agent/skills/data-validation-schemas/SKILL.md +333 -328
  96. package/.agent/skills/database-design/SKILL.md +247 -240
  97. package/.agent/skills/deployment-procedures/SKILL.md +172 -169
  98. package/.agent/skills/devops-engineer/SKILL.md +345 -345
  99. package/.agent/skills/devops-incident-responder/SKILL.md +143 -137
  100. package/.agent/skills/doc.md +209 -177
  101. package/.agent/skills/documentation-templates/SKILL.md +291 -279
  102. package/.agent/skills/edge-computing/SKILL.md +183 -181
  103. package/.agent/skills/error-resilience/SKILL.md +411 -428
  104. package/.agent/skills/extract-design-system/SKILL.md +160 -158
  105. package/.agent/skills/framer-motion-expert/SKILL.md +253 -244
  106. package/.agent/skills/frontend-design/SKILL.md +208 -201
  107. package/.agent/skills/frontend-security-expert/SKILL.md +16 -3
  108. package/.agent/skills/game-design-expert/SKILL.md +132 -129
  109. package/.agent/skills/game-engineering-expert/SKILL.md +148 -146
  110. package/.agent/skills/generative-ui-expert/SKILL.md +57 -1
  111. package/.agent/skills/geo-fundamentals/SKILL.md +148 -147
  112. package/.agent/skills/git-pro/SKILL.md +435 -0
  113. package/.agent/skills/github-operations/SKILL.md +335 -329
  114. package/.agent/skills/gsap-core/SKILL.md +319 -308
  115. package/.agent/skills/gsap-frameworks/SKILL.md +213 -207
  116. package/.agent/skills/gsap-performance/SKILL.md +139 -133
  117. package/.agent/skills/gsap-plugins/SKILL.md +486 -480
  118. package/.agent/skills/gsap-react/SKILL.md +202 -189
  119. package/.agent/skills/gsap-scrolltrigger/SKILL.md +357 -350
  120. package/.agent/skills/gsap-timeline/SKILL.md +165 -161
  121. package/.agent/skills/gsap-utils/SKILL.md +344 -338
  122. package/.agent/skills/harness-protocol/SKILL.md +48 -0
  123. package/.agent/skills/i18n-localization/SKILL.md +174 -163
  124. package/.agent/skills/intelligent-routing/SKILL.md +202 -246
  125. package/.agent/skills/knowledge-graph/SKILL.md +60 -52
  126. package/.agent/skills/lint-and-validate/SKILL.md +261 -261
  127. package/.agent/skills/llm-engineering/SKILL.md +400 -394
  128. package/.agent/skills/local-first/SKILL.md +178 -178
  129. package/.agent/skills/mcp-builder/SKILL.md +143 -142
  130. package/.agent/skills/mobile-design/SKILL.md +272 -263
  131. package/.agent/skills/monorepo-management/SKILL.md +335 -334
  132. package/.agent/skills/motion-engineering/SKILL.md +266 -234
  133. package/.agent/skills/nextjs-react-expert/SKILL.md +236 -234
  134. package/.agent/skills/nodejs-best-practices/SKILL.md +547 -548
  135. package/.agent/skills/observability/SKILL.md +343 -343
  136. package/.agent/skills/parallel-agents/SKILL.md +143 -146
  137. package/.agent/skills/performance-profiling/SKILL.md +259 -267
  138. package/.agent/skills/plan-writing/SKILL.md +150 -142
  139. package/.agent/skills/platform-engineer/SKILL.md +148 -147
  140. package/.agent/skills/playwright-best-practices/SKILL.md +188 -187
  141. package/.agent/skills/powershell-windows/SKILL.md +162 -162
  142. package/.agent/skills/project-idioms/SKILL.md +137 -137
  143. package/.agent/skills/python-patterns/SKILL.md +260 -259
  144. package/.agent/skills/python-pro/SKILL.md +324 -323
  145. package/.agent/skills/react-specialist/SKILL.md +305 -277
  146. package/.agent/skills/readme-builder/SKILL.md +310 -300
  147. package/.agent/skills/realtime-patterns/SKILL.md +323 -319
  148. package/.agent/skills/red-team-tactics/SKILL.md +231 -218
  149. package/.agent/skills/rust-pro/SKILL.md +671 -673
  150. package/.agent/skills/seo-fundamentals/SKILL.md +179 -179
  151. package/.agent/skills/server-management/SKILL.md +218 -214
  152. package/.agent/skills/shadcn-ui-expert/SKILL.md +231 -231
  153. package/.agent/skills/skill-creator/SKILL.md +87 -86
  154. package/.agent/skills/sql-pro/SKILL.md +629 -629
  155. package/.agent/skills/supabase-postgres-best-practices/SKILL.md +97 -97
  156. package/.agent/skills/swiftui-expert/SKILL.md +204 -201
  157. package/.agent/skills/system-design-pro/SKILL.md +345 -0
  158. package/.agent/skills/systematic-debugging/SKILL.md +153 -142
  159. package/.agent/skills/tailwind-patterns/SKILL.md +610 -566
  160. package/.agent/skills/tdd-workflow/SKILL.md +169 -161
  161. package/.agent/skills/test-result-analyzer/SKILL.md +313 -309
  162. package/.agent/skills/testing-patterns/SKILL.md +566 -579
  163. package/.agent/skills/trend-researcher/SKILL.md +243 -237
  164. package/.agent/skills/typescript-advanced/SKILL.md +336 -335
  165. package/.agent/skills/ui-ux-pro-max/SKILL.md +590 -562
  166. package/.agent/skills/ui-ux-researcher/SKILL.md +244 -244
  167. package/.agent/skills/vue-expert/SKILL.md +294 -275
  168. package/.agent/skills/vulnerability-scanner/SKILL.md +416 -404
  169. package/.agent/skills/web-accessibility-auditor/SKILL.md +219 -218
  170. package/.agent/skills/web-design-guidelines/SKILL.md +192 -186
  171. package/.agent/skills/webapp-testing/SKILL.md +167 -169
  172. package/.agent/skills/webgpu-performance/SKILL.md +56 -2
  173. package/.agent/skills/whimsy-injector/SKILL.md +346 -325
  174. package/.agent/skills/workflow-optimizer/SKILL.md +231 -229
  175. package/.agent/workflows/acf.md +141 -0
  176. package/.agent/workflows/api-tester.md +176 -151
  177. package/.agent/workflows/audit.md +150 -127
  178. package/.agent/workflows/brainstorm.md +134 -110
  179. package/.agent/workflows/changelog.md +140 -112
  180. package/.agent/workflows/create.md +168 -124
  181. package/.agent/workflows/debug.md +190 -165
  182. package/.agent/workflows/deploy.md +201 -180
  183. package/.agent/workflows/enhance.md +154 -128
  184. package/.agent/workflows/fix.md +136 -114
  185. package/.agent/workflows/generate.md +198 -183
  186. package/.agent/workflows/marathon.md +37 -11
  187. package/.agent/workflows/migrate.md +184 -160
  188. package/.agent/workflows/orchestrate.md +192 -168
  189. package/.agent/workflows/performance-benchmarker.md +135 -114
  190. package/.agent/workflows/plan.md +196 -173
  191. package/.agent/workflows/preview.md +103 -80
  192. package/.agent/workflows/refactor.md +192 -161
  193. package/.agent/workflows/review-ai.md +125 -101
  194. package/.agent/workflows/review.md +141 -116
  195. package/.agent/workflows/session.md +122 -94
  196. package/.agent/workflows/status.md +101 -79
  197. package/.agent/workflows/strengthen-skills.md +164 -138
  198. package/.agent/workflows/super-prompt.md +24 -0
  199. package/.agent/workflows/swarm.md +193 -179
  200. package/.agent/workflows/test.md +211 -189
  201. package/.agent/workflows/tribunal-backend.md +136 -105
  202. package/.agent/workflows/tribunal-database.md +122 -95
  203. package/.agent/workflows/tribunal-frontend.md +221 -96
  204. package/.agent/workflows/tribunal-full.md +129 -100
  205. package/.agent/workflows/tribunal-mobile.md +122 -95
  206. package/.agent/workflows/tribunal-performance.md +136 -110
  207. package/.agent/workflows/tribunal-speed.md +209 -183
  208. package/.agent/workflows/ui-ux-pro-max.md +145 -122
  209. package/README.md +107 -55
  210. package/mcp_config.json +1 -3
  211. package/package.json +94 -94
@@ -1,347 +1,345 @@
1
- ---
2
- name: observability
3
- description: Production observability mastery. Structured logging (Pino/Winston), OpenTelemetry tracing, metrics (Prometheus/Grafana), SLIs/SLOs/error budgets, distributed tracing, alerting design, health checks, and AI observability. Use when setting up monitoring, debugging production issues, or designing observable distributed systems.
4
- allowed-tools: Read, Write, Edit, Glob, Grep
5
- version: 2.0.0
6
- last-updated: 2026-04-01
7
- applies-to-model: gemini-2.5-pro, claude-3-7-sonnet
8
- ---
9
-
10
- # Observability — Production Monitoring Mastery
11
-
12
- ---
13
-
14
- ## The Three Pillars
15
-
16
- ```
17
- Logs → WHAT happened (structured events)
18
- Traces → WHERE it happened (request flow across services)
19
- Metrics → HOW MUCH is happening (counters, histograms, gauges)
20
-
21
- All three are needed. Logs alone are not observability.
22
- ```
23
-
24
- ---
25
-
26
- ## Structured Logging
27
-
28
- ```typescript
29
- import pino from "pino";
30
-
31
- // ✅ Structured JSON logging
32
- const logger = pino({
33
- level: process.env.LOG_LEVEL ?? "info",
34
- timestamp: pino.stdTimeFunctions.isoTime,
35
- ...(process.env.NODE_ENV === "development" && {
36
- transport: { target: "pino-pretty" },
37
- }),
38
- });
39
-
40
- // ✅ GOOD: Structured with context
41
- logger.info({ userId: user.id, action: "login", ip: req.ip }, "User logged in");
42
- logger.error({ err, orderId: order.id, paymentGateway: "stripe" }, "Payment failed");
43
- logger.warn({ queueDepth: 1500, threshold: 1000 }, "Queue depth exceeding threshold");
44
-
45
- // BAD: Unstructured string logging
46
- console.log("User " + user.id + " logged in from " + req.ip);
47
- console.log("Error: " + error.message);
48
-
49
- // ❌ HALLUCINATION TRAP: console.log is NOT production logging
50
- // - No severity levels (info/warn/error)
51
- // - No structured fields (can't search/filter)
52
- // - No timestamps in ISO format
53
- // - Can't be collected by log aggregators
54
- // ✅ Use Pino (Node.js) or structlog (Python) for production
55
- ```
56
-
57
- ### Log Levels
58
-
59
- ```
60
- fatal → App is crashing, immediate attention required
61
- error → Operation failed, needs investigation
62
- warn → Something unexpected, but app continues
63
- info → Business events (user login, order placed, deploy)
64
- debug → Technical details (query timing, cache hit/miss)
65
- trace → Verbose debugging (only in development)
66
-
67
- Rules:
68
- - Production default: info
69
- - Never log PII (names, emails, SSNs) at any level
70
- - Never log secrets (tokens, passwords, API keys)
71
- - Log request IDs for correlation
72
- - Log durations for performance tracking
73
- ```
74
-
75
- ### Request Context / Correlation
76
-
77
- ```typescript
78
- import { AsyncLocalStorage } from "node:async_hooks";
79
-
80
- const requestContext = new AsyncLocalStorage<{ requestId: string; userId?: string }>();
81
-
82
- // Middleware: set context per request
83
- app.use((req, res, next) => {
84
- const requestId = req.headers["x-request-id"]?.toString() ?? crypto.randomUUID();
85
- res.setHeader("x-request-id", requestId);
86
- requestContext.run({ requestId, userId: req.user?.id }, next);
87
- });
88
-
89
- // Child logger with context
90
- function getLogger() {
91
- const ctx = requestContext.getStore();
92
- return logger.child({
93
- requestId: ctx?.requestId,
94
- userId: ctx?.userId,
95
- });
96
- }
97
-
98
- // Every log from this request includes requestId and userId
99
- const log = getLogger();
100
- log.info("Processing order"); // { requestId: "abc-123", userId: "42", msg: "Processing order" }
101
- ```
102
-
103
- ---
104
-
105
- ## Distributed Tracing (OpenTelemetry)
106
-
107
- ```typescript
108
- import { NodeSDK } from "@opentelemetry/sdk-node";
109
- import { getNodeAutoInstrumentations } from "@opentelemetry/auto-instrumentations-node";
110
- import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http";
111
-
112
- // Initialize OpenTelemetry
113
- const sdk = new NodeSDK({
114
- traceExporter: new OTLPTraceExporter({
115
- url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT ?? "http://localhost:4318/v1/traces",
116
- }),
117
- instrumentations: [
118
- getNodeAutoInstrumentations({
119
- "@opentelemetry/instrumentation-http": { enabled: true },
120
- "@opentelemetry/instrumentation-express": { enabled: true },
121
- "@opentelemetry/instrumentation-pg": { enabled: true },
122
- "@opentelemetry/instrumentation-redis": { enabled: true },
123
- }),
124
- ],
125
- });
126
-
127
- sdk.start();
128
-
129
- // Manual span for custom business logic
130
- import { trace } from "@opentelemetry/api";
131
-
132
- const tracer = trace.getTracer("order-service");
133
-
134
- async function processOrder(order: Order) {
135
- return tracer.startActiveSpan("processOrder", async (span) => {
136
- try {
137
- span.setAttribute("order.id", order.id);
138
- span.setAttribute("order.total", order.total);
139
- span.setAttribute("order.items.count", order.items.length);
140
-
141
- const result = await executeOrder(order);
142
- span.setStatus({ code: SpanStatusCode.OK });
143
- return result;
144
- } catch (error) {
145
- span.setStatus({ code: SpanStatusCode.ERROR, message: error.message });
146
- span.recordException(error);
147
- throw error;
148
- } finally {
149
- span.end();
150
- }
151
- });
152
- }
153
- ```
154
-
155
- ---
156
-
157
- ## Metrics
158
-
159
- ```typescript
160
- import { metrics } from "@opentelemetry/api";
161
-
162
- const meter = metrics.getMeter("api-server");
163
-
164
- // Counter — things that only go up
165
- const requestCounter = meter.createCounter("http.requests.total", {
166
- description: "Total HTTP requests",
167
- });
168
-
169
- // Histogram — request durations
170
- const requestDuration = meter.createHistogram("http.request.duration_ms", {
171
- description: "HTTP request duration in milliseconds",
172
- unit: "ms",
173
- });
174
-
175
- // Gauge — current values
176
- const activeConnections = meter.createUpDownCounter("db.connections.active", {
177
- description: "Active database connections",
178
- });
179
-
180
- // Middleware to record metrics
181
- app.use((req, res, next) => {
182
- const start = performance.now();
183
- res.on("finish", () => {
184
- const duration = performance.now() - start;
185
- requestCounter.add(1, {
186
- method: req.method,
187
- path: req.route?.path ?? req.path,
188
- status: res.statusCode.toString(),
189
- });
190
- requestDuration.record(duration, {
191
- method: req.method,
192
- status: res.statusCode.toString(),
193
- });
194
- });
195
- next();
196
- });
197
- ```
198
-
199
- ### Key Metrics to Track
200
-
201
- ```
202
- RED method (for services):
203
- Rate → requests per second
204
- Errors → error rate (4xx, 5xx)
205
- Duration → latency percentiles (P50, P95, P99)
206
-
207
- USE method (for resources):
208
- Utilization → CPU %, memory %, disk %
209
- Saturation → queue depth, thread pool saturation
210
- Errors → disk failures, OOM kills
211
-
212
- Business metrics:
213
- - Sign-ups per hour
214
- - Orders processed per minute
215
- - Revenue per day
216
- - API calls per customer
217
- ```
218
-
219
- ---
220
-
221
- ## SLIs, SLOs & Error Budgets
222
-
223
- ```
224
- SLI (Service Level Indicator) → What you measure
225
- "99.2% of requests complete in <500ms"
226
-
227
- SLO (Service Level Objective) → Your target
228
- "99.9% of requests should complete in <500ms"
229
-
230
- SLA (Service Level Agreement) → Your contract (with penalties)
231
- "99.95% uptime or we refund 10%"
232
-
233
- Error Budget = 100% - SLO
234
- SLO: 99.9% → Error budget: 0.1% → 43 min downtime/month
235
- SLO: 99.5% → Error budget: 0.5% → 3.6 hours downtime/month
236
-
237
- Rules:
238
- - Burn error budget too fast → freeze deployments
239
- - Error budget remaining → ship features faster
240
- - Don't set SLOs you can't measure
241
- - SLOs should be slightly below actual performance
242
- ```
243
-
244
- ---
245
-
246
- ## Health Checks
247
-
248
- ```typescript
249
- // Liveness: Is the process running?
250
- app.get("/health/live", (req, res) => {
251
- res.status(200).json({ status: "ok" });
252
- });
253
-
254
- // Readiness: Can it accept traffic?
255
- app.get("/health/ready", async (req, res) => {
256
- try {
257
- await db.raw("SELECT 1"); // database check
258
- await redis.ping(); // cache check
259
- res.status(200).json({
260
- status: "ready",
261
- checks: { database: "ok", cache: "ok" },
262
- });
263
- } catch (error) {
264
- res.status(503).json({
265
- status: "not ready",
266
- checks: { database: error.message },
267
- });
268
- }
269
- });
270
-
271
- // ❌ HALLUCINATION TRAP: Liveness ≠ Readiness
272
- // Liveness fails → container restarts (only for unrecoverable states)
273
- // Readiness fails → stop sending traffic (temporary — DB down, etc.)
274
- // Making liveness check the DB → DB outage restarts all containers → cascade failure
275
- ```
276
-
277
- ---
278
-
279
- ## Alerting
280
-
281
- ```
282
- Alert design rules:
283
- 1. Alert on SYMPTOMS, not causes (high latency, not "CPU is 80%")
284
- 2. Every alert must have a runbook link
285
- 3. Every alert must be ACTIONABLE — if you can't do anything, it's a notification
286
- 4. Use severity levels:
287
- - Critical → page on-call (customer-facing outage)
288
- - Warning → Slack notification (degraded, not broken)
289
- - Info → dashboard only (awareness)
290
- 5. Avoid alert fatigue — fewer, meaningful alerts beat many noisy ones
291
- ```
292
-
293
- ---
294
-
295
-
296
- ---
297
-
298
-
299
-
300
- AI coding assistants often fall into specific bad habits when dealing with this domain. These are strictly forbidden:
301
-
302
- 1. **Over-engineering:** Proposing complex abstractions or distributed systems when a simpler approach suffices.
303
- 2. **Hallucinated Libraries/Methods:** Using non-existent methods or packages. Always `// VERIFY` or check `package.json` / `requirements.txt`.
304
- 3. **Skipping Edge Cases:** Writing the "happy path" and ignoring error handling, timeouts, or data validation.
305
- 4. **Context Amnesia:** Forgetting the user's constraints and offering generic advice instead of tailored solutions.
306
- 5. **Silent Degradation:** Catching and suppressing errors without logging or re-raising.
307
-
308
- ---
309
-
310
-
311
-
312
- **Slash command: `/review` or `/tribunal-full`**
313
- **Active reviewers: `logic-reviewer` · `security-auditor`**
314
-
315
- ### ❌ Forbidden AI Tropes
316
-
317
- 1. **Blind Assumptions:** Never make an assumption without documenting it clearly with `// VERIFY: [reason]`.
318
- 2. **Silent Degradation:** Catching and suppressing errors without logging or handling.
319
- 3. **Context Amnesia:** Forgetting the user's constraints and offering generic advice instead of tailored solutions.
320
-
321
-
322
-
323
- Review these questions before confirming output:
324
- ```
325
- ✅ Did I rely ONLY on real, verified tools and methods?
326
- ✅ Is this solution appropriately scoped to the user's constraints?
327
- ✅ Did I handle potential failure modes and edge cases?
328
- ✅ Have I avoided generic boilerplate that doesn't add value?
329
- ```
330
-
331
- ### 🛑 Verification-Before-Completion (VBC) Protocol
332
-
333
- **CRITICAL:** You must follow a strict "evidence-based closeout" state machine.
334
- - ❌ **Forbidden:** Declaring a task complete because the output "looks correct."
335
- - ✅ **Required:** You are explicitly forbidden from finalizing any task without providing **concrete evidence** (terminal output, passing tests, compile success, or equivalent proof) that your output works as intended.
336
-
337
-
338
- ## Pre-Flight Checklist
339
- - [ ] Have I reviewed the user's specific constraints and requests?
340
- - [ ] Have I checked the environment for relevant existing implementations?
341
-
342
- ## VBC Protocol (Verification-Before-Completion)
343
- You MUST verify existing code signatures and variables before attempting to modify or call them. No hallucination is permitted.
1
+ ---
2
+ name: observability
3
+ description: Production observability mastery. Structured logging (Pino/Winston), OpenTelemetry tracing, metrics (Prometheus/Grafana), SLIs/SLOs/error budgets, distributed tracing, alerting design, health checks, and AI observability. Use when setting up monitoring, debugging production issues, or designing observable distributed systems.
4
+ allowed-tools: Read, Write, Edit, Glob, Grep
5
+ version: 2.0.0
6
+ last-updated: 2026-04-01
7
+ applies-to-model: gemini-2.5-pro, claude-3-7-sonnet
8
+ routing:
9
+ domain: general
10
+ tier: basic
11
+ ---
12
+
13
+ # Observability — Production Monitoring Mastery
14
+
15
+ ---
16
+
17
+ ## The Three Pillars
18
+
19
+ ```
20
+ Logs → WHAT happened (structured events)
21
+ Traces → WHERE it happened (request flow across services)
22
+ Metrics → HOW MUCH is happening (counters, histograms, gauges)
23
+
24
+ All three are needed. Logs alone are not observability.
25
+ ```
26
+
27
+ ---
28
+
29
+ ## Structured Logging
30
+
31
+ ```typescript
32
+ import pino from "pino";
33
+
34
+ // ✅ Structured JSON logging
35
+ const logger = pino({
36
+ level: process.env.LOG_LEVEL ?? "info",
37
+ timestamp: pino.stdTimeFunctions.isoTime,
38
+ ...(process.env.NODE_ENV === "development" && {
39
+ transport: { target: "pino-pretty" },
40
+ }),
41
+ });
42
+
43
+ // GOOD: Structured with context
44
+ logger.info({ userId: user.id, action: "login", ip: req.ip }, "User logged in");
45
+ logger.error({ err, orderId: order.id, paymentGateway: "stripe" }, "Payment failed");
46
+ logger.warn({ queueDepth: 1500, threshold: 1000 }, "Queue depth exceeding threshold");
344
47
 
48
+ // ❌ BAD: Unstructured string logging
49
+ console.log("User " + user.id + " logged in from " + req.ip);
50
+ console.log("Error: " + error.message);
51
+
52
+ // ❌ HALLUCINATION TRAP: console.log is NOT production logging
53
+ // - No severity levels (info/warn/error)
54
+ // - No structured fields (can't search/filter)
55
+ // - No timestamps in ISO format
56
+ // - Can't be collected by log aggregators
57
+ // ✅ Use Pino (Node.js) or structlog (Python) for production
58
+ ```
59
+
60
+ ### Log Levels
61
+
62
+ ```
63
+ fatal → App is crashing, immediate attention required
64
+ error → Operation failed, needs investigation
65
+ warn → Something unexpected, but app continues
66
+ info → Business events (user login, order placed, deploy)
67
+ debug → Technical details (query timing, cache hit/miss)
68
+ trace → Verbose debugging (only in development)
69
+
70
+ Rules:
71
+ - Production default: info
72
+ - Never log PII (names, emails, SSNs) at any level
73
+ - Never log secrets (tokens, passwords, API keys)
74
+ - Log request IDs for correlation
75
+ - Log durations for performance tracking
76
+ ```
77
+
78
+ ### Request Context / Correlation
79
+
80
+ ```typescript
81
+ import { AsyncLocalStorage } from "node:async_hooks";
82
+
83
+ const requestContext = new AsyncLocalStorage<{ requestId: string; userId?: string }>();
84
+
85
+ // Middleware: set context per request
86
+ app.use((req, res, next) => {
87
+ const requestId = req.headers["x-request-id"]?.toString() ?? crypto.randomUUID();
88
+ res.setHeader("x-request-id", requestId);
89
+ requestContext.run({ requestId, userId: req.user?.id }, next);
90
+ });
91
+
92
+ // Child logger with context
93
+ function getLogger() {
94
+ const ctx = requestContext.getStore();
95
+ return logger.child({
96
+ requestId: ctx?.requestId,
97
+ userId: ctx?.userId,
98
+ });
99
+ }
100
+
101
+ // Every log from this request includes requestId and userId
102
+ const log = getLogger();
103
+ log.info("Processing order"); // { requestId: "abc-123", userId: "42", msg: "Processing order" }
104
+ ```
105
+
106
+ ---
107
+
108
+ ## Distributed Tracing (OpenTelemetry)
109
+
110
+ ```typescript
111
+ import { NodeSDK } from "@opentelemetry/sdk-node";
112
+ import { getNodeAutoInstrumentations } from "@opentelemetry/auto-instrumentations-node";
113
+ import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http";
114
+
115
+ // Initialize OpenTelemetry
116
+ const sdk = new NodeSDK({
117
+ traceExporter: new OTLPTraceExporter({
118
+ url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT ?? "http://localhost:4318/v1/traces",
119
+ }),
120
+ instrumentations: [
121
+ getNodeAutoInstrumentations({
122
+ "@opentelemetry/instrumentation-http": { enabled: true },
123
+ "@opentelemetry/instrumentation-express": { enabled: true },
124
+ "@opentelemetry/instrumentation-pg": { enabled: true },
125
+ "@opentelemetry/instrumentation-redis": { enabled: true },
126
+ }),
127
+ ],
128
+ });
129
+
130
+ sdk.start();
131
+
132
+ // Manual span for custom business logic
133
+ import { trace } from "@opentelemetry/api";
134
+
135
+ const tracer = trace.getTracer("order-service");
136
+
137
+ async function processOrder(order: Order) {
138
+ return tracer.startActiveSpan("processOrder", async (span) => {
139
+ try {
140
+ span.setAttribute("order.id", order.id);
141
+ span.setAttribute("order.total", order.total);
142
+ span.setAttribute("order.items.count", order.items.length);
143
+
144
+ const result = await executeOrder(order);
145
+ span.setStatus({ code: SpanStatusCode.OK });
146
+ return result;
147
+ } catch (error) {
148
+ span.setStatus({ code: SpanStatusCode.ERROR, message: error.message });
149
+ span.recordException(error);
150
+ throw error;
151
+ } finally {
152
+ span.end();
153
+ }
154
+ });
155
+ }
156
+ ```
157
+
158
+ ---
159
+
160
+ ## Metrics
161
+
162
+ ```typescript
163
+ import { metrics } from "@opentelemetry/api";
164
+
165
+ const meter = metrics.getMeter("api-server");
166
+
167
+ // Counter — things that only go up
168
+ const requestCounter = meter.createCounter("http.requests.total", {
169
+ description: "Total HTTP requests",
170
+ });
171
+
172
+ // Histogram — request durations
173
+ const requestDuration = meter.createHistogram("http.request.duration_ms", {
174
+ description: "HTTP request duration in milliseconds",
175
+ unit: "ms",
176
+ });
177
+
178
+ // Gauge — current values
179
+ const activeConnections = meter.createUpDownCounter("db.connections.active", {
180
+ description: "Active database connections",
181
+ });
182
+
183
+ // Middleware to record metrics
184
+ app.use((req, res, next) => {
185
+ const start = performance.now();
186
+ res.on("finish", () => {
187
+ const duration = performance.now() - start;
188
+ requestCounter.add(1, {
189
+ method: req.method,
190
+ path: req.route?.path ?? req.path,
191
+ status: res.statusCode.toString(),
192
+ });
193
+ requestDuration.record(duration, {
194
+ method: req.method,
195
+ status: res.statusCode.toString(),
196
+ });
197
+ });
198
+ next();
199
+ });
200
+ ```
201
+
202
+ ### Key Metrics to Track
203
+
204
+ ```
205
+ RED method (for services):
206
+ Rate → requests per second
207
+ Errors → error rate (4xx, 5xx)
208
+ Duration → latency percentiles (P50, P95, P99)
209
+
210
+ USE method (for resources):
211
+ Utilization → CPU %, memory %, disk %
212
+ Saturation → queue depth, thread pool saturation
213
+ Errors → disk failures, OOM kills
214
+
215
+ Business metrics:
216
+ - Sign-ups per hour
217
+ - Orders processed per minute
218
+ - Revenue per day
219
+ - API calls per customer
220
+ ```
221
+
222
+ ---
223
+
224
+ ## SLIs, SLOs & Error Budgets
225
+
226
+ ```
227
+ SLI (Service Level Indicator) → What you measure
228
+ "99.2% of requests complete in <500ms"
229
+
230
+ SLO (Service Level Objective) → Your target
231
+ "99.9% of requests should complete in <500ms"
232
+
233
+ SLA (Service Level Agreement) → Your contract (with penalties)
234
+ "99.95% uptime or we refund 10%"
235
+
236
+ Error Budget = 100% - SLO
237
+ SLO: 99.9% → Error budget: 0.1% → 43 min downtime/month
238
+ SLO: 99.5% → Error budget: 0.5% → 3.6 hours downtime/month
239
+
240
+ Rules:
241
+ - Burn error budget too fast → freeze deployments
242
+ - Error budget remaining → ship features faster
243
+ - Don't set SLOs you can't measure
244
+ - SLOs should be slightly below actual performance
245
+ ```
246
+
247
+ ---
248
+
249
+ ## Health Checks
250
+
251
+ ```typescript
252
+ // Liveness: Is the process running?
253
+ app.get("/health/live", (req, res) => {
254
+ res.status(200).json({ status: "ok" });
255
+ });
256
+
257
+ // Readiness: Can it accept traffic?
258
+ app.get("/health/ready", async (req, res) => {
259
+ try {
260
+ await db.raw("SELECT 1"); // database check
261
+ await redis.ping(); // cache check
262
+ res.status(200).json({
263
+ status: "ready",
264
+ checks: { database: "ok", cache: "ok" },
265
+ });
266
+ } catch (error) {
267
+ res.status(503).json({
268
+ status: "not ready",
269
+ checks: { database: error.message },
270
+ });
271
+ }
272
+ });
273
+
274
+ // ❌ HALLUCINATION TRAP: Liveness ≠ Readiness
275
+ // Liveness fails → container restarts (only for unrecoverable states)
276
+ // Readiness fails → stop sending traffic (temporary — DB down, etc.)
277
+ // Making liveness check the DB → DB outage restarts all containers → cascade failure
278
+ ```
279
+
280
+ ---
281
+
282
+ ## Alerting
283
+
284
+ ```
285
+ Alert design rules:
286
+ 1. Alert on SYMPTOMS, not causes (high latency, not "CPU is 80%")
287
+ 2. Every alert must have a runbook link
288
+ 3. Every alert must be ACTIONABLE — if you can't do anything, it's a notification
289
+ 4. Use severity levels:
290
+ - Critical → page on-call (customer-facing outage)
291
+ - Warning → Slack notification (degraded, not broken)
292
+ - Info → dashboard only (awareness)
293
+ 5. Avoid alert fatigue — fewer, meaningful alerts beat many noisy ones
294
+ ```
295
+
296
+ ---
297
+
298
+ ---
299
+
300
+ AI coding assistants often fall into specific bad habits when dealing with this domain. These are strictly forbidden:
301
+
302
+ 1. **Over-engineering:** Proposing complex abstractions or distributed systems when a simpler approach suffices.
303
+ 2. **Hallucinated Libraries/Methods:** Using non-existent methods or packages. Always `// VERIFY` or check `package.json` / `requirements.txt`.
304
+ 3. **Skipping Edge Cases:** Writing the "happy path" and ignoring error handling, timeouts, or data validation.
305
+ 4. **Context Amnesia:** Forgetting the user's constraints and offering generic advice instead of tailored solutions.
306
+ 5. **Silent Degradation:** Catching and suppressing errors without logging or re-raising.
307
+
308
+ ---
309
+
310
+ **Slash command: `/review` or `/tribunal-full`**
311
+ **Active reviewers: `logic-reviewer` · `security-auditor`**
312
+
313
+ ### ❌ Forbidden AI Tropes
314
+
315
+ 1. **Blind Assumptions:** Never make an assumption without documenting it clearly with `// VERIFY: [reason]`.
316
+ 2. **Silent Degradation:** Catching and suppressing errors without logging or handling.
317
+ 3. **Context Amnesia:** Forgetting the user's constraints and offering generic advice instead of tailored solutions.
318
+
319
+ Review these questions before confirming output:
320
+
321
+ ```
322
+ ✅ Did I rely ONLY on real, verified tools and methods?
323
+ ✅ Is this solution appropriately scoped to the user's constraints?
324
+ ✅ Did I handle potential failure modes and edge cases?
325
+ ✅ Have I avoided generic boilerplate that doesn't add value?
326
+ ```
327
+
328
+ ### 🛑 Verification-Before-Completion (VBC) Protocol
329
+
330
+ **CRITICAL:** You must follow a strict "evidence-based closeout" state machine.
331
+
332
+ - ❌ **Forbidden:** Declaring a task complete because the output "looks correct."
333
+ - ✅ **Required:** You are explicitly forbidden from finalizing any task without providing **concrete evidence** (terminal output, passing tests, compile success, or equivalent proof) that your output works as intended.
334
+
335
+ ## Pre-Flight Checklist
336
+
337
+ - [ ] Have I reviewed the user's specific constraints and requests?
338
+ - [ ] Have I checked the environment for relevant existing implementations?
339
+
340
+ ## VBC Protocol (Verification-Before-Completion)
341
+
342
+ You MUST verify existing code signatures and variables before attempting to modify or call them. No hallucination is permitted.
345
343
 
346
344
  ---
347
345
 
@@ -371,6 +369,7 @@ AI coding assistants often fall into specific bad habits when dealing with this
371
369
  ### ✅ Pre-Flight Self-Audit
372
370
 
373
371
  Review these questions before confirming output:
372
+
374
373
  ```
375
374
  ✅ Did I rely ONLY on real, verified tools and methods?
376
375
  ✅ Is this solution appropriately scoped to the user's constraints?
@@ -381,5 +380,6 @@ Review these questions before confirming output:
381
380
  ### 🛑 Verification-Before-Completion (VBC) Protocol
382
381
 
383
382
  **CRITICAL:** You must follow a strict "evidence-based closeout" state machine.
383
+
384
384
  - ❌ **Forbidden:** Declaring a task complete because the output "looks correct."
385
385
  - ✅ **Required:** You are explicitly forbidden from finalizing any task without providing **concrete evidence** (terminal output, passing tests, compile success, or equivalent proof) that your output works as intended.