opencode-skills-collection 4.0.68 → 4.0.69

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (309) hide show
  1. package/bundled-skills/.antigravity-install-manifest.json +266 -1
  2. package/bundled-skills/access-review/SKILL.md +394 -0
  3. package/bundled-skills/access-review/references/details.md +121 -0
  4. package/bundled-skills/agent-evals/SKILL.md +420 -0
  5. package/bundled-skills/agent-observability/SKILL.md +346 -0
  6. package/bundled-skills/agent-observability/references/details.md +786 -0
  7. package/bundled-skills/ai-agent-security/SKILL.md +393 -0
  8. package/bundled-skills/ai-agent-security/references/details.md +912 -0
  9. package/bundled-skills/ai-coding-agent-guardrails/SKILL.md +442 -0
  10. package/bundled-skills/ai-coding-agent-guardrails/references/details.md +753 -0
  11. package/bundled-skills/ai-inference-service-mesh/SKILL.md +449 -0
  12. package/bundled-skills/ai-pipeline-orchestration/SKILL.md +287 -0
  13. package/bundled-skills/ai-red-teaming/SKILL.md +409 -0
  14. package/bundled-skills/ai-security-hardening/SKILL.md +343 -0
  15. package/bundled-skills/ai-sre-incident-response/SKILL.md +336 -0
  16. package/bundled-skills/alerting-oncall/SKILL.md +458 -0
  17. package/bundled-skills/alerting-oncall/references/details.md +84 -0
  18. package/bundled-skills/apk-redteam-pipeline/SKILL.md +446 -0
  19. package/bundled-skills/argocd-gitops/SKILL.md +469 -0
  20. package/bundled-skills/arm-templates/SKILL.md +438 -0
  21. package/bundled-skills/arm-templates/references/details.md +64 -0
  22. package/bundled-skills/asset-inventory/SKILL.md +412 -0
  23. package/bundled-skills/asset-inventory/references/details.md +127 -0
  24. package/bundled-skills/audit-logging/SKILL.md +476 -0
  25. package/bundled-skills/aws-cloudtrail/SKILL.md +486 -0
  26. package/bundled-skills/aws-cost-optimization/SKILL.md +331 -0
  27. package/bundled-skills/aws-ec2/SKILL.md +426 -0
  28. package/bundled-skills/aws-ecs-fargate/SKILL.md +388 -0
  29. package/bundled-skills/aws-iam/SKILL.md +463 -0
  30. package/bundled-skills/aws-lambda/SKILL.md +428 -0
  31. package/bundled-skills/aws-rds/SKILL.md +380 -0
  32. package/bundled-skills/aws-s3/SKILL.md +434 -0
  33. package/bundled-skills/aws-secrets-manager/SKILL.md +486 -0
  34. package/bundled-skills/aws-vpc/SKILL.md +436 -0
  35. package/bundled-skills/azure-ai-document-intelligence-ts/SKILL.md +1 -1
  36. package/bundled-skills/azure-aks/SKILL.md +423 -0
  37. package/bundled-skills/azure-devops/SKILL.md +457 -0
  38. package/bundled-skills/azure-functions-devsec/SKILL.md +436 -0
  39. package/bundled-skills/azure-keyvault/SKILL.md +455 -0
  40. package/bundled-skills/azure-keyvault/references/details.md +83 -0
  41. package/bundled-skills/azure-monitor-audit/SKILL.md +379 -0
  42. package/bundled-skills/azure-networking/SKILL.md +448 -0
  43. package/bundled-skills/azure-networking/references/details.md +135 -0
  44. package/bundled-skills/azure-sql/SKILL.md +413 -0
  45. package/bundled-skills/azure-sql/references/details.md +113 -0
  46. package/bundled-skills/azure-vms/SKILL.md +402 -0
  47. package/bundled-skills/azure-vms/references/details.md +134 -0
  48. package/bundled-skills/backup-recovery/SKILL.md +388 -0
  49. package/bundled-skills/bb-methodology/SKILL.md +451 -0
  50. package/bundled-skills/bb-methodology/references/details.md +120 -0
  51. package/bundled-skills/block-storage/SKILL.md +371 -0
  52. package/bundled-skills/blue-green-deploy/SKILL.md +453 -0
  53. package/bundled-skills/blue-green-deploy/references/details.md +90 -0
  54. package/bundled-skills/bug-bounty/SKILL.md +447 -0
  55. package/bundled-skills/bug-bounty/references/details.md +1316 -0
  56. package/bundled-skills/bugcrowd-reporting/SKILL.md +351 -0
  57. package/bundled-skills/business-continuity/SKILL.md +463 -0
  58. package/bundled-skills/career-ops/SKILL.md +186 -0
  59. package/bundled-skills/cdn-setup/SKILL.md +374 -0
  60. package/bundled-skills/change-management/SKILL.md +438 -0
  61. package/bundled-skills/change-management/references/details.md +105 -0
  62. package/bundled-skills/circleci/SKILL.md +475 -0
  63. package/bundled-skills/cis-benchmarks/SKILL.md +150 -0
  64. package/bundled-skills/cloudflare-pages/SKILL.md +318 -0
  65. package/bundled-skills/cloudflare-r2/SKILL.md +353 -0
  66. package/bundled-skills/cloudflare-workers/SKILL.md +415 -0
  67. package/bundled-skills/cloudflare-zero-trust/SKILL.md +361 -0
  68. package/bundled-skills/cloudformation/SKILL.md +461 -0
  69. package/bundled-skills/constraint-driven-development/SKILL.md +335 -0
  70. package/bundled-skills/constraint-driven-development/references/floor-guard.md +99 -0
  71. package/bundled-skills/container-hardening/SKILL.md +126 -0
  72. package/bundled-skills/container-registries/SKILL.md +435 -0
  73. package/bundled-skills/container-scanning/SKILL.md +416 -0
  74. package/bundled-skills/convex-backend/SKILL.md +338 -0
  75. package/bundled-skills/dast-scanning/SKILL.md +437 -0
  76. package/bundled-skills/database-backups/SKILL.md +425 -0
  77. package/bundled-skills/datadog/SKILL.md +487 -0
  78. package/bundled-skills/dependency-scanning/SKILL.md +457 -0
  79. package/bundled-skills/devcontainers-nix/SKILL.md +416 -0
  80. package/bundled-skills/disaster-recovery/SKILL.md +374 -0
  81. package/bundled-skills/disaster-recovery/references/details.md +219 -0
  82. package/bundled-skills/dns-management/SKILL.md +375 -0
  83. package/bundled-skills/docker-compose/SKILL.md +482 -0
  84. package/bundled-skills/docker-management/SKILL.md +426 -0
  85. package/bundled-skills/ebpf-observability/SKILL.md +436 -0
  86. package/bundled-skills/ebpf-observability/references/details.md +542 -0
  87. package/bundled-skills/elk-stack/SKILL.md +487 -0
  88. package/bundled-skills/enterprise-vpn-attack/SKILL.md +395 -0
  89. package/bundled-skills/evidence-hygiene/SKILL.md +404 -0
  90. package/bundled-skills/feature-flags/SKILL.md +426 -0
  91. package/bundled-skills/feature-flags/references/details.md +86 -0
  92. package/bundled-skills/fedramp-compliance/SKILL.md +453 -0
  93. package/bundled-skills/firebase-app-platform/SKILL.md +381 -0
  94. package/bundled-skills/firewall-config/SKILL.md +479 -0
  95. package/bundled-skills/gcp-audit-logs/SKILL.md +452 -0
  96. package/bundled-skills/gcp-audit-logs/references/details.md +56 -0
  97. package/bundled-skills/gcp-cloud-functions/SKILL.md +284 -0
  98. package/bundled-skills/gcp-cloud-sql/SKILL.md +277 -0
  99. package/bundled-skills/gcp-compute/SKILL.md +319 -0
  100. package/bundled-skills/gcp-gke/SKILL.md +307 -0
  101. package/bundled-skills/gcp-networking/SKILL.md +293 -0
  102. package/bundled-skills/gcp-secret-manager/SKILL.md +421 -0
  103. package/bundled-skills/gcp-secret-manager/references/details.md +131 -0
  104. package/bundled-skills/gdpr-compliance/SKILL.md +451 -0
  105. package/bundled-skills/gdpr-compliance/references/details.md +145 -0
  106. package/bundled-skills/geo-audit/SKILL.md +368 -0
  107. package/bundled-skills/geo-brand-mentions/SKILL.md +68 -0
  108. package/bundled-skills/geo-brand-mentions/references/details.md +471 -0
  109. package/bundled-skills/geo-citability/SKILL.md +350 -0
  110. package/bundled-skills/geo-compare/SKILL.md +340 -0
  111. package/bundled-skills/geo-content/SKILL.md +383 -0
  112. package/bundled-skills/geo-crawlers/SKILL.md +408 -0
  113. package/bundled-skills/geo-llmstxt/SKILL.md +464 -0
  114. package/bundled-skills/geo-platform-optimizer/SKILL.md +314 -0
  115. package/bundled-skills/geo-proposal/SKILL.md +378 -0
  116. package/bundled-skills/geo-prospect/SKILL.md +225 -0
  117. package/bundled-skills/geo-report/SKILL.md +436 -0
  118. package/bundled-skills/geo-report-pdf/SKILL.md +157 -0
  119. package/bundled-skills/geo-schema/SKILL.md +408 -0
  120. package/bundled-skills/geo-technical/SKILL.md +78 -0
  121. package/bundled-skills/geo-technical/references/details.md +543 -0
  122. package/bundled-skills/git-workflow/SKILL.md +460 -0
  123. package/bundled-skills/github-actions/SKILL.md +368 -0
  124. package/bundled-skills/gitlab-ci/SKILL.md +340 -0
  125. package/bundled-skills/gpu-kubernetes-operations/SKILL.md +468 -0
  126. package/bundled-skills/gpu-server-management/SKILL.md +236 -0
  127. package/bundled-skills/hashicorp-vault/SKILL.md +408 -0
  128. package/bundled-skills/helm-charts/SKILL.md +469 -0
  129. package/bundled-skills/hipaa-compliance/SKILL.md +451 -0
  130. package/bundled-skills/hunt-aspnet/SKILL.md +321 -0
  131. package/bundled-skills/hunt-ato/SKILL.md +184 -0
  132. package/bundled-skills/hunt-auth-bypass/SKILL.md +426 -0
  133. package/bundled-skills/hunt-auth-bypass/references/details.md +80 -0
  134. package/bundled-skills/hunt-brute-force/SKILL.md +341 -0
  135. package/bundled-skills/hunt-business-logic/SKILL.md +281 -0
  136. package/bundled-skills/hunt-cache-poison/SKILL.md +382 -0
  137. package/bundled-skills/hunt-captcha-bypass/SKILL.md +136 -0
  138. package/bundled-skills/hunt-cicd/SKILL.md +311 -0
  139. package/bundled-skills/hunt-clickjacking/SKILL.md +110 -0
  140. package/bundled-skills/hunt-cors/SKILL.md +335 -0
  141. package/bundled-skills/hunt-dom/SKILL.md +323 -0
  142. package/bundled-skills/hunt-exceptional-conditions/SKILL.md +111 -0
  143. package/bundled-skills/hunt-file-upload/SKILL.md +202 -0
  144. package/bundled-skills/hunt-fintech-graphql/SKILL.md +289 -0
  145. package/bundled-skills/hunt-forgot-password/SKILL.md +114 -0
  146. package/bundled-skills/hunt-grpc/SKILL.md +317 -0
  147. package/bundled-skills/hunt-host-header/SKILL.md +309 -0
  148. package/bundled-skills/hunt-html-injection/SKILL.md +106 -0
  149. package/bundled-skills/hunt-http-smuggling/SKILL.md +129 -0
  150. package/bundled-skills/hunt-http-smuggling/references/phase2h-smuggling-cachepoison.md +177 -0
  151. package/bundled-skills/hunt-idor/SKILL.md +434 -0
  152. package/bundled-skills/hunt-jwt-crypto/SKILL.md +221 -0
  153. package/bundled-skills/hunt-k8s/SKILL.md +337 -0
  154. package/bundled-skills/hunt-laravel/SKILL.md +255 -0
  155. package/bundled-skills/hunt-ldap/SKILL.md +351 -0
  156. package/bundled-skills/hunt-lfi/SKILL.md +311 -0
  157. package/bundled-skills/hunt-llm-ai/SKILL.md +289 -0
  158. package/bundled-skills/hunt-mfa-bypass/SKILL.md +177 -0
  159. package/bundled-skills/hunt-misc/SKILL.md +378 -0
  160. package/bundled-skills/hunt-nextjs/SKILL.md +299 -0
  161. package/bundled-skills/hunt-nodejs/SKILL.md +263 -0
  162. package/bundled-skills/hunt-nosqli/SKILL.md +210 -0
  163. package/bundled-skills/hunt-ntlm-info/SKILL.md +314 -0
  164. package/bundled-skills/hunt-oauth/SKILL.md +459 -0
  165. package/bundled-skills/hunt-open-redirect/SKILL.md +223 -0
  166. package/bundled-skills/hunt-race-condition/SKILL.md +381 -0
  167. package/bundled-skills/hunt-race-condition/references/details.md +159 -0
  168. package/bundled-skills/hunt-rag-vector/SKILL.md +212 -0
  169. package/bundled-skills/hunt-rce/SKILL.md +444 -0
  170. package/bundled-skills/hunt-rce/references/details.md +110 -0
  171. package/bundled-skills/hunt-saml/SKILL.md +156 -0
  172. package/bundled-skills/hunt-session/SKILL.md +342 -0
  173. package/bundled-skills/hunt-shadow-api/SKILL.md +198 -0
  174. package/bundled-skills/hunt-source-leak/SKILL.md +345 -0
  175. package/bundled-skills/hunt-spa-api/SKILL.md +163 -0
  176. package/bundled-skills/hunt-springboot/SKILL.md +285 -0
  177. package/bundled-skills/hunt-sqli/SKILL.md +466 -0
  178. package/bundled-skills/hunt-ssrf/SKILL.md +396 -0
  179. package/bundled-skills/hunt-ssrf/references/details.md +179 -0
  180. package/bundled-skills/hunt-ssti/SKILL.md +163 -0
  181. package/bundled-skills/hunt-subdomain/SKILL.md +379 -0
  182. package/bundled-skills/hunt-tls-network/SKILL.md +399 -0
  183. package/bundled-skills/hunt-xxe/SKILL.md +466 -0
  184. package/bundled-skills/i-have-adhd/SKILL.md +170 -0
  185. package/bundled-skills/identity-access-management/SKILL.md +382 -0
  186. package/bundled-skills/identity-access-management/references/details.md +524 -0
  187. package/bundled-skills/incident-management/SKILL.md +484 -0
  188. package/bundled-skills/incident-response/SKILL.md +448 -0
  189. package/bundled-skills/incident-response/references/details.md +113 -0
  190. package/bundled-skills/interview-me/SKILL.md +248 -0
  191. package/bundled-skills/iso27001-compliance/SKILL.md +460 -0
  192. package/bundled-skills/jenkins/SKILL.md +462 -0
  193. package/bundled-skills/jev-use/SKILL.md +158 -0
  194. package/bundled-skills/kubernetes-hardening/SKILL.md +154 -0
  195. package/bundled-skills/kubernetes-ops/SKILL.md +449 -0
  196. package/bundled-skills/kubernetes-ops/references/details.md +108 -0
  197. package/bundled-skills/kustomize/SKILL.md +478 -0
  198. package/bundled-skills/linux-administration/SKILL.md +367 -0
  199. package/bundled-skills/linux-hardening/SKILL.md +154 -0
  200. package/bundled-skills/llm-app-security/SKILL.md +389 -0
  201. package/bundled-skills/llm-app-security/references/details.md +674 -0
  202. package/bundled-skills/llm-caching/SKILL.md +334 -0
  203. package/bundled-skills/llm-cost-optimization/SKILL.md +311 -0
  204. package/bundled-skills/llm-fine-tuning/SKILL.md +329 -0
  205. package/bundled-skills/llm-gateway/SKILL.md +282 -0
  206. package/bundled-skills/llm-inference-scaling/SKILL.md +286 -0
  207. package/bundled-skills/llmops-platform-engineering/SKILL.md +472 -0
  208. package/bundled-skills/load-balancing/SKILL.md +403 -0
  209. package/bundled-skills/loki-logging/SKILL.md +479 -0
  210. package/bundled-skills/m365-entra-attack/SKILL.md +423 -0
  211. package/bundled-skills/mac-mini-llm-lab/SKILL.md +350 -0
  212. package/bundled-skills/mcp-server-security/SKILL.md +356 -0
  213. package/bundled-skills/mcp-server-security/references/details.md +745 -0
  214. package/bundled-skills/mdm-device-management/SKILL.md +404 -0
  215. package/bundled-skills/mdm-device-management/references/details.md +410 -0
  216. package/bundled-skills/meme-coin-audit/SKILL.md +402 -0
  217. package/bundled-skills/mid-engagement-ir-detection/SKILL.md +377 -0
  218. package/bundled-skills/model-registry-governance/SKILL.md +452 -0
  219. package/bundled-skills/model-serving-kubernetes/SKILL.md +339 -0
  220. package/bundled-skills/model-supply-chain-security/SKILL.md +427 -0
  221. package/bundled-skills/mongodb/SKILL.md +436 -0
  222. package/bundled-skills/multi-tenant-llm-hosting/SKILL.md +435 -0
  223. package/bundled-skills/multi-tenant-llm-hosting/references/details.md +211 -0
  224. package/bundled-skills/mysql/SKILL.md +390 -0
  225. package/bundled-skills/new-relic/SKILL.md +472 -0
  226. package/bundled-skills/nfs-storage/SKILL.md +356 -0
  227. package/bundled-skills/object-storage/SKILL.md +378 -0
  228. package/bundled-skills/offensive-osint/SKILL.md +443 -0
  229. package/bundled-skills/okta-attack/SKILL.md +436 -0
  230. package/bundled-skills/ollama-stack/SKILL.md +379 -0
  231. package/bundled-skills/openclaw-deployment-hardening/SKILL.md +135 -0
  232. package/bundled-skills/openclaw-local-mac-mini/SKILL.md +426 -0
  233. package/bundled-skills/openclaw-local-mac-mini/references/details.md +221 -0
  234. package/bundled-skills/openclaw-security-hardening/SKILL.md +135 -0
  235. package/bundled-skills/openshift/SKILL.md +485 -0
  236. package/bundled-skills/opentelemetry/SKILL.md +438 -0
  237. package/bundled-skills/opentelemetry/references/details.md +78 -0
  238. package/bundled-skills/opentofu-migration/SKILL.md +349 -0
  239. package/bundled-skills/osint-methodology/SKILL.md +460 -0
  240. package/bundled-skills/osint-methodology/references/details.md +1350 -0
  241. package/bundled-skills/pci-dss-compliance/SKILL.md +446 -0
  242. package/bundled-skills/penetration-testing/SKILL.md +152 -0
  243. package/bundled-skills/performance-tuning/SKILL.md +381 -0
  244. package/bundled-skills/planetscale/SKILL.md +297 -0
  245. package/bundled-skills/platform-engineering/SKILL.md +348 -0
  246. package/bundled-skills/platform-engineering/references/details.md +944 -0
  247. package/bundled-skills/podman/SKILL.md +405 -0
  248. package/bundled-skills/policy-as-code/SKILL.md +434 -0
  249. package/bundled-skills/policy-as-code/references/details.md +204 -0
  250. package/bundled-skills/postgresql-devsec/SKILL.md +378 -0
  251. package/bundled-skills/prometheus-grafana/SKILL.md +469 -0
  252. package/bundled-skills/prompt-injection-defense/SKILL.md +483 -0
  253. package/bundled-skills/rag-infrastructure/SKILL.md +269 -0
  254. package/bundled-skills/rag-observability-evals/SKILL.md +444 -0
  255. package/bundled-skills/rag-observability-evals/references/details.md +92 -0
  256. package/bundled-skills/recon-scope-triage/SKILL.md +128 -0
  257. package/bundled-skills/redis/SKILL.md +421 -0
  258. package/bundled-skills/redteam-report-template/SKILL.md +370 -0
  259. package/bundled-skills/report-writing/SKILL.md +426 -0
  260. package/bundled-skills/report-writing/references/details.md +187 -0
  261. package/bundled-skills/reverse-proxy/SKILL.md +420 -0
  262. package/bundled-skills/runbook-creation/SKILL.md +438 -0
  263. package/bundled-skills/runbook-creation/references/details.md +71 -0
  264. package/bundled-skills/saas-security-posture/SKILL.md +415 -0
  265. package/bundled-skills/sast-scanning/SKILL.md +444 -0
  266. package/bundled-skills/sbom-supply-chain/SKILL.md +433 -0
  267. package/bundled-skills/security-arsenal/SKILL.md +446 -0
  268. package/bundled-skills/security-arsenal/references/details.md +540 -0
  269. package/bundled-skills/security-automation/SKILL.md +146 -0
  270. package/bundled-skills/semantic-versioning/SKILL.md +434 -0
  271. package/bundled-skills/semantic-versioning/references/details.md +83 -0
  272. package/bundled-skills/service-mesh/SKILL.md +422 -0
  273. package/bundled-skills/soc2-compliance/SKILL.md +409 -0
  274. package/bundled-skills/sops-encryption/SKILL.md +124 -0
  275. package/bundled-skills/sre-dashboards/SKILL.md +143 -0
  276. package/bundled-skills/ssh-configuration/SKILL.md +324 -0
  277. package/bundled-skills/ssl-tls-management/SKILL.md +428 -0
  278. package/bundled-skills/ssl-tls-management/references/details.md +99 -0
  279. package/bundled-skills/startup-it-troubleshooting/SKILL.md +415 -0
  280. package/bundled-skills/supply-chain-attack-recon/SKILL.md +453 -0
  281. package/bundled-skills/supply-chain-attack-recon/references/details.md +258 -0
  282. package/bundled-skills/systemd-services/SKILL.md +379 -0
  283. package/bundled-skills/terraform-aws/SKILL.md +125 -0
  284. package/bundled-skills/terraform-azure/SKILL.md +415 -0
  285. package/bundled-skills/terraform-azure/references/details.md +231 -0
  286. package/bundled-skills/terraform-gcp/SKILL.md +369 -0
  287. package/bundled-skills/threat-modeling/SKILL.md +487 -0
  288. package/bundled-skills/user-management/SKILL.md +383 -0
  289. package/bundled-skills/using-agent-skills/SKILL.md +220 -0
  290. package/bundled-skills/vector-database-ops/SKILL.md +300 -0
  291. package/bundled-skills/vendor-management/SKILL.md +439 -0
  292. package/bundled-skills/vendor-management/references/details.md +109 -0
  293. package/bundled-skills/vercel-deployments/SKILL.md +296 -0
  294. package/bundled-skills/vllm-server/SKILL.md +236 -0
  295. package/bundled-skills/vmware-vcenter-attack/SKILL.md +412 -0
  296. package/bundled-skills/vpn-setup/SKILL.md +452 -0
  297. package/bundled-skills/vulnerability-scanning/SKILL.md +448 -0
  298. package/bundled-skills/waf-setup/SKILL.md +354 -0
  299. package/bundled-skills/waf-setup/references/details.md +211 -0
  300. package/bundled-skills/web2-recon/SKILL.md +440 -0
  301. package/bundled-skills/web2-recon/references/details.md +319 -0
  302. package/bundled-skills/web3-audit/SKILL.md +445 -0
  303. package/bundled-skills/web3-audit/references/details.md +224 -0
  304. package/bundled-skills/windows-hardening/SKILL.md +454 -0
  305. package/bundled-skills/windows-hardening/references/details.md +204 -0
  306. package/bundled-skills/windows-server/SKILL.md +318 -0
  307. package/bundled-skills/zero-trust/SKILL.md +461 -0
  308. package/package.json +1 -1
  309. package/skills_index.json +6943 -323
@@ -0,0 +1,343 @@
1
+ ---
2
+ name: ai-security-hardening
3
+ description: Harden AI/LLM deployments against prompt injection, data exfiltration,
4
+ model theft, and supply chain attacks.
5
+ category: security
6
+ risk: safe
7
+ source: https://github.com/BagelHole/DevOps-Security-Agent-Skills
8
+ source_repo: BagelHole/DevOps-Security-Agent-Skills
9
+ source_type: community
10
+ date_added: '2026-09-20'
11
+ license: MIT
12
+ license_source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
13
+ compatibility: Requires the relevant security tooling (scanners, vault CLIs) and an
14
+ authorized scope for any active assessment. Docs-only; helper scripts and templates
15
+ not bundled.
16
+ metadata:
17
+ author: devops-skills
18
+ version: '1.0'
19
+ ---
20
+
21
+ # AI Security Hardening
22
+
23
+ Secure LLM and AI systems against prompt injection, jailbreaks, data leakage, and supply chain threats in production environments.
24
+
25
+ ## When to Use This Skill
26
+
27
+ Use this skill when:
28
+ - Deploying an LLM-powered application handling sensitive user data
29
+ - Protecting against prompt injection attacks in AI agents
30
+ - Implementing output filtering and content moderation
31
+ - Securing model weights and API endpoints from theft
32
+ - Achieving SOC2 or ISO 27001 compliance for AI systems
33
+
34
+ ## AI-Specific Threat Model
35
+
36
+ ```
37
+ Threat Risk Control
38
+ ─────────────────────────────────────────────────────────────────────
39
+ Prompt injection System prompt override Input sanitization, separate context
40
+ Data exfiltration PII in model outputs Output filtering, DLP scanning
41
+ Jailbreaking Policy bypass Content moderation, guardrails
42
+ Model theft Weight extraction via API Rate limiting, access controls
43
+ Training data poisoning Backdoored fine-tuned model Dataset validation, provenance
44
+ Supply chain attack Malicious model weights Signature verification, scanning
45
+ Insecure output XSS/SQLi from LLM response Output encoding, parameterized queries
46
+ ```
47
+
48
+ ## Prompt Injection Defense
49
+
50
+ ```python
51
+ import re
52
+ from typing import Optional
53
+
54
+ INJECTION_PATTERNS = [
55
+ r"ignore\s+(all\s+)?(previous|prior|above)\s+instructions",
56
+ r"you\s+are\s+now\s+",
57
+ r"new\s+instructions?:",
58
+ r"system\s+prompt",
59
+ r"forget\s+everything",
60
+ r"act\s+as\s+",
61
+ r"jailbreak",
62
+ r"dan\s+mode",
63
+ r"<\s*system\s*>",
64
+ r"\[INST\]",
65
+ ]
66
+
67
+ def detect_prompt_injection(user_input: str) -> tuple[bool, Optional[str]]:
68
+ """Return (is_suspicious, matched_pattern)."""
69
+ normalized = user_input.lower().strip()
70
+ for pattern in INJECTION_PATTERNS:
71
+ if re.search(pattern, normalized, re.IGNORECASE):
72
+ return True, pattern
73
+ return False, None
74
+
75
+ def sanitize_user_input(user_input: str, max_length: int = 4000) -> str:
76
+ """Sanitize input before passing to LLM."""
77
+ # Truncate
78
+ user_input = user_input[:max_length]
79
+
80
+ # Remove null bytes and control characters
81
+ user_input = re.sub(r'[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]', '', user_input)
82
+
83
+ # Check for injection
84
+ suspicious, pattern = detect_prompt_injection(user_input)
85
+ if suspicious:
86
+ raise ValueError(f"Potential prompt injection detected: {pattern}")
87
+
88
+ return user_input
89
+ ```
90
+
91
+ ## Guardrails with NeMo Guardrails
92
+
93
+ ```python
94
+ # guardrails.yaml
95
+ from nemoguardrails import RailsConfig, LLMRails
96
+
97
+ config = RailsConfig.from_path("./guardrails-config")
98
+ rails = LLMRails(config)
99
+
100
+ async def safe_llm_call(user_message: str) -> str:
101
+ response = await rails.generate_async(
102
+ messages=[{"role": "user", "content": user_message}]
103
+ )
104
+ return response["content"]
105
+ ```
106
+
107
+ ```yaml
108
+ # guardrails-config/config.yml
109
+ models:
110
+ - type: main
111
+ engine: openai
112
+ model: gpt-4o-mini
113
+
114
+ rails:
115
+ input:
116
+ flows:
117
+ - check jailbreak
118
+ - check sensitive data
119
+ output:
120
+ flows:
121
+ - check output for PII
122
+ - check output for harmful content
123
+ ```
124
+
125
+ ## Output Filtering & PII Scrubbing
126
+
127
+ ```python
128
+ import re
129
+ from presidio_analyzer import AnalyzerEngine
130
+ from presidio_anonymizer import AnonymizerEngine
131
+
132
+ analyzer = AnalyzerEngine()
133
+ anonymizer = AnonymizerEngine()
134
+
135
+ PII_ENTITIES = ["PERSON", "EMAIL_ADDRESS", "PHONE_NUMBER", "CREDIT_CARD",
136
+ "US_SSN", "IBAN_CODE", "IP_ADDRESS", "LOCATION"]
137
+
138
+ def scrub_pii_from_output(text: str) -> str:
139
+ """Remove PII from LLM output before returning to user."""
140
+ results = analyzer.analyze(text=text, entities=PII_ENTITIES, language="en")
141
+ if not results:
142
+ return text
143
+ anonymized = anonymizer.anonymize(text=text, analyzer_results=results)
144
+ return anonymized.text
145
+
146
+ def validate_output_safety(output: str) -> bool:
147
+ """Check output doesn't contain prompt injection artifacts."""
148
+ dangerous_patterns = [
149
+ r"<\s*script\s*>", # XSS
150
+ r"javascript:", # XSS
151
+ r";\s*(DROP|DELETE|INSERT)",# SQLi
152
+ r"\$\{.*\}", # template injection
153
+ r"`.*`", # command injection in some contexts
154
+ ]
155
+ for pattern in dangerous_patterns:
156
+ if re.search(pattern, output, re.IGNORECASE):
157
+ return False
158
+ return True
159
+ ```
160
+
161
+ ## API Security for LLM Endpoints
162
+
163
+ ```python
164
+ from fastapi import FastAPI, HTTPException, Depends, Request
165
+ from fastapi.security import HTTPBearer, HTTPAuthorizationCredentials
166
+ import jwt
167
+ import time
168
+ from collections import defaultdict
169
+
170
+ app = FastAPI()
171
+ security = HTTPBearer()
172
+
173
+ # Rate limiting (per API key)
174
+ request_counts = defaultdict(list)
175
+
176
+ def rate_limit(api_key: str, max_requests: int = 100, window_seconds: int = 60):
177
+ now = time.time()
178
+ requests = request_counts[api_key]
179
+ # Remove old requests outside window
180
+ request_counts[api_key] = [t for t in requests if now - t < window_seconds]
181
+ if len(request_counts[api_key]) >= max_requests:
182
+ raise HTTPException(status_code=429, detail="Rate limit exceeded")
183
+ request_counts[api_key].append(now)
184
+
185
+ async def verify_token(
186
+ credentials: HTTPAuthorizationCredentials = Depends(security)
187
+ ) -> dict:
188
+ try:
189
+ payload = jwt.decode(credentials.credentials, SECRET_KEY, algorithms=["HS256"])
190
+ rate_limit(payload["sub"])
191
+ return payload
192
+ except jwt.ExpiredSignatureError:
193
+ raise HTTPException(status_code=401, detail="Token expired")
194
+ except jwt.InvalidTokenError:
195
+ raise HTTPException(status_code=401, detail="Invalid token")
196
+
197
+ @app.post("/v1/chat/completions")
198
+ async def chat(request: Request, token: dict = Depends(verify_token)):
199
+ body = await request.json()
200
+
201
+ # Input validation
202
+ user_msg = body.get("messages", [{}])[-1].get("content", "")
203
+ try:
204
+ safe_input = sanitize_user_input(user_msg)
205
+ except ValueError as e:
206
+ raise HTTPException(status_code=400, detail=str(e))
207
+
208
+ # Call LLM and scrub output
209
+ response = await call_llm(safe_input, token["scope"])
210
+ response["choices"][0]["message"]["content"] = scrub_pii_from_output(
211
+ response["choices"][0]["message"]["content"]
212
+ )
213
+ return response
214
+ ```
215
+
216
+ ## Model Weight Security
217
+
218
+ ```bash
219
+ # Verify model weights with SHA-256 hash before loading
220
+ MODEL_DIR="./models/llama-3.1-8b"
221
+ EXPECTED_HASH="sha256:abc123..."
222
+
223
+ # Generate hash of downloaded model
224
+ actual_hash=$(find "$MODEL_DIR" -name "*.safetensors" | sort | xargs sha256sum | sha256sum)
225
+ echo "Model hash: $actual_hash"
226
+
227
+ # Compare (automate in CI/CD)
228
+ if [ "$actual_hash" != "$EXPECTED_HASH" ]; then
229
+ echo "ERROR: Model hash mismatch — possible tampering!"
230
+ exit 1
231
+ fi
232
+
233
+ # Scan model files for embedded malware (ModelScan)
234
+ pip install modelscan
235
+ modelscan scan -p "$MODEL_DIR"
236
+ ```
237
+
238
+ ## Network Isolation for AI Services
239
+
240
+ ```yaml
241
+ # Kubernetes NetworkPolicy — isolate LLM API
242
+ apiVersion: networking.k8s.io/v1
243
+ kind: NetworkPolicy
244
+ metadata:
245
+ name: llm-api-isolation
246
+ namespace: ai-services
247
+ spec:
248
+ podSelector:
249
+ matchLabels:
250
+ app: vllm
251
+ policyTypes:
252
+ - Ingress
253
+ - Egress
254
+ ingress:
255
+ - from:
256
+ - namespaceSelector:
257
+ matchLabels:
258
+ name: backend # only backend can call LLM
259
+ ports:
260
+ - protocol: TCP
261
+ port: 8000
262
+ egress:
263
+ - to:
264
+ - namespaceSelector:
265
+ matchLabels:
266
+ name: monitoring # metrics only
267
+ ports:
268
+ - protocol: TCP
269
+ port: 9090
270
+ # Block egress to internet — prevent data exfiltration
271
+ # (allow only internal cluster traffic)
272
+ ```
273
+
274
+ ## Audit Logging
275
+
276
+ ```python
277
+ import structlog
278
+ from datetime import datetime, timezone
279
+
280
+ audit_log = structlog.get_logger("ai.audit")
281
+
282
+ def log_llm_interaction(
283
+ user_id: str,
284
+ session_id: str,
285
+ model: str,
286
+ prompt_tokens: int,
287
+ completion_tokens: int,
288
+ was_filtered: bool,
289
+ injection_detected: bool,
290
+ ):
291
+ audit_log.info(
292
+ "llm_interaction",
293
+ timestamp=datetime.now(timezone.utc).isoformat(),
294
+ user_id=user_id,
295
+ session_id=session_id,
296
+ model=model,
297
+ prompt_tokens=prompt_tokens,
298
+ completion_tokens=completion_tokens,
299
+ was_filtered=was_filtered,
300
+ injection_detected=injection_detected,
301
+ # DO NOT log prompt/completion content — PII risk
302
+ )
303
+ ```
304
+
305
+ ## Common Issues
306
+
307
+ | Issue | Cause | Fix |
308
+ |-------|-------|-----|
309
+ | False positive injection blocks | Overly broad regex | Tune patterns; use ML-based classifier for high-traffic |
310
+ | PII in model outputs | Model trained on PII data | Add Presidio scrubbing to output layer |
311
+ | API key leakage | Keys in logs or responses | Mask keys in logging; use vault for key storage |
312
+ | Model weight tampering | Unverified downloads | Always verify SHA-256; use `modelscan` |
313
+ | Rate limit bypass | Per-IP not per-user | Rate limit on authenticated user ID, not IP |
314
+
315
+ ## Best Practices
316
+
317
+ - Never log raw prompts or completions — they may contain PII or sensitive data.
318
+ - Treat LLM output as untrusted input — always encode before rendering in HTML.
319
+ - Use network policies to prevent LLM pods from making outbound internet calls.
320
+ - Rotate API keys quarterly; use short-lived JWT tokens for service-to-service auth.
321
+ - Run `modelscan` on any model downloaded from the internet before serving.
322
+
323
+ ## Related Skills
324
+
325
+ - hashicorp-vault (`hashicorp-vault`) - Secrets management for API keys
326
+ - [network-security] (../../network/) - Network-level controls
327
+ - linux-hardening (`linux-hardening`) - Host hardening
328
+ - agent-observability (`agent-observability`) - AI audit logging
329
+ - llm-gateway (`llm-gateway`) - Centralized access control
330
+
331
+ ## Limitations
332
+
333
+ - Apply guidance only within authorized scope; test destructive steps in non-production first.
334
+ - Docs-only import: upstream scripts and templates not bundled.
335
+
336
+ ### Example
337
+
338
+ ```bash
339
+ # Read-only first: inventory before any active step.
340
+ which <tool> && <tool> --help | head -n 20
341
+ ```
342
+
343
+ > Adapted from [BagelHole/DevOps-Security-Agent-Skills](https://github.com/BagelHole/DevOps-Security-Agent-Skills) (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.
@@ -0,0 +1,336 @@
1
+ ---
2
+ name: ai-sre-incident-response
3
+ description: Build AI-focused SRE incident response practices for LLM outages, degraded
4
+ quality, runaway cost events, and safety regressions.
5
+ category: devops
6
+ risk: critical
7
+ source: https://github.com/BagelHole/DevOps-Security-Agent-Skills
8
+ source_repo: BagelHole/DevOps-Security-Agent-Skills
9
+ source_type: community
10
+ date_added: '2026-09-20'
11
+ license: MIT
12
+ license_source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
13
+ compatibility: Requires the relevant platform CLIs (kubectl, helm, terraform, git,
14
+ CI runners) and authorized access to the target environment. Docs-only; helper scripts
15
+ and templates not bundled.
16
+ metadata:
17
+ author: devops-skills
18
+ version: '1.0'
19
+ ---
20
+
21
+ # AI SRE Incident Response
22
+
23
+ Apply SRE rigor to AI systems where incidents include quality regressions, unsafe outputs, and budget explosions.
24
+
25
+ ## When to Use This Skill
26
+
27
+ - An LLM endpoint begins returning degraded or hallucinated answers
28
+ - Token spend spikes beyond budget thresholds
29
+ - A model provider goes down and traffic must fail over
30
+ - Safety guardrails fire at abnormal rates
31
+ - A new model deployment causes latency or accuracy regression
32
+
33
+ ## Prerequisites
34
+
35
+ - Prometheus and Alertmanager deployed with scrape targets for AI services
36
+ - Grafana dashboards for golden signals (latency, error rate, cost, quality)
37
+ - On-call rotation configured in PagerDuty, Opsgenie, or equivalent
38
+ - Runbook repository accessible to responders
39
+ - Rollback mechanism for model and prompt versions (GitOps or feature flags)
40
+
41
+ ## AI Incident Classes
42
+
43
+ - **Availability incident**: model/provider unavailable, timeout storm.
44
+ - **Quality incident**: answer accuracy or tool success drops below SLO.
45
+ - **Safety incident**: harmful or policy-violating outputs increase.
46
+ - **Cost incident**: unexpected token or provider spend spike.
47
+
48
+ ## Severity Framework
49
+
50
+ | Severity | Criteria | Response Time | Notification |
51
+ |----------|----------|---------------|--------------|
52
+ | SEV1 | User-facing outage, compliance risk, data leak | 5 min | Page on-call + incident commander |
53
+ | SEV2 | Major degradation in key flows | 15 min | Page on-call |
54
+ | SEV3 | Limited impact or internal-only issue | 1 hour | Slack alert |
55
+ | SEV4 | Cosmetic or low-priority regression | Next business day | Ticket |
56
+
57
+ ## Golden Signals for AI Services
58
+
59
+ - Request success rate
60
+ - Latency (queue + generation + tool execution)
61
+ - Hallucination/groundedness proxy metrics
62
+ - Cost per minute and per tenant
63
+ - Guardrail violation rate
64
+
65
+ ## Prometheus Alert Rules
66
+
67
+ ```yaml
68
+ # prometheus-ai-alerts.yaml
69
+ groups:
70
+ - name: ai-service-alerts
71
+ rules:
72
+ - alert: ModelEndpointDown
73
+ expr: up{job="llm-inference"} == 0
74
+ for: 2m
75
+ labels:
76
+ severity: sev1
77
+ annotations:
78
+ summary: "LLM inference endpoint {{ $labels.instance }} is down"
79
+ runbook_url: "https://runbooks.internal/ai/model-outage"
80
+
81
+ - alert: HighHallucinationRate
82
+ expr: |
83
+ rate(llm_hallucination_detected_total[10m])
84
+ / rate(llm_requests_total[10m]) > 0.15
85
+ for: 5m
86
+ labels:
87
+ severity: sev2
88
+ annotations:
89
+ summary: "Hallucination rate above 15% for {{ $labels.model }}"
90
+ runbook_url: "https://runbooks.internal/ai/quality-regression"
91
+
92
+ - alert: TokenCostExplosion
93
+ expr: |
94
+ sum(rate(llm_token_cost_dollars[5m])) by (tenant)
95
+ > 0.50
96
+ for: 3m
97
+ labels:
98
+ severity: sev2
99
+ annotations:
100
+ summary: "Token spend exceeds $0.50/min for tenant {{ $labels.tenant }}"
101
+ runbook_url: "https://runbooks.internal/ai/cost-spike"
102
+
103
+ - alert: LatencyP95Exceeded
104
+ expr: |
105
+ histogram_quantile(0.95,
106
+ rate(llm_request_duration_seconds_bucket[5m])
107
+ ) > 5
108
+ for: 5m
109
+ labels:
110
+ severity: sev2
111
+ annotations:
112
+ summary: "LLM p95 latency exceeds 5s for {{ $labels.service }}"
113
+
114
+ - alert: GuardrailViolationSpike
115
+ expr: |
116
+ rate(llm_guardrail_violations_total[10m])
117
+ / rate(llm_requests_total[10m]) > 0.05
118
+ for: 5m
119
+ labels:
120
+ severity: sev1
121
+ annotations:
122
+ summary: "Guardrail violations above 5% for {{ $labels.model }}"
123
+ runbook_url: "https://runbooks.internal/ai/safety-incident"
124
+
125
+ - alert: ModelQualityDrop
126
+ expr: |
127
+ llm_eval_score{metric="groundedness"} < 0.70
128
+ for: 10m
129
+ labels:
130
+ severity: sev2
131
+ annotations:
132
+ summary: "Groundedness score dropped below 0.70 for {{ $labels.model }}"
133
+
134
+ - alert: ProviderErrorRateHigh
135
+ expr: |
136
+ rate(llm_provider_errors_total[5m])
137
+ / rate(llm_provider_requests_total[5m]) > 0.10
138
+ for: 3m
139
+ labels:
140
+ severity: sev2
141
+ annotations:
142
+ summary: "Provider {{ $labels.provider }} error rate above 10%"
143
+ ```
144
+
145
+ ## Response Playbooks
146
+
147
+ ### Model Outage Runbook
148
+
149
+ ```text
150
+ TRIGGER: ModelEndpointDown fires for > 2 minutes
151
+ RESPONDER: On-call AI platform engineer
152
+
153
+ 1. Acknowledge alert in PagerDuty.
154
+ 2. Check provider status page (e.g., status.openai.com).
155
+ 3. Verify network connectivity:
156
+ curl -s -o /dev/null -w "%{http_code}" https://api.provider.com/health
157
+ 4. If provider is down:
158
+ a. Enable fallback model route in gateway config.
159
+ b. kubectl set env deployment/llm-gateway FALLBACK_ENABLED=true
160
+ c. Verify fallback traffic is flowing via Grafana dashboard.
161
+ 5. If self-hosted model is down:
162
+ a. Check pod status: kubectl get pods -l app=llm-inference -n ai
163
+ b. Check GPU health: kubectl logs -l app=llm-inference --tail=50
164
+ c. Restart if OOM: kubectl rollout restart deployment/llm-inference -n ai
165
+ 6. Freeze all deployments:
166
+ kubectl annotate deployment --all deploy-freeze=true -n ai
167
+ 7. Communicate ETA in #incident-channel.
168
+ 8. When resolved, unfreeze and run smoke tests.
169
+ ```
170
+
171
+ ### Quality Regression Runbook (Hallucination Spike)
172
+
173
+ ```text
174
+ TRIGGER: HighHallucinationRate or ModelQualityDrop fires
175
+ RESPONDER: On-call AI engineer + ML lead
176
+
177
+ 1. Acknowledge alert. Open incident ticket.
178
+ 2. Identify scope:
179
+ - Which model version? Check deployment metadata.
180
+ - Which routes/tenants affected? Filter by labels in Grafana.
181
+ 3. Check recent changes:
182
+ - Model version promotion in last 24h?
183
+ - Prompt template changes in last 24h?
184
+ - Retrieval index rebuild in last 24h?
185
+ 4. If recent model change:
186
+ kubectl rollout undo deployment/llm-inference -n ai
187
+ 5. If recent prompt change:
188
+ git revert <commit> && git push # triggers GitOps redeploy
189
+ 6. Increase trace sampling to 100% for affected route:
190
+ kubectl set env deployment/llm-gateway TRACE_SAMPLE_RATE=1.0
191
+ 7. Run offline eval suite against current production:
192
+ python run_evals.py --target prod --suite quality --compare baseline
193
+ 8. Confirm metrics return to baseline before closing.
194
+ ```
195
+
196
+ ### Token Cost Explosion Runbook
197
+
198
+ ```text
199
+ TRIGGER: TokenCostExplosion fires
200
+ RESPONDER: On-call platform engineer
201
+
202
+ 1. Identify top consumers:
203
+ Query: topk(10, sum(rate(llm_token_cost_dollars[15m])) by (tenant, model, route))
204
+ 2. Check for runaway loops:
205
+ - Agent retry storms (exponential token growth per request)
206
+ - Missing max_tokens caps on new routes
207
+ - Cache bypass due to config change
208
+ 3. Apply immediate caps:
209
+ kubectl patch configmap llm-quotas -n ai --patch '
210
+ data:
211
+ max_tokens_per_request: "4096"
212
+ rpm_limit: "60"
213
+ '
214
+ 4. Enable semantic cache if disabled:
215
+ kubectl set env deployment/llm-gateway CACHE_ENABLED=true
216
+ 5. Route traffic to cheaper model tier:
217
+ kubectl set env deployment/llm-gateway DEFAULT_MODEL=gpt-4o-mini
218
+ 6. Notify affected tenants of temporary limits.
219
+ 7. Open postmortem with cost attribution analysis.
220
+ ```
221
+
222
+ ## Escalation Procedures
223
+
224
+ ```text
225
+ Level 1 (0-15 min): On-call AI platform engineer
226
+ Level 2 (15-30 min): AI platform team lead + affected product owner
227
+ Level 3 (30-60 min): Engineering director + security (if safety incident)
228
+ Level 4 (60+ min): VP Engineering + legal (if compliance/data incident)
229
+
230
+ Safety incidents always start at Level 2 minimum.
231
+ Provider-side incidents: open support ticket immediately at Level 1.
232
+ ```
233
+
234
+ ## Detection Queries (PromQL)
235
+
236
+ ```promql
237
+ # Request success rate by model
238
+ 1 - (
239
+ sum(rate(llm_requests_total{status="error"}[5m])) by (model)
240
+ / sum(rate(llm_requests_total[5m])) by (model)
241
+ )
242
+
243
+ # Cost per successful answer
244
+ sum(rate(llm_token_cost_dollars[5m])) by (route)
245
+ / sum(rate(llm_requests_total{status="success"}[5m])) by (route)
246
+
247
+ # Hallucination rate trend (1h window, 5m steps)
248
+ rate(llm_hallucination_detected_total[1h])
249
+ / rate(llm_requests_total[1h])
250
+
251
+ # Latency breakdown by stage
252
+ histogram_quantile(0.95, rate(llm_retrieval_duration_seconds_bucket[5m]))
253
+ histogram_quantile(0.95, rate(llm_generation_duration_seconds_bucket[5m]))
254
+ histogram_quantile(0.95, rate(llm_tool_execution_duration_seconds_bucket[5m]))
255
+
256
+ # Tenant cost leaderboard
257
+ topk(10, sum(rate(llm_token_cost_dollars[1h])) by (tenant))
258
+ ```
259
+
260
+ ## Postmortem Requirements
261
+
262
+ - Timeline with detector and responder timestamps
263
+ - Blast radius by tenant and feature
264
+ - Missed signals and alert tuning actions
265
+ - Concrete hardening tasks with owners and due dates
266
+ - Cost impact (dollars, tokens, affected requests)
267
+ - Customer communication log
268
+
269
+ ## Postmortem Template
270
+
271
+ ```markdown
272
+ ## Incident Summary
273
+ - **Severity**: SEVx
274
+ - **Duration**: start_time - end_time (Xh Ym)
275
+ - **Detection**: How was it detected? (alert / customer report / manual)
276
+ - **Impact**: X tenants, Y requests, $Z cost
277
+
278
+ ## Timeline
279
+ | Time (UTC) | Event |
280
+ |------------|-------|
281
+ | HH:MM | Alert fired |
282
+ | HH:MM | Responder acknowledged |
283
+ | HH:MM | Root cause identified |
284
+ | HH:MM | Mitigation applied |
285
+ | HH:MM | Incident resolved |
286
+
287
+ ## Root Cause
288
+ [Description]
289
+
290
+ ## Action Items
291
+ | Action | Owner | Due Date | Status |
292
+ |--------|-------|----------|--------|
293
+ | Tune alert threshold | @engineer | YYYY-MM-DD | Open |
294
+ | Add fallback route | @platform | YYYY-MM-DD | Open |
295
+ ```
296
+
297
+ ## Chaos Engineering for AI Systems
298
+
299
+ Regularly test incident readiness:
300
+
301
+ - **Provider failover drill**: block provider API at network level, verify fallback activates within SLO.
302
+ - **Model rollback drill**: deploy known-bad model version, verify automated quality gate catches it.
303
+ - **Cost cap drill**: simulate runaway token usage, verify quotas trigger before budget threshold.
304
+ - **Cache failure drill**: disable semantic cache, verify system degrades gracefully.
305
+
306
+ ## Troubleshooting
307
+
308
+ | Symptom | Check | Fix |
309
+ |---------|-------|-----|
310
+ | All requests timing out | Provider status page, DNS resolution | Enable fallback provider |
311
+ | Gradual quality decline | Recent model/prompt deployments | Roll back to last known good |
312
+ | Sudden cost spike | Per-tenant token usage dashboard | Apply emergency token caps |
313
+ | Guardrail violations spike | Model version, prompt injection logs | Enable stricter input filtering |
314
+ | Intermittent 503 errors | Pod restarts, GPU OOM events | Increase memory limits or reduce batch size |
315
+
316
+ ## Related Skills
317
+
318
+ - incident-response (`incident-response`) - Standard incident process and evidence
319
+ - alerting-oncall (`alerting-oncall`) - Paging and escalation policy
320
+ - llm-cost-optimization (`llm-cost-optimization`) - Spend controls and efficiency patterns
321
+ - agent-observability (`agent-observability`) - Instrument requests, traces, and costs
322
+ - rag-observability-evals (`rag-observability-evals`) - RAG quality monitoring
323
+
324
+ ## Limitations
325
+
326
+ - Guidance executes against real environments: confirm target, blast radius, and rollback plan before applying anything.
327
+ - Never deploy to production without explicit approval. Docs-only import: upstream scripts and templates not bundled.
328
+
329
+ ### Example
330
+
331
+ ```bash
332
+ git status && git diff --stat
333
+ kubectl diff -f manifest.yaml
334
+ ```
335
+
336
+ > Adapted from [BagelHole/DevOps-Security-Agent-Skills](https://github.com/BagelHole/DevOps-Security-Agent-Skills) (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.