opencode-skills-collection 4.0.68 → 4.0.69
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bundled-skills/.antigravity-install-manifest.json +266 -1
- package/bundled-skills/access-review/SKILL.md +394 -0
- package/bundled-skills/access-review/references/details.md +121 -0
- package/bundled-skills/agent-evals/SKILL.md +420 -0
- package/bundled-skills/agent-observability/SKILL.md +346 -0
- package/bundled-skills/agent-observability/references/details.md +786 -0
- package/bundled-skills/ai-agent-security/SKILL.md +393 -0
- package/bundled-skills/ai-agent-security/references/details.md +912 -0
- package/bundled-skills/ai-coding-agent-guardrails/SKILL.md +442 -0
- package/bundled-skills/ai-coding-agent-guardrails/references/details.md +753 -0
- package/bundled-skills/ai-inference-service-mesh/SKILL.md +449 -0
- package/bundled-skills/ai-pipeline-orchestration/SKILL.md +287 -0
- package/bundled-skills/ai-red-teaming/SKILL.md +409 -0
- package/bundled-skills/ai-security-hardening/SKILL.md +343 -0
- package/bundled-skills/ai-sre-incident-response/SKILL.md +336 -0
- package/bundled-skills/alerting-oncall/SKILL.md +458 -0
- package/bundled-skills/alerting-oncall/references/details.md +84 -0
- package/bundled-skills/apk-redteam-pipeline/SKILL.md +446 -0
- package/bundled-skills/argocd-gitops/SKILL.md +469 -0
- package/bundled-skills/arm-templates/SKILL.md +438 -0
- package/bundled-skills/arm-templates/references/details.md +64 -0
- package/bundled-skills/asset-inventory/SKILL.md +412 -0
- package/bundled-skills/asset-inventory/references/details.md +127 -0
- package/bundled-skills/audit-logging/SKILL.md +476 -0
- package/bundled-skills/aws-cloudtrail/SKILL.md +486 -0
- package/bundled-skills/aws-cost-optimization/SKILL.md +331 -0
- package/bundled-skills/aws-ec2/SKILL.md +426 -0
- package/bundled-skills/aws-ecs-fargate/SKILL.md +388 -0
- package/bundled-skills/aws-iam/SKILL.md +463 -0
- package/bundled-skills/aws-lambda/SKILL.md +428 -0
- package/bundled-skills/aws-rds/SKILL.md +380 -0
- package/bundled-skills/aws-s3/SKILL.md +434 -0
- package/bundled-skills/aws-secrets-manager/SKILL.md +486 -0
- package/bundled-skills/aws-vpc/SKILL.md +436 -0
- package/bundled-skills/azure-ai-document-intelligence-ts/SKILL.md +1 -1
- package/bundled-skills/azure-aks/SKILL.md +423 -0
- package/bundled-skills/azure-devops/SKILL.md +457 -0
- package/bundled-skills/azure-functions-devsec/SKILL.md +436 -0
- package/bundled-skills/azure-keyvault/SKILL.md +455 -0
- package/bundled-skills/azure-keyvault/references/details.md +83 -0
- package/bundled-skills/azure-monitor-audit/SKILL.md +379 -0
- package/bundled-skills/azure-networking/SKILL.md +448 -0
- package/bundled-skills/azure-networking/references/details.md +135 -0
- package/bundled-skills/azure-sql/SKILL.md +413 -0
- package/bundled-skills/azure-sql/references/details.md +113 -0
- package/bundled-skills/azure-vms/SKILL.md +402 -0
- package/bundled-skills/azure-vms/references/details.md +134 -0
- package/bundled-skills/backup-recovery/SKILL.md +388 -0
- package/bundled-skills/bb-methodology/SKILL.md +451 -0
- package/bundled-skills/bb-methodology/references/details.md +120 -0
- package/bundled-skills/block-storage/SKILL.md +371 -0
- package/bundled-skills/blue-green-deploy/SKILL.md +453 -0
- package/bundled-skills/blue-green-deploy/references/details.md +90 -0
- package/bundled-skills/bug-bounty/SKILL.md +447 -0
- package/bundled-skills/bug-bounty/references/details.md +1316 -0
- package/bundled-skills/bugcrowd-reporting/SKILL.md +351 -0
- package/bundled-skills/business-continuity/SKILL.md +463 -0
- package/bundled-skills/career-ops/SKILL.md +186 -0
- package/bundled-skills/cdn-setup/SKILL.md +374 -0
- package/bundled-skills/change-management/SKILL.md +438 -0
- package/bundled-skills/change-management/references/details.md +105 -0
- package/bundled-skills/circleci/SKILL.md +475 -0
- package/bundled-skills/cis-benchmarks/SKILL.md +150 -0
- package/bundled-skills/cloudflare-pages/SKILL.md +318 -0
- package/bundled-skills/cloudflare-r2/SKILL.md +353 -0
- package/bundled-skills/cloudflare-workers/SKILL.md +415 -0
- package/bundled-skills/cloudflare-zero-trust/SKILL.md +361 -0
- package/bundled-skills/cloudformation/SKILL.md +461 -0
- package/bundled-skills/constraint-driven-development/SKILL.md +335 -0
- package/bundled-skills/constraint-driven-development/references/floor-guard.md +99 -0
- package/bundled-skills/container-hardening/SKILL.md +126 -0
- package/bundled-skills/container-registries/SKILL.md +435 -0
- package/bundled-skills/container-scanning/SKILL.md +416 -0
- package/bundled-skills/convex-backend/SKILL.md +338 -0
- package/bundled-skills/dast-scanning/SKILL.md +437 -0
- package/bundled-skills/database-backups/SKILL.md +425 -0
- package/bundled-skills/datadog/SKILL.md +487 -0
- package/bundled-skills/dependency-scanning/SKILL.md +457 -0
- package/bundled-skills/devcontainers-nix/SKILL.md +416 -0
- package/bundled-skills/disaster-recovery/SKILL.md +374 -0
- package/bundled-skills/disaster-recovery/references/details.md +219 -0
- package/bundled-skills/dns-management/SKILL.md +375 -0
- package/bundled-skills/docker-compose/SKILL.md +482 -0
- package/bundled-skills/docker-management/SKILL.md +426 -0
- package/bundled-skills/ebpf-observability/SKILL.md +436 -0
- package/bundled-skills/ebpf-observability/references/details.md +542 -0
- package/bundled-skills/elk-stack/SKILL.md +487 -0
- package/bundled-skills/enterprise-vpn-attack/SKILL.md +395 -0
- package/bundled-skills/evidence-hygiene/SKILL.md +404 -0
- package/bundled-skills/feature-flags/SKILL.md +426 -0
- package/bundled-skills/feature-flags/references/details.md +86 -0
- package/bundled-skills/fedramp-compliance/SKILL.md +453 -0
- package/bundled-skills/firebase-app-platform/SKILL.md +381 -0
- package/bundled-skills/firewall-config/SKILL.md +479 -0
- package/bundled-skills/gcp-audit-logs/SKILL.md +452 -0
- package/bundled-skills/gcp-audit-logs/references/details.md +56 -0
- package/bundled-skills/gcp-cloud-functions/SKILL.md +284 -0
- package/bundled-skills/gcp-cloud-sql/SKILL.md +277 -0
- package/bundled-skills/gcp-compute/SKILL.md +319 -0
- package/bundled-skills/gcp-gke/SKILL.md +307 -0
- package/bundled-skills/gcp-networking/SKILL.md +293 -0
- package/bundled-skills/gcp-secret-manager/SKILL.md +421 -0
- package/bundled-skills/gcp-secret-manager/references/details.md +131 -0
- package/bundled-skills/gdpr-compliance/SKILL.md +451 -0
- package/bundled-skills/gdpr-compliance/references/details.md +145 -0
- package/bundled-skills/geo-audit/SKILL.md +368 -0
- package/bundled-skills/geo-brand-mentions/SKILL.md +68 -0
- package/bundled-skills/geo-brand-mentions/references/details.md +471 -0
- package/bundled-skills/geo-citability/SKILL.md +350 -0
- package/bundled-skills/geo-compare/SKILL.md +340 -0
- package/bundled-skills/geo-content/SKILL.md +383 -0
- package/bundled-skills/geo-crawlers/SKILL.md +408 -0
- package/bundled-skills/geo-llmstxt/SKILL.md +464 -0
- package/bundled-skills/geo-platform-optimizer/SKILL.md +314 -0
- package/bundled-skills/geo-proposal/SKILL.md +378 -0
- package/bundled-skills/geo-prospect/SKILL.md +225 -0
- package/bundled-skills/geo-report/SKILL.md +436 -0
- package/bundled-skills/geo-report-pdf/SKILL.md +157 -0
- package/bundled-skills/geo-schema/SKILL.md +408 -0
- package/bundled-skills/geo-technical/SKILL.md +78 -0
- package/bundled-skills/geo-technical/references/details.md +543 -0
- package/bundled-skills/git-workflow/SKILL.md +460 -0
- package/bundled-skills/github-actions/SKILL.md +368 -0
- package/bundled-skills/gitlab-ci/SKILL.md +340 -0
- package/bundled-skills/gpu-kubernetes-operations/SKILL.md +468 -0
- package/bundled-skills/gpu-server-management/SKILL.md +236 -0
- package/bundled-skills/hashicorp-vault/SKILL.md +408 -0
- package/bundled-skills/helm-charts/SKILL.md +469 -0
- package/bundled-skills/hipaa-compliance/SKILL.md +451 -0
- package/bundled-skills/hunt-aspnet/SKILL.md +321 -0
- package/bundled-skills/hunt-ato/SKILL.md +184 -0
- package/bundled-skills/hunt-auth-bypass/SKILL.md +426 -0
- package/bundled-skills/hunt-auth-bypass/references/details.md +80 -0
- package/bundled-skills/hunt-brute-force/SKILL.md +341 -0
- package/bundled-skills/hunt-business-logic/SKILL.md +281 -0
- package/bundled-skills/hunt-cache-poison/SKILL.md +382 -0
- package/bundled-skills/hunt-captcha-bypass/SKILL.md +136 -0
- package/bundled-skills/hunt-cicd/SKILL.md +311 -0
- package/bundled-skills/hunt-clickjacking/SKILL.md +110 -0
- package/bundled-skills/hunt-cors/SKILL.md +335 -0
- package/bundled-skills/hunt-dom/SKILL.md +323 -0
- package/bundled-skills/hunt-exceptional-conditions/SKILL.md +111 -0
- package/bundled-skills/hunt-file-upload/SKILL.md +202 -0
- package/bundled-skills/hunt-fintech-graphql/SKILL.md +289 -0
- package/bundled-skills/hunt-forgot-password/SKILL.md +114 -0
- package/bundled-skills/hunt-grpc/SKILL.md +317 -0
- package/bundled-skills/hunt-host-header/SKILL.md +309 -0
- package/bundled-skills/hunt-html-injection/SKILL.md +106 -0
- package/bundled-skills/hunt-http-smuggling/SKILL.md +129 -0
- package/bundled-skills/hunt-http-smuggling/references/phase2h-smuggling-cachepoison.md +177 -0
- package/bundled-skills/hunt-idor/SKILL.md +434 -0
- package/bundled-skills/hunt-jwt-crypto/SKILL.md +221 -0
- package/bundled-skills/hunt-k8s/SKILL.md +337 -0
- package/bundled-skills/hunt-laravel/SKILL.md +255 -0
- package/bundled-skills/hunt-ldap/SKILL.md +351 -0
- package/bundled-skills/hunt-lfi/SKILL.md +311 -0
- package/bundled-skills/hunt-llm-ai/SKILL.md +289 -0
- package/bundled-skills/hunt-mfa-bypass/SKILL.md +177 -0
- package/bundled-skills/hunt-misc/SKILL.md +378 -0
- package/bundled-skills/hunt-nextjs/SKILL.md +299 -0
- package/bundled-skills/hunt-nodejs/SKILL.md +263 -0
- package/bundled-skills/hunt-nosqli/SKILL.md +210 -0
- package/bundled-skills/hunt-ntlm-info/SKILL.md +314 -0
- package/bundled-skills/hunt-oauth/SKILL.md +459 -0
- package/bundled-skills/hunt-open-redirect/SKILL.md +223 -0
- package/bundled-skills/hunt-race-condition/SKILL.md +381 -0
- package/bundled-skills/hunt-race-condition/references/details.md +159 -0
- package/bundled-skills/hunt-rag-vector/SKILL.md +212 -0
- package/bundled-skills/hunt-rce/SKILL.md +444 -0
- package/bundled-skills/hunt-rce/references/details.md +110 -0
- package/bundled-skills/hunt-saml/SKILL.md +156 -0
- package/bundled-skills/hunt-session/SKILL.md +342 -0
- package/bundled-skills/hunt-shadow-api/SKILL.md +198 -0
- package/bundled-skills/hunt-source-leak/SKILL.md +345 -0
- package/bundled-skills/hunt-spa-api/SKILL.md +163 -0
- package/bundled-skills/hunt-springboot/SKILL.md +285 -0
- package/bundled-skills/hunt-sqli/SKILL.md +466 -0
- package/bundled-skills/hunt-ssrf/SKILL.md +396 -0
- package/bundled-skills/hunt-ssrf/references/details.md +179 -0
- package/bundled-skills/hunt-ssti/SKILL.md +163 -0
- package/bundled-skills/hunt-subdomain/SKILL.md +379 -0
- package/bundled-skills/hunt-tls-network/SKILL.md +399 -0
- package/bundled-skills/hunt-xxe/SKILL.md +466 -0
- package/bundled-skills/i-have-adhd/SKILL.md +170 -0
- package/bundled-skills/identity-access-management/SKILL.md +382 -0
- package/bundled-skills/identity-access-management/references/details.md +524 -0
- package/bundled-skills/incident-management/SKILL.md +484 -0
- package/bundled-skills/incident-response/SKILL.md +448 -0
- package/bundled-skills/incident-response/references/details.md +113 -0
- package/bundled-skills/interview-me/SKILL.md +248 -0
- package/bundled-skills/iso27001-compliance/SKILL.md +460 -0
- package/bundled-skills/jenkins/SKILL.md +462 -0
- package/bundled-skills/jev-use/SKILL.md +158 -0
- package/bundled-skills/kubernetes-hardening/SKILL.md +154 -0
- package/bundled-skills/kubernetes-ops/SKILL.md +449 -0
- package/bundled-skills/kubernetes-ops/references/details.md +108 -0
- package/bundled-skills/kustomize/SKILL.md +478 -0
- package/bundled-skills/linux-administration/SKILL.md +367 -0
- package/bundled-skills/linux-hardening/SKILL.md +154 -0
- package/bundled-skills/llm-app-security/SKILL.md +389 -0
- package/bundled-skills/llm-app-security/references/details.md +674 -0
- package/bundled-skills/llm-caching/SKILL.md +334 -0
- package/bundled-skills/llm-cost-optimization/SKILL.md +311 -0
- package/bundled-skills/llm-fine-tuning/SKILL.md +329 -0
- package/bundled-skills/llm-gateway/SKILL.md +282 -0
- package/bundled-skills/llm-inference-scaling/SKILL.md +286 -0
- package/bundled-skills/llmops-platform-engineering/SKILL.md +472 -0
- package/bundled-skills/load-balancing/SKILL.md +403 -0
- package/bundled-skills/loki-logging/SKILL.md +479 -0
- package/bundled-skills/m365-entra-attack/SKILL.md +423 -0
- package/bundled-skills/mac-mini-llm-lab/SKILL.md +350 -0
- package/bundled-skills/mcp-server-security/SKILL.md +356 -0
- package/bundled-skills/mcp-server-security/references/details.md +745 -0
- package/bundled-skills/mdm-device-management/SKILL.md +404 -0
- package/bundled-skills/mdm-device-management/references/details.md +410 -0
- package/bundled-skills/meme-coin-audit/SKILL.md +402 -0
- package/bundled-skills/mid-engagement-ir-detection/SKILL.md +377 -0
- package/bundled-skills/model-registry-governance/SKILL.md +452 -0
- package/bundled-skills/model-serving-kubernetes/SKILL.md +339 -0
- package/bundled-skills/model-supply-chain-security/SKILL.md +427 -0
- package/bundled-skills/mongodb/SKILL.md +436 -0
- package/bundled-skills/multi-tenant-llm-hosting/SKILL.md +435 -0
- package/bundled-skills/multi-tenant-llm-hosting/references/details.md +211 -0
- package/bundled-skills/mysql/SKILL.md +390 -0
- package/bundled-skills/new-relic/SKILL.md +472 -0
- package/bundled-skills/nfs-storage/SKILL.md +356 -0
- package/bundled-skills/object-storage/SKILL.md +378 -0
- package/bundled-skills/offensive-osint/SKILL.md +443 -0
- package/bundled-skills/okta-attack/SKILL.md +436 -0
- package/bundled-skills/ollama-stack/SKILL.md +379 -0
- package/bundled-skills/openclaw-deployment-hardening/SKILL.md +135 -0
- package/bundled-skills/openclaw-local-mac-mini/SKILL.md +426 -0
- package/bundled-skills/openclaw-local-mac-mini/references/details.md +221 -0
- package/bundled-skills/openclaw-security-hardening/SKILL.md +135 -0
- package/bundled-skills/openshift/SKILL.md +485 -0
- package/bundled-skills/opentelemetry/SKILL.md +438 -0
- package/bundled-skills/opentelemetry/references/details.md +78 -0
- package/bundled-skills/opentofu-migration/SKILL.md +349 -0
- package/bundled-skills/osint-methodology/SKILL.md +460 -0
- package/bundled-skills/osint-methodology/references/details.md +1350 -0
- package/bundled-skills/pci-dss-compliance/SKILL.md +446 -0
- package/bundled-skills/penetration-testing/SKILL.md +152 -0
- package/bundled-skills/performance-tuning/SKILL.md +381 -0
- package/bundled-skills/planetscale/SKILL.md +297 -0
- package/bundled-skills/platform-engineering/SKILL.md +348 -0
- package/bundled-skills/platform-engineering/references/details.md +944 -0
- package/bundled-skills/podman/SKILL.md +405 -0
- package/bundled-skills/policy-as-code/SKILL.md +434 -0
- package/bundled-skills/policy-as-code/references/details.md +204 -0
- package/bundled-skills/postgresql-devsec/SKILL.md +378 -0
- package/bundled-skills/prometheus-grafana/SKILL.md +469 -0
- package/bundled-skills/prompt-injection-defense/SKILL.md +483 -0
- package/bundled-skills/rag-infrastructure/SKILL.md +269 -0
- package/bundled-skills/rag-observability-evals/SKILL.md +444 -0
- package/bundled-skills/rag-observability-evals/references/details.md +92 -0
- package/bundled-skills/recon-scope-triage/SKILL.md +128 -0
- package/bundled-skills/redis/SKILL.md +421 -0
- package/bundled-skills/redteam-report-template/SKILL.md +370 -0
- package/bundled-skills/report-writing/SKILL.md +426 -0
- package/bundled-skills/report-writing/references/details.md +187 -0
- package/bundled-skills/reverse-proxy/SKILL.md +420 -0
- package/bundled-skills/runbook-creation/SKILL.md +438 -0
- package/bundled-skills/runbook-creation/references/details.md +71 -0
- package/bundled-skills/saas-security-posture/SKILL.md +415 -0
- package/bundled-skills/sast-scanning/SKILL.md +444 -0
- package/bundled-skills/sbom-supply-chain/SKILL.md +433 -0
- package/bundled-skills/security-arsenal/SKILL.md +446 -0
- package/bundled-skills/security-arsenal/references/details.md +540 -0
- package/bundled-skills/security-automation/SKILL.md +146 -0
- package/bundled-skills/semantic-versioning/SKILL.md +434 -0
- package/bundled-skills/semantic-versioning/references/details.md +83 -0
- package/bundled-skills/service-mesh/SKILL.md +422 -0
- package/bundled-skills/soc2-compliance/SKILL.md +409 -0
- package/bundled-skills/sops-encryption/SKILL.md +124 -0
- package/bundled-skills/sre-dashboards/SKILL.md +143 -0
- package/bundled-skills/ssh-configuration/SKILL.md +324 -0
- package/bundled-skills/ssl-tls-management/SKILL.md +428 -0
- package/bundled-skills/ssl-tls-management/references/details.md +99 -0
- package/bundled-skills/startup-it-troubleshooting/SKILL.md +415 -0
- package/bundled-skills/supply-chain-attack-recon/SKILL.md +453 -0
- package/bundled-skills/supply-chain-attack-recon/references/details.md +258 -0
- package/bundled-skills/systemd-services/SKILL.md +379 -0
- package/bundled-skills/terraform-aws/SKILL.md +125 -0
- package/bundled-skills/terraform-azure/SKILL.md +415 -0
- package/bundled-skills/terraform-azure/references/details.md +231 -0
- package/bundled-skills/terraform-gcp/SKILL.md +369 -0
- package/bundled-skills/threat-modeling/SKILL.md +487 -0
- package/bundled-skills/user-management/SKILL.md +383 -0
- package/bundled-skills/using-agent-skills/SKILL.md +220 -0
- package/bundled-skills/vector-database-ops/SKILL.md +300 -0
- package/bundled-skills/vendor-management/SKILL.md +439 -0
- package/bundled-skills/vendor-management/references/details.md +109 -0
- package/bundled-skills/vercel-deployments/SKILL.md +296 -0
- package/bundled-skills/vllm-server/SKILL.md +236 -0
- package/bundled-skills/vmware-vcenter-attack/SKILL.md +412 -0
- package/bundled-skills/vpn-setup/SKILL.md +452 -0
- package/bundled-skills/vulnerability-scanning/SKILL.md +448 -0
- package/bundled-skills/waf-setup/SKILL.md +354 -0
- package/bundled-skills/waf-setup/references/details.md +211 -0
- package/bundled-skills/web2-recon/SKILL.md +440 -0
- package/bundled-skills/web2-recon/references/details.md +319 -0
- package/bundled-skills/web3-audit/SKILL.md +445 -0
- package/bundled-skills/web3-audit/references/details.md +224 -0
- package/bundled-skills/windows-hardening/SKILL.md +454 -0
- package/bundled-skills/windows-hardening/references/details.md +204 -0
- package/bundled-skills/windows-server/SKILL.md +318 -0
- package/bundled-skills/zero-trust/SKILL.md +461 -0
- package/package.json +1 -1
- package/skills_index.json +6943 -323
|
@@ -0,0 +1,343 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: ai-security-hardening
|
|
3
|
+
description: Harden AI/LLM deployments against prompt injection, data exfiltration,
|
|
4
|
+
model theft, and supply chain attacks.
|
|
5
|
+
category: security
|
|
6
|
+
risk: safe
|
|
7
|
+
source: https://github.com/BagelHole/DevOps-Security-Agent-Skills
|
|
8
|
+
source_repo: BagelHole/DevOps-Security-Agent-Skills
|
|
9
|
+
source_type: community
|
|
10
|
+
date_added: '2026-09-20'
|
|
11
|
+
license: MIT
|
|
12
|
+
license_source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
|
|
13
|
+
compatibility: Requires the relevant security tooling (scanners, vault CLIs) and an
|
|
14
|
+
authorized scope for any active assessment. Docs-only; helper scripts and templates
|
|
15
|
+
not bundled.
|
|
16
|
+
metadata:
|
|
17
|
+
author: devops-skills
|
|
18
|
+
version: '1.0'
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
# AI Security Hardening
|
|
22
|
+
|
|
23
|
+
Secure LLM and AI systems against prompt injection, jailbreaks, data leakage, and supply chain threats in production environments.
|
|
24
|
+
|
|
25
|
+
## When to Use This Skill
|
|
26
|
+
|
|
27
|
+
Use this skill when:
|
|
28
|
+
- Deploying an LLM-powered application handling sensitive user data
|
|
29
|
+
- Protecting against prompt injection attacks in AI agents
|
|
30
|
+
- Implementing output filtering and content moderation
|
|
31
|
+
- Securing model weights and API endpoints from theft
|
|
32
|
+
- Achieving SOC2 or ISO 27001 compliance for AI systems
|
|
33
|
+
|
|
34
|
+
## AI-Specific Threat Model
|
|
35
|
+
|
|
36
|
+
```
|
|
37
|
+
Threat Risk Control
|
|
38
|
+
─────────────────────────────────────────────────────────────────────
|
|
39
|
+
Prompt injection System prompt override Input sanitization, separate context
|
|
40
|
+
Data exfiltration PII in model outputs Output filtering, DLP scanning
|
|
41
|
+
Jailbreaking Policy bypass Content moderation, guardrails
|
|
42
|
+
Model theft Weight extraction via API Rate limiting, access controls
|
|
43
|
+
Training data poisoning Backdoored fine-tuned model Dataset validation, provenance
|
|
44
|
+
Supply chain attack Malicious model weights Signature verification, scanning
|
|
45
|
+
Insecure output XSS/SQLi from LLM response Output encoding, parameterized queries
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
## Prompt Injection Defense
|
|
49
|
+
|
|
50
|
+
```python
|
|
51
|
+
import re
|
|
52
|
+
from typing import Optional
|
|
53
|
+
|
|
54
|
+
INJECTION_PATTERNS = [
|
|
55
|
+
r"ignore\s+(all\s+)?(previous|prior|above)\s+instructions",
|
|
56
|
+
r"you\s+are\s+now\s+",
|
|
57
|
+
r"new\s+instructions?:",
|
|
58
|
+
r"system\s+prompt",
|
|
59
|
+
r"forget\s+everything",
|
|
60
|
+
r"act\s+as\s+",
|
|
61
|
+
r"jailbreak",
|
|
62
|
+
r"dan\s+mode",
|
|
63
|
+
r"<\s*system\s*>",
|
|
64
|
+
r"\[INST\]",
|
|
65
|
+
]
|
|
66
|
+
|
|
67
|
+
def detect_prompt_injection(user_input: str) -> tuple[bool, Optional[str]]:
|
|
68
|
+
"""Return (is_suspicious, matched_pattern)."""
|
|
69
|
+
normalized = user_input.lower().strip()
|
|
70
|
+
for pattern in INJECTION_PATTERNS:
|
|
71
|
+
if re.search(pattern, normalized, re.IGNORECASE):
|
|
72
|
+
return True, pattern
|
|
73
|
+
return False, None
|
|
74
|
+
|
|
75
|
+
def sanitize_user_input(user_input: str, max_length: int = 4000) -> str:
|
|
76
|
+
"""Sanitize input before passing to LLM."""
|
|
77
|
+
# Truncate
|
|
78
|
+
user_input = user_input[:max_length]
|
|
79
|
+
|
|
80
|
+
# Remove null bytes and control characters
|
|
81
|
+
user_input = re.sub(r'[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]', '', user_input)
|
|
82
|
+
|
|
83
|
+
# Check for injection
|
|
84
|
+
suspicious, pattern = detect_prompt_injection(user_input)
|
|
85
|
+
if suspicious:
|
|
86
|
+
raise ValueError(f"Potential prompt injection detected: {pattern}")
|
|
87
|
+
|
|
88
|
+
return user_input
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
## Guardrails with NeMo Guardrails
|
|
92
|
+
|
|
93
|
+
```python
|
|
94
|
+
# guardrails.yaml
|
|
95
|
+
from nemoguardrails import RailsConfig, LLMRails
|
|
96
|
+
|
|
97
|
+
config = RailsConfig.from_path("./guardrails-config")
|
|
98
|
+
rails = LLMRails(config)
|
|
99
|
+
|
|
100
|
+
async def safe_llm_call(user_message: str) -> str:
|
|
101
|
+
response = await rails.generate_async(
|
|
102
|
+
messages=[{"role": "user", "content": user_message}]
|
|
103
|
+
)
|
|
104
|
+
return response["content"]
|
|
105
|
+
```
|
|
106
|
+
|
|
107
|
+
```yaml
|
|
108
|
+
# guardrails-config/config.yml
|
|
109
|
+
models:
|
|
110
|
+
- type: main
|
|
111
|
+
engine: openai
|
|
112
|
+
model: gpt-4o-mini
|
|
113
|
+
|
|
114
|
+
rails:
|
|
115
|
+
input:
|
|
116
|
+
flows:
|
|
117
|
+
- check jailbreak
|
|
118
|
+
- check sensitive data
|
|
119
|
+
output:
|
|
120
|
+
flows:
|
|
121
|
+
- check output for PII
|
|
122
|
+
- check output for harmful content
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
## Output Filtering & PII Scrubbing
|
|
126
|
+
|
|
127
|
+
```python
|
|
128
|
+
import re
|
|
129
|
+
from presidio_analyzer import AnalyzerEngine
|
|
130
|
+
from presidio_anonymizer import AnonymizerEngine
|
|
131
|
+
|
|
132
|
+
analyzer = AnalyzerEngine()
|
|
133
|
+
anonymizer = AnonymizerEngine()
|
|
134
|
+
|
|
135
|
+
PII_ENTITIES = ["PERSON", "EMAIL_ADDRESS", "PHONE_NUMBER", "CREDIT_CARD",
|
|
136
|
+
"US_SSN", "IBAN_CODE", "IP_ADDRESS", "LOCATION"]
|
|
137
|
+
|
|
138
|
+
def scrub_pii_from_output(text: str) -> str:
|
|
139
|
+
"""Remove PII from LLM output before returning to user."""
|
|
140
|
+
results = analyzer.analyze(text=text, entities=PII_ENTITIES, language="en")
|
|
141
|
+
if not results:
|
|
142
|
+
return text
|
|
143
|
+
anonymized = anonymizer.anonymize(text=text, analyzer_results=results)
|
|
144
|
+
return anonymized.text
|
|
145
|
+
|
|
146
|
+
def validate_output_safety(output: str) -> bool:
|
|
147
|
+
"""Check output doesn't contain prompt injection artifacts."""
|
|
148
|
+
dangerous_patterns = [
|
|
149
|
+
r"<\s*script\s*>", # XSS
|
|
150
|
+
r"javascript:", # XSS
|
|
151
|
+
r";\s*(DROP|DELETE|INSERT)",# SQLi
|
|
152
|
+
r"\$\{.*\}", # template injection
|
|
153
|
+
r"`.*`", # command injection in some contexts
|
|
154
|
+
]
|
|
155
|
+
for pattern in dangerous_patterns:
|
|
156
|
+
if re.search(pattern, output, re.IGNORECASE):
|
|
157
|
+
return False
|
|
158
|
+
return True
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
## API Security for LLM Endpoints
|
|
162
|
+
|
|
163
|
+
```python
|
|
164
|
+
from fastapi import FastAPI, HTTPException, Depends, Request
|
|
165
|
+
from fastapi.security import HTTPBearer, HTTPAuthorizationCredentials
|
|
166
|
+
import jwt
|
|
167
|
+
import time
|
|
168
|
+
from collections import defaultdict
|
|
169
|
+
|
|
170
|
+
app = FastAPI()
|
|
171
|
+
security = HTTPBearer()
|
|
172
|
+
|
|
173
|
+
# Rate limiting (per API key)
|
|
174
|
+
request_counts = defaultdict(list)
|
|
175
|
+
|
|
176
|
+
def rate_limit(api_key: str, max_requests: int = 100, window_seconds: int = 60):
|
|
177
|
+
now = time.time()
|
|
178
|
+
requests = request_counts[api_key]
|
|
179
|
+
# Remove old requests outside window
|
|
180
|
+
request_counts[api_key] = [t for t in requests if now - t < window_seconds]
|
|
181
|
+
if len(request_counts[api_key]) >= max_requests:
|
|
182
|
+
raise HTTPException(status_code=429, detail="Rate limit exceeded")
|
|
183
|
+
request_counts[api_key].append(now)
|
|
184
|
+
|
|
185
|
+
async def verify_token(
|
|
186
|
+
credentials: HTTPAuthorizationCredentials = Depends(security)
|
|
187
|
+
) -> dict:
|
|
188
|
+
try:
|
|
189
|
+
payload = jwt.decode(credentials.credentials, SECRET_KEY, algorithms=["HS256"])
|
|
190
|
+
rate_limit(payload["sub"])
|
|
191
|
+
return payload
|
|
192
|
+
except jwt.ExpiredSignatureError:
|
|
193
|
+
raise HTTPException(status_code=401, detail="Token expired")
|
|
194
|
+
except jwt.InvalidTokenError:
|
|
195
|
+
raise HTTPException(status_code=401, detail="Invalid token")
|
|
196
|
+
|
|
197
|
+
@app.post("/v1/chat/completions")
|
|
198
|
+
async def chat(request: Request, token: dict = Depends(verify_token)):
|
|
199
|
+
body = await request.json()
|
|
200
|
+
|
|
201
|
+
# Input validation
|
|
202
|
+
user_msg = body.get("messages", [{}])[-1].get("content", "")
|
|
203
|
+
try:
|
|
204
|
+
safe_input = sanitize_user_input(user_msg)
|
|
205
|
+
except ValueError as e:
|
|
206
|
+
raise HTTPException(status_code=400, detail=str(e))
|
|
207
|
+
|
|
208
|
+
# Call LLM and scrub output
|
|
209
|
+
response = await call_llm(safe_input, token["scope"])
|
|
210
|
+
response["choices"][0]["message"]["content"] = scrub_pii_from_output(
|
|
211
|
+
response["choices"][0]["message"]["content"]
|
|
212
|
+
)
|
|
213
|
+
return response
|
|
214
|
+
```
|
|
215
|
+
|
|
216
|
+
## Model Weight Security
|
|
217
|
+
|
|
218
|
+
```bash
|
|
219
|
+
# Verify model weights with SHA-256 hash before loading
|
|
220
|
+
MODEL_DIR="./models/llama-3.1-8b"
|
|
221
|
+
EXPECTED_HASH="sha256:abc123..."
|
|
222
|
+
|
|
223
|
+
# Generate hash of downloaded model
|
|
224
|
+
actual_hash=$(find "$MODEL_DIR" -name "*.safetensors" | sort | xargs sha256sum | sha256sum)
|
|
225
|
+
echo "Model hash: $actual_hash"
|
|
226
|
+
|
|
227
|
+
# Compare (automate in CI/CD)
|
|
228
|
+
if [ "$actual_hash" != "$EXPECTED_HASH" ]; then
|
|
229
|
+
echo "ERROR: Model hash mismatch — possible tampering!"
|
|
230
|
+
exit 1
|
|
231
|
+
fi
|
|
232
|
+
|
|
233
|
+
# Scan model files for embedded malware (ModelScan)
|
|
234
|
+
pip install modelscan
|
|
235
|
+
modelscan scan -p "$MODEL_DIR"
|
|
236
|
+
```
|
|
237
|
+
|
|
238
|
+
## Network Isolation for AI Services
|
|
239
|
+
|
|
240
|
+
```yaml
|
|
241
|
+
# Kubernetes NetworkPolicy — isolate LLM API
|
|
242
|
+
apiVersion: networking.k8s.io/v1
|
|
243
|
+
kind: NetworkPolicy
|
|
244
|
+
metadata:
|
|
245
|
+
name: llm-api-isolation
|
|
246
|
+
namespace: ai-services
|
|
247
|
+
spec:
|
|
248
|
+
podSelector:
|
|
249
|
+
matchLabels:
|
|
250
|
+
app: vllm
|
|
251
|
+
policyTypes:
|
|
252
|
+
- Ingress
|
|
253
|
+
- Egress
|
|
254
|
+
ingress:
|
|
255
|
+
- from:
|
|
256
|
+
- namespaceSelector:
|
|
257
|
+
matchLabels:
|
|
258
|
+
name: backend # only backend can call LLM
|
|
259
|
+
ports:
|
|
260
|
+
- protocol: TCP
|
|
261
|
+
port: 8000
|
|
262
|
+
egress:
|
|
263
|
+
- to:
|
|
264
|
+
- namespaceSelector:
|
|
265
|
+
matchLabels:
|
|
266
|
+
name: monitoring # metrics only
|
|
267
|
+
ports:
|
|
268
|
+
- protocol: TCP
|
|
269
|
+
port: 9090
|
|
270
|
+
# Block egress to internet — prevent data exfiltration
|
|
271
|
+
# (allow only internal cluster traffic)
|
|
272
|
+
```
|
|
273
|
+
|
|
274
|
+
## Audit Logging
|
|
275
|
+
|
|
276
|
+
```python
|
|
277
|
+
import structlog
|
|
278
|
+
from datetime import datetime, timezone
|
|
279
|
+
|
|
280
|
+
audit_log = structlog.get_logger("ai.audit")
|
|
281
|
+
|
|
282
|
+
def log_llm_interaction(
|
|
283
|
+
user_id: str,
|
|
284
|
+
session_id: str,
|
|
285
|
+
model: str,
|
|
286
|
+
prompt_tokens: int,
|
|
287
|
+
completion_tokens: int,
|
|
288
|
+
was_filtered: bool,
|
|
289
|
+
injection_detected: bool,
|
|
290
|
+
):
|
|
291
|
+
audit_log.info(
|
|
292
|
+
"llm_interaction",
|
|
293
|
+
timestamp=datetime.now(timezone.utc).isoformat(),
|
|
294
|
+
user_id=user_id,
|
|
295
|
+
session_id=session_id,
|
|
296
|
+
model=model,
|
|
297
|
+
prompt_tokens=prompt_tokens,
|
|
298
|
+
completion_tokens=completion_tokens,
|
|
299
|
+
was_filtered=was_filtered,
|
|
300
|
+
injection_detected=injection_detected,
|
|
301
|
+
# DO NOT log prompt/completion content — PII risk
|
|
302
|
+
)
|
|
303
|
+
```
|
|
304
|
+
|
|
305
|
+
## Common Issues
|
|
306
|
+
|
|
307
|
+
| Issue | Cause | Fix |
|
|
308
|
+
|-------|-------|-----|
|
|
309
|
+
| False positive injection blocks | Overly broad regex | Tune patterns; use ML-based classifier for high-traffic |
|
|
310
|
+
| PII in model outputs | Model trained on PII data | Add Presidio scrubbing to output layer |
|
|
311
|
+
| API key leakage | Keys in logs or responses | Mask keys in logging; use vault for key storage |
|
|
312
|
+
| Model weight tampering | Unverified downloads | Always verify SHA-256; use `modelscan` |
|
|
313
|
+
| Rate limit bypass | Per-IP not per-user | Rate limit on authenticated user ID, not IP |
|
|
314
|
+
|
|
315
|
+
## Best Practices
|
|
316
|
+
|
|
317
|
+
- Never log raw prompts or completions — they may contain PII or sensitive data.
|
|
318
|
+
- Treat LLM output as untrusted input — always encode before rendering in HTML.
|
|
319
|
+
- Use network policies to prevent LLM pods from making outbound internet calls.
|
|
320
|
+
- Rotate API keys quarterly; use short-lived JWT tokens for service-to-service auth.
|
|
321
|
+
- Run `modelscan` on any model downloaded from the internet before serving.
|
|
322
|
+
|
|
323
|
+
## Related Skills
|
|
324
|
+
|
|
325
|
+
- hashicorp-vault (`hashicorp-vault`) - Secrets management for API keys
|
|
326
|
+
- [network-security] (../../network/) - Network-level controls
|
|
327
|
+
- linux-hardening (`linux-hardening`) - Host hardening
|
|
328
|
+
- agent-observability (`agent-observability`) - AI audit logging
|
|
329
|
+
- llm-gateway (`llm-gateway`) - Centralized access control
|
|
330
|
+
|
|
331
|
+
## Limitations
|
|
332
|
+
|
|
333
|
+
- Apply guidance only within authorized scope; test destructive steps in non-production first.
|
|
334
|
+
- Docs-only import: upstream scripts and templates not bundled.
|
|
335
|
+
|
|
336
|
+
### Example
|
|
337
|
+
|
|
338
|
+
```bash
|
|
339
|
+
# Read-only first: inventory before any active step.
|
|
340
|
+
which <tool> && <tool> --help | head -n 20
|
|
341
|
+
```
|
|
342
|
+
|
|
343
|
+
> Adapted from [BagelHole/DevOps-Security-Agent-Skills](https://github.com/BagelHole/DevOps-Security-Agent-Skills) (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.
|
|
@@ -0,0 +1,336 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: ai-sre-incident-response
|
|
3
|
+
description: Build AI-focused SRE incident response practices for LLM outages, degraded
|
|
4
|
+
quality, runaway cost events, and safety regressions.
|
|
5
|
+
category: devops
|
|
6
|
+
risk: critical
|
|
7
|
+
source: https://github.com/BagelHole/DevOps-Security-Agent-Skills
|
|
8
|
+
source_repo: BagelHole/DevOps-Security-Agent-Skills
|
|
9
|
+
source_type: community
|
|
10
|
+
date_added: '2026-09-20'
|
|
11
|
+
license: MIT
|
|
12
|
+
license_source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
|
|
13
|
+
compatibility: Requires the relevant platform CLIs (kubectl, helm, terraform, git,
|
|
14
|
+
CI runners) and authorized access to the target environment. Docs-only; helper scripts
|
|
15
|
+
and templates not bundled.
|
|
16
|
+
metadata:
|
|
17
|
+
author: devops-skills
|
|
18
|
+
version: '1.0'
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
# AI SRE Incident Response
|
|
22
|
+
|
|
23
|
+
Apply SRE rigor to AI systems where incidents include quality regressions, unsafe outputs, and budget explosions.
|
|
24
|
+
|
|
25
|
+
## When to Use This Skill
|
|
26
|
+
|
|
27
|
+
- An LLM endpoint begins returning degraded or hallucinated answers
|
|
28
|
+
- Token spend spikes beyond budget thresholds
|
|
29
|
+
- A model provider goes down and traffic must fail over
|
|
30
|
+
- Safety guardrails fire at abnormal rates
|
|
31
|
+
- A new model deployment causes latency or accuracy regression
|
|
32
|
+
|
|
33
|
+
## Prerequisites
|
|
34
|
+
|
|
35
|
+
- Prometheus and Alertmanager deployed with scrape targets for AI services
|
|
36
|
+
- Grafana dashboards for golden signals (latency, error rate, cost, quality)
|
|
37
|
+
- On-call rotation configured in PagerDuty, Opsgenie, or equivalent
|
|
38
|
+
- Runbook repository accessible to responders
|
|
39
|
+
- Rollback mechanism for model and prompt versions (GitOps or feature flags)
|
|
40
|
+
|
|
41
|
+
## AI Incident Classes
|
|
42
|
+
|
|
43
|
+
- **Availability incident**: model/provider unavailable, timeout storm.
|
|
44
|
+
- **Quality incident**: answer accuracy or tool success drops below SLO.
|
|
45
|
+
- **Safety incident**: harmful or policy-violating outputs increase.
|
|
46
|
+
- **Cost incident**: unexpected token or provider spend spike.
|
|
47
|
+
|
|
48
|
+
## Severity Framework
|
|
49
|
+
|
|
50
|
+
| Severity | Criteria | Response Time | Notification |
|
|
51
|
+
|----------|----------|---------------|--------------|
|
|
52
|
+
| SEV1 | User-facing outage, compliance risk, data leak | 5 min | Page on-call + incident commander |
|
|
53
|
+
| SEV2 | Major degradation in key flows | 15 min | Page on-call |
|
|
54
|
+
| SEV3 | Limited impact or internal-only issue | 1 hour | Slack alert |
|
|
55
|
+
| SEV4 | Cosmetic or low-priority regression | Next business day | Ticket |
|
|
56
|
+
|
|
57
|
+
## Golden Signals for AI Services
|
|
58
|
+
|
|
59
|
+
- Request success rate
|
|
60
|
+
- Latency (queue + generation + tool execution)
|
|
61
|
+
- Hallucination/groundedness proxy metrics
|
|
62
|
+
- Cost per minute and per tenant
|
|
63
|
+
- Guardrail violation rate
|
|
64
|
+
|
|
65
|
+
## Prometheus Alert Rules
|
|
66
|
+
|
|
67
|
+
```yaml
|
|
68
|
+
# prometheus-ai-alerts.yaml
|
|
69
|
+
groups:
|
|
70
|
+
- name: ai-service-alerts
|
|
71
|
+
rules:
|
|
72
|
+
- alert: ModelEndpointDown
|
|
73
|
+
expr: up{job="llm-inference"} == 0
|
|
74
|
+
for: 2m
|
|
75
|
+
labels:
|
|
76
|
+
severity: sev1
|
|
77
|
+
annotations:
|
|
78
|
+
summary: "LLM inference endpoint {{ $labels.instance }} is down"
|
|
79
|
+
runbook_url: "https://runbooks.internal/ai/model-outage"
|
|
80
|
+
|
|
81
|
+
- alert: HighHallucinationRate
|
|
82
|
+
expr: |
|
|
83
|
+
rate(llm_hallucination_detected_total[10m])
|
|
84
|
+
/ rate(llm_requests_total[10m]) > 0.15
|
|
85
|
+
for: 5m
|
|
86
|
+
labels:
|
|
87
|
+
severity: sev2
|
|
88
|
+
annotations:
|
|
89
|
+
summary: "Hallucination rate above 15% for {{ $labels.model }}"
|
|
90
|
+
runbook_url: "https://runbooks.internal/ai/quality-regression"
|
|
91
|
+
|
|
92
|
+
- alert: TokenCostExplosion
|
|
93
|
+
expr: |
|
|
94
|
+
sum(rate(llm_token_cost_dollars[5m])) by (tenant)
|
|
95
|
+
> 0.50
|
|
96
|
+
for: 3m
|
|
97
|
+
labels:
|
|
98
|
+
severity: sev2
|
|
99
|
+
annotations:
|
|
100
|
+
summary: "Token spend exceeds $0.50/min for tenant {{ $labels.tenant }}"
|
|
101
|
+
runbook_url: "https://runbooks.internal/ai/cost-spike"
|
|
102
|
+
|
|
103
|
+
- alert: LatencyP95Exceeded
|
|
104
|
+
expr: |
|
|
105
|
+
histogram_quantile(0.95,
|
|
106
|
+
rate(llm_request_duration_seconds_bucket[5m])
|
|
107
|
+
) > 5
|
|
108
|
+
for: 5m
|
|
109
|
+
labels:
|
|
110
|
+
severity: sev2
|
|
111
|
+
annotations:
|
|
112
|
+
summary: "LLM p95 latency exceeds 5s for {{ $labels.service }}"
|
|
113
|
+
|
|
114
|
+
- alert: GuardrailViolationSpike
|
|
115
|
+
expr: |
|
|
116
|
+
rate(llm_guardrail_violations_total[10m])
|
|
117
|
+
/ rate(llm_requests_total[10m]) > 0.05
|
|
118
|
+
for: 5m
|
|
119
|
+
labels:
|
|
120
|
+
severity: sev1
|
|
121
|
+
annotations:
|
|
122
|
+
summary: "Guardrail violations above 5% for {{ $labels.model }}"
|
|
123
|
+
runbook_url: "https://runbooks.internal/ai/safety-incident"
|
|
124
|
+
|
|
125
|
+
- alert: ModelQualityDrop
|
|
126
|
+
expr: |
|
|
127
|
+
llm_eval_score{metric="groundedness"} < 0.70
|
|
128
|
+
for: 10m
|
|
129
|
+
labels:
|
|
130
|
+
severity: sev2
|
|
131
|
+
annotations:
|
|
132
|
+
summary: "Groundedness score dropped below 0.70 for {{ $labels.model }}"
|
|
133
|
+
|
|
134
|
+
- alert: ProviderErrorRateHigh
|
|
135
|
+
expr: |
|
|
136
|
+
rate(llm_provider_errors_total[5m])
|
|
137
|
+
/ rate(llm_provider_requests_total[5m]) > 0.10
|
|
138
|
+
for: 3m
|
|
139
|
+
labels:
|
|
140
|
+
severity: sev2
|
|
141
|
+
annotations:
|
|
142
|
+
summary: "Provider {{ $labels.provider }} error rate above 10%"
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
## Response Playbooks
|
|
146
|
+
|
|
147
|
+
### Model Outage Runbook
|
|
148
|
+
|
|
149
|
+
```text
|
|
150
|
+
TRIGGER: ModelEndpointDown fires for > 2 minutes
|
|
151
|
+
RESPONDER: On-call AI platform engineer
|
|
152
|
+
|
|
153
|
+
1. Acknowledge alert in PagerDuty.
|
|
154
|
+
2. Check provider status page (e.g., status.openai.com).
|
|
155
|
+
3. Verify network connectivity:
|
|
156
|
+
curl -s -o /dev/null -w "%{http_code}" https://api.provider.com/health
|
|
157
|
+
4. If provider is down:
|
|
158
|
+
a. Enable fallback model route in gateway config.
|
|
159
|
+
b. kubectl set env deployment/llm-gateway FALLBACK_ENABLED=true
|
|
160
|
+
c. Verify fallback traffic is flowing via Grafana dashboard.
|
|
161
|
+
5. If self-hosted model is down:
|
|
162
|
+
a. Check pod status: kubectl get pods -l app=llm-inference -n ai
|
|
163
|
+
b. Check GPU health: kubectl logs -l app=llm-inference --tail=50
|
|
164
|
+
c. Restart if OOM: kubectl rollout restart deployment/llm-inference -n ai
|
|
165
|
+
6. Freeze all deployments:
|
|
166
|
+
kubectl annotate deployment --all deploy-freeze=true -n ai
|
|
167
|
+
7. Communicate ETA in #incident-channel.
|
|
168
|
+
8. When resolved, unfreeze and run smoke tests.
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
### Quality Regression Runbook (Hallucination Spike)
|
|
172
|
+
|
|
173
|
+
```text
|
|
174
|
+
TRIGGER: HighHallucinationRate or ModelQualityDrop fires
|
|
175
|
+
RESPONDER: On-call AI engineer + ML lead
|
|
176
|
+
|
|
177
|
+
1. Acknowledge alert. Open incident ticket.
|
|
178
|
+
2. Identify scope:
|
|
179
|
+
- Which model version? Check deployment metadata.
|
|
180
|
+
- Which routes/tenants affected? Filter by labels in Grafana.
|
|
181
|
+
3. Check recent changes:
|
|
182
|
+
- Model version promotion in last 24h?
|
|
183
|
+
- Prompt template changes in last 24h?
|
|
184
|
+
- Retrieval index rebuild in last 24h?
|
|
185
|
+
4. If recent model change:
|
|
186
|
+
kubectl rollout undo deployment/llm-inference -n ai
|
|
187
|
+
5. If recent prompt change:
|
|
188
|
+
git revert <commit> && git push # triggers GitOps redeploy
|
|
189
|
+
6. Increase trace sampling to 100% for affected route:
|
|
190
|
+
kubectl set env deployment/llm-gateway TRACE_SAMPLE_RATE=1.0
|
|
191
|
+
7. Run offline eval suite against current production:
|
|
192
|
+
python run_evals.py --target prod --suite quality --compare baseline
|
|
193
|
+
8. Confirm metrics return to baseline before closing.
|
|
194
|
+
```
|
|
195
|
+
|
|
196
|
+
### Token Cost Explosion Runbook
|
|
197
|
+
|
|
198
|
+
```text
|
|
199
|
+
TRIGGER: TokenCostExplosion fires
|
|
200
|
+
RESPONDER: On-call platform engineer
|
|
201
|
+
|
|
202
|
+
1. Identify top consumers:
|
|
203
|
+
Query: topk(10, sum(rate(llm_token_cost_dollars[15m])) by (tenant, model, route))
|
|
204
|
+
2. Check for runaway loops:
|
|
205
|
+
- Agent retry storms (exponential token growth per request)
|
|
206
|
+
- Missing max_tokens caps on new routes
|
|
207
|
+
- Cache bypass due to config change
|
|
208
|
+
3. Apply immediate caps:
|
|
209
|
+
kubectl patch configmap llm-quotas -n ai --patch '
|
|
210
|
+
data:
|
|
211
|
+
max_tokens_per_request: "4096"
|
|
212
|
+
rpm_limit: "60"
|
|
213
|
+
'
|
|
214
|
+
4. Enable semantic cache if disabled:
|
|
215
|
+
kubectl set env deployment/llm-gateway CACHE_ENABLED=true
|
|
216
|
+
5. Route traffic to cheaper model tier:
|
|
217
|
+
kubectl set env deployment/llm-gateway DEFAULT_MODEL=gpt-4o-mini
|
|
218
|
+
6. Notify affected tenants of temporary limits.
|
|
219
|
+
7. Open postmortem with cost attribution analysis.
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
## Escalation Procedures
|
|
223
|
+
|
|
224
|
+
```text
|
|
225
|
+
Level 1 (0-15 min): On-call AI platform engineer
|
|
226
|
+
Level 2 (15-30 min): AI platform team lead + affected product owner
|
|
227
|
+
Level 3 (30-60 min): Engineering director + security (if safety incident)
|
|
228
|
+
Level 4 (60+ min): VP Engineering + legal (if compliance/data incident)
|
|
229
|
+
|
|
230
|
+
Safety incidents always start at Level 2 minimum.
|
|
231
|
+
Provider-side incidents: open support ticket immediately at Level 1.
|
|
232
|
+
```
|
|
233
|
+
|
|
234
|
+
## Detection Queries (PromQL)
|
|
235
|
+
|
|
236
|
+
```promql
|
|
237
|
+
# Request success rate by model
|
|
238
|
+
1 - (
|
|
239
|
+
sum(rate(llm_requests_total{status="error"}[5m])) by (model)
|
|
240
|
+
/ sum(rate(llm_requests_total[5m])) by (model)
|
|
241
|
+
)
|
|
242
|
+
|
|
243
|
+
# Cost per successful answer
|
|
244
|
+
sum(rate(llm_token_cost_dollars[5m])) by (route)
|
|
245
|
+
/ sum(rate(llm_requests_total{status="success"}[5m])) by (route)
|
|
246
|
+
|
|
247
|
+
# Hallucination rate trend (1h window, 5m steps)
|
|
248
|
+
rate(llm_hallucination_detected_total[1h])
|
|
249
|
+
/ rate(llm_requests_total[1h])
|
|
250
|
+
|
|
251
|
+
# Latency breakdown by stage
|
|
252
|
+
histogram_quantile(0.95, rate(llm_retrieval_duration_seconds_bucket[5m]))
|
|
253
|
+
histogram_quantile(0.95, rate(llm_generation_duration_seconds_bucket[5m]))
|
|
254
|
+
histogram_quantile(0.95, rate(llm_tool_execution_duration_seconds_bucket[5m]))
|
|
255
|
+
|
|
256
|
+
# Tenant cost leaderboard
|
|
257
|
+
topk(10, sum(rate(llm_token_cost_dollars[1h])) by (tenant))
|
|
258
|
+
```
|
|
259
|
+
|
|
260
|
+
## Postmortem Requirements
|
|
261
|
+
|
|
262
|
+
- Timeline with detector and responder timestamps
|
|
263
|
+
- Blast radius by tenant and feature
|
|
264
|
+
- Missed signals and alert tuning actions
|
|
265
|
+
- Concrete hardening tasks with owners and due dates
|
|
266
|
+
- Cost impact (dollars, tokens, affected requests)
|
|
267
|
+
- Customer communication log
|
|
268
|
+
|
|
269
|
+
## Postmortem Template
|
|
270
|
+
|
|
271
|
+
```markdown
|
|
272
|
+
## Incident Summary
|
|
273
|
+
- **Severity**: SEVx
|
|
274
|
+
- **Duration**: start_time - end_time (Xh Ym)
|
|
275
|
+
- **Detection**: How was it detected? (alert / customer report / manual)
|
|
276
|
+
- **Impact**: X tenants, Y requests, $Z cost
|
|
277
|
+
|
|
278
|
+
## Timeline
|
|
279
|
+
| Time (UTC) | Event |
|
|
280
|
+
|------------|-------|
|
|
281
|
+
| HH:MM | Alert fired |
|
|
282
|
+
| HH:MM | Responder acknowledged |
|
|
283
|
+
| HH:MM | Root cause identified |
|
|
284
|
+
| HH:MM | Mitigation applied |
|
|
285
|
+
| HH:MM | Incident resolved |
|
|
286
|
+
|
|
287
|
+
## Root Cause
|
|
288
|
+
[Description]
|
|
289
|
+
|
|
290
|
+
## Action Items
|
|
291
|
+
| Action | Owner | Due Date | Status |
|
|
292
|
+
|--------|-------|----------|--------|
|
|
293
|
+
| Tune alert threshold | @engineer | YYYY-MM-DD | Open |
|
|
294
|
+
| Add fallback route | @platform | YYYY-MM-DD | Open |
|
|
295
|
+
```
|
|
296
|
+
|
|
297
|
+
## Chaos Engineering for AI Systems
|
|
298
|
+
|
|
299
|
+
Regularly test incident readiness:
|
|
300
|
+
|
|
301
|
+
- **Provider failover drill**: block provider API at network level, verify fallback activates within SLO.
|
|
302
|
+
- **Model rollback drill**: deploy known-bad model version, verify automated quality gate catches it.
|
|
303
|
+
- **Cost cap drill**: simulate runaway token usage, verify quotas trigger before budget threshold.
|
|
304
|
+
- **Cache failure drill**: disable semantic cache, verify system degrades gracefully.
|
|
305
|
+
|
|
306
|
+
## Troubleshooting
|
|
307
|
+
|
|
308
|
+
| Symptom | Check | Fix |
|
|
309
|
+
|---------|-------|-----|
|
|
310
|
+
| All requests timing out | Provider status page, DNS resolution | Enable fallback provider |
|
|
311
|
+
| Gradual quality decline | Recent model/prompt deployments | Roll back to last known good |
|
|
312
|
+
| Sudden cost spike | Per-tenant token usage dashboard | Apply emergency token caps |
|
|
313
|
+
| Guardrail violations spike | Model version, prompt injection logs | Enable stricter input filtering |
|
|
314
|
+
| Intermittent 503 errors | Pod restarts, GPU OOM events | Increase memory limits or reduce batch size |
|
|
315
|
+
|
|
316
|
+
## Related Skills
|
|
317
|
+
|
|
318
|
+
- incident-response (`incident-response`) - Standard incident process and evidence
|
|
319
|
+
- alerting-oncall (`alerting-oncall`) - Paging and escalation policy
|
|
320
|
+
- llm-cost-optimization (`llm-cost-optimization`) - Spend controls and efficiency patterns
|
|
321
|
+
- agent-observability (`agent-observability`) - Instrument requests, traces, and costs
|
|
322
|
+
- rag-observability-evals (`rag-observability-evals`) - RAG quality monitoring
|
|
323
|
+
|
|
324
|
+
## Limitations
|
|
325
|
+
|
|
326
|
+
- Guidance executes against real environments: confirm target, blast radius, and rollback plan before applying anything.
|
|
327
|
+
- Never deploy to production without explicit approval. Docs-only import: upstream scripts and templates not bundled.
|
|
328
|
+
|
|
329
|
+
### Example
|
|
330
|
+
|
|
331
|
+
```bash
|
|
332
|
+
git status && git diff --stat
|
|
333
|
+
kubectl diff -f manifest.yaml
|
|
334
|
+
```
|
|
335
|
+
|
|
336
|
+
> Adapted from [BagelHole/DevOps-Security-Agent-Skills](https://github.com/BagelHole/DevOps-Security-Agent-Skills) (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.
|