opencode-skills-collection 4.0.68 → 4.0.69

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (309) hide show
  1. package/bundled-skills/.antigravity-install-manifest.json +266 -1
  2. package/bundled-skills/access-review/SKILL.md +394 -0
  3. package/bundled-skills/access-review/references/details.md +121 -0
  4. package/bundled-skills/agent-evals/SKILL.md +420 -0
  5. package/bundled-skills/agent-observability/SKILL.md +346 -0
  6. package/bundled-skills/agent-observability/references/details.md +786 -0
  7. package/bundled-skills/ai-agent-security/SKILL.md +393 -0
  8. package/bundled-skills/ai-agent-security/references/details.md +912 -0
  9. package/bundled-skills/ai-coding-agent-guardrails/SKILL.md +442 -0
  10. package/bundled-skills/ai-coding-agent-guardrails/references/details.md +753 -0
  11. package/bundled-skills/ai-inference-service-mesh/SKILL.md +449 -0
  12. package/bundled-skills/ai-pipeline-orchestration/SKILL.md +287 -0
  13. package/bundled-skills/ai-red-teaming/SKILL.md +409 -0
  14. package/bundled-skills/ai-security-hardening/SKILL.md +343 -0
  15. package/bundled-skills/ai-sre-incident-response/SKILL.md +336 -0
  16. package/bundled-skills/alerting-oncall/SKILL.md +458 -0
  17. package/bundled-skills/alerting-oncall/references/details.md +84 -0
  18. package/bundled-skills/apk-redteam-pipeline/SKILL.md +446 -0
  19. package/bundled-skills/argocd-gitops/SKILL.md +469 -0
  20. package/bundled-skills/arm-templates/SKILL.md +438 -0
  21. package/bundled-skills/arm-templates/references/details.md +64 -0
  22. package/bundled-skills/asset-inventory/SKILL.md +412 -0
  23. package/bundled-skills/asset-inventory/references/details.md +127 -0
  24. package/bundled-skills/audit-logging/SKILL.md +476 -0
  25. package/bundled-skills/aws-cloudtrail/SKILL.md +486 -0
  26. package/bundled-skills/aws-cost-optimization/SKILL.md +331 -0
  27. package/bundled-skills/aws-ec2/SKILL.md +426 -0
  28. package/bundled-skills/aws-ecs-fargate/SKILL.md +388 -0
  29. package/bundled-skills/aws-iam/SKILL.md +463 -0
  30. package/bundled-skills/aws-lambda/SKILL.md +428 -0
  31. package/bundled-skills/aws-rds/SKILL.md +380 -0
  32. package/bundled-skills/aws-s3/SKILL.md +434 -0
  33. package/bundled-skills/aws-secrets-manager/SKILL.md +486 -0
  34. package/bundled-skills/aws-vpc/SKILL.md +436 -0
  35. package/bundled-skills/azure-ai-document-intelligence-ts/SKILL.md +1 -1
  36. package/bundled-skills/azure-aks/SKILL.md +423 -0
  37. package/bundled-skills/azure-devops/SKILL.md +457 -0
  38. package/bundled-skills/azure-functions-devsec/SKILL.md +436 -0
  39. package/bundled-skills/azure-keyvault/SKILL.md +455 -0
  40. package/bundled-skills/azure-keyvault/references/details.md +83 -0
  41. package/bundled-skills/azure-monitor-audit/SKILL.md +379 -0
  42. package/bundled-skills/azure-networking/SKILL.md +448 -0
  43. package/bundled-skills/azure-networking/references/details.md +135 -0
  44. package/bundled-skills/azure-sql/SKILL.md +413 -0
  45. package/bundled-skills/azure-sql/references/details.md +113 -0
  46. package/bundled-skills/azure-vms/SKILL.md +402 -0
  47. package/bundled-skills/azure-vms/references/details.md +134 -0
  48. package/bundled-skills/backup-recovery/SKILL.md +388 -0
  49. package/bundled-skills/bb-methodology/SKILL.md +451 -0
  50. package/bundled-skills/bb-methodology/references/details.md +120 -0
  51. package/bundled-skills/block-storage/SKILL.md +371 -0
  52. package/bundled-skills/blue-green-deploy/SKILL.md +453 -0
  53. package/bundled-skills/blue-green-deploy/references/details.md +90 -0
  54. package/bundled-skills/bug-bounty/SKILL.md +447 -0
  55. package/bundled-skills/bug-bounty/references/details.md +1316 -0
  56. package/bundled-skills/bugcrowd-reporting/SKILL.md +351 -0
  57. package/bundled-skills/business-continuity/SKILL.md +463 -0
  58. package/bundled-skills/career-ops/SKILL.md +186 -0
  59. package/bundled-skills/cdn-setup/SKILL.md +374 -0
  60. package/bundled-skills/change-management/SKILL.md +438 -0
  61. package/bundled-skills/change-management/references/details.md +105 -0
  62. package/bundled-skills/circleci/SKILL.md +475 -0
  63. package/bundled-skills/cis-benchmarks/SKILL.md +150 -0
  64. package/bundled-skills/cloudflare-pages/SKILL.md +318 -0
  65. package/bundled-skills/cloudflare-r2/SKILL.md +353 -0
  66. package/bundled-skills/cloudflare-workers/SKILL.md +415 -0
  67. package/bundled-skills/cloudflare-zero-trust/SKILL.md +361 -0
  68. package/bundled-skills/cloudformation/SKILL.md +461 -0
  69. package/bundled-skills/constraint-driven-development/SKILL.md +335 -0
  70. package/bundled-skills/constraint-driven-development/references/floor-guard.md +99 -0
  71. package/bundled-skills/container-hardening/SKILL.md +126 -0
  72. package/bundled-skills/container-registries/SKILL.md +435 -0
  73. package/bundled-skills/container-scanning/SKILL.md +416 -0
  74. package/bundled-skills/convex-backend/SKILL.md +338 -0
  75. package/bundled-skills/dast-scanning/SKILL.md +437 -0
  76. package/bundled-skills/database-backups/SKILL.md +425 -0
  77. package/bundled-skills/datadog/SKILL.md +487 -0
  78. package/bundled-skills/dependency-scanning/SKILL.md +457 -0
  79. package/bundled-skills/devcontainers-nix/SKILL.md +416 -0
  80. package/bundled-skills/disaster-recovery/SKILL.md +374 -0
  81. package/bundled-skills/disaster-recovery/references/details.md +219 -0
  82. package/bundled-skills/dns-management/SKILL.md +375 -0
  83. package/bundled-skills/docker-compose/SKILL.md +482 -0
  84. package/bundled-skills/docker-management/SKILL.md +426 -0
  85. package/bundled-skills/ebpf-observability/SKILL.md +436 -0
  86. package/bundled-skills/ebpf-observability/references/details.md +542 -0
  87. package/bundled-skills/elk-stack/SKILL.md +487 -0
  88. package/bundled-skills/enterprise-vpn-attack/SKILL.md +395 -0
  89. package/bundled-skills/evidence-hygiene/SKILL.md +404 -0
  90. package/bundled-skills/feature-flags/SKILL.md +426 -0
  91. package/bundled-skills/feature-flags/references/details.md +86 -0
  92. package/bundled-skills/fedramp-compliance/SKILL.md +453 -0
  93. package/bundled-skills/firebase-app-platform/SKILL.md +381 -0
  94. package/bundled-skills/firewall-config/SKILL.md +479 -0
  95. package/bundled-skills/gcp-audit-logs/SKILL.md +452 -0
  96. package/bundled-skills/gcp-audit-logs/references/details.md +56 -0
  97. package/bundled-skills/gcp-cloud-functions/SKILL.md +284 -0
  98. package/bundled-skills/gcp-cloud-sql/SKILL.md +277 -0
  99. package/bundled-skills/gcp-compute/SKILL.md +319 -0
  100. package/bundled-skills/gcp-gke/SKILL.md +307 -0
  101. package/bundled-skills/gcp-networking/SKILL.md +293 -0
  102. package/bundled-skills/gcp-secret-manager/SKILL.md +421 -0
  103. package/bundled-skills/gcp-secret-manager/references/details.md +131 -0
  104. package/bundled-skills/gdpr-compliance/SKILL.md +451 -0
  105. package/bundled-skills/gdpr-compliance/references/details.md +145 -0
  106. package/bundled-skills/geo-audit/SKILL.md +368 -0
  107. package/bundled-skills/geo-brand-mentions/SKILL.md +68 -0
  108. package/bundled-skills/geo-brand-mentions/references/details.md +471 -0
  109. package/bundled-skills/geo-citability/SKILL.md +350 -0
  110. package/bundled-skills/geo-compare/SKILL.md +340 -0
  111. package/bundled-skills/geo-content/SKILL.md +383 -0
  112. package/bundled-skills/geo-crawlers/SKILL.md +408 -0
  113. package/bundled-skills/geo-llmstxt/SKILL.md +464 -0
  114. package/bundled-skills/geo-platform-optimizer/SKILL.md +314 -0
  115. package/bundled-skills/geo-proposal/SKILL.md +378 -0
  116. package/bundled-skills/geo-prospect/SKILL.md +225 -0
  117. package/bundled-skills/geo-report/SKILL.md +436 -0
  118. package/bundled-skills/geo-report-pdf/SKILL.md +157 -0
  119. package/bundled-skills/geo-schema/SKILL.md +408 -0
  120. package/bundled-skills/geo-technical/SKILL.md +78 -0
  121. package/bundled-skills/geo-technical/references/details.md +543 -0
  122. package/bundled-skills/git-workflow/SKILL.md +460 -0
  123. package/bundled-skills/github-actions/SKILL.md +368 -0
  124. package/bundled-skills/gitlab-ci/SKILL.md +340 -0
  125. package/bundled-skills/gpu-kubernetes-operations/SKILL.md +468 -0
  126. package/bundled-skills/gpu-server-management/SKILL.md +236 -0
  127. package/bundled-skills/hashicorp-vault/SKILL.md +408 -0
  128. package/bundled-skills/helm-charts/SKILL.md +469 -0
  129. package/bundled-skills/hipaa-compliance/SKILL.md +451 -0
  130. package/bundled-skills/hunt-aspnet/SKILL.md +321 -0
  131. package/bundled-skills/hunt-ato/SKILL.md +184 -0
  132. package/bundled-skills/hunt-auth-bypass/SKILL.md +426 -0
  133. package/bundled-skills/hunt-auth-bypass/references/details.md +80 -0
  134. package/bundled-skills/hunt-brute-force/SKILL.md +341 -0
  135. package/bundled-skills/hunt-business-logic/SKILL.md +281 -0
  136. package/bundled-skills/hunt-cache-poison/SKILL.md +382 -0
  137. package/bundled-skills/hunt-captcha-bypass/SKILL.md +136 -0
  138. package/bundled-skills/hunt-cicd/SKILL.md +311 -0
  139. package/bundled-skills/hunt-clickjacking/SKILL.md +110 -0
  140. package/bundled-skills/hunt-cors/SKILL.md +335 -0
  141. package/bundled-skills/hunt-dom/SKILL.md +323 -0
  142. package/bundled-skills/hunt-exceptional-conditions/SKILL.md +111 -0
  143. package/bundled-skills/hunt-file-upload/SKILL.md +202 -0
  144. package/bundled-skills/hunt-fintech-graphql/SKILL.md +289 -0
  145. package/bundled-skills/hunt-forgot-password/SKILL.md +114 -0
  146. package/bundled-skills/hunt-grpc/SKILL.md +317 -0
  147. package/bundled-skills/hunt-host-header/SKILL.md +309 -0
  148. package/bundled-skills/hunt-html-injection/SKILL.md +106 -0
  149. package/bundled-skills/hunt-http-smuggling/SKILL.md +129 -0
  150. package/bundled-skills/hunt-http-smuggling/references/phase2h-smuggling-cachepoison.md +177 -0
  151. package/bundled-skills/hunt-idor/SKILL.md +434 -0
  152. package/bundled-skills/hunt-jwt-crypto/SKILL.md +221 -0
  153. package/bundled-skills/hunt-k8s/SKILL.md +337 -0
  154. package/bundled-skills/hunt-laravel/SKILL.md +255 -0
  155. package/bundled-skills/hunt-ldap/SKILL.md +351 -0
  156. package/bundled-skills/hunt-lfi/SKILL.md +311 -0
  157. package/bundled-skills/hunt-llm-ai/SKILL.md +289 -0
  158. package/bundled-skills/hunt-mfa-bypass/SKILL.md +177 -0
  159. package/bundled-skills/hunt-misc/SKILL.md +378 -0
  160. package/bundled-skills/hunt-nextjs/SKILL.md +299 -0
  161. package/bundled-skills/hunt-nodejs/SKILL.md +263 -0
  162. package/bundled-skills/hunt-nosqli/SKILL.md +210 -0
  163. package/bundled-skills/hunt-ntlm-info/SKILL.md +314 -0
  164. package/bundled-skills/hunt-oauth/SKILL.md +459 -0
  165. package/bundled-skills/hunt-open-redirect/SKILL.md +223 -0
  166. package/bundled-skills/hunt-race-condition/SKILL.md +381 -0
  167. package/bundled-skills/hunt-race-condition/references/details.md +159 -0
  168. package/bundled-skills/hunt-rag-vector/SKILL.md +212 -0
  169. package/bundled-skills/hunt-rce/SKILL.md +444 -0
  170. package/bundled-skills/hunt-rce/references/details.md +110 -0
  171. package/bundled-skills/hunt-saml/SKILL.md +156 -0
  172. package/bundled-skills/hunt-session/SKILL.md +342 -0
  173. package/bundled-skills/hunt-shadow-api/SKILL.md +198 -0
  174. package/bundled-skills/hunt-source-leak/SKILL.md +345 -0
  175. package/bundled-skills/hunt-spa-api/SKILL.md +163 -0
  176. package/bundled-skills/hunt-springboot/SKILL.md +285 -0
  177. package/bundled-skills/hunt-sqli/SKILL.md +466 -0
  178. package/bundled-skills/hunt-ssrf/SKILL.md +396 -0
  179. package/bundled-skills/hunt-ssrf/references/details.md +179 -0
  180. package/bundled-skills/hunt-ssti/SKILL.md +163 -0
  181. package/bundled-skills/hunt-subdomain/SKILL.md +379 -0
  182. package/bundled-skills/hunt-tls-network/SKILL.md +399 -0
  183. package/bundled-skills/hunt-xxe/SKILL.md +466 -0
  184. package/bundled-skills/i-have-adhd/SKILL.md +170 -0
  185. package/bundled-skills/identity-access-management/SKILL.md +382 -0
  186. package/bundled-skills/identity-access-management/references/details.md +524 -0
  187. package/bundled-skills/incident-management/SKILL.md +484 -0
  188. package/bundled-skills/incident-response/SKILL.md +448 -0
  189. package/bundled-skills/incident-response/references/details.md +113 -0
  190. package/bundled-skills/interview-me/SKILL.md +248 -0
  191. package/bundled-skills/iso27001-compliance/SKILL.md +460 -0
  192. package/bundled-skills/jenkins/SKILL.md +462 -0
  193. package/bundled-skills/jev-use/SKILL.md +158 -0
  194. package/bundled-skills/kubernetes-hardening/SKILL.md +154 -0
  195. package/bundled-skills/kubernetes-ops/SKILL.md +449 -0
  196. package/bundled-skills/kubernetes-ops/references/details.md +108 -0
  197. package/bundled-skills/kustomize/SKILL.md +478 -0
  198. package/bundled-skills/linux-administration/SKILL.md +367 -0
  199. package/bundled-skills/linux-hardening/SKILL.md +154 -0
  200. package/bundled-skills/llm-app-security/SKILL.md +389 -0
  201. package/bundled-skills/llm-app-security/references/details.md +674 -0
  202. package/bundled-skills/llm-caching/SKILL.md +334 -0
  203. package/bundled-skills/llm-cost-optimization/SKILL.md +311 -0
  204. package/bundled-skills/llm-fine-tuning/SKILL.md +329 -0
  205. package/bundled-skills/llm-gateway/SKILL.md +282 -0
  206. package/bundled-skills/llm-inference-scaling/SKILL.md +286 -0
  207. package/bundled-skills/llmops-platform-engineering/SKILL.md +472 -0
  208. package/bundled-skills/load-balancing/SKILL.md +403 -0
  209. package/bundled-skills/loki-logging/SKILL.md +479 -0
  210. package/bundled-skills/m365-entra-attack/SKILL.md +423 -0
  211. package/bundled-skills/mac-mini-llm-lab/SKILL.md +350 -0
  212. package/bundled-skills/mcp-server-security/SKILL.md +356 -0
  213. package/bundled-skills/mcp-server-security/references/details.md +745 -0
  214. package/bundled-skills/mdm-device-management/SKILL.md +404 -0
  215. package/bundled-skills/mdm-device-management/references/details.md +410 -0
  216. package/bundled-skills/meme-coin-audit/SKILL.md +402 -0
  217. package/bundled-skills/mid-engagement-ir-detection/SKILL.md +377 -0
  218. package/bundled-skills/model-registry-governance/SKILL.md +452 -0
  219. package/bundled-skills/model-serving-kubernetes/SKILL.md +339 -0
  220. package/bundled-skills/model-supply-chain-security/SKILL.md +427 -0
  221. package/bundled-skills/mongodb/SKILL.md +436 -0
  222. package/bundled-skills/multi-tenant-llm-hosting/SKILL.md +435 -0
  223. package/bundled-skills/multi-tenant-llm-hosting/references/details.md +211 -0
  224. package/bundled-skills/mysql/SKILL.md +390 -0
  225. package/bundled-skills/new-relic/SKILL.md +472 -0
  226. package/bundled-skills/nfs-storage/SKILL.md +356 -0
  227. package/bundled-skills/object-storage/SKILL.md +378 -0
  228. package/bundled-skills/offensive-osint/SKILL.md +443 -0
  229. package/bundled-skills/okta-attack/SKILL.md +436 -0
  230. package/bundled-skills/ollama-stack/SKILL.md +379 -0
  231. package/bundled-skills/openclaw-deployment-hardening/SKILL.md +135 -0
  232. package/bundled-skills/openclaw-local-mac-mini/SKILL.md +426 -0
  233. package/bundled-skills/openclaw-local-mac-mini/references/details.md +221 -0
  234. package/bundled-skills/openclaw-security-hardening/SKILL.md +135 -0
  235. package/bundled-skills/openshift/SKILL.md +485 -0
  236. package/bundled-skills/opentelemetry/SKILL.md +438 -0
  237. package/bundled-skills/opentelemetry/references/details.md +78 -0
  238. package/bundled-skills/opentofu-migration/SKILL.md +349 -0
  239. package/bundled-skills/osint-methodology/SKILL.md +460 -0
  240. package/bundled-skills/osint-methodology/references/details.md +1350 -0
  241. package/bundled-skills/pci-dss-compliance/SKILL.md +446 -0
  242. package/bundled-skills/penetration-testing/SKILL.md +152 -0
  243. package/bundled-skills/performance-tuning/SKILL.md +381 -0
  244. package/bundled-skills/planetscale/SKILL.md +297 -0
  245. package/bundled-skills/platform-engineering/SKILL.md +348 -0
  246. package/bundled-skills/platform-engineering/references/details.md +944 -0
  247. package/bundled-skills/podman/SKILL.md +405 -0
  248. package/bundled-skills/policy-as-code/SKILL.md +434 -0
  249. package/bundled-skills/policy-as-code/references/details.md +204 -0
  250. package/bundled-skills/postgresql-devsec/SKILL.md +378 -0
  251. package/bundled-skills/prometheus-grafana/SKILL.md +469 -0
  252. package/bundled-skills/prompt-injection-defense/SKILL.md +483 -0
  253. package/bundled-skills/rag-infrastructure/SKILL.md +269 -0
  254. package/bundled-skills/rag-observability-evals/SKILL.md +444 -0
  255. package/bundled-skills/rag-observability-evals/references/details.md +92 -0
  256. package/bundled-skills/recon-scope-triage/SKILL.md +128 -0
  257. package/bundled-skills/redis/SKILL.md +421 -0
  258. package/bundled-skills/redteam-report-template/SKILL.md +370 -0
  259. package/bundled-skills/report-writing/SKILL.md +426 -0
  260. package/bundled-skills/report-writing/references/details.md +187 -0
  261. package/bundled-skills/reverse-proxy/SKILL.md +420 -0
  262. package/bundled-skills/runbook-creation/SKILL.md +438 -0
  263. package/bundled-skills/runbook-creation/references/details.md +71 -0
  264. package/bundled-skills/saas-security-posture/SKILL.md +415 -0
  265. package/bundled-skills/sast-scanning/SKILL.md +444 -0
  266. package/bundled-skills/sbom-supply-chain/SKILL.md +433 -0
  267. package/bundled-skills/security-arsenal/SKILL.md +446 -0
  268. package/bundled-skills/security-arsenal/references/details.md +540 -0
  269. package/bundled-skills/security-automation/SKILL.md +146 -0
  270. package/bundled-skills/semantic-versioning/SKILL.md +434 -0
  271. package/bundled-skills/semantic-versioning/references/details.md +83 -0
  272. package/bundled-skills/service-mesh/SKILL.md +422 -0
  273. package/bundled-skills/soc2-compliance/SKILL.md +409 -0
  274. package/bundled-skills/sops-encryption/SKILL.md +124 -0
  275. package/bundled-skills/sre-dashboards/SKILL.md +143 -0
  276. package/bundled-skills/ssh-configuration/SKILL.md +324 -0
  277. package/bundled-skills/ssl-tls-management/SKILL.md +428 -0
  278. package/bundled-skills/ssl-tls-management/references/details.md +99 -0
  279. package/bundled-skills/startup-it-troubleshooting/SKILL.md +415 -0
  280. package/bundled-skills/supply-chain-attack-recon/SKILL.md +453 -0
  281. package/bundled-skills/supply-chain-attack-recon/references/details.md +258 -0
  282. package/bundled-skills/systemd-services/SKILL.md +379 -0
  283. package/bundled-skills/terraform-aws/SKILL.md +125 -0
  284. package/bundled-skills/terraform-azure/SKILL.md +415 -0
  285. package/bundled-skills/terraform-azure/references/details.md +231 -0
  286. package/bundled-skills/terraform-gcp/SKILL.md +369 -0
  287. package/bundled-skills/threat-modeling/SKILL.md +487 -0
  288. package/bundled-skills/user-management/SKILL.md +383 -0
  289. package/bundled-skills/using-agent-skills/SKILL.md +220 -0
  290. package/bundled-skills/vector-database-ops/SKILL.md +300 -0
  291. package/bundled-skills/vendor-management/SKILL.md +439 -0
  292. package/bundled-skills/vendor-management/references/details.md +109 -0
  293. package/bundled-skills/vercel-deployments/SKILL.md +296 -0
  294. package/bundled-skills/vllm-server/SKILL.md +236 -0
  295. package/bundled-skills/vmware-vcenter-attack/SKILL.md +412 -0
  296. package/bundled-skills/vpn-setup/SKILL.md +452 -0
  297. package/bundled-skills/vulnerability-scanning/SKILL.md +448 -0
  298. package/bundled-skills/waf-setup/SKILL.md +354 -0
  299. package/bundled-skills/waf-setup/references/details.md +211 -0
  300. package/bundled-skills/web2-recon/SKILL.md +440 -0
  301. package/bundled-skills/web2-recon/references/details.md +319 -0
  302. package/bundled-skills/web3-audit/SKILL.md +445 -0
  303. package/bundled-skills/web3-audit/references/details.md +224 -0
  304. package/bundled-skills/windows-hardening/SKILL.md +454 -0
  305. package/bundled-skills/windows-hardening/references/details.md +204 -0
  306. package/bundled-skills/windows-server/SKILL.md +318 -0
  307. package/bundled-skills/zero-trust/SKILL.md +461 -0
  308. package/package.json +1 -1
  309. package/skills_index.json +6943 -323
@@ -0,0 +1,472 @@
1
+ ---
2
+ name: llmops-platform-engineering
3
+ description: Build production LLMOps platforms with CI/CD, model promotion workflows,
4
+ evaluation gates, rollback, and governance across cloud and self-hosted inference.
5
+ category: devops
6
+ risk: critical
7
+ source: https://github.com/BagelHole/DevOps-Security-Agent-Skills
8
+ source_repo: BagelHole/DevOps-Security-Agent-Skills
9
+ source_type: community
10
+ date_added: '2026-09-20'
11
+ license: MIT
12
+ license_source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
13
+ compatibility: Requires the relevant platform CLIs (kubectl, helm, terraform, git,
14
+ CI runners) and authorized access to the target environment. Docs-only; helper scripts
15
+ and templates not bundled.
16
+ metadata:
17
+ author: devops-skills
18
+ version: '1.0'
19
+ ---
20
+
21
+ # LLMOps Platform Engineering
22
+
23
+ Design and operate an internal LLM platform that supports rapid experimentation without compromising reliability, cost, or compliance.
24
+
25
+ ## When to Use This Skill
26
+
27
+ - Building an internal platform for teams to deploy and manage LLM-powered features
28
+ - Designing CI/CD pipelines that include model evaluation gates
29
+ - Setting up A/B testing infrastructure for model versions
30
+ - Creating Kubernetes-based model serving infrastructure
31
+ - Establishing governance workflows for model promotion
32
+
33
+ ## Prerequisites
34
+
35
+ - Kubernetes cluster with GPU node pools (or cloud inference API access)
36
+ - Container registry (Harbor, ECR, GCR, or ACR)
37
+ - CI/CD system (GitHub Actions, GitLab CI, or Argo Workflows)
38
+ - Observability stack (Prometheus + Grafana + OpenTelemetry)
39
+ - Model registry (MLflow or custom metadata store)
40
+
41
+ ## Outcomes
42
+
43
+ - Standardized path from experiment to production
44
+ - Safe model rollout with quality and safety gates
45
+ - Repeatable infra modules for inference, vector DB, and observability
46
+ - Clear ownership model across platform, app, and security teams
47
+
48
+ ## Reference Architecture
49
+
50
+ 1. **Control Plane**: model registry, prompt/version catalog, policy checks, eval pipeline.
51
+ 2. **Data Plane**: inference gateway, vector database, cache, feature store.
52
+ 3. **Ops Plane**: telemetry, alerting, SLO dashboards, cost analytics.
53
+ 4. **Security Plane**: IAM boundaries, secret rotation, content filters, audit logs.
54
+
55
+ ## Model Promotion Pipeline
56
+
57
+ ```yaml
58
+ # .github/workflows/model-promotion.yaml
59
+ name: Model Promotion Pipeline
60
+ on:
61
+ workflow_dispatch:
62
+ inputs:
63
+ model_name:
64
+ description: "Model identifier"
65
+ required: true
66
+ model_version:
67
+ description: "Model version to promote"
68
+ required: true
69
+ target_env:
70
+ description: "Target environment"
71
+ required: true
72
+ type: choice
73
+ options: [staging, production]
74
+
75
+ jobs:
76
+ evaluate:
77
+ runs-on: ubuntu-latest
78
+ steps:
79
+ - uses: actions/checkout@v4
80
+
81
+ - name: Run quality evaluation suite
82
+ run: |
83
+ python -m evals.run \
84
+ --model "${{ inputs.model_name }}:${{ inputs.model_version }}" \
85
+ --suite quality \
86
+ --output results/quality.json
87
+
88
+ - name: Run safety evaluation suite
89
+ run: |
90
+ python -m evals.run \
91
+ --model "${{ inputs.model_name }}:${{ inputs.model_version }}" \
92
+ --suite safety \
93
+ --output results/safety.json
94
+
95
+ - name: Run latency benchmark
96
+ run: |
97
+ python -m evals.benchmark \
98
+ --model "${{ inputs.model_name }}:${{ inputs.model_version }}" \
99
+ --concurrent-users 50 \
100
+ --duration 300 \
101
+ --output results/latency.json
102
+
103
+ - name: Gate check - quality
104
+ run: |
105
+ python -m evals.gate_check \
106
+ --results results/quality.json \
107
+ --threshold-file thresholds/quality.yaml
108
+
109
+ - name: Gate check - safety
110
+ run: |
111
+ python -m evals.gate_check \
112
+ --results results/safety.json \
113
+ --threshold-file thresholds/safety.yaml
114
+
115
+ - name: Gate check - latency
116
+ run: |
117
+ python -m evals.gate_check \
118
+ --results results/latency.json \
119
+ --threshold-file thresholds/latency.yaml
120
+
121
+ - name: Upload eval evidence
122
+ uses: actions/upload-artifact@v4
123
+ with:
124
+ name: eval-results-${{ inputs.model_version }}
125
+ path: results/
126
+
127
+ approve:
128
+ needs: evaluate
129
+ runs-on: ubuntu-latest
130
+ environment: ${{ inputs.target_env }}
131
+ steps:
132
+ - name: Record approval
133
+ run: |
134
+ echo "Approved by: ${{ github.actor }}"
135
+ echo "Model: ${{ inputs.model_name }}:${{ inputs.model_version }}"
136
+ echo "Target: ${{ inputs.target_env }}"
137
+ echo "Time: $(date -u +%Y-%m-%dT%H:%M:%SZ)"
138
+
139
+ deploy:
140
+ needs: approve
141
+ runs-on: ubuntu-latest
142
+ steps:
143
+ - uses: actions/checkout@v4
144
+
145
+ - name: Deploy canary
146
+ run: |
147
+ kubectl set image deployment/${{ inputs.model_name }}-canary \
148
+ model=${{ inputs.model_name }}:${{ inputs.model_version }} \
149
+ -n ai-${{ inputs.target_env }}
150
+
151
+ - name: Wait for canary validation (15 min)
152
+ run: |
153
+ python -m canary.validate \
154
+ --deployment ${{ inputs.model_name }}-canary \
155
+ --namespace ai-${{ inputs.target_env }} \
156
+ --duration 900 \
157
+ --quality-threshold 0.85 \
158
+ --error-rate-threshold 0.02
159
+
160
+ - name: Promote to full rollout
161
+ run: |
162
+ kubectl set image deployment/${{ inputs.model_name }} \
163
+ model=${{ inputs.model_name }}:${{ inputs.model_version }} \
164
+ -n ai-${{ inputs.target_env }}
165
+ kubectl rollout status deployment/${{ inputs.model_name }} \
166
+ -n ai-${{ inputs.target_env }} --timeout=300s
167
+ ```
168
+
169
+ ## Evaluation Gate Thresholds
170
+
171
+ ```yaml
172
+ # thresholds/quality.yaml
173
+ gates:
174
+ groundedness:
175
+ metric: groundedness_score
176
+ min: 0.85
177
+ comparison: gte
178
+ task_success:
179
+ metric: task_success_rate
180
+ min: 0.90
181
+ comparison: gte
182
+ hallucination:
183
+ metric: hallucination_rate
184
+ max: 0.08
185
+ comparison: lte
186
+ regression:
187
+ metric: quality_delta_vs_baseline
188
+ min: -0.02
189
+ comparison: gte
190
+ description: "Must not regress more than 2% vs current production"
191
+
192
+ # thresholds/latency.yaml
193
+ gates:
194
+ p50_latency:
195
+ metric: latency_p50_ms
196
+ max: 800
197
+ comparison: lte
198
+ p95_latency:
199
+ metric: latency_p95_ms
200
+ max: 2000
201
+ comparison: lte
202
+ p99_latency:
203
+ metric: latency_p99_ms
204
+ max: 5000
205
+ comparison: lte
206
+ throughput:
207
+ metric: requests_per_second
208
+ min: 50
209
+ comparison: gte
210
+ ```
211
+
212
+ ## A/B Testing Configuration
213
+
214
+ ```yaml
215
+ # ab-test-config.yaml
216
+ apiVersion: gateway.ai/v1
217
+ kind: ABTest
218
+ metadata:
219
+ name: model-comparison-q1
220
+ namespace: ai-production
221
+ spec:
222
+ duration: 7d
223
+ traffic_split:
224
+ control:
225
+ model: gpt-4o-2024-08-06
226
+ weight: 70
227
+ treatment:
228
+ model: gpt-4o-2025-01-15
229
+ weight: 30
230
+ metrics:
231
+ primary:
232
+ - task_success_rate
233
+ - user_satisfaction_score
234
+ secondary:
235
+ - latency_p95
236
+ - cost_per_request
237
+ - hallucination_rate
238
+ guardrails:
239
+ auto_rollback_if:
240
+ - metric: task_success_rate
241
+ threshold: 0.80
242
+ window: 1h
243
+ - metric: hallucination_rate
244
+ threshold: 0.15
245
+ window: 30m
246
+ assignment:
247
+ strategy: sticky_user
248
+ hash_key: user_id
249
+ ```
250
+
251
+ ## Kubernetes Model Serving Deployment
252
+
253
+ ```yaml
254
+ # model-serving-deployment.yaml
255
+ apiVersion: apps/v1
256
+ kind: Deployment
257
+ metadata:
258
+ name: llm-inference
259
+ namespace: ai-production
260
+ labels:
261
+ app: llm-inference
262
+ model: gpt-4o
263
+ version: "2025-01"
264
+ spec:
265
+ replicas: 3
266
+ strategy:
267
+ type: RollingUpdate
268
+ rollingUpdate:
269
+ maxSurge: 1
270
+ maxUnavailable: 0
271
+ selector:
272
+ matchLabels:
273
+ app: llm-inference
274
+ template:
275
+ metadata:
276
+ labels:
277
+ app: llm-inference
278
+ model: gpt-4o
279
+ annotations:
280
+ prometheus.io/scrape: "true"
281
+ prometheus.io/port: "8080"
282
+ prometheus.io/path: "/metrics"
283
+ spec:
284
+ topologySpreadConstraints:
285
+ - maxSkew: 1
286
+ topologyKey: topology.kubernetes.io/zone
287
+ whenUnsatisfiable: DoNotSchedule
288
+ labelSelector:
289
+ matchLabels:
290
+ app: llm-inference
291
+ containers:
292
+ - name: model
293
+ image: registry.internal/vllm-server:0.4.1
294
+ args:
295
+ - "--model=/models/current"
296
+ - "--tensor-parallel-size=1"
297
+ - "--max-model-len=8192"
298
+ - "--gpu-memory-utilization=0.90"
299
+ ports:
300
+ - containerPort: 8000
301
+ name: inference
302
+ - containerPort: 8080
303
+ name: metrics
304
+ resources:
305
+ requests:
306
+ cpu: "4"
307
+ memory: "16Gi"
308
+ nvidia.com/gpu: "1"
309
+ limits:
310
+ cpu: "8"
311
+ memory: "32Gi"
312
+ nvidia.com/gpu: "1"
313
+ readinessProbe:
314
+ httpGet:
315
+ path: /health
316
+ port: 8000
317
+ initialDelaySeconds: 60
318
+ periodSeconds: 10
319
+ livenessProbe:
320
+ httpGet:
321
+ path: /health
322
+ port: 8000
323
+ initialDelaySeconds: 120
324
+ periodSeconds: 30
325
+ volumeMounts:
326
+ - name: model-weights
327
+ mountPath: /models
328
+ readOnly: true
329
+ - name: config
330
+ mountPath: /etc/vllm
331
+ volumes:
332
+ - name: model-weights
333
+ persistentVolumeClaim:
334
+ claimName: model-weights-pvc
335
+ - name: config
336
+ configMap:
337
+ name: vllm-config
338
+ tolerations:
339
+ - key: nvidia.com/gpu
340
+ operator: Exists
341
+ effect: NoSchedule
342
+ nodeSelector:
343
+ gpu-type: a100
344
+ ---
345
+ apiVersion: v1
346
+ kind: Service
347
+ metadata:
348
+ name: llm-inference
349
+ namespace: ai-production
350
+ spec:
351
+ selector:
352
+ app: llm-inference
353
+ ports:
354
+ - name: inference
355
+ port: 8000
356
+ targetPort: 8000
357
+ - name: metrics
358
+ port: 8080
359
+ targetPort: 8080
360
+ ---
361
+ apiVersion: autoscaling/v2
362
+ kind: HorizontalPodAutoscaler
363
+ metadata:
364
+ name: llm-inference-hpa
365
+ namespace: ai-production
366
+ spec:
367
+ scaleTargetRef:
368
+ apiVersion: apps/v1
369
+ kind: Deployment
370
+ name: llm-inference
371
+ minReplicas: 2
372
+ maxReplicas: 10
373
+ metrics:
374
+ - type: Pods
375
+ pods:
376
+ metric:
377
+ name: llm_queue_depth
378
+ target:
379
+ type: AverageValue
380
+ averageValue: "5"
381
+ - type: Pods
382
+ pods:
383
+ metric:
384
+ name: gpu_utilization_percent
385
+ target:
386
+ type: AverageValue
387
+ averageValue: "75"
388
+ behavior:
389
+ scaleUp:
390
+ stabilizationWindowSeconds: 60
391
+ policies:
392
+ - type: Pods
393
+ value: 2
394
+ periodSeconds: 120
395
+ scaleDown:
396
+ stabilizationWindowSeconds: 300
397
+ policies:
398
+ - type: Pods
399
+ value: 1
400
+ periodSeconds: 300
401
+ ```
402
+
403
+ ## CI/CD Design for AI Services
404
+
405
+ - Build immutable containers with pinned dependencies and model hashes.
406
+ - Use environment promotion: `dev -> stage -> prod`.
407
+ - Fail deployment if:
408
+ - regression evals drop below baseline,
409
+ - safety tests exceed risk threshold,
410
+ - p95 latency exceeds SLO budget.
411
+ - Store deployment evidence for audits (commit SHA, eval report, approver).
412
+
413
+ ## Operational SLOs
414
+
415
+ | Signal | Target | Measurement Window |
416
+ |--------|--------|--------------------|
417
+ | Availability | 99.9% | 30-day rolling |
418
+ | p95 Latency | < 1200ms | 5-min buckets |
419
+ | Cost per request | < $0.05 | 1-hour average |
420
+ | Task success rate | > 90% | 24-hour rolling |
421
+ | Groundedness | > 85% | 24-hour rolling |
422
+
423
+ ## Platform Guardrails
424
+
425
+ - Enforce tenant quotas and model allow-lists.
426
+ - Require structured output contracts for automation paths.
427
+ - Default to low-risk model settings for critical workflows.
428
+ - Disable unconstrained tool execution in production.
429
+
430
+ ## Tooling Stack (Example)
431
+
432
+ | Layer | Tools |
433
+ |-------|-------|
434
+ | Orchestration | Argo Workflows, GitHub Actions, Airflow |
435
+ | Model Registry | MLflow, custom metadata DB |
436
+ | Gateway | LiteLLM, Envoy-based API gateway |
437
+ | Observability | OpenTelemetry + Prometheus + Grafana + Langfuse |
438
+ | Policy | OPA/Rego for deployment and runtime checks |
439
+ | Evaluation | RAGAS, custom eval harness, Promptfoo |
440
+ | Serving | vLLM, TGI, Triton Inference Server |
441
+
442
+ ## Troubleshooting
443
+
444
+ | Issue | Diagnosis | Resolution |
445
+ |-------|-----------|------------|
446
+ | Canary fails quality gate | Compare eval results with baseline | Adjust model config or revert version |
447
+ | Deployment stuck in rollout | Check pod events and resource quotas | Fix resource limits or node availability |
448
+ | A/B test shows no significant difference | Verify traffic split and sample size | Extend test duration or increase treatment weight |
449
+ | Model cold start too slow | Large model weight download | Use pre-cached PVCs or init containers |
450
+ | Eval pipeline flaky | Non-deterministic model outputs | Set temperature=0 for evals, increase sample size |
451
+
452
+ ## Related Skills
453
+
454
+ - ai-pipeline-orchestration (`ai-pipeline-orchestration`) - Orchestrate ingestion and inference workflows
455
+ - agent-evals (`agent-evals`) - Build evaluation gates for releases
456
+ - llm-gateway (`llm-gateway`) - Route and control LLM traffic
457
+ - model-registry-governance (`model-registry-governance`) - Model lifecycle and approval workflows
458
+ - ai-sre-incident-response (`ai-sre-incident-response`) - AI-specific incident response
459
+
460
+ ## Limitations
461
+
462
+ - Guidance executes against real environments: confirm target, blast radius, and rollback plan before applying anything.
463
+ - Never deploy to production without explicit approval. Docs-only import: upstream scripts and templates not bundled.
464
+
465
+ ### Example
466
+
467
+ ```bash
468
+ git status && git diff --stat
469
+ kubectl diff -f manifest.yaml
470
+ ```
471
+
472
+ > Adapted from [BagelHole/DevOps-Security-Agent-Skills](https://github.com/BagelHole/DevOps-Security-Agent-Skills) (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.