opencode-skills-collection 4.0.68 → 4.0.69

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (309) hide show
  1. package/bundled-skills/.antigravity-install-manifest.json +266 -1
  2. package/bundled-skills/access-review/SKILL.md +394 -0
  3. package/bundled-skills/access-review/references/details.md +121 -0
  4. package/bundled-skills/agent-evals/SKILL.md +420 -0
  5. package/bundled-skills/agent-observability/SKILL.md +346 -0
  6. package/bundled-skills/agent-observability/references/details.md +786 -0
  7. package/bundled-skills/ai-agent-security/SKILL.md +393 -0
  8. package/bundled-skills/ai-agent-security/references/details.md +912 -0
  9. package/bundled-skills/ai-coding-agent-guardrails/SKILL.md +442 -0
  10. package/bundled-skills/ai-coding-agent-guardrails/references/details.md +753 -0
  11. package/bundled-skills/ai-inference-service-mesh/SKILL.md +449 -0
  12. package/bundled-skills/ai-pipeline-orchestration/SKILL.md +287 -0
  13. package/bundled-skills/ai-red-teaming/SKILL.md +409 -0
  14. package/bundled-skills/ai-security-hardening/SKILL.md +343 -0
  15. package/bundled-skills/ai-sre-incident-response/SKILL.md +336 -0
  16. package/bundled-skills/alerting-oncall/SKILL.md +458 -0
  17. package/bundled-skills/alerting-oncall/references/details.md +84 -0
  18. package/bundled-skills/apk-redteam-pipeline/SKILL.md +446 -0
  19. package/bundled-skills/argocd-gitops/SKILL.md +469 -0
  20. package/bundled-skills/arm-templates/SKILL.md +438 -0
  21. package/bundled-skills/arm-templates/references/details.md +64 -0
  22. package/bundled-skills/asset-inventory/SKILL.md +412 -0
  23. package/bundled-skills/asset-inventory/references/details.md +127 -0
  24. package/bundled-skills/audit-logging/SKILL.md +476 -0
  25. package/bundled-skills/aws-cloudtrail/SKILL.md +486 -0
  26. package/bundled-skills/aws-cost-optimization/SKILL.md +331 -0
  27. package/bundled-skills/aws-ec2/SKILL.md +426 -0
  28. package/bundled-skills/aws-ecs-fargate/SKILL.md +388 -0
  29. package/bundled-skills/aws-iam/SKILL.md +463 -0
  30. package/bundled-skills/aws-lambda/SKILL.md +428 -0
  31. package/bundled-skills/aws-rds/SKILL.md +380 -0
  32. package/bundled-skills/aws-s3/SKILL.md +434 -0
  33. package/bundled-skills/aws-secrets-manager/SKILL.md +486 -0
  34. package/bundled-skills/aws-vpc/SKILL.md +436 -0
  35. package/bundled-skills/azure-ai-document-intelligence-ts/SKILL.md +1 -1
  36. package/bundled-skills/azure-aks/SKILL.md +423 -0
  37. package/bundled-skills/azure-devops/SKILL.md +457 -0
  38. package/bundled-skills/azure-functions-devsec/SKILL.md +436 -0
  39. package/bundled-skills/azure-keyvault/SKILL.md +455 -0
  40. package/bundled-skills/azure-keyvault/references/details.md +83 -0
  41. package/bundled-skills/azure-monitor-audit/SKILL.md +379 -0
  42. package/bundled-skills/azure-networking/SKILL.md +448 -0
  43. package/bundled-skills/azure-networking/references/details.md +135 -0
  44. package/bundled-skills/azure-sql/SKILL.md +413 -0
  45. package/bundled-skills/azure-sql/references/details.md +113 -0
  46. package/bundled-skills/azure-vms/SKILL.md +402 -0
  47. package/bundled-skills/azure-vms/references/details.md +134 -0
  48. package/bundled-skills/backup-recovery/SKILL.md +388 -0
  49. package/bundled-skills/bb-methodology/SKILL.md +451 -0
  50. package/bundled-skills/bb-methodology/references/details.md +120 -0
  51. package/bundled-skills/block-storage/SKILL.md +371 -0
  52. package/bundled-skills/blue-green-deploy/SKILL.md +453 -0
  53. package/bundled-skills/blue-green-deploy/references/details.md +90 -0
  54. package/bundled-skills/bug-bounty/SKILL.md +447 -0
  55. package/bundled-skills/bug-bounty/references/details.md +1316 -0
  56. package/bundled-skills/bugcrowd-reporting/SKILL.md +351 -0
  57. package/bundled-skills/business-continuity/SKILL.md +463 -0
  58. package/bundled-skills/career-ops/SKILL.md +186 -0
  59. package/bundled-skills/cdn-setup/SKILL.md +374 -0
  60. package/bundled-skills/change-management/SKILL.md +438 -0
  61. package/bundled-skills/change-management/references/details.md +105 -0
  62. package/bundled-skills/circleci/SKILL.md +475 -0
  63. package/bundled-skills/cis-benchmarks/SKILL.md +150 -0
  64. package/bundled-skills/cloudflare-pages/SKILL.md +318 -0
  65. package/bundled-skills/cloudflare-r2/SKILL.md +353 -0
  66. package/bundled-skills/cloudflare-workers/SKILL.md +415 -0
  67. package/bundled-skills/cloudflare-zero-trust/SKILL.md +361 -0
  68. package/bundled-skills/cloudformation/SKILL.md +461 -0
  69. package/bundled-skills/constraint-driven-development/SKILL.md +335 -0
  70. package/bundled-skills/constraint-driven-development/references/floor-guard.md +99 -0
  71. package/bundled-skills/container-hardening/SKILL.md +126 -0
  72. package/bundled-skills/container-registries/SKILL.md +435 -0
  73. package/bundled-skills/container-scanning/SKILL.md +416 -0
  74. package/bundled-skills/convex-backend/SKILL.md +338 -0
  75. package/bundled-skills/dast-scanning/SKILL.md +437 -0
  76. package/bundled-skills/database-backups/SKILL.md +425 -0
  77. package/bundled-skills/datadog/SKILL.md +487 -0
  78. package/bundled-skills/dependency-scanning/SKILL.md +457 -0
  79. package/bundled-skills/devcontainers-nix/SKILL.md +416 -0
  80. package/bundled-skills/disaster-recovery/SKILL.md +374 -0
  81. package/bundled-skills/disaster-recovery/references/details.md +219 -0
  82. package/bundled-skills/dns-management/SKILL.md +375 -0
  83. package/bundled-skills/docker-compose/SKILL.md +482 -0
  84. package/bundled-skills/docker-management/SKILL.md +426 -0
  85. package/bundled-skills/ebpf-observability/SKILL.md +436 -0
  86. package/bundled-skills/ebpf-observability/references/details.md +542 -0
  87. package/bundled-skills/elk-stack/SKILL.md +487 -0
  88. package/bundled-skills/enterprise-vpn-attack/SKILL.md +395 -0
  89. package/bundled-skills/evidence-hygiene/SKILL.md +404 -0
  90. package/bundled-skills/feature-flags/SKILL.md +426 -0
  91. package/bundled-skills/feature-flags/references/details.md +86 -0
  92. package/bundled-skills/fedramp-compliance/SKILL.md +453 -0
  93. package/bundled-skills/firebase-app-platform/SKILL.md +381 -0
  94. package/bundled-skills/firewall-config/SKILL.md +479 -0
  95. package/bundled-skills/gcp-audit-logs/SKILL.md +452 -0
  96. package/bundled-skills/gcp-audit-logs/references/details.md +56 -0
  97. package/bundled-skills/gcp-cloud-functions/SKILL.md +284 -0
  98. package/bundled-skills/gcp-cloud-sql/SKILL.md +277 -0
  99. package/bundled-skills/gcp-compute/SKILL.md +319 -0
  100. package/bundled-skills/gcp-gke/SKILL.md +307 -0
  101. package/bundled-skills/gcp-networking/SKILL.md +293 -0
  102. package/bundled-skills/gcp-secret-manager/SKILL.md +421 -0
  103. package/bundled-skills/gcp-secret-manager/references/details.md +131 -0
  104. package/bundled-skills/gdpr-compliance/SKILL.md +451 -0
  105. package/bundled-skills/gdpr-compliance/references/details.md +145 -0
  106. package/bundled-skills/geo-audit/SKILL.md +368 -0
  107. package/bundled-skills/geo-brand-mentions/SKILL.md +68 -0
  108. package/bundled-skills/geo-brand-mentions/references/details.md +471 -0
  109. package/bundled-skills/geo-citability/SKILL.md +350 -0
  110. package/bundled-skills/geo-compare/SKILL.md +340 -0
  111. package/bundled-skills/geo-content/SKILL.md +383 -0
  112. package/bundled-skills/geo-crawlers/SKILL.md +408 -0
  113. package/bundled-skills/geo-llmstxt/SKILL.md +464 -0
  114. package/bundled-skills/geo-platform-optimizer/SKILL.md +314 -0
  115. package/bundled-skills/geo-proposal/SKILL.md +378 -0
  116. package/bundled-skills/geo-prospect/SKILL.md +225 -0
  117. package/bundled-skills/geo-report/SKILL.md +436 -0
  118. package/bundled-skills/geo-report-pdf/SKILL.md +157 -0
  119. package/bundled-skills/geo-schema/SKILL.md +408 -0
  120. package/bundled-skills/geo-technical/SKILL.md +78 -0
  121. package/bundled-skills/geo-technical/references/details.md +543 -0
  122. package/bundled-skills/git-workflow/SKILL.md +460 -0
  123. package/bundled-skills/github-actions/SKILL.md +368 -0
  124. package/bundled-skills/gitlab-ci/SKILL.md +340 -0
  125. package/bundled-skills/gpu-kubernetes-operations/SKILL.md +468 -0
  126. package/bundled-skills/gpu-server-management/SKILL.md +236 -0
  127. package/bundled-skills/hashicorp-vault/SKILL.md +408 -0
  128. package/bundled-skills/helm-charts/SKILL.md +469 -0
  129. package/bundled-skills/hipaa-compliance/SKILL.md +451 -0
  130. package/bundled-skills/hunt-aspnet/SKILL.md +321 -0
  131. package/bundled-skills/hunt-ato/SKILL.md +184 -0
  132. package/bundled-skills/hunt-auth-bypass/SKILL.md +426 -0
  133. package/bundled-skills/hunt-auth-bypass/references/details.md +80 -0
  134. package/bundled-skills/hunt-brute-force/SKILL.md +341 -0
  135. package/bundled-skills/hunt-business-logic/SKILL.md +281 -0
  136. package/bundled-skills/hunt-cache-poison/SKILL.md +382 -0
  137. package/bundled-skills/hunt-captcha-bypass/SKILL.md +136 -0
  138. package/bundled-skills/hunt-cicd/SKILL.md +311 -0
  139. package/bundled-skills/hunt-clickjacking/SKILL.md +110 -0
  140. package/bundled-skills/hunt-cors/SKILL.md +335 -0
  141. package/bundled-skills/hunt-dom/SKILL.md +323 -0
  142. package/bundled-skills/hunt-exceptional-conditions/SKILL.md +111 -0
  143. package/bundled-skills/hunt-file-upload/SKILL.md +202 -0
  144. package/bundled-skills/hunt-fintech-graphql/SKILL.md +289 -0
  145. package/bundled-skills/hunt-forgot-password/SKILL.md +114 -0
  146. package/bundled-skills/hunt-grpc/SKILL.md +317 -0
  147. package/bundled-skills/hunt-host-header/SKILL.md +309 -0
  148. package/bundled-skills/hunt-html-injection/SKILL.md +106 -0
  149. package/bundled-skills/hunt-http-smuggling/SKILL.md +129 -0
  150. package/bundled-skills/hunt-http-smuggling/references/phase2h-smuggling-cachepoison.md +177 -0
  151. package/bundled-skills/hunt-idor/SKILL.md +434 -0
  152. package/bundled-skills/hunt-jwt-crypto/SKILL.md +221 -0
  153. package/bundled-skills/hunt-k8s/SKILL.md +337 -0
  154. package/bundled-skills/hunt-laravel/SKILL.md +255 -0
  155. package/bundled-skills/hunt-ldap/SKILL.md +351 -0
  156. package/bundled-skills/hunt-lfi/SKILL.md +311 -0
  157. package/bundled-skills/hunt-llm-ai/SKILL.md +289 -0
  158. package/bundled-skills/hunt-mfa-bypass/SKILL.md +177 -0
  159. package/bundled-skills/hunt-misc/SKILL.md +378 -0
  160. package/bundled-skills/hunt-nextjs/SKILL.md +299 -0
  161. package/bundled-skills/hunt-nodejs/SKILL.md +263 -0
  162. package/bundled-skills/hunt-nosqli/SKILL.md +210 -0
  163. package/bundled-skills/hunt-ntlm-info/SKILL.md +314 -0
  164. package/bundled-skills/hunt-oauth/SKILL.md +459 -0
  165. package/bundled-skills/hunt-open-redirect/SKILL.md +223 -0
  166. package/bundled-skills/hunt-race-condition/SKILL.md +381 -0
  167. package/bundled-skills/hunt-race-condition/references/details.md +159 -0
  168. package/bundled-skills/hunt-rag-vector/SKILL.md +212 -0
  169. package/bundled-skills/hunt-rce/SKILL.md +444 -0
  170. package/bundled-skills/hunt-rce/references/details.md +110 -0
  171. package/bundled-skills/hunt-saml/SKILL.md +156 -0
  172. package/bundled-skills/hunt-session/SKILL.md +342 -0
  173. package/bundled-skills/hunt-shadow-api/SKILL.md +198 -0
  174. package/bundled-skills/hunt-source-leak/SKILL.md +345 -0
  175. package/bundled-skills/hunt-spa-api/SKILL.md +163 -0
  176. package/bundled-skills/hunt-springboot/SKILL.md +285 -0
  177. package/bundled-skills/hunt-sqli/SKILL.md +466 -0
  178. package/bundled-skills/hunt-ssrf/SKILL.md +396 -0
  179. package/bundled-skills/hunt-ssrf/references/details.md +179 -0
  180. package/bundled-skills/hunt-ssti/SKILL.md +163 -0
  181. package/bundled-skills/hunt-subdomain/SKILL.md +379 -0
  182. package/bundled-skills/hunt-tls-network/SKILL.md +399 -0
  183. package/bundled-skills/hunt-xxe/SKILL.md +466 -0
  184. package/bundled-skills/i-have-adhd/SKILL.md +170 -0
  185. package/bundled-skills/identity-access-management/SKILL.md +382 -0
  186. package/bundled-skills/identity-access-management/references/details.md +524 -0
  187. package/bundled-skills/incident-management/SKILL.md +484 -0
  188. package/bundled-skills/incident-response/SKILL.md +448 -0
  189. package/bundled-skills/incident-response/references/details.md +113 -0
  190. package/bundled-skills/interview-me/SKILL.md +248 -0
  191. package/bundled-skills/iso27001-compliance/SKILL.md +460 -0
  192. package/bundled-skills/jenkins/SKILL.md +462 -0
  193. package/bundled-skills/jev-use/SKILL.md +158 -0
  194. package/bundled-skills/kubernetes-hardening/SKILL.md +154 -0
  195. package/bundled-skills/kubernetes-ops/SKILL.md +449 -0
  196. package/bundled-skills/kubernetes-ops/references/details.md +108 -0
  197. package/bundled-skills/kustomize/SKILL.md +478 -0
  198. package/bundled-skills/linux-administration/SKILL.md +367 -0
  199. package/bundled-skills/linux-hardening/SKILL.md +154 -0
  200. package/bundled-skills/llm-app-security/SKILL.md +389 -0
  201. package/bundled-skills/llm-app-security/references/details.md +674 -0
  202. package/bundled-skills/llm-caching/SKILL.md +334 -0
  203. package/bundled-skills/llm-cost-optimization/SKILL.md +311 -0
  204. package/bundled-skills/llm-fine-tuning/SKILL.md +329 -0
  205. package/bundled-skills/llm-gateway/SKILL.md +282 -0
  206. package/bundled-skills/llm-inference-scaling/SKILL.md +286 -0
  207. package/bundled-skills/llmops-platform-engineering/SKILL.md +472 -0
  208. package/bundled-skills/load-balancing/SKILL.md +403 -0
  209. package/bundled-skills/loki-logging/SKILL.md +479 -0
  210. package/bundled-skills/m365-entra-attack/SKILL.md +423 -0
  211. package/bundled-skills/mac-mini-llm-lab/SKILL.md +350 -0
  212. package/bundled-skills/mcp-server-security/SKILL.md +356 -0
  213. package/bundled-skills/mcp-server-security/references/details.md +745 -0
  214. package/bundled-skills/mdm-device-management/SKILL.md +404 -0
  215. package/bundled-skills/mdm-device-management/references/details.md +410 -0
  216. package/bundled-skills/meme-coin-audit/SKILL.md +402 -0
  217. package/bundled-skills/mid-engagement-ir-detection/SKILL.md +377 -0
  218. package/bundled-skills/model-registry-governance/SKILL.md +452 -0
  219. package/bundled-skills/model-serving-kubernetes/SKILL.md +339 -0
  220. package/bundled-skills/model-supply-chain-security/SKILL.md +427 -0
  221. package/bundled-skills/mongodb/SKILL.md +436 -0
  222. package/bundled-skills/multi-tenant-llm-hosting/SKILL.md +435 -0
  223. package/bundled-skills/multi-tenant-llm-hosting/references/details.md +211 -0
  224. package/bundled-skills/mysql/SKILL.md +390 -0
  225. package/bundled-skills/new-relic/SKILL.md +472 -0
  226. package/bundled-skills/nfs-storage/SKILL.md +356 -0
  227. package/bundled-skills/object-storage/SKILL.md +378 -0
  228. package/bundled-skills/offensive-osint/SKILL.md +443 -0
  229. package/bundled-skills/okta-attack/SKILL.md +436 -0
  230. package/bundled-skills/ollama-stack/SKILL.md +379 -0
  231. package/bundled-skills/openclaw-deployment-hardening/SKILL.md +135 -0
  232. package/bundled-skills/openclaw-local-mac-mini/SKILL.md +426 -0
  233. package/bundled-skills/openclaw-local-mac-mini/references/details.md +221 -0
  234. package/bundled-skills/openclaw-security-hardening/SKILL.md +135 -0
  235. package/bundled-skills/openshift/SKILL.md +485 -0
  236. package/bundled-skills/opentelemetry/SKILL.md +438 -0
  237. package/bundled-skills/opentelemetry/references/details.md +78 -0
  238. package/bundled-skills/opentofu-migration/SKILL.md +349 -0
  239. package/bundled-skills/osint-methodology/SKILL.md +460 -0
  240. package/bundled-skills/osint-methodology/references/details.md +1350 -0
  241. package/bundled-skills/pci-dss-compliance/SKILL.md +446 -0
  242. package/bundled-skills/penetration-testing/SKILL.md +152 -0
  243. package/bundled-skills/performance-tuning/SKILL.md +381 -0
  244. package/bundled-skills/planetscale/SKILL.md +297 -0
  245. package/bundled-skills/platform-engineering/SKILL.md +348 -0
  246. package/bundled-skills/platform-engineering/references/details.md +944 -0
  247. package/bundled-skills/podman/SKILL.md +405 -0
  248. package/bundled-skills/policy-as-code/SKILL.md +434 -0
  249. package/bundled-skills/policy-as-code/references/details.md +204 -0
  250. package/bundled-skills/postgresql-devsec/SKILL.md +378 -0
  251. package/bundled-skills/prometheus-grafana/SKILL.md +469 -0
  252. package/bundled-skills/prompt-injection-defense/SKILL.md +483 -0
  253. package/bundled-skills/rag-infrastructure/SKILL.md +269 -0
  254. package/bundled-skills/rag-observability-evals/SKILL.md +444 -0
  255. package/bundled-skills/rag-observability-evals/references/details.md +92 -0
  256. package/bundled-skills/recon-scope-triage/SKILL.md +128 -0
  257. package/bundled-skills/redis/SKILL.md +421 -0
  258. package/bundled-skills/redteam-report-template/SKILL.md +370 -0
  259. package/bundled-skills/report-writing/SKILL.md +426 -0
  260. package/bundled-skills/report-writing/references/details.md +187 -0
  261. package/bundled-skills/reverse-proxy/SKILL.md +420 -0
  262. package/bundled-skills/runbook-creation/SKILL.md +438 -0
  263. package/bundled-skills/runbook-creation/references/details.md +71 -0
  264. package/bundled-skills/saas-security-posture/SKILL.md +415 -0
  265. package/bundled-skills/sast-scanning/SKILL.md +444 -0
  266. package/bundled-skills/sbom-supply-chain/SKILL.md +433 -0
  267. package/bundled-skills/security-arsenal/SKILL.md +446 -0
  268. package/bundled-skills/security-arsenal/references/details.md +540 -0
  269. package/bundled-skills/security-automation/SKILL.md +146 -0
  270. package/bundled-skills/semantic-versioning/SKILL.md +434 -0
  271. package/bundled-skills/semantic-versioning/references/details.md +83 -0
  272. package/bundled-skills/service-mesh/SKILL.md +422 -0
  273. package/bundled-skills/soc2-compliance/SKILL.md +409 -0
  274. package/bundled-skills/sops-encryption/SKILL.md +124 -0
  275. package/bundled-skills/sre-dashboards/SKILL.md +143 -0
  276. package/bundled-skills/ssh-configuration/SKILL.md +324 -0
  277. package/bundled-skills/ssl-tls-management/SKILL.md +428 -0
  278. package/bundled-skills/ssl-tls-management/references/details.md +99 -0
  279. package/bundled-skills/startup-it-troubleshooting/SKILL.md +415 -0
  280. package/bundled-skills/supply-chain-attack-recon/SKILL.md +453 -0
  281. package/bundled-skills/supply-chain-attack-recon/references/details.md +258 -0
  282. package/bundled-skills/systemd-services/SKILL.md +379 -0
  283. package/bundled-skills/terraform-aws/SKILL.md +125 -0
  284. package/bundled-skills/terraform-azure/SKILL.md +415 -0
  285. package/bundled-skills/terraform-azure/references/details.md +231 -0
  286. package/bundled-skills/terraform-gcp/SKILL.md +369 -0
  287. package/bundled-skills/threat-modeling/SKILL.md +487 -0
  288. package/bundled-skills/user-management/SKILL.md +383 -0
  289. package/bundled-skills/using-agent-skills/SKILL.md +220 -0
  290. package/bundled-skills/vector-database-ops/SKILL.md +300 -0
  291. package/bundled-skills/vendor-management/SKILL.md +439 -0
  292. package/bundled-skills/vendor-management/references/details.md +109 -0
  293. package/bundled-skills/vercel-deployments/SKILL.md +296 -0
  294. package/bundled-skills/vllm-server/SKILL.md +236 -0
  295. package/bundled-skills/vmware-vcenter-attack/SKILL.md +412 -0
  296. package/bundled-skills/vpn-setup/SKILL.md +452 -0
  297. package/bundled-skills/vulnerability-scanning/SKILL.md +448 -0
  298. package/bundled-skills/waf-setup/SKILL.md +354 -0
  299. package/bundled-skills/waf-setup/references/details.md +211 -0
  300. package/bundled-skills/web2-recon/SKILL.md +440 -0
  301. package/bundled-skills/web2-recon/references/details.md +319 -0
  302. package/bundled-skills/web3-audit/SKILL.md +445 -0
  303. package/bundled-skills/web3-audit/references/details.md +224 -0
  304. package/bundled-skills/windows-hardening/SKILL.md +454 -0
  305. package/bundled-skills/windows-hardening/references/details.md +204 -0
  306. package/bundled-skills/windows-server/SKILL.md +318 -0
  307. package/bundled-skills/zero-trust/SKILL.md +461 -0
  308. package/package.json +1 -1
  309. package/skills_index.json +6943 -323
@@ -0,0 +1,435 @@
1
+ ---
2
+ name: multi-tenant-llm-hosting
3
+ description: Design secure, multi-tenant LLM hosting platforms with tenant isolation,
4
+ quotas, billing attribution, noisy-neighbor protection, and per-tenant policy controls.
5
+ category: devops
6
+ risk: critical
7
+ source: https://github.com/BagelHole/DevOps-Security-Agent-Skills
8
+ source_repo: BagelHole/DevOps-Security-Agent-Skills
9
+ source_type: community
10
+ date_added: '2026-09-20'
11
+ license: MIT
12
+ license_source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
13
+ compatibility: Requires the relevant OS/platform tooling and privileged access where
14
+ noted. Docs-only; helper scripts and templates not bundled.
15
+ metadata:
16
+ author: devops-skills
17
+ version: '1.0'
18
+ ---
19
+
20
+ # Multi-Tenant LLM Hosting
21
+
22
+ Host many teams/customers on shared inference infrastructure without sacrificing security, performance, or cost governance.
23
+
24
+ ## Prerequisites
25
+
26
+ - Kubernetes cluster with GPU node pools
27
+ - API gateway or LLM gateway (LiteLLM, Envoy, Kong)
28
+ - Prometheus + Grafana for per-tenant observability
29
+ - Redis or equivalent for rate limiting state
30
+ - Billing system or cost attribution database
31
+
32
+ ## Isolation Model
33
+
34
+ - Strong tenant identity on every request
35
+ - Per-tenant API keys and scoped model access
36
+ - Namespace or workload isolation for high-risk tenants
37
+ - Strict data retention and log partitioning controls
38
+
39
+ ## vLLM Multi-Model Serving
40
+
41
+ ```yaml
42
+ # vllm-deployment.yaml - Multi-model serving with vLLM
43
+ apiVersion: apps/v1
44
+ kind: Deployment
45
+ metadata:
46
+ name: vllm-gpt4o-equivalent
47
+ namespace: llm-serving
48
+ labels:
49
+ app: vllm
50
+ model-tier: premium
51
+ spec:
52
+ replicas: 3
53
+ selector:
54
+ matchLabels:
55
+ app: vllm
56
+ model-tier: premium
57
+ template:
58
+ metadata:
59
+ labels:
60
+ app: vllm
61
+ model-tier: premium
62
+ annotations:
63
+ prometheus.io/scrape: "true"
64
+ prometheus.io/port: "8080"
65
+ spec:
66
+ containers:
67
+ - name: vllm
68
+ image: vllm/vllm-openai:v0.4.1
69
+ args:
70
+ - "--model=/models/llama-3.1-70b"
71
+ - "--tensor-parallel-size=2"
72
+ - "--max-model-len=8192"
73
+ - "--gpu-memory-utilization=0.90"
74
+ - "--max-num-seqs=128"
75
+ - "--enable-prefix-caching"
76
+ ports:
77
+ - containerPort: 8000
78
+ name: inference
79
+ - containerPort: 8080
80
+ name: metrics
81
+ resources:
82
+ requests:
83
+ nvidia.com/gpu: 2
84
+ cpu: "8"
85
+ memory: "64Gi"
86
+ limits:
87
+ nvidia.com/gpu: 2
88
+ cpu: "16"
89
+ memory: "128Gi"
90
+ volumeMounts:
91
+ - name: model-weights
92
+ mountPath: /models
93
+ readOnly: true
94
+ volumes:
95
+ - name: model-weights
96
+ persistentVolumeClaim:
97
+ claimName: premium-model-weights
98
+ tolerations:
99
+ - key: nvidia.com/gpu
100
+ operator: Exists
101
+ effect: NoSchedule
102
+ nodeSelector:
103
+ gpu-type: a100
104
+ ---
105
+ apiVersion: apps/v1
106
+ kind: Deployment
107
+ metadata:
108
+ name: vllm-economy
109
+ namespace: llm-serving
110
+ labels:
111
+ app: vllm
112
+ model-tier: economy
113
+ spec:
114
+ replicas: 2
115
+ selector:
116
+ matchLabels:
117
+ app: vllm
118
+ model-tier: economy
119
+ template:
120
+ metadata:
121
+ labels:
122
+ app: vllm
123
+ model-tier: economy
124
+ spec:
125
+ containers:
126
+ - name: vllm
127
+ image: vllm/vllm-openai:v0.4.1
128
+ args:
129
+ - "--model=/models/llama-3.1-8b"
130
+ - "--max-model-len=4096"
131
+ - "--gpu-memory-utilization=0.85"
132
+ - "--max-num-seqs=256"
133
+ - "--enable-prefix-caching"
134
+ ports:
135
+ - containerPort: 8000
136
+ name: inference
137
+ - containerPort: 8080
138
+ name: metrics
139
+ resources:
140
+ requests:
141
+ nvidia.com/gpu: 1
142
+ cpu: "4"
143
+ memory: "32Gi"
144
+ limits:
145
+ nvidia.com/gpu: 1
146
+ cpu: "8"
147
+ memory: "64Gi"
148
+ volumeMounts:
149
+ - name: model-weights
150
+ mountPath: /models
151
+ readOnly: true
152
+ volumes:
153
+ - name: model-weights
154
+ persistentVolumeClaim:
155
+ claimName: economy-model-weights
156
+ tolerations:
157
+ - key: nvidia.com/gpu
158
+ operator: Exists
159
+ effect: NoSchedule
160
+ ```
161
+
162
+ ## Per-Tenant Quota Configuration
163
+
164
+ ```yaml
165
+ # tenant-quotas-configmap.yaml
166
+ apiVersion: v1
167
+ kind: ConfigMap
168
+ metadata:
169
+ name: tenant-quotas
170
+ namespace: llm-serving
171
+ data:
172
+ quotas.yaml: |
173
+ tenants:
174
+ acme-corp:
175
+ tier: enterprise
176
+ models_allowed:
177
+ - llama-3.1-70b
178
+ - llama-3.1-8b
179
+ - nomic-embed-text
180
+ rate_limits:
181
+ requests_per_minute: 300
182
+ tokens_per_minute: 500000
183
+ concurrent_requests: 50
184
+ budget:
185
+ daily_limit_usd: 500.00
186
+ monthly_limit_usd: 10000.00
187
+ alert_threshold_percent: 80
188
+ priority: high
189
+
190
+ startup-xyz:
191
+ tier: standard
192
+ models_allowed:
193
+ - llama-3.1-8b
194
+ - nomic-embed-text
195
+ rate_limits:
196
+ requests_per_minute: 60
197
+ tokens_per_minute: 100000
198
+ concurrent_requests: 10
199
+ budget:
200
+ daily_limit_usd: 50.00
201
+ monthly_limit_usd: 1000.00
202
+ alert_threshold_percent: 80
203
+ priority: medium
204
+
205
+ internal-dev:
206
+ tier: free
207
+ models_allowed:
208
+ - llama-3.1-8b
209
+ rate_limits:
210
+ requests_per_minute: 20
211
+ tokens_per_minute: 50000
212
+ concurrent_requests: 5
213
+ budget:
214
+ daily_limit_usd: 10.00
215
+ monthly_limit_usd: 200.00
216
+ alert_threshold_percent: 90
217
+ priority: low
218
+ ```
219
+
220
+ ## Namespace Isolation for High-Risk Tenants
221
+
222
+ ```yaml
223
+ # tenant-namespace.yaml
224
+ apiVersion: v1
225
+ kind: Namespace
226
+ metadata:
227
+ name: tenant-acme-corp
228
+ labels:
229
+ tenant: acme-corp
230
+ isolation: strict
231
+ ---
232
+ apiVersion: networking.k8s.io/v1
233
+ kind: NetworkPolicy
234
+ metadata:
235
+ name: tenant-isolation
236
+ namespace: tenant-acme-corp
237
+ spec:
238
+ podSelector: {}
239
+ policyTypes:
240
+ - Ingress
241
+ - Egress
242
+ ingress:
243
+ - from:
244
+ - namespaceSelector:
245
+ matchLabels:
246
+ name: llm-gateway
247
+ egress:
248
+ - to:
249
+ - namespaceSelector:
250
+ matchLabels:
251
+ name: llm-serving
252
+ ports:
253
+ - port: 8000
254
+ protocol: TCP
255
+ - to:
256
+ - namespaceSelector:
257
+ matchLabels:
258
+ name: kube-dns
259
+ ports:
260
+ - port: 53
261
+ protocol: UDP
262
+ ---
263
+ apiVersion: v1
264
+ kind: ResourceQuota
265
+ metadata:
266
+ name: tenant-quota
267
+ namespace: tenant-acme-corp
268
+ spec:
269
+ hard:
270
+ requests.cpu: "16"
271
+ requests.memory: "64Gi"
272
+ limits.cpu: "32"
273
+ limits.memory: "128Gi"
274
+ requests.nvidia.com/gpu: "4"
275
+ pods: "20"
276
+ ```
277
+
278
+ ## Request Routing and Rate Limiting
279
+
280
+ ```python
281
+ # gateway_router.py
282
+ """Multi-tenant request router with rate limiting and model routing."""
283
+ import time
284
+ import json
285
+ import redis
286
+ from fastapi import FastAPI, HTTPException, Header, Request
287
+ from typing import Optional
288
+ import httpx
289
+ import yaml
290
+
291
+ app = FastAPI()
292
+ redis_client = redis.Redis(host="redis", port=6379, decode_responses=True)
293
+
294
+ # Load tenant config
295
+ with open("/etc/config/quotas.yaml") as f:
296
+ TENANT_CONFIG = yaml.safe_load(f)["tenants"]
297
+
298
+ MODEL_ENDPOINTS = {
299
+ "llama-3.1-70b": "http://vllm-gpt4o-equivalent:8000",
300
+ "llama-3.1-8b": "http://vllm-economy:8000",
301
+ "nomic-embed-text": "http://embedding-service:8000",
302
+ }
303
+
304
+ def check_rate_limit(tenant_id: str, config: dict) -> bool:
305
+ """Check and update rate limit for a tenant."""
306
+ key = f"ratelimit:{tenant_id}:{int(time.time() // 60)}"
307
+ current = redis_client.incr(key)
308
+ if current == 1:
309
+ redis_client.expire(key, 120)
310
+ return current <= config["rate_limits"]["requests_per_minute"]
311
+
312
+ def check_concurrent(tenant_id: str, config: dict) -> bool:
313
+ """Check concurrent request limit."""
314
+ key = f"concurrent:{tenant_id}"
315
+ current = int(redis_client.get(key) or 0)
316
+ return current < config["rate_limits"]["concurrent_requests"]
317
+
318
+ def check_budget(tenant_id: str, config: dict) -> bool:
319
+ """Check if tenant is within daily budget."""
320
+ key = f"spend:{tenant_id}:{time.strftime('%Y-%m-%d')}"
321
+ current_spend = float(redis_client.get(key) or 0)
322
+ return current_spend < config["budget"]["daily_limit_usd"]
323
+
324
+ def record_usage(tenant_id: str, model: str, prompt_tokens: int, completion_tokens: int):
325
+ """Record token usage and cost for billing."""
326
+ # Cost rates per 1K tokens
327
+ rates = {
328
+ "llama-3.1-70b": {"prompt": 0.004, "completion": 0.012},
329
+ "llama-3.1-8b": {"prompt": 0.0005, "completion": 0.0015},
330
+ "nomic-embed-text": {"prompt": 0.0001, "completion": 0.0},
331
+ }
332
+ rate = rates.get(model, {"prompt": 0.001, "completion": 0.003})
333
+ cost = (prompt_tokens * rate["prompt"] + completion_tokens * rate["completion"]) / 1000
334
+
335
+ # Update daily spend
336
+ spend_key = f"spend:{tenant_id}:{time.strftime('%Y-%m-%d')}"
337
+ redis_client.incrbyfloat(spend_key, cost)
338
+ redis_client.expire(spend_key, 172800)
339
+
340
+ # Record for billing export
341
+ billing_key = f"billing:{tenant_id}:{time.strftime('%Y-%m')}"
342
+ redis_client.rpush(billing_key, json.dumps({
343
+ "timestamp": time.time(),
344
+ "model": model,
345
+ "prompt_tokens": prompt_tokens,
346
+ "completion_tokens": completion_tokens,
347
+ "cost_usd": cost,
348
+ }))
349
+
350
+ @app.post("/v1/chat/completions")
351
+ async def chat_completions(
352
+ request: Request,
353
+ x_tenant_id: str = Header(...),
354
+ x_api_key: str = Header(...),
355
+ ):
356
+ """Route chat completion request with tenant controls."""
357
+ if x_tenant_id not in TENANT_CONFIG:
358
+ raise HTTPException(status_code=403, detail="Unknown tenant")
359
+
360
+ config = TENANT_CONFIG[x_tenant_id]
361
+ body = await request.json()
362
+ model = body.get("model", "llama-3.1-8b")
363
+
364
+ # Check model access
365
+ if model not in config["models_allowed"]:
366
+ raise HTTPException(status_code=403, detail=f"Model {model} not allowed for tenant")
367
+
368
+ # Check rate limit
369
+ if not check_rate_limit(x_tenant_id, config):
370
+ raise HTTPException(status_code=429, detail="Rate limit exceeded")
371
+
372
+ # Check concurrent requests
373
+ if not check_concurrent(x_tenant_id, config):
374
+ raise HTTPException(status_code=429, detail="Concurrent request limit exceeded")
375
+
376
+ # Check budget
377
+ if not check_budget(x_tenant_id, config):
378
+ raise HTTPException(status_code=402, detail="Daily budget exceeded")
379
+
380
+ # Route to model endpoint
381
+ endpoint = MODEL_ENDPOINTS.get(model)
382
+ if not endpoint:
383
+ raise HTTPException(status_code=404, detail=f"Model {model} not available")
384
+
385
+ # Track concurrent requests
386
+ concurrent_key = f"concurrent:{x_tenant_id}"
387
+ redis_client.incr(concurrent_key)
388
+
389
+ try:
390
+ async with httpx.AsyncClient(timeout=120.0) as client:
391
+ response = await client.post(
392
+ f"{endpoint}/v1/chat/completions",
393
+ json=body,
394
+ headers={"Content-Type": "application/json"},
395
+ )
396
+ result = response.json()
397
+
398
+ # Record usage
399
+ usage = result.get("usage", {})
400
+ record_usage(
401
+ x_tenant_id, model,
402
+ usage.get("prompt_tokens", 0),
403
+ usage.get("completion_tokens", 0),
404
+ )
405
+
406
+ return result
407
+ finally:
408
+ redis_client.decr(concurrent_key)
409
+ ```
410
+
411
+
412
+ ## Contents
413
+
414
+ - [Rate Limiting with Envoy](references/details.md)
415
+ - [Billing Integration](references/details.md)
416
+ - [Noisy-Neighbor Controls](references/details.md)
417
+ - [Per-Tenant Monitoring](references/details.md)
418
+ - [Security Baseline](references/details.md)
419
+ - [Operational Runbook](references/details.md)
420
+ - [Troubleshooting](references/details.md)
421
+ - [Related Skills](references/details.md)
422
+
423
+ ## When to Use This Skill
424
+
425
+ - Building an internal LLM platform shared by multiple teams
426
+ - Hosting LLM inference for external customers with isolation requirements
427
+ - Implementing per-tenant quotas, billing, and rate limiting
428
+ - Designing request routing for multi-model, multi-tenant environments
429
+ - Preventing noisy-neighbor issues on shared GPU infrastructure
430
+
431
+ ## Limitations
432
+
433
+ - Infrastructure commands can disrupt services: confirm target host/scope and have backups/snapshots before mutating state.
434
+ - Docs-only import: upstream scripts and templates not bundled.
435
+
@@ -0,0 +1,211 @@
1
+ # Details (moved from SKILL.md)
2
+
3
+ > Extended reference content for `multi-tenant-llm-hosting`, kept under `references/` so the entrypoint stays within the audit budget.
4
+
5
+ ## Rate Limiting with Envoy
6
+
7
+ ```yaml
8
+ # envoy-ratelimit.yaml
9
+ apiVersion: v1
10
+ kind: ConfigMap
11
+ metadata:
12
+ name: envoy-ratelimit-config
13
+ namespace: llm-serving
14
+ data:
15
+ config.yaml: |
16
+ domain: llm-gateway
17
+ descriptors:
18
+ # Per-tenant rate limits
19
+ - key: tenant_id
20
+ value: acme-corp
21
+ rate_limit:
22
+ unit: minute
23
+ requests_per_unit: 300
24
+ - key: tenant_id
25
+ value: startup-xyz
26
+ rate_limit:
27
+ unit: minute
28
+ requests_per_unit: 60
29
+ - key: tenant_id
30
+ value: internal-dev
31
+ rate_limit:
32
+ unit: minute
33
+ requests_per_unit: 20
34
+
35
+ # Global rate limit as safety net
36
+ - key: global
37
+ rate_limit:
38
+ unit: second
39
+ requests_per_unit: 100
40
+ ```
41
+
42
+
43
+ ## Billing Integration
44
+
45
+ ```python
46
+ # billing_export.py
47
+ """Export tenant usage data for billing systems."""
48
+ import redis
49
+ import json
50
+ from datetime import datetime, timedelta
51
+ from typing import Dict, List
52
+
53
+ redis_client = redis.Redis(host="redis", port=6379, decode_responses=True)
54
+
55
+ def generate_tenant_invoice(tenant_id: str, month: str) -> Dict:
56
+ """Generate monthly invoice for a tenant."""
57
+ billing_key = f"billing:{tenant_id}:{month}"
58
+ records = redis_client.lrange(billing_key, 0, -1)
59
+
60
+ usage_by_model = {}
61
+ total_cost = 0.0
62
+ total_requests = 0
63
+
64
+ for record_json in records:
65
+ record = json.loads(record_json)
66
+ model = record["model"]
67
+
68
+ if model not in usage_by_model:
69
+ usage_by_model[model] = {
70
+ "requests": 0,
71
+ "prompt_tokens": 0,
72
+ "completion_tokens": 0,
73
+ "cost_usd": 0.0,
74
+ }
75
+
76
+ usage_by_model[model]["requests"] += 1
77
+ usage_by_model[model]["prompt_tokens"] += record["prompt_tokens"]
78
+ usage_by_model[model]["completion_tokens"] += record["completion_tokens"]
79
+ usage_by_model[model]["cost_usd"] += record["cost_usd"]
80
+
81
+ total_cost += record["cost_usd"]
82
+ total_requests += 1
83
+
84
+ return {
85
+ "tenant_id": tenant_id,
86
+ "billing_period": month,
87
+ "generated_at": datetime.utcnow().isoformat(),
88
+ "summary": {
89
+ "total_requests": total_requests,
90
+ "total_cost_usd": round(total_cost, 4),
91
+ },
92
+ "usage_by_model": usage_by_model,
93
+ }
94
+
95
+ def get_tenant_spend_today(tenant_id: str) -> float:
96
+ """Get current day spend for budget alerts."""
97
+ key = f"spend:{tenant_id}:{datetime.utcnow().strftime('%Y-%m-%d')}"
98
+ return float(redis_client.get(key) or 0)
99
+ ```
100
+
101
+
102
+ ## Noisy-Neighbor Controls
103
+
104
+ - Per-tenant RPM/TPM limits
105
+ - Concurrency caps and queue isolation
106
+ - Fair scheduling with weighted priority classes
107
+ - Backpressure and graceful degradation policies
108
+
109
+ ```yaml
110
+ # priority-classes.yaml
111
+ apiVersion: scheduling.k8s.io/v1
112
+ kind: PriorityClass
113
+ metadata:
114
+ name: tenant-enterprise
115
+ value: 1000
116
+ globalDefault: false
117
+ description: "Enterprise tenant workloads"
118
+ ---
119
+ apiVersion: scheduling.k8s.io/v1
120
+ kind: PriorityClass
121
+ metadata:
122
+ name: tenant-standard
123
+ value: 500
124
+ globalDefault: false
125
+ description: "Standard tenant workloads"
126
+ ---
127
+ apiVersion: scheduling.k8s.io/v1
128
+ kind: PriorityClass
129
+ metadata:
130
+ name: tenant-free
131
+ value: 100
132
+ globalDefault: false
133
+ description: "Free tier tenant workloads"
134
+ ```
135
+
136
+
137
+ ## Per-Tenant Monitoring
138
+
139
+ ```yaml
140
+ # tenant-alerts.yaml
141
+ groups:
142
+ - name: tenant-alerts
143
+ rules:
144
+ - alert: TenantBudgetWarning
145
+ expr: |
146
+ llm_tenant_daily_spend_usd
147
+ / llm_tenant_daily_budget_usd > 0.80
148
+ for: 5m
149
+ labels:
150
+ severity: warning
151
+ annotations:
152
+ summary: "Tenant {{ $labels.tenant }} at 80% of daily budget"
153
+
154
+ - alert: TenantRateLimitHitting
155
+ expr: |
156
+ rate(llm_rate_limit_rejections_total[5m]) > 1
157
+ for: 5m
158
+ labels:
159
+ severity: info
160
+ annotations:
161
+ summary: "Tenant {{ $labels.tenant }} hitting rate limits"
162
+
163
+ - alert: TenantErrorRateHigh
164
+ expr: |
165
+ rate(llm_tenant_errors_total[5m])
166
+ / rate(llm_tenant_requests_total[5m]) > 0.10
167
+ for: 5m
168
+ labels:
169
+ severity: warning
170
+ annotations:
171
+ summary: "Tenant {{ $labels.tenant }} error rate above 10%"
172
+ ```
173
+
174
+
175
+ ## Security Baseline
176
+
177
+ - Encrypt data in transit and at rest.
178
+ - Disallow cross-tenant cache leakage.
179
+ - Restrict debug data access by role.
180
+ - Audit all privileged administrative actions.
181
+
182
+
183
+ ## Operational Runbook
184
+
185
+ 1. Onboard tenant with policy template.
186
+ 2. Issue virtual key and quota profile.
187
+ 3. Validate observability and billing tags.
188
+ 4. Run tenant-specific load/safety tests.
189
+ 5. Enable production traffic with canary limits.
190
+
191
+
192
+ ## Troubleshooting
193
+
194
+ | Symptom | Check | Fix |
195
+ |---------|-------|-----|
196
+ | Tenant getting 429 errors | Rate limit counters in Redis | Increase RPM/TPM limits or upgrade tier |
197
+ | One tenant slowing others | Concurrent request counts per tenant | Reduce concurrency cap for offending tenant |
198
+ | Billing data missing | Redis billing keys and export job logs | Check billing export CronJob and Redis connectivity |
199
+ | Tenant cannot access model | Tenant config in ConfigMap | Add model to `models_allowed` list |
200
+ | Cross-tenant data leakage | Cache key prefixes and namespace isolation | Ensure cache keys include tenant_id prefix |
201
+ | Budget alerts not firing | Prometheus scrape targets and alert rules | Verify metric export and Alertmanager config |
202
+
203
+
204
+ ## Related Skills
205
+
206
+ - llm-gateway (`llm-gateway`) - Key management and traffic routing
207
+ - llm-cost-optimization (`llm-cost-optimization`) - Cost controls and optimization tactics
208
+ - zero-trust (`zero-trust`) - Identity-centric network and access patterns
209
+ - gpu-kubernetes-operations (`gpu-kubernetes-operations`) - GPU cluster management
210
+ - llm-inference-scaling (`llm-inference-scaling`) - Autoscaling inference workloads
211
+