@doist/doistbot-cli 1.0.12 β 1.0.14
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/config.js +1 -1
- package/package.json +5 -5
- package/sandbox/dist/core/pi.js +2 -2
- package/sandbox/dist/core/repository-map.js +1 -1
- package/sandbox/dist/tasks/chat/chat.js +1 -1
- package/sandbox/dist/tasks/issue-fix-retry/fix-retry.js +5 -1
- package/sandbox/dist/tasks/issue-summarize/context.js +1 -0
- package/sandbox/dist/tasks/issue-summarize/model.js +24 -7
- package/sandbox/dist/tasks/issue-summarize/summarize.js +6 -4
- package/sandbox/dist/tasks/issue-triage/assignee-review.js +114 -0
- package/sandbox/dist/tasks/issue-triage/fix-attempt.js +16 -11
- package/sandbox/dist/tasks/issue-triage/fix-dispatch.js +2 -0
- package/sandbox/dist/tasks/issue-triage/fix-loop/effects.js +21 -19
- package/sandbox/dist/tasks/issue-triage/fix-loop/runners.js +32 -31
- package/sandbox/dist/tasks/issue-triage/github-request.js +2 -0
- package/sandbox/dist/tasks/issue-triage/hero-group-map.js +41 -3
- package/sandbox/dist/tasks/issue-triage/markers.js +1 -0
- package/sandbox/dist/tasks/issue-triage/pr-creator.js +59 -1
- package/sandbox/dist/tasks/issue-triage/pr-readiness.js +25 -0
- package/sandbox/dist/tasks/issue-triage/resolution.js +3 -0
- package/sandbox/dist/tasks/issue-triage/triage.js +5 -1
- package/sandbox/dist/tasks/review/thinking-level.js +1 -1
- package/sandbox/src/review/prompts/ai-internal-tools.md +3 -3
- package/sandbox/src/review/prompts/backend-repository-standards.md +102 -0
- package/sandbox/src/review/prompts/outline-map.json +14 -10
- package/sandbox/src/review/prompts/service-production-readiness.md +2 -2
- package/sandbox/src/review/prompts/which-cloud-and-where-to-deploy.md +0 -344
|
@@ -1,344 +0,0 @@
|
|
|
1
|
-
# Platform Standard: Which Cloud & Where to Deploy
|
|
2
|
-
|
|
3
|
-
| π Document Type | Platform Standard |
|
|
4
|
-
| :--------------- | :----------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
5
|
-
| π― Status | Approved |
|
|
6
|
-
| π€ Owner | Luciano Facchinelli |
|
|
7
|
-
| π
Last Reviewed | 2026-03-06 |
|
|
8
|
-
| ποΈ References | [AWS Well-Architected](https://aws.amazon.com/architecture/well-architected/), [GCP Architecture Framework](https://cloud.google.com/architecture/framework) |
|
|
9
|
-
| π Version | 1.0 |
|
|
10
|
-
|
|
11
|
-
# Workload types
|
|
12
|
-
|
|
13
|
-
## Remote State Services
|
|
14
|
-
|
|
15
|
-
- **Definition**: Services that either store state in external systems (databases, caches, object storage) and don't require persistent local disk or doesn't have a state at all.
|
|
16
|
-
Containers are disposable and can be recreated without data loss.
|
|
17
|
-
- **Typical use cases**:
|
|
18
|
-
- REST APIs
|
|
19
|
-
- Web applications
|
|
20
|
-
- Backend microservices
|
|
21
|
-
|
|
22
|
-
## No State / Local State Services
|
|
23
|
-
|
|
24
|
-
- **Definition**: Services that require persistent local disk storage and expect data to survive container restarts. These need node affinity and persistent volumes.
|
|
25
|
-
- **Typical use cases**:
|
|
26
|
-
- Analytics database (i.e ClickHouse)
|
|
27
|
-
- Observability tool (i.e Langfuse)
|
|
28
|
-
- Self-hosted databases
|
|
29
|
-
- Message brokers with persistence
|
|
30
|
-
- Runs a database or database-like system
|
|
31
|
-
- Benefits from node affinity control for performance
|
|
32
|
-
|
|
33
|
-
## Event-Driven / Short-Lived Tasks
|
|
34
|
-
|
|
35
|
-
- **Definition**: Sporadic, short-duration tasks triggered by events that must complete within 15 minutes
|
|
36
|
-
- **Typical use cases**:
|
|
37
|
-
- Webhooks
|
|
38
|
-
- Scheduled jobs
|
|
39
|
-
- Queue processors
|
|
40
|
-
- Glue code between services
|
|
41
|
-
- File processing triggers
|
|
42
|
-
- Code must be small and optimized to reduce consumption from a resource perspective.
|
|
43
|
-
- Cold starts can impact latency for infrequent calls
|
|
44
|
-
- Not suitable for long-running or high-memory processes
|
|
45
|
-
|
|
46
|
-
## Preview / Ephemeral Environments
|
|
47
|
-
|
|
48
|
-
- **Definition**: Temporary, short-lived environments designed for testing, demos, and validation with a clear expiration
|
|
49
|
-
- **Typical use cases**:
|
|
50
|
-
- PR preview deployments
|
|
51
|
-
- Demo environments
|
|
52
|
-
- Temporary test instances
|
|
53
|
-
- Feature branch testing
|
|
54
|
-
- Excellent developer experience for fast iterations
|
|
55
|
-
|
|
56
|
-
## Internal Tools (Low Traffic)
|
|
57
|
-
|
|
58
|
-
- **Definition**: Internal-facing utilities and dashboards with low, predictable usage patterns accessed only by team members
|
|
59
|
-
- **Typical use cases**:
|
|
60
|
-
- Admin dashboards
|
|
61
|
-
- Internal utilities
|
|
62
|
-
- Analytics dashboards
|
|
63
|
-
- Configuration management UIs
|
|
64
|
-
- Reporting tools
|
|
65
|
-
- Low and predictable traffic patterns
|
|
66
|
-
- Users are internal team members only
|
|
67
|
-
- May or may not require access to production resources
|
|
68
|
-
- Simpler reliability requirements than user-facing services
|
|
69
|
-
|
|
70
|
-
## Batch Processing / ML Training
|
|
71
|
-
|
|
72
|
-
- **Definition**: Large-scale data processing or machine learning workloads that run for extended periods, often on a schedule or ad-hoc basis
|
|
73
|
-
- **Typical use cases**:
|
|
74
|
-
- Large batch data jobs
|
|
75
|
-
- ML model training
|
|
76
|
-
- Data pipeline processing
|
|
77
|
-
- One-off data migrations
|
|
78
|
-
- ETL jobs
|
|
79
|
-
- Variable duration β can run for hours or days
|
|
80
|
-
- Often requires burst compute capacity
|
|
81
|
-
- May need GPU resources for ML workloads
|
|
82
|
-
- Clear distinction between experimentation and production
|
|
83
|
-
- **Note: Requires case-by-case evaluation based on requirements**
|
|
84
|
-
|
|
85
|
-
# 1\. Purpose
|
|
86
|
-
|
|
87
|
-
"Where should I deploy this?" and "Can I use GCP for this?" shouldn't require a Platform team consultation. This standard provides clear defaults so teams can move fast while maintaining operational consistency.
|
|
88
|
-
|
|
89
|
-
These two decisions are tightly coupledβyour cloud choice often constrains your compute options, and vice versa. We're treating them as one standard.
|
|
90
|
-
|
|
91
|
-
## 2\. Scope
|
|
92
|
-
|
|
93
|
-
This standard applies to:
|
|
94
|
-
|
|
95
|
-
- All new services, tools, and workloads
|
|
96
|
-
- Migrations of existing services to new compute platforms
|
|
97
|
-
- Any infrastructure that runs production or production-adjacent workloads
|
|
98
|
-
|
|
99
|
-
This does **not** prescribe:
|
|
100
|
-
|
|
101
|
-
- How to configure specific services (separate runbooks)
|
|
102
|
-
- Database or storage choices (separate standard)
|
|
103
|
-
|
|
104
|
-
---
|
|
105
|
-
|
|
106
|
-
# Part A: Which Cloud
|
|
107
|
-
|
|
108
|
-
## 3A. Standard
|
|
109
|
-
|
|
110
|
-
### Default: AWS
|
|
111
|
-
|
|
112
|
-
**All new production workloads MUST be deployed to AWS unless there's a documented exception.**
|
|
113
|
-
|
|
114
|
-
AWS is our primary cloud. This means:
|
|
115
|
-
|
|
116
|
-
- All production databases, queues, and storage
|
|
117
|
-
- All user-facing services
|
|
118
|
-
- All internal services that modifies production data
|
|
119
|
-
|
|
120
|
-
### When GCP is Acceptable
|
|
121
|
-
|
|
122
|
-
GCP may be used for:
|
|
123
|
-
|
|
124
|
-
| Use Case | Rationale | Examples |
|
|
125
|
-
| :-------------------------- | :----------------------------------------------- | :------------------------------------------ |
|
|
126
|
-
| **BigQuery analytics** | Best-in-class data warehouse, existing pipelines | Data team experiments, analytics dashboards |
|
|
127
|
-
| **ML/AI experimentation** | Vertex AI, Colab integration | Prototype ML models before productionizing |
|
|
128
|
-
| **One-off data processing** | Cost-effective for burst compute | Large batch jobs with clear end dates |
|
|
129
|
-
| **Cloud Run previews** | Fast, cheap preview environments | PR preview deployments for web apps |
|
|
130
|
-
|
|
131
|
-
GCP is **NOT acceptable** for:
|
|
132
|
-
|
|
133
|
-
- Services that need to access production databases directly
|
|
134
|
-
- Anything requiring low-latency connections to AWS resources
|
|
135
|
-
- User-facing production services
|
|
136
|
-
- Long-running local state services
|
|
137
|
-
|
|
138
|
-
### Multi-Cloud \= Complexity Tax
|
|
139
|
-
|
|
140
|
-
Every service in GCP means:
|
|
141
|
-
|
|
142
|
-
- Separate IAM, networking, and monitoring
|
|
143
|
-
- Cross-cloud latency and data transfer costs
|
|
144
|
-
- Split operational knowledge on the team
|
|
145
|
-
- Harder incident response
|
|
146
|
-
|
|
147
|
-
The bar for GCP should be: "AWS genuinely can't do this well" β not "GCP has a nicer UI for this."
|
|
148
|
-
|
|
149
|
-
## 4A. Rationale
|
|
150
|
-
|
|
151
|
-
**Why default to AWS?**
|
|
152
|
-
|
|
153
|
-
- Operational consistency: one set of tools, dashboards, runbooks
|
|
154
|
-
- Network topology: VPCs, peering, and security groups are already configured
|
|
155
|
-
- Cost visibility: unified billing, tagging, and attribution
|
|
156
|
-
- Team knowledge: everyone knows AWS; GCP expertise is spotty
|
|
157
|
-
|
|
158
|
-
**Why allow GCP at all?**
|
|
159
|
-
|
|
160
|
-
- BigQuery is genuinely better for our analytics use cases
|
|
161
|
-
- Cloud Run is excellent for preview environments (fast, cheap, ephemeral)
|
|
162
|
-
- Forcing everything into AWS would be dogmatic, not practical
|
|
163
|
-
|
|
164
|
-
---
|
|
165
|
-
|
|
166
|
-
# Part B: Where to Deploy (Compute)
|
|
167
|
-
|
|
168
|
-
## 3B. Standard
|
|
169
|
-
|
|
170
|
-
### Decision Tree
|
|
171
|
-
|
|
172
|
-
```mermaid
|
|
173
|
-
%%{init: {'flowchart': {'curve': 'linear'}}}%%
|
|
174
|
-
flowchart TB
|
|
175
|
-
subgraph PartA["Part A: Cloud Selection"]
|
|
176
|
-
AWS[/"β
AWS<br>(Primary Cloud)"/]
|
|
177
|
-
Q1{"Which Cloud?"}
|
|
178
|
-
GCP_Check{"GCP Use Case?"}
|
|
179
|
-
GCP[/"βοΈ GCP Acceptable"/]
|
|
180
|
-
end
|
|
181
|
-
subgraph GCP_No["π« GCP NOT Acceptable For"]
|
|
182
|
-
direction TB
|
|
183
|
-
N1["Production DB Access"]
|
|
184
|
-
N2["Low-latency to AWS"]
|
|
185
|
-
N3["User-facing Production"]
|
|
186
|
-
N4["Long-running local state"]
|
|
187
|
-
end
|
|
188
|
-
subgraph PartB["Part B: Compute Platform Selection"]
|
|
189
|
-
ECS[/"π¦ ECS Fargate<br>(Default)"/]
|
|
190
|
-
Q2{"What Type<br>of Workload?"}
|
|
191
|
-
EKS[/"βΈοΈ EKS Kubernetes"/]
|
|
192
|
-
Lambda[/"β‘ Lambda"/]
|
|
193
|
-
Platform[/"π€ Consult Platform Team"/]
|
|
194
|
-
CloudRun[/"π Cloud Run"/]
|
|
195
|
-
Q3{"GCP Compute?"}
|
|
196
|
-
end
|
|
197
|
-
subgraph Examples["π Examples"]
|
|
198
|
-
direction TB
|
|
199
|
-
Ex1["REST API β ECS Fargate"]
|
|
200
|
-
Ex2["ClickHouse β EKS"]
|
|
201
|
-
Ex3["Langfuse β EKS"]
|
|
202
|
-
Ex4["Webhooks β Lambda"]
|
|
203
|
-
Ex5["PR Previews β Cloud Run"]
|
|
204
|
-
end
|
|
205
|
-
subgraph Exceptions["β οΈ Exception Process"]
|
|
206
|
-
direction TB
|
|
207
|
-
E1["1. Document rationale"]
|
|
208
|
-
E2["2. Post in Platform channel"]
|
|
209
|
-
E3["3. Get Platform approval"]
|
|
210
|
-
end
|
|
211
|
-
Start(["π New Service/Workload"]) --> Q1
|
|
212
|
-
Q1 -- Default --> AWS
|
|
213
|
-
Q1 -- Exception Cases --> GCP_Check
|
|
214
|
-
GCP_Check -- BigQuery Analytics --> GCP
|
|
215
|
-
GCP_Check -- ML/AI Experimentation --> GCP
|
|
216
|
-
GCP_Check -- "One-off Data Processing" --> GCP
|
|
217
|
-
GCP_Check -- Preview Environments --> GCP
|
|
218
|
-
GCP_Check -- None of above --> AWS
|
|
219
|
-
GCP -. Check Restrictions .-> GCP_No
|
|
220
|
-
AWS --> Q2
|
|
221
|
-
GCP --> Q3
|
|
222
|
-
Q2 -- remote state<br>(API, Web App) --> ECS
|
|
223
|
-
Q2 -- local state<br>(Persistent Storage) --> EKS
|
|
224
|
-
Q2 -- "Event-driven<br>(under 15 min)" --> Lambda
|
|
225
|
-
Q2 -- Batch/ML Training --> Platform
|
|
226
|
-
Q3 -- Previews/Internal Tools --> CloudRun
|
|
227
|
-
ECS -.-> Ex1
|
|
228
|
-
EKS -.-> Ex2 & Ex3
|
|
229
|
-
Lambda -.-> Ex4
|
|
230
|
-
CloudRun -.-> Ex5
|
|
231
|
-
E1 --> E2
|
|
232
|
-
E2 --> E3
|
|
233
|
-
Q2 -- Doesn't fit? --> Exceptions
|
|
234
|
-
|
|
235
|
-
AWS:::aws
|
|
236
|
-
Q1:::decision
|
|
237
|
-
GCP_Check:::decision
|
|
238
|
-
GCP:::gcp
|
|
239
|
-
ECS:::aws
|
|
240
|
-
Q2:::decision
|
|
241
|
-
EKS:::aws
|
|
242
|
-
Lambda:::aws
|
|
243
|
-
CloudRun:::gcp
|
|
244
|
-
Q3:::decision
|
|
245
|
-
classDef aws fill:#FF9900,stroke:#232F3E,color:#232F3E
|
|
246
|
-
classDef gcp fill:#4285F4,stroke:#174EA6,color:white
|
|
247
|
-
classDef decision fill:#f9f,stroke:#333,stroke-width:2px
|
|
248
|
-
classDef compute fill:#90EE90,stroke:#228B22
|
|
249
|
-
```
|
|
250
|
-
|
|
251
|
-
### Compute Platform Characteristics
|
|
252
|
-
|
|
253
|
-
| Platform | Best For | Avoid When | Notes |
|
|
254
|
-
| :------------------- | :----------------------------------------- | :------------------------------------------- | :----------------------------------------------------- |
|
|
255
|
-
| **ECS Fargate** | remote state services, APIs, web apps | Need persistent local storage, need GPU | Our default. Good tooling, well-understood. |
|
|
256
|
-
| **EKS (Kubernetes)** | local state workloads, complex deployments | Simple local state services (overkill) | Higher operational overhead. Worth it for local state. |
|
|
257
|
-
| **Lambda** | Event-driven, short tasks, glue code | Long-running processes, high-cpu needs | Cold starts matter. Keep functions small. |
|
|
258
|
-
| **Cloud Run** | Previews, internal tools, experiments | Production services, AWS-dependent workloads | Great DX, but it's GCP (see Part A). |
|
|
259
|
-
|
|
260
|
-
## 4B. Rationale
|
|
261
|
-
|
|
262
|
-
**Why ECS as default?**
|
|
263
|
-
|
|
264
|
-
- Simpler than Kubernetes for most use cases
|
|
265
|
-
- Fargate \= no EC2 instance management
|
|
266
|
-
- Good integration with AWS services (ALB, CloudWatch, Secrets Manager)
|
|
267
|
-
- Lower cognitive overhead for teams
|
|
268
|
-
|
|
269
|
-
**Why Kubernetes for local state?**
|
|
270
|
-
|
|
271
|
-
- local stateSets, persistent volumes, and operators
|
|
272
|
-
- Better control over scheduling and node affinity
|
|
273
|
-
- Community ecosystem for databases and local state apps
|
|
274
|
-
- We're investing in EKS anyway; leverage it where it shines
|
|
275
|
-
|
|
276
|
-
**Why Lambda for event-driven?**
|
|
277
|
-
|
|
278
|
-
- Pay-per-invocation is cost-effective for sporadic workloads
|
|
279
|
-
- Built-in scaling, no capacity planning
|
|
280
|
-
- Native integration with SQS, SNS, EventBridge
|
|
281
|
-
- Forces good practices (small, focused functions)
|
|
282
|
-
|
|
283
|
-
**Why Cloud Run for previews?**
|
|
284
|
-
|
|
285
|
-
- Deploys in seconds, scales to zero
|
|
286
|
-
- Cheap for low-traffic ephemeral environments
|
|
287
|
-
- Great developer experience
|
|
288
|
-
- Isolated from production (GCP \= natural boundary)
|
|
289
|
-
|
|
290
|
-
---
|
|
291
|
-
|
|
292
|
-
## 5\. How to Apply
|
|
293
|
-
|
|
294
|
-
### Example: New API Service
|
|
295
|
-
|
|
296
|
-
You're building a new internal API that serves data to the mobile apps.
|
|
297
|
-
|
|
298
|
-
1. **Which cloud?** β AWS (it's production, needs DB access)
|
|
299
|
-
2. **What compute?** β ECS Fargate (local state API)
|
|
300
|
-
3. **How to deploy?** β CloudFormation stack, GitHub Actions workflow
|
|
301
|
-
|
|
302
|
-
### Example: Analytics Dashboard
|
|
303
|
-
|
|
304
|
-
You're building an internal dashboard that queries BigQuery and displays charts.
|
|
305
|
-
|
|
306
|
-
1. **Which cloud?** β GCP (BigQuery-native, no prod DB access)
|
|
307
|
-
2. **What compute?** β Cloud Run (internal tool, low traffic)
|
|
308
|
-
3. **How to deploy?** β Cloud Run from container registry
|
|
309
|
-
|
|
310
|
-
### Example: New Observability Tool (like Langfuse)
|
|
311
|
-
|
|
312
|
-
You're deploying a self-hosted observability tool that needs persistent storage.
|
|
313
|
-
|
|
314
|
-
1. **Which cloud?** β AWS (needs to ingest data from AWS services)
|
|
315
|
-
2. **What compute?** β EKS (local state, needs persistent volumes)
|
|
316
|
-
3. **How to deploy?** β Helm chart, ArgoCD
|
|
317
|
-
|
|
318
|
-
## 6\. Exceptions
|
|
319
|
-
|
|
320
|
-
### Requesting an Exception
|
|
321
|
-
|
|
322
|
-
If your use case doesn't fit the decision tree:
|
|
323
|
-
|
|
324
|
-
1. Document: What you're building, why the default doesn't work
|
|
325
|
-
2. Post in Platform channel
|
|
326
|
-
3. Get sign-off before proceeding
|
|
327
|
-
|
|
328
|
-
Common valid exceptions:
|
|
329
|
-
|
|
330
|
-
- Vendor-specific requirements (e.g., "only runs on GCP")
|
|
331
|
-
- Cost optimization for specific workload patterns
|
|
332
|
-
- Experimentation with explicit time bounds
|
|
333
|
-
|
|
334
|
-
### Grandfathered Services
|
|
335
|
-
|
|
336
|
-
We should apply the [Gartnerβs TIME framework](https://unstoppablesoftware.com/understanding-gartners-time-model-maximizing-business-value-with-software-portfolio-management/#:~:text=1%2E%20Tolerate%20%28Low%20Value%2C%20Low%20Cost%2FRisk) when deciding. After that, we may realize some existing services don't follow this standard. We're not migrating them unless there's a compelling reason. Document them, but don't use them as precedent for new services.
|
|
337
|
-
|
|
338
|
-
## 7\. References
|
|
339
|
-
|
|
340
|
-
- [ECS Service Setup Runbook](https://file+.vscode-resource.vscode-cdn.net/Users/arod/work/doist/twist-threads/work/20260114-platform-standards-analysis/TBD)
|
|
341
|
-
- [EKS Deployment Guide](https://file+.vscode-resource.vscode-cdn.net/Users/arod/work/doist/twist-threads/work/20260114-platform-standards-analysis/TBD)
|
|
342
|
-
- [Lambda Best Practices](https://file+.vscode-resource.vscode-cdn.net/Users/arod/work/doist/twist-threads/work/20260114-platform-standards-analysis/TBD)
|
|
343
|
-
- [Cloud Run Setup for Previews](https://file+.vscode-resource.vscode-cdn.net/Users/arod/work/doist/twist-threads/work/20260114-platform-standards-analysis/TBD)
|
|
344
|
-
- \[Cost Tagging Standard\](TBD \- Group 3\)
|