bmad-method-quarkus 1.0.4 → 1.0.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/src/bmm-skills/agents/bmad-quarkus-build/SKILL.md +25 -5
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-error-handling-i18n/SKILL.md +6 -2
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-hexagonal-core/SKILL.md +106 -38
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-kafka-messaging/SKILL.md +42 -12
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-observability-otel/SKILL.md +191 -13
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-openapi-tmforum/SKILL.md +7 -8
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-security-standards/SKILL.md +132 -0
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-sql-jdbc-agroal/SKILL.md +30 -13
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-temporal-workflows/SKILL.md +616 -0
- package/src/commands/quarkus-all.md +6 -5
|
@@ -0,0 +1,616 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: quarkus-temporal-workflows
|
|
3
|
+
description: Standard for durable orchestration and distributed sagas in Quarkus native services using the Temporal Java SDK — plain `io.temporal:temporal-sdk` (1.27.0+) with NO Quarkiverse extension, manual CDI bootstrap (`@Produces WorkflowClient` + `@Observes StartupEvent` starting the `WorkerFactory`), the `io.temporal.workflow.Saga` compensation pattern, Java records as `@RegisterForReflection` payloads, the GraalVM Netty/gRPC native fix, and `TestWorkflowEnvironment` in-memory testing. Workflows are invoked predominantly over gRPC — inbound through a `*GrpcService` that starts the workflow without waiting, outbound from activities through `common/client` beans — so it also covers the deadline-vs-activity-timeout ordering and the stacked-retry trap. Use this skill whenever the user mentions Temporal, temporal-sdk, workflow orchestration, durable execution, saga, compensation, compensating transaction, activity/activities, WorkflowClient, WorkerFactory, task queue, workflow id, signal/query, `@WorkflowInterface`/`@ActivityInterface`, TestWorkflowEnvironment, long-running or multi-step business processes spanning several slices or services, or "roll back what already committed in another service". Covers BOTH topologies and says which to pick — a standalone **worker-only project** (the default — an independent deployable named with the `-worker` suffix, with no REST layer, no vertical slices and no datasource, where the REST/OpenAPI/slice/SQL standards do not apply) and an embedded worker inside a normal slice app. Use it as well whenever the user asks about a Temporal-only project, an orchestrator service, a worker deployable, or whether a workflow project needs REST endpoints or slices. Applies ONLY to projects that use Temporal; it never applies to a single-slice operation that one `Handler` and one database transaction already cover.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Temporal Durable Orchestration Standard (Quarkus)
|
|
7
|
+
|
|
8
|
+
Applies to any Quarkus backend project **that uses Temporal**; project directives (CLAUDE.md, ADRs, explicit instructions) override these defaults where they conflict.
|
|
9
|
+
|
|
10
|
+
Temporal is for the work a single `@Transactional Handler.process()` cannot make atomic: a business process that spans several slices, several services, or a span of time (minutes to months), where a failure halfway through must **undo** what already committed elsewhere. A JTA transaction already gives you all-or-nothing inside one database — reach for Temporal only when the operation crosses that boundary. Wrapping a one-slice insert in a workflow buys nothing and costs a worker, a task queue and a replay contract.
|
|
11
|
+
|
|
12
|
+
**Two topologies, and §3 is the first decision you make.** The default is a **worker-only project**: an independent deployable that hosts workflows and activities and has no REST layer, no slices and no database — the shape to reach for when the process spans services. The alternative is an **embedded worker** inside a normal slice-architecture app. §3 gives the layout for each and, for worker-only projects, states which house standards go quiet: no `Resource`, no `<Slice>Handler`/`<Slice>Sql`, no OpenAPI, usually no datasource.
|
|
13
|
+
|
|
14
|
+
**Scope inside a slice architecture (embedded only):** Temporal is quarantined to `orchestration/` and `common/temporal`, the same way Mutiny is quarantined to `*GrpcService` and `common/client` (see quarkus-hexagonal-core skill). No `io.temporal.*` type ever appears in a slice folder, a `Handler`, a `Sql` class or a slice DTO. A slice stays fully testable and deployable with Temporal absent from the classpath.
|
|
15
|
+
|
|
16
|
+
## 1. Dependencies — plain SDK, no Quarkiverse extension
|
|
17
|
+
|
|
18
|
+
```xml
|
|
19
|
+
<properties>
|
|
20
|
+
<temporal.version>1.27.0</temporal.version>
|
|
21
|
+
</properties>
|
|
22
|
+
|
|
23
|
+
<dependencies>
|
|
24
|
+
<dependency>
|
|
25
|
+
<groupId>io.temporal</groupId>
|
|
26
|
+
<artifactId>temporal-sdk</artifactId>
|
|
27
|
+
<version>${temporal.version}</version>
|
|
28
|
+
</dependency>
|
|
29
|
+
<dependency>
|
|
30
|
+
<groupId>io.temporal</groupId>
|
|
31
|
+
<artifactId>temporal-opentracing</artifactId> <!-- trace propagation — see §10 -->
|
|
32
|
+
<version>${temporal.version}</version>
|
|
33
|
+
</dependency>
|
|
34
|
+
<dependency>
|
|
35
|
+
<groupId>io.temporal</groupId>
|
|
36
|
+
<artifactId>temporal-testing</artifactId>
|
|
37
|
+
<version>${temporal.version}</version>
|
|
38
|
+
<scope>test</scope>
|
|
39
|
+
</dependency>
|
|
40
|
+
</dependencies>
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
- **`io.temporal:temporal-sdk` 1.27.0 or newer.** Pin an explicit version — the Quarkus BOM does not manage Temporal, so nothing else will.
|
|
44
|
+
- **Do not use a Quarkiverse Temporal extension.** The SDK is wired by hand (§2). The trade is explicit and worth naming: no extension means no build-time wiring, so the CDI bootstrap and the native-image flags (§8) are *yours* to maintain — and in exchange the SDK version moves independently of the Quarkus platform release, which is what a durable-execution dependency needs.
|
|
45
|
+
- `temporal-testing` is `test` scope only. It must never reach the runtime classpath.
|
|
46
|
+
- The code you write follows the patterns in Temporal's official Java samples — interface + implementation pair, activity stubs built inside the workflow, `Saga` for compensation. Do not invent a house abstraction over the SDK; a wrapper around `Workflow.newActivityStub` hides the options that make a workflow correct.
|
|
47
|
+
|
|
48
|
+
## 2. Bootstrap: one CDI bean owns the client and the worker
|
|
49
|
+
|
|
50
|
+
There is no extension to start the worker, so exactly one `@ApplicationScoped` bean in `common/temporal` produces the `WorkflowClient` and starts the `WorkerFactory` on `StartupEvent`.
|
|
51
|
+
|
|
52
|
+
```java
|
|
53
|
+
// common/temporal
|
|
54
|
+
@ApplicationScoped
|
|
55
|
+
public class TemporalBootstrap {
|
|
56
|
+
|
|
57
|
+
@Inject TemporalConfig config;
|
|
58
|
+
|
|
59
|
+
private WorkerFactory factory;
|
|
60
|
+
|
|
61
|
+
@Produces
|
|
62
|
+
@ApplicationScoped
|
|
63
|
+
public WorkflowServiceStubs serviceStubs() {
|
|
64
|
+
return WorkflowServiceStubs.newServiceStubs(
|
|
65
|
+
WorkflowServiceStubsOptions.newBuilder()
|
|
66
|
+
.setTarget(config.target())
|
|
67
|
+
.build());
|
|
68
|
+
}
|
|
69
|
+
|
|
70
|
+
@Produces
|
|
71
|
+
@ApplicationScoped
|
|
72
|
+
public WorkflowClient workflowClient(WorkflowServiceStubs stubs) {
|
|
73
|
+
return WorkflowClient.newInstance(stubs,
|
|
74
|
+
WorkflowClientOptions.newBuilder()
|
|
75
|
+
.setNamespace(config.namespace())
|
|
76
|
+
.build());
|
|
77
|
+
}
|
|
78
|
+
|
|
79
|
+
void onStart(@Observes StartupEvent event,
|
|
80
|
+
WorkflowClient client,
|
|
81
|
+
PartyOnboardingActivitiesImpl activities) {
|
|
82
|
+
factory = WorkerFactory.newInstance(client);
|
|
83
|
+
Worker worker = factory.newWorker(config.taskQueue());
|
|
84
|
+
worker.registerWorkflowImplementationTypes(PartyOnboardingWorkflowImpl.class); // by CLASS
|
|
85
|
+
worker.registerActivitiesImplementations(activities); // by INSTANCE
|
|
86
|
+
factory.start();
|
|
87
|
+
}
|
|
88
|
+
|
|
89
|
+
void onStop(@Observes ShutdownEvent event, WorkflowServiceStubs stubs) {
|
|
90
|
+
if (factory != null) {
|
|
91
|
+
factory.shutdown(); // stop polling
|
|
92
|
+
factory.awaitTermination(20, TimeUnit.SECONDS); // drain in-flight tasks
|
|
93
|
+
}
|
|
94
|
+
stubs.shutdown();
|
|
95
|
+
}
|
|
96
|
+
}
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
The two registration calls are **not** symmetrical, and the asymmetry is the rule that trips people:
|
|
100
|
+
|
|
101
|
+
| Registration | Form | Why |
|
|
102
|
+
|---|---|---|
|
|
103
|
+
| Workflows | `registerWorkflowImplementationTypes(XImpl.class)` — a **class** | Temporal instantiates a fresh workflow object per execution and replays it. **CDI cannot inject into a workflow implementation**; a field injected once would not survive replay. Its only dependencies are activity stubs it builds itself (§4). |
|
|
104
|
+
| Activities | `registerActivitiesImplementations(bean)` — an **instance** | Activities are ordinary side-effecting beans. Pass the CDI-managed instance so `@Inject`ed `Handler`s are wired normally. |
|
|
105
|
+
|
|
106
|
+
Rules:
|
|
107
|
+
- **One bootstrap bean per app.** A second one registering a second worker on the same task queue is a duplicate-poller bug, not scaling — scale with replicas.
|
|
108
|
+
- `onStop` matters: `factory.shutdown()` + `awaitTermination` lets in-flight activity tasks finish instead of being abandoned to a timeout-and-retry cycle. Budget it inside the pod's grace period alongside the telemetry flush (see quarkus-observability-otel skill, "The shutdown budget").
|
|
109
|
+
- Connection settings come from `@ConfigMapping` (`<Area>Config`, per the hexagonal-core naming table), never from literals:
|
|
110
|
+
|
|
111
|
+
```java
|
|
112
|
+
@ConfigMapping(prefix = "temporal")
|
|
113
|
+
public interface TemporalConfig {
|
|
114
|
+
String target();
|
|
115
|
+
String namespace();
|
|
116
|
+
String taskQueue();
|
|
117
|
+
}
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
```properties
|
|
121
|
+
temporal.target=${TEMPORAL_TARGET:localhost:7233}
|
|
122
|
+
temporal.namespace=${TEMPORAL_NAMESPACE:default}
|
|
123
|
+
temporal.task-queue=alva.customer.party-onboarding.v1
|
|
124
|
+
```
|
|
125
|
+
|
|
126
|
+
The endpoint, namespace and any mTLS certificate path are environment-specific: `.env` locally, platform vault in deployed environments, `${VAR}` in properties — see quarkus-security-standards skill §1–2. A Temporal Cloud client certificate is a secret and is never committed.
|
|
127
|
+
|
|
128
|
+
## 3. Topology: a worker-only project (default), or an embedded worker
|
|
129
|
+
|
|
130
|
+
Decide this **before** the first file, and record it in the app's `README.md` — the two shapes have different layouts and suspend different house rules.
|
|
131
|
+
|
|
132
|
+
### A. Worker-only project — the default
|
|
133
|
+
|
|
134
|
+
A deployable that hosts workflows and activities **and nothing else**: no REST, no slices, and normally no database. Every side effect is a remote call to the service that owns that data (§7). This is the right shape whenever the process spans services, for two reasons that matter in production:
|
|
135
|
+
|
|
136
|
+
- Workflow code is **replay-sensitive** (§4). Keeping it in its own deployable means a workflow change is released on its own cadence instead of riding along with an unrelated business-service deploy — and a business service can be redeployed freely without touching in-flight executions.
|
|
137
|
+
- Workers scale on **task-queue backlog**, not on request rate. A pod sized for orchestration has nothing in common with one sized for an API.
|
|
138
|
+
|
|
139
|
+
```
|
|
140
|
+
apps/customer-onboarding-worker/ # worker-only deployable — note the -worker suffix
|
|
141
|
+
├── src/main/java/com/alva/customer/onboarding/
|
|
142
|
+
│ ├── common/
|
|
143
|
+
│ │ ├── temporal/TemporalConfig.java # @ConfigMapping
|
|
144
|
+
│ │ ├── temporal/TemporalBootstrap.java # §2 — the only class touching WorkerFactory
|
|
145
|
+
│ │ ├── client/PartyRegistrar.java # gRPC beans — how activities reach other services
|
|
146
|
+
│ │ ├── client/WalletOpener.java
|
|
147
|
+
│ │ └── exception/BusinessException.java # the house error taxonomy still applies (§9)
|
|
148
|
+
│ └── orchestration/
|
|
149
|
+
│ └── party_onboarding/ # snake_case, capability-named
|
|
150
|
+
│ ├── dto/
|
|
151
|
+
│ │ ├── PartyOnboardingRequestDto.java # record + @RegisterForReflection
|
|
152
|
+
│ │ └── PartyOnboardingResultDto.java
|
|
153
|
+
│ ├── PartyOnboardingWorkflow.java # @WorkflowInterface — the contract
|
|
154
|
+
│ ├── PartyOnboardingWorkflowImpl.java # the saga; deterministic, zero I/O
|
|
155
|
+
│ ├── PartyOnboardingActivities.java # @ActivityInterface
|
|
156
|
+
│ ├── PartyOnboardingActivitiesImpl.java # injects common/client beans — no local Handler
|
|
157
|
+
│ └── README.md # steps, compensations, timeouts, task queue
|
|
158
|
+
├── src/main/proto/ # ONLY if this app owns the trigger RPC (below)
|
|
159
|
+
├── Dockerfile · service.yaml · README.md
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
There is no `Resource`, no `@Path`, no slice folder, no `Sql` class and no `db/` migration. **Do not scaffold them "for later"** — an unused datasource is a connection pool, a credential and a readiness dependency bought for nothing.
|
|
163
|
+
|
|
164
|
+
**The deployable carries the `-worker` suffix, never `-ms`** — `{module}-{capability}-worker`, e.g. `customer-onboarding-worker` (see the suffix table in the quarkus-hexagonal-core skill). The suffix is a deploy-time contract, and this kind of app contradicts everything `-ms` implies: no ingress, no HTTP route to probe, no API to publish, and scaling driven by task-queue backlog instead of request rate. Naming it `-ms` would hand the platform a service it would try to route to and scale on RPS it never receives. Its `service.yaml` matches: no provided APIs, task queues listed instead.
|
|
165
|
+
|
|
166
|
+
**Which monorepo owns it:** the domain that owns the business **process**, even when most steps call other domains. Cross-domain calls go through contracts published in `contracts/`, exactly like any other inter-domain call.
|
|
167
|
+
|
|
168
|
+
#### What does *not* apply in a worker-only project
|
|
169
|
+
|
|
170
|
+
Marcus reads eight universal standards; several of them have nothing to govern here. Applying them anyway produces a REST layer nobody calls and a database nobody reads.
|
|
171
|
+
|
|
172
|
+
| Standard | Status in a worker-only project |
|
|
173
|
+
|---|---|
|
|
174
|
+
| `quarkus-openapi-tmforum` | **Does not apply.** No `Resource`, no OpenAPI document, no Swagger. The only HTTP surface is `/q/health` and `/q/metrics`. |
|
|
175
|
+
| Vertical slices (`quarkus-hexagonal-core`) | **Does not apply.** No slice folders, no `<Slice>Handler`, no `<Slice>Sql`. `orchestration/` + `common/` *is* the structure. |
|
|
176
|
+
| `quarkus-sql-jdbc-agroal` | **Usually does not apply** — Temporal holds the process state, so there is no datasource. If the app genuinely owns a table (an idempotency/dedup ledger), its SQL follows that skill, with one adjustment: there is no `Handler` below the activity, so **the activity itself plays the Handler role** — it carries the `@Transactional`, opens the single `Connection` and passes it to every `Sql` call. That is the only case where an activity is transactional. |
|
|
177
|
+
| `quarkus-grpc-services` | **Applies fully** — it is how activities reach every other service (§7), and how the trigger RPC is exposed when this app owns one. |
|
|
178
|
+
| `quarkus-error-handling-i18n` | Error codes and the `BusinessException` taxonomy apply (§9). The REST edge mapper is absent; the gRPC interceptor applies only if the app exposes an RPC. |
|
|
179
|
+
| `quarkus-observability-otel` | **Applies fully, and matters more** — there is no request log to fall back on when a workflow misbehaves. |
|
|
180
|
+
| `quarkus-security-standards` | **Applies fully** — Temporal endpoint, namespace and mTLS certificates are secrets (§2). |
|
|
181
|
+
| `quarkus-kafka-messaging` | Applies when a workflow is triggered by, or emits, domain events. |
|
|
182
|
+
| Rest of `quarkus-hexagonal-core` | Naming, `apps/` layout, Java 25, native build + Dockerfile, `service.yaml` + README: **all unchanged**. |
|
|
183
|
+
|
|
184
|
+
ArchUnit ships a **subset**: `bannedSuffixes` (with the `..orchestration..` `*Impl` carve-out below), `reactiveIsQuarantined`, `noOrm`, the DTO rules, plus `workflowsDoNoIo`. Drop `slicesAreIndependent`, `onlyHandlersTouchSql`, `transactionalOnlyOnProcess`, `handlersExposeOnlyProcess`, `restResources` and `transportHasNoJdbc` — with no slices to match they can never fail, and a rule that cannot fail is noise that makes the suite look stronger than it is.
|
|
185
|
+
|
|
186
|
+
#### Who starts the workflow
|
|
187
|
+
|
|
188
|
+
- **Default: the calling service does.** The worker exposes **no inbound API at all** — it polls its task queue and nothing else. The service that owns the triggering business event holds its own `*Starter` bean (§6); the workflow id and the payload record are the contract between them. Nothing to route, nothing to secure at the edge.
|
|
189
|
+
- **Alternative: the worker owns the trigger RPC.** Then it has `src/main/proto/` and one `*GrpcService`, and that adapter calls the `*Starter` **directly — with no `Handler` in between.** This is a deliberate deviation from "a transport adapter always delegates to a `Handler`": a `Handler` exists to hold business logic and a transaction, and a trigger RPC has neither. An empty pass-through `Handler` would be pure ceremony. The adapter still does only proto↔DTO translation plus the start call.
|
|
190
|
+
|
|
191
|
+
Readiness reflects the **worker**, not a route: can it reach the Temporal frontend, and is the factory polling. Scale on task-queue backlog (§10), never on RPS.
|
|
192
|
+
|
|
193
|
+
### B. Embedded worker — the exception
|
|
194
|
+
|
|
195
|
+
A normal `-ms` app that also hosts a worker — and it **stays `-ms`**, because it still serves an API; the `-worker` suffix is reserved for deployables that serve none. Use it only when every step of the saga is a capability **this same app already owns**, and Temporal is there for durability, retries or long waits rather than for crossing a service boundary. The cost is that the app now has two runtime roles and a replay-sensitive workflow tied to the service's release cadence.
|
|
196
|
+
|
|
197
|
+
```
|
|
198
|
+
src/main/java/com/<company>/<module>/
|
|
199
|
+
├── common/temporal/… # as above, plus PartyOnboardingStarter (§6)
|
|
200
|
+
├── orchestration/
|
|
201
|
+
│ └── party_onboarding/… # same contents; activities call local Handlers
|
|
202
|
+
└── create_party_individual/ # slices — untouched, no Temporal import
|
|
203
|
+
```
|
|
204
|
+
|
|
205
|
+
Here the dependency direction is deliberately **one-way**: `orchestration/` may depend on slice `Handler`s and slice DTOs; a slice may never depend on `orchestration/`. A slice that imports a workflow has been made un-deployable without a Temporal cluster. Extend `slicesAreIndependent` with the single asymmetric exemption shown below, and keep every other slice rule intact — in this topology they all still bite.
|
|
206
|
+
|
|
207
|
+
### Naming
|
|
208
|
+
|
|
209
|
+
| Artifact | Convention | Example |
|
|
210
|
+
|---|---|---|
|
|
211
|
+
| Workflow interface | `<Capability>Workflow`, `@WorkflowInterface` | `PartyOnboardingWorkflow` |
|
|
212
|
+
| Workflow implementation | `<Capability>WorkflowImpl` | `PartyOnboardingWorkflowImpl` |
|
|
213
|
+
| Activity interface | `<Capability>Activities`, `@ActivityInterface` | `PartyOnboardingActivities` |
|
|
214
|
+
| Activity implementation | `<Capability>ActivitiesImpl` | `PartyOnboardingActivitiesImpl` |
|
|
215
|
+
| Workflow method | imperative verb, `@WorkflowMethod` | `onboard` |
|
|
216
|
+
| Signal / Query | `<verb>` / `get<Noun>` | `approve` / `getStatus` |
|
|
217
|
+
| Starter bean | capability noun in `common/temporal` | `PartyOnboardingStarter` |
|
|
218
|
+
| Payload DTO | `*Dto` **record** in `dto/` | `PartyOnboardingRequestDto` |
|
|
219
|
+
| Task queue | `<org>.<module>.<capability-kebab>.v<major>` | `alva.customer.party-onboarding.v1` |
|
|
220
|
+
| Workflow id | `<capability-kebab>-<business-id>` — deterministic | `party-onboarding-9f3e…` |
|
|
221
|
+
|
|
222
|
+
**`*Impl` is banned everywhere else in this codebase and sanctioned here.** The SDK builds its client stubs from the annotated interface, so the interface/implementation pair is structural, not stylistic, and the official samples name it `*Impl`. Carve it out of the ArchUnit `bannedSuffixes` rule **by package**, exactly as `common/util/StringUtils` and the `Jdbc` helper are carved out — never by weakening the rule:
|
|
223
|
+
|
|
224
|
+
```java
|
|
225
|
+
@ArchTest
|
|
226
|
+
static final ArchRule bannedSuffixesOutsideOrchestration = noClasses()
|
|
227
|
+
.that().resideOutsideOfPackage("..orchestration..")
|
|
228
|
+
.should().haveSimpleNameEndingWith("Impl"); // …plus the existing suffixes
|
|
229
|
+
|
|
230
|
+
// Temporal stays quarantined
|
|
231
|
+
@ArchTest
|
|
232
|
+
static final ArchRule temporalIsQuarantined = noClasses()
|
|
233
|
+
.that().resideOutsideOfPackage("..orchestration..")
|
|
234
|
+
.and().resideOutsideOfPackage("..common.temporal..")
|
|
235
|
+
.should().dependOnClassesThat().resideInAPackage("io.temporal..");
|
|
236
|
+
|
|
237
|
+
// A workflow implementation performs no I/O of its own (§4)
|
|
238
|
+
@ArchTest
|
|
239
|
+
static final ArchRule workflowsDoNoIo = noClasses()
|
|
240
|
+
.that().haveSimpleNameEndingWith("WorkflowImpl")
|
|
241
|
+
.should().dependOnClassesThat().resideInAnyPackage(
|
|
242
|
+
"java.sql..", "javax.sql..", "jakarta.ws.rs..", "io.quarkus.grpc..");
|
|
243
|
+
```
|
|
244
|
+
|
|
245
|
+
Extend the existing `slicesAreIndependent` rule with **one** asymmetric exemption — orchestration may reach into slices, never the reverse:
|
|
246
|
+
|
|
247
|
+
```java
|
|
248
|
+
.ignoreDependency(resideInAPackage("..orchestration.."), alwaysTrue())
|
|
249
|
+
```
|
|
250
|
+
|
|
251
|
+
## 4. The workflow implementation: deterministic, and that is a hard constraint
|
|
252
|
+
|
|
253
|
+
Temporal replays workflow code from its event history after every worker restart. Replay must produce the identical sequence of commands, so a workflow implementation is **pure orchestration**: it calls activity stubs and does nothing else.
|
|
254
|
+
|
|
255
|
+
Banned inside a `*WorkflowImpl`, with the replacement the SDK provides:
|
|
256
|
+
|
|
257
|
+
| Never | Always |
|
|
258
|
+
|---|---|
|
|
259
|
+
| `System.currentTimeMillis()`, `LocalDateTime.now()` | `Workflow.currentTimeMillis()` |
|
|
260
|
+
| `UUID.randomUUID()`, `Math.random()` | `Workflow.randomUUID()`, `Workflow.newRandom()` |
|
|
261
|
+
| `Thread.sleep(…)` | `Workflow.sleep(Duration…)` |
|
|
262
|
+
| `new Thread(…)`, executor services | `Async.function(…)` / `Async.procedure(…)`, scoped with `Workflow.newCancellationScope` |
|
|
263
|
+
| JDBC, HTTP, gRPC, Kafka, file I/O | an **activity** |
|
|
264
|
+
| `Logger.getLogger(…)` | `Workflow.getLogger(…)` — suppresses duplicate lines on replay |
|
|
265
|
+
| iteration over `HashMap`/`HashSet` | `LinkedHashMap`/`TreeMap`, or sort explicitly |
|
|
266
|
+
|
|
267
|
+
`@Inject` does not work here either (§2) — the workflow's only collaborators are its activity stubs. A workflow that needs configuration receives it as a **workflow argument**, never read from `ConfigProvider` at replay time.
|
|
268
|
+
|
|
269
|
+
**Versioning is the production footgun.** Editing a workflow implementation while executions of the old code are still running causes a non-deterministic replay and a stuck workflow. Any change to the *sequence* of activity calls requires `Workflow.getVersion("addWallet", DEFAULT_VERSION, 1)` branching, or a new task queue (`…v2`) drained in parallel. Changing an activity's *body* is always safe; changing the workflow's call order never is.
|
|
270
|
+
|
|
271
|
+
## 5. The saga: `io.temporal.workflow.Saga` with explicit compensations
|
|
272
|
+
|
|
273
|
+
```java
|
|
274
|
+
public class PartyOnboardingWorkflowImpl implements PartyOnboardingWorkflow {
|
|
275
|
+
|
|
276
|
+
private static final Logger log = Workflow.getLogger(PartyOnboardingWorkflowImpl.class);
|
|
277
|
+
|
|
278
|
+
private final PartyOnboardingActivities activities = Workflow.newActivityStub(
|
|
279
|
+
PartyOnboardingActivities.class,
|
|
280
|
+
ActivityOptions.newBuilder()
|
|
281
|
+
.setStartToCloseTimeout(Duration.ofSeconds(30))
|
|
282
|
+
.setScheduleToCloseTimeout(Duration.ofMinutes(10)) // the outer bound — never omit
|
|
283
|
+
.setRetryOptions(RetryOptions.newBuilder()
|
|
284
|
+
.setInitialInterval(Duration.ofSeconds(1))
|
|
285
|
+
.setBackoffCoefficient(2.0)
|
|
286
|
+
.setMaximumAttempts(5)
|
|
287
|
+
.setDoNotRetry("BusinessException") // matches the FAILURE TYPE — see §9
|
|
288
|
+
.build())
|
|
289
|
+
.build());
|
|
290
|
+
|
|
291
|
+
@Override
|
|
292
|
+
public PartyOnboardingResultDto onboard(PartyOnboardingRequestDto request) {
|
|
293
|
+
Saga saga = new Saga(new Saga.Options.Builder()
|
|
294
|
+
.setParallelCompensation(false) // compensate in reverse order, one at a time
|
|
295
|
+
.build());
|
|
296
|
+
try {
|
|
297
|
+
String partyId = activities.createParty(request);
|
|
298
|
+
saga.addCompensation(activities::deleteParty, partyId);
|
|
299
|
+
|
|
300
|
+
String walletId = activities.openWallet(partyId, request.currency());
|
|
301
|
+
saga.addCompensation(activities::closeWallet, walletId);
|
|
302
|
+
|
|
303
|
+
activities.sendWelcomeNotification(partyId);
|
|
304
|
+
return new PartyOnboardingResultDto(partyId, walletId);
|
|
305
|
+
|
|
306
|
+
} catch (TemporalFailure e) { // NOT just ActivityFailure — see below
|
|
307
|
+
log.warn("Onboarding failed, compensating: {}", e.getMessage());
|
|
308
|
+
saga.compensate();
|
|
309
|
+
throw e; // the workflow still fails — do not swallow
|
|
310
|
+
}
|
|
311
|
+
}
|
|
312
|
+
}
|
|
313
|
+
```
|
|
314
|
+
|
|
315
|
+
Rules:
|
|
316
|
+
- **Catch `TemporalFailure`, not `ActivityFailure`.** `ActivityFailure` misses `ChildWorkflowFailure`, `CanceledFailure` and every other `TemporalFailure` subtype — and a miss means the compensation chain silently never runs, which is the exact failure the saga exists to prevent.
|
|
317
|
+
- **Decide `Saga.Options.setContinueWithError` deliberately.** It defaults to `false`, so the *first* failing compensation aborts the rest of the chain and leaves the earlier steps uncompensated. Set it to `true` when the compensations are independent and you would rather attempt all of them; leave it `false` only when a failed compensation genuinely makes the remaining ones unsafe.
|
|
318
|
+
- **Register the compensation *after* the forward activity returns**, using its result. Registering before means compensating something that never happened.
|
|
319
|
+
- Compensations run in **reverse registration order** (LIFO). Keep `setParallelCompensation(false)` unless the steps are provably independent — parallel compensation of dependent resources reintroduces the ordering bug the saga exists to avoid.
|
|
320
|
+
- **`saga.compensate()` then rethrow.** A saga that compensates and returns normally reports success for a business process that did not happen.
|
|
321
|
+
- A compensation is itself an activity, so it is retried like any other — and it must be **idempotent**: `deleteParty` on an already-deleted party succeeds quietly rather than failing the compensation chain.
|
|
322
|
+
- No compensation is possible for some steps (an email that went out). Order the workflow so irreversible steps come **last**, after everything reversible has succeeded.
|
|
323
|
+
|
|
324
|
+
## 6. Activities: the only side-effecting layer
|
|
325
|
+
|
|
326
|
+
An activity implementation is thin: extract the arguments, perform **exactly one** side effect, return a plain value. **No business logic, no orchestration, no `@Transactional` of its own** — the transaction always lives one level down. What it delegates *to* is the one thing the topology (§3) changes:
|
|
327
|
+
|
|
328
|
+
```java
|
|
329
|
+
// A. Worker-only project (default): the side effect is a remote call
|
|
330
|
+
@ApplicationScoped
|
|
331
|
+
public class PartyOnboardingActivitiesImpl implements PartyOnboardingActivities {
|
|
332
|
+
|
|
333
|
+
@Inject PartyRegistrar partyRegistrar; // common/client — gRPC, awaits internally (§7)
|
|
334
|
+
@Inject WalletOpener walletOpener;
|
|
335
|
+
|
|
336
|
+
@Override
|
|
337
|
+
public String createParty(PartyOnboardingRequestDto request) {
|
|
338
|
+
return partyRegistrar.register(request.tenantId(), request.givenName(), request.familyName());
|
|
339
|
+
}
|
|
340
|
+
}
|
|
341
|
+
```
|
|
342
|
+
|
|
343
|
+
```java
|
|
344
|
+
// B. Embedded worker: the side effect is a Handler in this same app
|
|
345
|
+
@Inject CreatePartyIndividualHandler createPartyHandler;
|
|
346
|
+
|
|
347
|
+
@Override
|
|
348
|
+
public String createParty(PartyOnboardingRequestDto request) {
|
|
349
|
+
return createPartyHandler.process(toSliceRequest(request)).getPartyId();
|
|
350
|
+
}
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
Either way the activity is a transport adapter in the sense the hexagonal-core skill defines, and the `@Transactional` boundary is the remote service's `Handler` (A) or the local one (B) — never here. The single exception is a worker-only app that owns a table of its own (§3): with no `Handler` beneath it, that activity plays the `Handler` role and carries the transaction itself.
|
|
354
|
+
|
|
355
|
+
- **Activities are at-least-once.** Temporal retries on worker crash, timeout and failure, so every activity must be idempotent. Where the underlying operation is not naturally idempotent, dedupe on a business key or on `Activity.getExecutionContext().getInfo().getWorkflowId()`.
|
|
356
|
+
- Every activity needs a `setScheduleToCloseTimeout` (or a `setMaximumAttempts`). An activity with unlimited retries and no outer bound does not fail — it hangs forever, in production and in your tests.
|
|
357
|
+
- Long-running activities heartbeat (`Activity.getExecutionContext().heartbeat(progress)`) so a dead worker is detected in seconds rather than at the start-to-close timeout.
|
|
358
|
+
|
|
359
|
+
### Starting a workflow
|
|
360
|
+
|
|
361
|
+
The `*Starter` bean below is the same in both topologies — what differs is **which deployable owns it**:
|
|
362
|
+
|
|
363
|
+
- **Worker-only (A):** it lives in the *calling* service, which therefore carries `temporal-sdk` and a `common/temporal` package of its own (client only — it registers no worker). The worker project itself does not start workflows; it runs them.
|
|
364
|
+
- **Embedded (B):** it lives in this app, and a slice `Handler` injects it.
|
|
365
|
+
|
|
366
|
+
In the embedded case a `Handler` must start the workflow through this bean and **never** by injecting `WorkflowClient` into the slice — that would drag `io.temporal` past the §3 quarantine:
|
|
367
|
+
|
|
368
|
+
```java
|
|
369
|
+
// common/temporal
|
|
370
|
+
@ApplicationScoped
|
|
371
|
+
public class PartyOnboardingStarter {
|
|
372
|
+
|
|
373
|
+
@Inject WorkflowClient client;
|
|
374
|
+
@Inject TemporalConfig config;
|
|
375
|
+
|
|
376
|
+
public void start(String partyId, PartyOnboardingRequestDto request) {
|
|
377
|
+
PartyOnboardingWorkflow workflow = client.newWorkflowStub(
|
|
378
|
+
PartyOnboardingWorkflow.class,
|
|
379
|
+
WorkflowOptions.newBuilder()
|
|
380
|
+
.setTaskQueue(config.taskQueue())
|
|
381
|
+
.setWorkflowId("party-onboarding-" + partyId) // deterministic = deduplicated
|
|
382
|
+
.setWorkflowIdReusePolicy(
|
|
383
|
+
WorkflowIdReusePolicy.WORKFLOW_ID_REUSE_POLICY_ALLOW_DUPLICATE_FAILED_ONLY)
|
|
384
|
+
.build());
|
|
385
|
+
WorkflowClient.start(workflow::onboard, request); // async — never block the Handler
|
|
386
|
+
}
|
|
387
|
+
}
|
|
388
|
+
```
|
|
389
|
+
|
|
390
|
+
**Starting a workflow inside a database transaction is a dual write.** If `process()` rolls back after the start call succeeded, the workflow orchestrates an entity that does not exist. This is the same hazard the outbox solves for Kafka (see quarkus-kafka-messaging skill), and it gets the same two answers:
|
|
391
|
+
|
|
392
|
+
1. **Deterministic workflow id derived from the business entity** (above) — a retry of the whole operation cannot start a second workflow, so the failure mode is recoverable rather than duplicated. This is mandatory, not optional. Use `ALLOW_DUPLICATE_FAILED_ONLY`, **not** `REJECT_DUPLICATE`: the latter blocks reuse even after the run closed, so a legitimate re-run of that business id becomes permanently impossible. The starter must also catch `WorkflowExecutionAlreadyStarted` and treat it as success — that exception *is* the deduplication working, not an error to propagate.
|
|
393
|
+
2. **Start after commit**, the way `OutboxDispatcher.dispatchAfterCommit(eventId)` already does it — or record an outbox row and let the relay start the workflow. Never call `WorkflowClient.start` in the middle of `execution()` and hope the transaction commits.
|
|
394
|
+
|
|
395
|
+
## 7. Transport: the invocations are gRPC
|
|
396
|
+
|
|
397
|
+
Orchestration is internal by nature — a saga exists because a process crosses services — so **gRPC is the predominant transport at both ends of a workflow**, per the house rule that internal synchronous calls are always gRPC and never internal REST (see quarkus-grpc-services and quarkus-hexagonal-core skills). North-bound REST triggers exist (a customer action that kicks off onboarding) but are the minority; the default assumption for a Temporal capability is gRPC in, gRPC out.
|
|
398
|
+
|
|
399
|
+
**A. Worker-only project** — the calling service starts the workflow; the worker only runs it and calls back out over gRPC:
|
|
400
|
+
|
|
401
|
+
```
|
|
402
|
+
calling service (-ms) worker project (-worker) other services
|
|
403
|
+
┌───────────────────────┐ ┌────────────────────────┐
|
|
404
|
+
│ GrpcService → Handler │ │ PartyOnboardingWorkflowImpl │ (no I/O here)
|
|
405
|
+
│ ↓ │ │ ↓ │
|
|
406
|
+
│ <Capability>Starter ─┼── Temporal ──┼→ PartyOnboardingActivitiesImpl │
|
|
407
|
+
└───────────────────────┘ task queue │ ↓ │
|
|
408
|
+
│ common/client bean ───┼─ gRPC ─→ party-ms, wallet-ms
|
|
409
|
+
└────────────────────────┘
|
|
410
|
+
```
|
|
411
|
+
|
|
412
|
+
**B. Embedded worker** — everything in one deployable, the activity reaching a local `Handler`:
|
|
413
|
+
|
|
414
|
+
```
|
|
415
|
+
<Proto>GrpcService → Handler.process() → <Capability>Starter → Temporal
|
|
416
|
+
(slice folder) (slice folder) (common/temporal)
|
|
417
|
+
↓
|
|
418
|
+
PartyOnboardingWorkflowImpl (no I/O — orchestration only)
|
|
419
|
+
↓
|
|
420
|
+
PartyOnboardingActivitiesImpl → slice Handler
|
|
421
|
+
```
|
|
422
|
+
|
|
423
|
+
Nothing about the inbound chain is special-cased for Temporal: the `*GrpcService` delegates to a `Handler`, and the `Handler` injects the `*Starter`. **A `*GrpcService` never injects `WorkflowClient` directly** — that would put `io.temporal` inside a slice folder and break the §3 quarantine. The one exception is the trigger RPC of a worker-only project, which has no `Handler` to delegate to and calls the `*Starter` itself (§3).
|
|
424
|
+
|
|
425
|
+
### Inbound: the gRPC call starts the workflow, it does not wait for it
|
|
426
|
+
|
|
427
|
+
A gRPC client carries a deadline measured in seconds (`quarkus.grpc.clients.<name>.deadline=2s`); a saga runs for minutes, days or months. So the RPC that triggers a workflow **returns as soon as the workflow is started** — the workflow id is the response — and never blocks on completion:
|
|
428
|
+
|
|
429
|
+
- Start asynchronously (`WorkflowClient.start(...)`, as in §6), return the workflow id to the caller.
|
|
430
|
+
- The caller learns the outcome by polling a `@QueryMethod`, by consuming the domain event the workflow's final activity emits (outbox → Kafka, per the messaging skill), or by its own callback RPC.
|
|
431
|
+
- Blocking the RPC on `workflow.onboard(request)` ties a durable execution to a 2-second deadline: the deadline fires, the caller retries, and the only thing the deterministic workflow id saved you from was a second workflow. The first one is still running, unobserved.
|
|
432
|
+
|
|
433
|
+
### Outbound: activities call other services through `common/client`
|
|
434
|
+
|
|
435
|
+
An activity that reaches another service uses a capability-named bean from `common/client` — the same bean a `Handler` would use, never a raw `@GrpcClient` stub and never an internal REST client:
|
|
436
|
+
|
|
437
|
+
```java
|
|
438
|
+
@ApplicationScoped
|
|
439
|
+
public class PartyOnboardingActivitiesImpl implements PartyOnboardingActivities {
|
|
440
|
+
|
|
441
|
+
@Inject CredentialsValidator credentialsValidator; // common/client — awaits internally
|
|
442
|
+
|
|
443
|
+
@Override
|
|
444
|
+
public boolean validateCredentials(String username, String password) {
|
|
445
|
+
return credentialsValidator.validate(username, password); // plain types in, plain types out
|
|
446
|
+
}
|
|
447
|
+
}
|
|
448
|
+
```
|
|
449
|
+
|
|
450
|
+
Because those beans `await()` internally and expose plain types, the `reactiveIsQuarantined` ArchUnit rule keeps passing unchanged — no Mutiny ever reaches `orchestration/`, exactly as none reaches a `Handler`.
|
|
451
|
+
|
|
452
|
+
**Only activities may call out. A workflow never holds a client**, gRPC or otherwise (§4).
|
|
453
|
+
|
|
454
|
+
### Two timeout layers, and the order between them is not optional
|
|
455
|
+
|
|
456
|
+
```
|
|
457
|
+
grpc deadline < activity startToCloseTimeout < activity scheduleToCloseTimeout
|
|
458
|
+
2s 30s 10m
|
|
459
|
+
```
|
|
460
|
+
|
|
461
|
+
Get this backwards and the activity times out while the gRPC call is still in flight: Temporal schedules a retry, and now two identical calls are hitting a downstream that is already struggling. The gRPC deadline must be the *inner* bound so the call is dead before Temporal reacts to it.
|
|
462
|
+
|
|
463
|
+
**Do not stack retries.** Temporal's activity retry policy and the gRPC/mesh retry policy multiply: 5 activity attempts × 3 mesh retries = 15 calls, sent to the exact service that is failing. Temporal owns the retry — its attempts are durable, backed off, and visible in the workflow history, where a mesh retry is invisible. Keep the client/mesh layer at a single retry for connection-level `UNAVAILABLE` blips at most, and let the activity policy do the rest.
|
|
464
|
+
|
|
465
|
+
### Mapping gRPC failures onto the retry policy
|
|
466
|
+
|
|
467
|
+
A `StatusRuntimeException` from a downstream is classified at the activity boundary, the same way §9 classifies the house exceptions:
|
|
468
|
+
|
|
469
|
+
| gRPC status from the downstream | Temporal treatment |
|
|
470
|
+
|---|---|
|
|
471
|
+
| `INVALID_ARGUMENT`, `FAILED_PRECONDITION`, `NOT_FOUND`, `ALREADY_EXISTS`, `PERMISSION_DENIED` | **non-retryable** — translate to the slice's `BusinessException` code, then to a non-retryable `ApplicationFailure` |
|
|
472
|
+
| `UNAVAILABLE`, `DEADLINE_EXCEEDED`, `RESOURCE_EXHAUSTED`, `ABORTED` | retryable — this is what the activity retry policy is for |
|
|
473
|
+
| `INTERNAL`, `UNKNOWN` | retryable, but bounded by `scheduleToCloseTimeout` |
|
|
474
|
+
|
|
475
|
+
The downstream already maps its `BusinessException` codes onto gRPC statuses on the way out (see the grpc skill's error-mapping table), so this table is that one read in reverse — keep the two consistent when either changes.
|
|
476
|
+
|
|
477
|
+
### The SDK's own transport
|
|
478
|
+
|
|
479
|
+
The worker and the client talk to the Temporal service over gRPC as well, on shaded Netty. That is not trivia: it is why the native build needs the run-time initialization flag in §8, why a Temporal Cloud connection is mTLS with a client certificate handled as a secret (§2), and why `WorkflowServiceStubsOptions` carries its own RPC timeout and keepalive settings, separate from every timeout above.
|
|
480
|
+
|
|
481
|
+
## 8. Payloads and native image
|
|
482
|
+
|
|
483
|
+
**Payloads are Java records, always annotated `@RegisterForReflection`.** Records are immutable, which is what replay safety wants, and the SDK's default `DataConverter` serializes them with Jackson — which needs reflection metadata that GraalVM will not infer.
|
|
484
|
+
|
|
485
|
+
```java
|
|
486
|
+
@RegisterForReflection
|
|
487
|
+
public record PartyOnboardingRequestDto(
|
|
488
|
+
String tenantId,
|
|
489
|
+
String givenName,
|
|
490
|
+
String familyName,
|
|
491
|
+
String currency) {}
|
|
492
|
+
```
|
|
493
|
+
|
|
494
|
+
This is a deliberate local deviation from the slice DTO convention (Lombok `@Data` classes, per hexagonal-core): a Temporal payload is serialized into an immutable event history, so a mutable DTO with setters is the wrong shape. The `*Dto` suffix and the `dto/` package still apply, so the existing `dtosStayInDtoPackage` ArchUnit rule keeps passing.
|
|
495
|
+
|
|
496
|
+
**The native build needs the gRPC/Netty run-time initialization flag.** Temporal talks gRPC over shaded Netty, which fails at image-build time without it:
|
|
497
|
+
|
|
498
|
+
```properties
|
|
499
|
+
quarkus.native.additional-build-args=--initialize-at-run-time=io.grpc.netty.shaded.io.netty
|
|
500
|
+
```
|
|
501
|
+
|
|
502
|
+
- This property is a **comma-separated list**: if the app already sets it, append rather than overwrite — a silently replaced flag is a native failure two builds later.
|
|
503
|
+
- The SDK builds workflow and activity stubs with JDK dynamic proxies, which is precisely the pattern hexagonal-core tells you to avoid. Here it is unavoidable, so it is a *verified* exception, not an assumed one: if the native build or the binary fails on a missing proxy, register the interface with Quarkus's `@RegisterForProxy`. (The old `-H:DynamicProxyConfigurationResources` flag is gone on GraalVM for JDK 21+ — proxies are handled through reachability metadata now; do not reach for it.) Either way, **a Temporal app is not "native-ready" until `@QuarkusIntegrationTest` has run green against the built binary** — the JVM-mode test proves nothing about this.
|
|
504
|
+
|
|
505
|
+
## 9. Errors and retries
|
|
506
|
+
|
|
507
|
+
Temporal's retry policy is the automatic remediation this codebase otherwise has to hand-write — but only if failures are classified correctly. The house exception taxonomy (see quarkus-error-handling-i18n and quarkus-sql-jdbc-agroal skills) maps cleanly:
|
|
508
|
+
|
|
509
|
+
| Exception from the `Handler` | Temporal treatment | Why |
|
|
510
|
+
|---|---|---|
|
|
511
|
+
| `BusinessException` (validation, business rule) | **non-retryable** — `setDoNotRetry(...)` | The input will not become valid on attempt 4. Retrying burns the retry budget and delays compensation. |
|
|
512
|
+
| `TransientPersistenceException`, `PersistenceTimeoutException` (deadlock, `40001`, `57014`) | retryable — the default | Exactly what the retry policy exists for. |
|
|
513
|
+
| `StaleVersionException` (optimistic lock) | retryable, small `maximumAttempts` | A genuine concurrent-update race that usually clears. |
|
|
514
|
+
| `PersistenceException` (infrastructure) | retryable, capped by `scheduleToCloseTimeout` | Recoverable, but must not retry forever. |
|
|
515
|
+
|
|
516
|
+
Make the classification explicit at the activity boundary rather than relying on class-name matching alone:
|
|
517
|
+
|
|
518
|
+
```java
|
|
519
|
+
catch (BusinessException e) {
|
|
520
|
+
// The 2nd argument is the failure TYPE and must match setDoNotRetry("BusinessException") in §5 —
|
|
521
|
+
// passing e.code() there instead would leave the declarative rule matching nothing.
|
|
522
|
+
// The error code travels as a detail, so the caller still gets it.
|
|
523
|
+
throw ApplicationFailure.newNonRetryableFailure(e.getMessage(), "BusinessException", e.code());
|
|
524
|
+
}
|
|
525
|
+
```
|
|
526
|
+
|
|
527
|
+
`newNonRetryableFailure` already marks the failure non-retryable on its own, so the `setDoNotRetry` entry in §5 is the declarative belt to this braces — but only while the two strings agree. Keep them in sync, or the retry policy silently stops matching.
|
|
528
|
+
|
|
529
|
+
The failure the client ultimately sees is still a house error code: `WorkflowFailedException` surfacing at the REST edge is translated by `GlobalExceptionHandler` into the catalogued `<MOD>-<HTTP>-<seq>` for that capability, localized as usual. A raw Temporal stack trace never reaches an API consumer.
|
|
530
|
+
|
|
531
|
+
## 10. Observability
|
|
532
|
+
|
|
533
|
+
- Propagate trace context into workflows and activities with `OpenTracingClientInterceptor` / `OpenTracingWorkerInterceptor`, registered on `WorkflowClientOptions` and `WorkerFactoryOptions` in the bootstrap bean. They come from `io.temporal:temporal-opentracing` (declared in §1) and need `io.opentelemetry:opentelemetry-opentracing-shim` to bridge onto the OTel SDK this platform uses — there is no OTel-native interceptor module, which is why a tracing-only dependency shows up in an otherwise minimal pom. Without them a workflow is a hole in the trace: the REST span ends, and the activity spans belong to no one.
|
|
534
|
+
- Feed SDK metrics into the mandatory Micrometer registry (see quarkus-observability-otel skill) via `MicrometerClientStatsReporter` on `WorkflowServiceStubsOptions.setMetricsScope(...)`. Workflow task latency and activity failure rates are the client-side signals that a worker fleet is unhealthy. **Task-queue backlog is not among them** — it is server-side state, read from `DescribeTaskQueue` or the Temporal server's own metrics, so wire the autoscaling signal there rather than expecting it in `/q/metrics`.
|
|
535
|
+
- Log inside workflow code only through `Workflow.getLogger(...)`; a plain logger re-emits every line on every replay and makes the log unreadable exactly when you need it.
|
|
536
|
+
- Workflow ids and task queues are bounded, business-meaningful values — good span attributes. The workflow *arguments* are not: they carry the same PII prohibition as everything else.
|
|
537
|
+
|
|
538
|
+
## 11. Testing: `TestWorkflowEnvironment`, in memory
|
|
539
|
+
|
|
540
|
+
Saga behaviour is tested with Temporal's in-memory environment — plain JUnit 5 + Mockito, **no `@QuarkusTest`**, consistent with how `Handler`s are tested (see hexagonal-core testing table). The environment skips time, so a `Workflow.sleep(Duration.ofDays(30))` resolves instantly.
|
|
541
|
+
|
|
542
|
+
```java
|
|
543
|
+
class PartyOnboardingWorkflowTest {
|
|
544
|
+
|
|
545
|
+
private static final String TASK_QUEUE = "test-queue";
|
|
546
|
+
|
|
547
|
+
private TestWorkflowEnvironment testEnv;
|
|
548
|
+
private Worker worker;
|
|
549
|
+
private PartyOnboardingActivities activities;
|
|
550
|
+
|
|
551
|
+
@BeforeEach
|
|
552
|
+
void setUp() {
|
|
553
|
+
testEnv = TestWorkflowEnvironment.newInstance();
|
|
554
|
+
worker = testEnv.newWorker(TASK_QUEUE);
|
|
555
|
+
worker.registerWorkflowImplementationTypes(PartyOnboardingWorkflowImpl.class);
|
|
556
|
+
// withoutAnnotations() is REQUIRED: Mockito copies @ActivityInterface/@ActivityMethod
|
|
557
|
+
// onto the mock class and the worker rejects the registration without it.
|
|
558
|
+
activities = mock(PartyOnboardingActivities.class, withSettings().withoutAnnotations());
|
|
559
|
+
worker.registerActivitiesImplementations(activities);
|
|
560
|
+
}
|
|
561
|
+
|
|
562
|
+
@AfterEach
|
|
563
|
+
void tearDown() {
|
|
564
|
+
testEnv.close();
|
|
565
|
+
}
|
|
566
|
+
|
|
567
|
+
@Test
|
|
568
|
+
void shouldCompensateCreatedPartyWhenWalletFails() {
|
|
569
|
+
when(activities.createParty(any())).thenReturn("party-1");
|
|
570
|
+
when(activities.openWallet(any(), any())).thenThrow(new IllegalStateException("wallet down"));
|
|
571
|
+
testEnv.start();
|
|
572
|
+
|
|
573
|
+
PartyOnboardingWorkflow workflow = testEnv.getWorkflowClient().newWorkflowStub(
|
|
574
|
+
PartyOnboardingWorkflow.class,
|
|
575
|
+
WorkflowOptions.newBuilder().setTaskQueue(TASK_QUEUE).build());
|
|
576
|
+
|
|
577
|
+
assertThatThrownBy(() -> workflow.onboard(request()))
|
|
578
|
+
.isInstanceOf(WorkflowException.class);
|
|
579
|
+
|
|
580
|
+
verify(activities).deleteParty("party-1"); // the compensation ran
|
|
581
|
+
verify(activities, never()).closeWallet(any()); // the wallet was never created
|
|
582
|
+
}
|
|
583
|
+
}
|
|
584
|
+
```
|
|
585
|
+
|
|
586
|
+
Mandatory scenarios per workflow — the compensation paths are the point of the whole exercise:
|
|
587
|
+
|
|
588
|
+
| # | Scenario | Assert |
|
|
589
|
+
|---|---|---|
|
|
590
|
+
| 1 | Happy path | result DTO correct; activities invoked in order; **no** compensation invoked |
|
|
591
|
+
| 2 | Failure at step *n* | compensations for steps `1..n-1` ran, in reverse order; none for `n..end`; the workflow still fails |
|
|
592
|
+
| 3 | Non-retryable business failure | the activity ran **once** (`verify(activities, times(1))`) — proves `setDoNotRetry` is wired |
|
|
593
|
+
| 4 | Compensation is idempotent | a compensation invoked twice leaves the same end state |
|
|
594
|
+
| 5 | Signal / timer paths, where present | time-skipping drives the timer; assert the state the signal produced |
|
|
595
|
+
|
|
596
|
+
- Mockito stubs the **activity interface**, so a workflow test never touches a database or a slice `Handler`. Slice behaviour is already covered by that slice's own `HandlerTest`.
|
|
597
|
+
- Give mocked failing activities a bounded retry policy (or a non-retryable failure), or the test exercises the full retry schedule before failing.
|
|
598
|
+
- Add a `WorkflowReplayer` test against a committed history JSON for any workflow already running in production — it fails the build on a change that would break in-flight executions (§4), which is the one bug integration tests cannot find.
|
|
599
|
+
|
|
600
|
+
## Checklist for a new workflow
|
|
601
|
+
|
|
602
|
+
1. The operation genuinely crosses a slice, service or time boundary — a single `Handler` + one transaction cannot do it.
|
|
603
|
+
2. Topology chosen and written in the app README: **worker-only** (default — deployable named `{module}-{capability}-worker`, no REST, no slices, no datasource, activities call remote services) or **embedded** (stays `-ms`). In a worker-only project, none of the REST/slice/SQL standards were scaffolded and the ArchUnit suite ships the documented subset.
|
|
604
|
+
3. `temporal-sdk` (1.27.0+) compile scope, `temporal-testing` test scope, explicit version, no Quarkiverse extension.
|
|
605
|
+
4. Exactly one `TemporalBootstrap` in `common/temporal`: `@Produces WorkflowClient`, `@Observes StartupEvent` → register (workflows **by class**, activities **by instance**) → `factory.start()`; `@Observes ShutdownEvent` → `shutdown()` + `awaitTermination`.
|
|
606
|
+
5. Endpoint, namespace and certificates come from `${VAR}` / `.env` / vault — never literals (security skill §1).
|
|
607
|
+
6. Workflow lives in `orchestration/<capability>/`; in an embedded app, no `io.temporal` import inside any slice; ArchUnit quarantine + `*Impl` carve-out in place.
|
|
608
|
+
7. `*WorkflowImpl` is deterministic: no clock, no random, no threads, no I/O, no `@Inject`, no unordered-collection iteration.
|
|
609
|
+
8. Saga compensations registered **after** each forward step, LIFO, idempotent; irreversible steps ordered last; `compensate()` then rethrow.
|
|
610
|
+
9. Every activity: idempotent, `scheduleToCloseTimeout` set, `BusinessException` non-retryable, delegating to a `Handler` with no logic of its own.
|
|
611
|
+
10. Workflow started with a deterministic workflow id, after commit — never mid-transaction.
|
|
612
|
+
11. Transport is gRPC in and out: the trigger reaches the `*Starter` through a `Handler` (embedded) or directly from the worker's own `*GrpcService` (worker-only), it **returns the workflow id without waiting**, and every outbound call goes through a `common/client` bean.
|
|
613
|
+
12. Timeouts ordered `grpc deadline < startToCloseTimeout < scheduleToCloseTimeout`, and retries are not stacked — Temporal owns the retry, the mesh keeps at most one.
|
|
614
|
+
13. Payload records annotated `@RegisterForReflection`; `--initialize-at-run-time=io.grpc.netty.shaded.io.netty` appended (not overwriting) `quarkus.native.additional-build-args`; `@QuarkusIntegrationTest` green against the native binary.
|
|
615
|
+
14. `TestWorkflowEnvironment` tests cover happy path **and** every compensation path; tracing interceptors and Micrometer metrics scope registered.
|
|
616
|
+
15. `orchestration/<capability>/README.md` documents the steps, their compensations, the timeouts and the task queue.
|