bmad-method-quarkus 1.0.5 → 1.0.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/src/bmm-skills/agents/bmad-quarkus-build/SKILL.md +7 -3
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-error-handling-i18n/SKILL.md +6 -2
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-grpc-services/SKILL.md +1 -0
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-hexagonal-core/SKILL.md +70 -18
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-kafka-messaging/SKILL.md +5 -3
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-observability-otel/SKILL.md +191 -13
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-openapi-tmforum/SKILL.md +6 -7
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-security-standards/SKILL.md +132 -0
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-sql-jdbc-agroal/SKILL.md +5 -2
- package/src/bmm-skills/agents/bmad-quarkus-build/skills/quarkus-temporal-workflows/SKILL.md +592 -0
- package/src/commands/quarkus-all.md +6 -5
|
@@ -0,0 +1,592 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: quarkus-temporal-workflows
|
|
3
|
+
description: Standard for durable orchestration and distributed sagas in Quarkus native services using Temporal via the `io.quarkiverse.temporal:quarkus-temporal` extension — the plain `io.temporal:temporal-sdk` with manual wiring cannot produce a working native image. Covers `quarkus.temporal.*` config and auto-discovery (no bootstrap bean), the `io.temporal.workflow.Saga` compensation pattern, `@RegisterForReflection` record payloads, gRPC at both ends, non-retryable `ApplicationFailure`, and `TestWorkflowEnvironment` testing. Defines two topologies — a worker-only deployable (default, `-worker` suffix, no REST, no slices, no datasource) and an embedded worker in a slice app. Use when the user mentions Temporal, quarkus-temporal, workflow orchestration, durable execution, saga, compensation, activities, WorkflowClient, WorkerFactory, task queue, workflow id, `@WorkflowInterface`/`@ActivityInterface`, TestWorkflowEnvironment, a worker-only or orchestrator project, or a long-running process spanning services.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Temporal Durable Orchestration Standard (Quarkus)
|
|
7
|
+
|
|
8
|
+
Applies to any Quarkus backend project **that uses Temporal**; project directives (CLAUDE.md, ADRs, explicit instructions) override these defaults where they conflict.
|
|
9
|
+
|
|
10
|
+
Temporal is for the work a single `@Transactional Handler.process()` cannot make atomic: a business process that spans several slices, several services, or a span of time (minutes to months), where a failure halfway through must **undo** what already committed elsewhere. A JTA transaction already gives you all-or-nothing inside one database — reach for Temporal only when the operation crosses that boundary. Wrapping a one-slice insert in a workflow buys nothing and costs a worker, a task queue and a replay contract.
|
|
11
|
+
|
|
12
|
+
**Two topologies, and §3 is the first decision you make.** The default is a **worker-only project**: an independent deployable that hosts workflows and activities and has no REST layer, no slices and no database — the shape to reach for when the process spans services. The alternative is an **embedded worker** inside a normal slice-architecture app. §3 gives the layout for each and, for worker-only projects, states which house standards go quiet: no `Resource`, no `<Slice>Handler`/`<Slice>Sql`, no OpenAPI, usually no datasource.
|
|
13
|
+
|
|
14
|
+
**Scope inside a slice architecture (embedded only):** Temporal is quarantined to `orchestration/` and `common/temporal`, the same way Mutiny is quarantined to `*GrpcService` and `common/client` (see quarkus-hexagonal-core skill). No `io.temporal.*` type ever appears in a slice folder, a `Handler`, a `Sql` class or a slice DTO. A slice stays fully testable and deployable with Temporal absent from the classpath.
|
|
15
|
+
|
|
16
|
+
## 1. Dependencies — the Quarkiverse extension, not the plain SDK
|
|
17
|
+
|
|
18
|
+
```xml
|
|
19
|
+
<properties>
|
|
20
|
+
<quarkus.platform.version>3.33.3</quarkus.platform.version>
|
|
21
|
+
<quarkus-temporal.version>0.6.0</quarkus-temporal.version>
|
|
22
|
+
<!-- Must track the version quarkus-temporal itself pins (0.6.0 -> 1.38.0):
|
|
23
|
+
temporal-testing is the plain SDK's own artifact and has to match the
|
|
24
|
+
SDK the extension pulls in transitively. -->
|
|
25
|
+
<temporal.version>1.38.0</temporal.version>
|
|
26
|
+
</properties>
|
|
27
|
+
|
|
28
|
+
<dependencies>
|
|
29
|
+
<dependency>
|
|
30
|
+
<groupId>io.quarkiverse.temporal</groupId>
|
|
31
|
+
<artifactId>quarkus-temporal</artifactId>
|
|
32
|
+
<version>${quarkus-temporal.version}</version>
|
|
33
|
+
</dependency>
|
|
34
|
+
<dependency>
|
|
35
|
+
<groupId>io.temporal</groupId>
|
|
36
|
+
<artifactId>temporal-testing</artifactId>
|
|
37
|
+
<version>${temporal.version}</version>
|
|
38
|
+
<scope>test</scope>
|
|
39
|
+
</dependency>
|
|
40
|
+
</dependencies>
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
**Use `io.quarkiverse.temporal:quarkus-temporal`. Do not wire `io.temporal:temporal-sdk` by hand — it cannot produce a working native image.** This reverses earlier guidance, and the reason is specific enough to be worth recording so nobody re-litigates it: the manual channel-building path hits a hard GraalVM `BUILD_TIME`/`RUN_TIME` policy conflict on `io.grpc.netty.shaded.io.netty.buffer.PooledByteBufAllocator` that **no `--initialize-at-run-time` / `--initialize-at-build-time` combination can override**. The extension ships GraalVM substitution classes for the shaded gRPC/Netty classes, which is what makes the binary build at all. Since the native binary is the delivery artifact (see quarkus-hexagonal-core skill), a JVM-only approach is not an option to fall back on.
|
|
44
|
+
|
|
45
|
+
- The extension pulls `temporal-sdk` transitively — do not declare it yourself, or the two versions will drift.
|
|
46
|
+
- `temporal-testing` stays `test` scope and must match the SDK version the extension pins. Bump both together.
|
|
47
|
+
- `quarkus-smallrye-health` is still required even though the worker serves no REST (see quarkus-hexagonal-core skill, README §11) — liveness/readiness are not optional because the app has no HTTP surface of its own.
|
|
48
|
+
- Everything you write is still plain Temporal SDK code following the official Java samples — `@WorkflowInterface`/`@ActivityInterface` pairs, activity stubs built inside the workflow, `Saga` for compensation. The extension owns the *lifecycle*, not your workflow code. Do not build a house abstraction over `Workflow.newActivityStub`; it hides the options that make a workflow correct.
|
|
49
|
+
|
|
50
|
+
## 2. Configuration and lifecycle — owned by the extension
|
|
51
|
+
|
|
52
|
+
There is **no composition root to write**. The extension owns the `WorkflowServiceStubs` / `WorkflowClient` / `WorkerFactory` lifecycle and **auto-discovers** your workflow and activity implementations from their `@WorkflowInterface` / `@ActivityInterface` contracts. Everything is configuration:
|
|
53
|
+
|
|
54
|
+
```properties
|
|
55
|
+
quarkus.temporal.connection.target=${TEMPORAL_SERVICE_TARGET:127.0.0.1:7233}
|
|
56
|
+
quarkus.temporal.worker.task-queue=alva.customer.party-onboarding.v1
|
|
57
|
+
# Blocks shutdown until in-flight activity tasks finish, or this timeout hits
|
|
58
|
+
quarkus.temporal.termination-timeout=20s
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
- **Do not hand-write a `TemporalBootstrap`/`TemporalLifecycle` bean, a `@Produces WorkflowClient`, or a `WorkerFactory.start()` in a `StartupEvent` observer.** The extension already did it; a second registration on the same task queue is a duplicate poller. If you are migrating from the old manual wiring, delete the bean and the `@ConfigMapping` that fed it.
|
|
62
|
+
- `quarkus.temporal.termination-timeout` replaces the manual `factory.shutdown()` + `awaitTermination` drain. Budget it inside the pod's grace period alongside the telemetry flush (see quarkus-observability-otel skill, "The shutdown budget").
|
|
63
|
+
- The task queue is configured **and** referenced from code, so pin it once and keep the two in sync — a constant on the workflow interface, mirrored by the property:
|
|
64
|
+
|
|
65
|
+
```java
|
|
66
|
+
@WorkflowInterface
|
|
67
|
+
public interface PartyOnboardingWorkflow {
|
|
68
|
+
|
|
69
|
+
/** Must match quarkus.temporal.worker.task-queue in application.properties. */
|
|
70
|
+
String TASK_QUEUE = "alva.customer.party-onboarding.v1";
|
|
71
|
+
|
|
72
|
+
@WorkflowMethod
|
|
73
|
+
PartyOnboardingResultDto onboard(PartyOnboardingRequestDto request);
|
|
74
|
+
}
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
- Inject `WorkflowClient` wherever you start a workflow (§6); the extension produces it.
|
|
78
|
+
- The endpoint, namespace and any mTLS certificate path are environment-specific — `.env` locally, platform vault in deployed environments, `${VAR}` in properties (see quarkus-security-standards skill §1–2). A Temporal Cloud client certificate is a secret and is never committed.
|
|
79
|
+
|
|
80
|
+
**What has not changed:** a workflow implementation is still instantiated fresh per execution and replayed, so **CDI cannot inject into it** — its only collaborators are the activity stubs it builds itself (§4). Activity implementations, by contrast, are ordinary `@ApplicationScoped` beans and are injected normally.
|
|
81
|
+
|
|
82
|
+
## 3. Topology: a worker-only project (default), or an embedded worker
|
|
83
|
+
|
|
84
|
+
Decide this **before** the first file, and record it in the app's `README.md` — the two shapes have different layouts and suspend different house rules.
|
|
85
|
+
|
|
86
|
+
### A. Worker-only project — the default
|
|
87
|
+
|
|
88
|
+
A deployable that hosts workflows and activities **and nothing else**: no REST, no slices, and normally no database. Every side effect is a remote call to the service that owns that data (§7). This is the right shape whenever the process spans services, for two reasons that matter in production:
|
|
89
|
+
|
|
90
|
+
- Workflow code is **replay-sensitive** (§4). Keeping it in its own deployable means a workflow change is released on its own cadence instead of riding along with an unrelated business-service deploy — and a business service can be redeployed freely without touching in-flight executions.
|
|
91
|
+
- Workers scale on **task-queue backlog**, not on request rate. A pod sized for orchestration has nothing in common with one sized for an API.
|
|
92
|
+
|
|
93
|
+
```
|
|
94
|
+
apps/customer-onboarding-worker/ # worker-only deployable — note the -worker suffix
|
|
95
|
+
├── src/main/java/com/alva/customer/onboarding/
|
|
96
|
+
│ ├── common/
|
|
97
|
+
│ │ │ # no TemporalConfig, no bootstrap bean — §2
|
|
98
|
+
│ │ ├── client/PartyRegistrar.java # gRPC beans — how activities reach other services
|
|
99
|
+
│ │ ├── client/WalletOpener.java
|
|
100
|
+
│ │ └── exception/BusinessException.java # the house error taxonomy still applies (§9)
|
|
101
|
+
│ └── orchestration/
|
|
102
|
+
│ └── party_onboarding/ # snake_case, capability-named
|
|
103
|
+
│ ├── dto/
|
|
104
|
+
│ │ ├── PartyOnboardingRequestDto.java # record + @RegisterForReflection
|
|
105
|
+
│ │ └── PartyOnboardingResultDto.java
|
|
106
|
+
│ ├── PartyOnboardingWorkflow.java # @WorkflowInterface — the contract
|
|
107
|
+
│ ├── PartyOnboardingWorkflowImpl.java # the saga; deterministic, zero I/O
|
|
108
|
+
│ ├── PartyOnboardingActivities.java # @ActivityInterface
|
|
109
|
+
│ ├── PartyOnboardingActivitiesImpl.java # injects common/client beans — no local Handler
|
|
110
|
+
│ └── README.md # steps, compensations, timeouts, task queue
|
|
111
|
+
├── src/main/proto/ # ONLY if this app owns the trigger RPC (below)
|
|
112
|
+
├── Dockerfile · service.yaml · README.md
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
There is no `Resource`, no `@Path`, no slice folder, no `Sql` class and no `db/` migration. **Do not scaffold them "for later"** — an unused datasource is a connection pool, a credential and a readiness dependency bought for nothing.
|
|
116
|
+
|
|
117
|
+
**The deployable carries the `-worker` suffix, never `-ms`** — `{module}-{capability}-worker`, e.g. `customer-onboarding-worker` (see the suffix table in the quarkus-hexagonal-core skill). The suffix is a deploy-time contract, and this kind of app contradicts everything `-ms` implies: no ingress, no HTTP route to probe, no API to publish, and scaling driven by task-queue backlog instead of request rate. Naming it `-ms` would hand the platform a service it would try to route to and scale on RPS it never receives. Its `service.yaml` matches: no provided APIs, task queues listed instead.
|
|
118
|
+
|
|
119
|
+
**Which monorepo owns it:** the domain that owns the business **process**, even when most steps call other domains. Cross-domain calls go through contracts published in `contracts/`, exactly like any other inter-domain call.
|
|
120
|
+
|
|
121
|
+
#### What does *not* apply in a worker-only project
|
|
122
|
+
|
|
123
|
+
Marcus reads eight universal standards; several of them have nothing to govern here. Applying them anyway produces a REST layer nobody calls and a database nobody reads.
|
|
124
|
+
|
|
125
|
+
| Standard | Status in a worker-only project |
|
|
126
|
+
|---|---|
|
|
127
|
+
| `quarkus-openapi-tmforum` | **Does not apply.** No `Resource`, no OpenAPI document, no Swagger. The only HTTP surface is `/q/health` and `/q/metrics`. |
|
|
128
|
+
| Vertical slices (`quarkus-hexagonal-core`) | **Does not apply.** No slice folders, no `<Slice>Handler`, no `<Slice>Sql`. `orchestration/` + `common/` *is* the structure. |
|
|
129
|
+
| `quarkus-sql-jdbc-agroal` | **Usually does not apply** — Temporal holds the process state, so there is no datasource. If the app genuinely owns a table (an idempotency/dedup ledger), its SQL follows that skill, with one adjustment: there is no `Handler` below the activity, so **the activity itself plays the Handler role** — it carries the `@Transactional`, opens the single `Connection` and passes it to every `Sql` call. That is the only case where an activity is transactional. |
|
|
130
|
+
| `quarkus-grpc-services` | **Applies fully** — it is how activities reach every other service (§7), and how the trigger RPC is exposed when this app owns one. |
|
|
131
|
+
| `quarkus-error-handling-i18n` | Error codes and the `BusinessException` taxonomy apply (§9). The REST edge mapper is absent; the gRPC interceptor applies only if the app exposes an RPC. |
|
|
132
|
+
| `quarkus-observability-otel` | **Applies fully, and matters more** — there is no request log to fall back on when a workflow misbehaves. |
|
|
133
|
+
| `quarkus-security-standards` | **Applies fully** — Temporal endpoint, namespace and mTLS certificates are secrets (§2). |
|
|
134
|
+
| `quarkus-kafka-messaging` | Applies when a workflow is triggered by, or emits, domain events. |
|
|
135
|
+
| Rest of `quarkus-hexagonal-core` | Naming, `apps/` layout, Java 25, native build + Dockerfile, `service.yaml` + README: **all unchanged**. |
|
|
136
|
+
|
|
137
|
+
ArchUnit ships a **subset**, and the two lists below partition hexagonal-core's canonical suite — every rule it defines is either kept or explicitly dropped, so nothing is left ambiguous.
|
|
138
|
+
|
|
139
|
+
**Keep:** `bannedSuffixes` (with the `..orchestration..` `*Impl` carve-out below), `serviceSuffixIsReserved` and `grpcServices` (they bite the moment the worker owns its trigger RPC), `reactiveIsQuarantined`, `noOrm`, `businessExceptionIsTheOnlyOne`, `dtosStayInDtoPackage`, `dtosHaveNoLogic`, plus this skill's own `workflowsDoNoIo`.
|
|
140
|
+
|
|
141
|
+
**Drop:** `slicesAreIndependent`, `onlyHandlersTouchSql`, `transactionalOnlyOnProcess`, `handlersExposeOnlyProcess`, `handlerKnowsNoTransport`, `crossCuttingWritesAreCentralised`, `restResources` and `transportHasNoJdbc` — with no slices, no `Handler` and no `Sql` to match, they can never fail, and a rule that cannot fail is noise that makes the suite look stronger than it is.
|
|
142
|
+
|
|
143
|
+
`temporalIsQuarantined` is **dropped too, and only here**: in a worker-only app the whole codebase is Temporal, so the rule would forbid what the app exists to do. It is mandatory in the embedded topology (B).
|
|
144
|
+
|
|
145
|
+
#### Who starts the workflow
|
|
146
|
+
|
|
147
|
+
- **Default: the calling service does.** The worker exposes **no inbound API at all** — it polls its task queue and nothing else. The service that owns the triggering business event holds its own `*Starter` bean (§6); the workflow id and the payload record are the contract between them. Nothing to route, nothing to secure at the edge.
|
|
148
|
+
- **Alternative: the worker owns the trigger RPC.** Then it has `src/main/proto/` and one `*GrpcService`, and that adapter calls the `*Starter` **directly — with no `Handler` in between.** This is a deliberate deviation from "a transport adapter always delegates to a `Handler`": a `Handler` exists to hold business logic and a transaction, and a trigger RPC has neither. An empty pass-through `Handler` would be pure ceremony. The adapter still does only proto↔DTO translation plus the start call.
|
|
149
|
+
|
|
150
|
+
Readiness reflects the **worker**, not a route: can it reach the Temporal frontend, and is the factory polling. Scale on task-queue backlog (§10), never on RPS.
|
|
151
|
+
|
|
152
|
+
### B. Embedded worker — the exception
|
|
153
|
+
|
|
154
|
+
A normal `-ms` app that also hosts a worker — and it **stays `-ms`**, because it still serves an API; the `-worker` suffix is reserved for deployables that serve none. Use it only when every step of the saga is a capability **this same app already owns**, and Temporal is there for durability, retries or long waits rather than for crossing a service boundary. The cost is that the app now has two runtime roles and a replay-sensitive workflow tied to the service's release cadence.
|
|
155
|
+
|
|
156
|
+
```
|
|
157
|
+
src/main/java/com/<company>/<module>/
|
|
158
|
+
├── common/temporal/PartyOnboardingStarter.java # §6 — the only Temporal-aware class outside orchestration/
|
|
159
|
+
├── orchestration/
|
|
160
|
+
│ └── party_onboarding/… # same contents; activities call local Handlers
|
|
161
|
+
└── create_party_individual/ # slices — untouched, no Temporal import
|
|
162
|
+
```
|
|
163
|
+
|
|
164
|
+
Here the dependency direction is deliberately **one-way**: `orchestration/` may depend on slice `Handler`s and slice DTOs; a slice may never depend on `orchestration/`. A slice that imports a workflow has been made un-deployable without a Temporal cluster. Extend `slicesAreIndependent` with the single asymmetric exemption shown below, and keep every other slice rule intact — in this topology they all still bite.
|
|
165
|
+
|
|
166
|
+
### Naming
|
|
167
|
+
|
|
168
|
+
| Artifact | Convention | Example |
|
|
169
|
+
|---|---|---|
|
|
170
|
+
| Workflow interface | `<Capability>Workflow`, `@WorkflowInterface` | `PartyOnboardingWorkflow` |
|
|
171
|
+
| Workflow implementation | `<Capability>WorkflowImpl` | `PartyOnboardingWorkflowImpl` |
|
|
172
|
+
| Activity interface | `<Capability>Activities`, `@ActivityInterface` | `PartyOnboardingActivities` |
|
|
173
|
+
| Activity implementation | `<Capability>ActivitiesImpl` | `PartyOnboardingActivitiesImpl` |
|
|
174
|
+
| Workflow method | imperative verb, `@WorkflowMethod` | `onboard` |
|
|
175
|
+
| Signal / Query | `<verb>` / `get<Noun>` | `approve` / `getStatus` |
|
|
176
|
+
| Starter bean | capability noun in `common/temporal` | `PartyOnboardingStarter` |
|
|
177
|
+
| Payload DTO | `*Dto` **record** in `dto/` | `PartyOnboardingRequestDto` |
|
|
178
|
+
| Task queue | `<org>.<module>.<capability-kebab>.v<major>` | `alva.customer.party-onboarding.v1` |
|
|
179
|
+
| Workflow id | `<capability-kebab>-<business-id>` — deterministic | `party-onboarding-9f3e…` |
|
|
180
|
+
|
|
181
|
+
**`*Impl` is banned everywhere else in this codebase and sanctioned here.** The SDK builds its client stubs from the annotated interface, so the interface/implementation pair is structural, not stylistic, and the official samples name it `*Impl`. Carve it out of the ArchUnit `bannedSuffixes` rule **by package**, exactly as `common/util/StringUtils` and the `Jdbc` helper are carved out — never by weakening the rule:
|
|
182
|
+
|
|
183
|
+
```java
|
|
184
|
+
@ArchTest
|
|
185
|
+
static final ArchRule bannedSuffixesOutsideOrchestration = noClasses()
|
|
186
|
+
.that().resideOutsideOfPackage("..orchestration..")
|
|
187
|
+
.should().haveSimpleNameEndingWith("Impl"); // …plus the existing suffixes
|
|
188
|
+
|
|
189
|
+
// Temporal stays quarantined
|
|
190
|
+
@ArchTest
|
|
191
|
+
static final ArchRule temporalIsQuarantined = noClasses()
|
|
192
|
+
.that().resideOutsideOfPackage("..orchestration..")
|
|
193
|
+
.and().resideOutsideOfPackage("..common.temporal..")
|
|
194
|
+
.should().dependOnClassesThat().resideInAPackage("io.temporal..");
|
|
195
|
+
|
|
196
|
+
// A workflow implementation performs no I/O of its own (§4)
|
|
197
|
+
@ArchTest
|
|
198
|
+
static final ArchRule workflowsDoNoIo = noClasses()
|
|
199
|
+
.that().haveSimpleNameEndingWith("WorkflowImpl")
|
|
200
|
+
.should().dependOnClassesThat().resideInAnyPackage(
|
|
201
|
+
"java.sql..", "javax.sql..", "jakarta.ws.rs..", "io.quarkus.grpc..");
|
|
202
|
+
```
|
|
203
|
+
|
|
204
|
+
Extend the existing `slicesAreIndependent` rule with **one** asymmetric exemption — orchestration may reach into slices, never the reverse:
|
|
205
|
+
|
|
206
|
+
```java
|
|
207
|
+
.ignoreDependency(resideInAPackage("..orchestration.."), alwaysTrue())
|
|
208
|
+
```
|
|
209
|
+
|
|
210
|
+
## 4. The workflow implementation: deterministic, and that is a hard constraint
|
|
211
|
+
|
|
212
|
+
Temporal replays workflow code from its event history after every worker restart. Replay must produce the identical sequence of commands, so a workflow implementation is **pure orchestration**: it calls activity stubs and does nothing else.
|
|
213
|
+
|
|
214
|
+
Banned inside a `*WorkflowImpl`, with the replacement the SDK provides:
|
|
215
|
+
|
|
216
|
+
| Never | Always |
|
|
217
|
+
|---|---|
|
|
218
|
+
| `System.currentTimeMillis()`, `LocalDateTime.now()` | `Workflow.currentTimeMillis()` |
|
|
219
|
+
| `UUID.randomUUID()`, `Math.random()` | `Workflow.randomUUID()`, `Workflow.newRandom()` |
|
|
220
|
+
| `Thread.sleep(…)` | `Workflow.sleep(Duration…)` |
|
|
221
|
+
| `new Thread(…)`, executor services | `Async.function(…)` / `Async.procedure(…)`, scoped with `Workflow.newCancellationScope` |
|
|
222
|
+
| JDBC, HTTP, gRPC, Kafka, file I/O | an **activity** |
|
|
223
|
+
| `Logger.getLogger(…)` | `Workflow.getLogger(…)` — suppresses duplicate lines on replay |
|
|
224
|
+
| iteration over `HashMap`/`HashSet` | `LinkedHashMap`/`TreeMap`, or sort explicitly |
|
|
225
|
+
|
|
226
|
+
`@Inject` does not work here either (§2) — the workflow's only collaborators are its activity stubs. A workflow that needs configuration receives it as a **workflow argument**, never read from `ConfigProvider` at replay time.
|
|
227
|
+
|
|
228
|
+
**Versioning is the production footgun.** Editing a workflow implementation while executions of the old code are still running causes a non-deterministic replay and a stuck workflow. Any change to the *sequence* of activity calls requires `Workflow.getVersion("addWallet", DEFAULT_VERSION, 1)` branching, or a new task queue (`…v2`) drained in parallel. Changing an activity's *body* is always safe; changing the workflow's call order never is.
|
|
229
|
+
|
|
230
|
+
## 5. The saga: `io.temporal.workflow.Saga` with explicit compensations
|
|
231
|
+
|
|
232
|
+
```java
|
|
233
|
+
public class PartyOnboardingWorkflowImpl implements PartyOnboardingWorkflow {
|
|
234
|
+
|
|
235
|
+
private static final Logger log = Workflow.getLogger(PartyOnboardingWorkflowImpl.class);
|
|
236
|
+
|
|
237
|
+
private final PartyOnboardingActivities activities = Workflow.newActivityStub(
|
|
238
|
+
PartyOnboardingActivities.class,
|
|
239
|
+
ActivityOptions.newBuilder()
|
|
240
|
+
.setStartToCloseTimeout(Duration.ofSeconds(30))
|
|
241
|
+
.setScheduleToCloseTimeout(Duration.ofMinutes(10)) // the outer bound — never omit
|
|
242
|
+
.setRetryOptions(RetryOptions.newBuilder()
|
|
243
|
+
.setInitialInterval(Duration.ofSeconds(1))
|
|
244
|
+
.setMaximumInterval(Duration.ofSeconds(30))
|
|
245
|
+
.setBackoffCoefficient(2.0)
|
|
246
|
+
.build()) // no setDoNotRetry — see §9
|
|
247
|
+
.build());
|
|
248
|
+
|
|
249
|
+
@Override
|
|
250
|
+
public PartyOnboardingResultDto onboard(PartyOnboardingRequestDto request) {
|
|
251
|
+
Saga saga = new Saga(new Saga.Options.Builder()
|
|
252
|
+
.setParallelCompensation(false) // compensate in reverse order, one at a time
|
|
253
|
+
.setContinueWithError(true) // attempt every compensation — see below
|
|
254
|
+
.build());
|
|
255
|
+
try {
|
|
256
|
+
String partyId = activities.createParty(request);
|
|
257
|
+
saga.addCompensation(activities::deleteParty, partyId);
|
|
258
|
+
|
|
259
|
+
String walletId = activities.openWallet(partyId, request.currency());
|
|
260
|
+
saga.addCompensation(activities::closeWallet, walletId);
|
|
261
|
+
|
|
262
|
+
activities.sendWelcomeNotification(partyId);
|
|
263
|
+
return new PartyOnboardingResultDto(partyId, walletId);
|
|
264
|
+
|
|
265
|
+
} catch (RuntimeException e) { // deliberately broad — see below
|
|
266
|
+
log.warn("Onboarding failed, compensating: {}", e.getMessage());
|
|
267
|
+
saga.compensate();
|
|
268
|
+
throw e; // the workflow still fails — do not swallow
|
|
269
|
+
}
|
|
270
|
+
}
|
|
271
|
+
}
|
|
272
|
+
```
|
|
273
|
+
|
|
274
|
+
Rules:
|
|
275
|
+
- **Catch `RuntimeException`, not `ActivityFailure`.** `ActivityFailure` misses `ChildWorkflowFailure`, `CanceledFailure` and every other `TemporalFailure` subtype, and a miss means the compensation chain silently never runs — the exact failure the saga exists to prevent. `RuntimeException` is deliberately broader than even `TemporalFailure`: anything that escapes the forward path should compensate.
|
|
276
|
+
- **`setContinueWithError(true)` is the default here.** Out of the box it is `false`, which makes the *first* failing compensation abort the rest of the chain and leave earlier steps uncompensated — almost never what you want, since the compensations are usually independent deletes. Leave it `false` only when a failed compensation genuinely makes the remaining ones unsafe.
|
|
277
|
+
- **A step whose effect is an idempotent, resubmittable correction needs no compensation at all.** Register one only for steps that created something. Say so in a comment where you skip it, or the next reader will assume it was forgotten.
|
|
278
|
+
- **Register the compensation *after* the forward activity returns**, using its result. Registering before means compensating something that never happened.
|
|
279
|
+
- Compensations run in **reverse registration order** (LIFO). Keep `setParallelCompensation(false)` unless the steps are provably independent — parallel compensation of dependent resources reintroduces the ordering bug the saga exists to avoid.
|
|
280
|
+
- **`saga.compensate()` then rethrow.** A saga that compensates and returns normally reports success for a business process that did not happen.
|
|
281
|
+
- A compensation is itself an activity, so it is retried like any other — and it must be **idempotent**: `deleteParty` on an already-deleted party succeeds quietly rather than failing the compensation chain.
|
|
282
|
+
- No compensation is possible for some steps (an email that went out). Order the workflow so irreversible steps come **last**, after everything reversible has succeeded.
|
|
283
|
+
|
|
284
|
+
## 6. Activities: the only side-effecting layer
|
|
285
|
+
|
|
286
|
+
An activity implementation is thin: extract the arguments, perform **exactly one** side effect, return a plain value. **No business logic, no orchestration, no `@Transactional` of its own** — the transaction always lives one level down. What it delegates *to* is the one thing the topology (§3) changes:
|
|
287
|
+
|
|
288
|
+
```java
|
|
289
|
+
// A. Worker-only project (default): the side effect is a remote call
|
|
290
|
+
@ApplicationScoped
|
|
291
|
+
public class PartyOnboardingActivitiesImpl implements PartyOnboardingActivities {
|
|
292
|
+
|
|
293
|
+
@Inject PartyRegistrar partyRegistrar; // common/client — gRPC, awaits internally (§7)
|
|
294
|
+
@Inject WalletOpener walletOpener;
|
|
295
|
+
|
|
296
|
+
@Override
|
|
297
|
+
public String createParty(PartyOnboardingRequestDto request) {
|
|
298
|
+
return partyRegistrar.register(request.tenantId(), request.givenName(), request.familyName());
|
|
299
|
+
}
|
|
300
|
+
}
|
|
301
|
+
```
|
|
302
|
+
|
|
303
|
+
```java
|
|
304
|
+
// B. Embedded worker: the side effect is a Handler in this same app
|
|
305
|
+
@Inject CreatePartyIndividualHandler createPartyHandler;
|
|
306
|
+
|
|
307
|
+
@Override
|
|
308
|
+
public String createParty(PartyOnboardingRequestDto request) {
|
|
309
|
+
return createPartyHandler.process(toSliceRequest(request)).getPartyId();
|
|
310
|
+
}
|
|
311
|
+
```
|
|
312
|
+
|
|
313
|
+
Either way the activity is a transport adapter in the sense the hexagonal-core skill defines, and the `@Transactional` boundary is the remote service's `Handler` (A) or the local one (B) — never here. The single exception is a worker-only app that owns a table of its own (§3): with no `Handler` beneath it, that activity plays the `Handler` role and carries the transaction itself.
|
|
314
|
+
|
|
315
|
+
- **Activities are at-least-once.** Temporal retries on worker crash, timeout and failure, so every activity must be idempotent. Where the underlying operation is not naturally idempotent, dedupe on a business key or on `Activity.getExecutionContext().getInfo().getWorkflowId()`.
|
|
316
|
+
- Every activity needs a `setScheduleToCloseTimeout` (or a `setMaximumAttempts`). An activity with unlimited retries and no outer bound does not fail — it hangs forever, in production and in your tests.
|
|
317
|
+
- Long-running activities heartbeat (`Activity.getExecutionContext().heartbeat(progress)`) so a dead worker is detected in seconds rather than at the start-to-close timeout.
|
|
318
|
+
|
|
319
|
+
### Starting a workflow
|
|
320
|
+
|
|
321
|
+
The `*Starter` bean below is the same in both topologies — what differs is **which deployable owns it**:
|
|
322
|
+
|
|
323
|
+
- **Worker-only (A):** the `*Starter` lives in the *calling* service, which therefore carries `quarkus-temporal` of its own — configured as a **client without a worker**, since it starts workflows rather than running them. That falls out of auto-discovery on its own: the calling service contains no `@WorkflowInterface`/`@ActivityInterface` implementations, so the extension has nothing to register and starts no poller — just set `quarkus.temporal.connection.target` and leave the worker's task queue unset. Never register a worker on the orchestrator's task queue from a service that cannot execute the workflow. When the worker app owns its own trigger RPC (§3), its `*GrpcService` injects `WorkflowClient` and starts the workflow inline — no separate bean.
|
|
324
|
+
- **Embedded (B):** it lives in this app, and a slice `Handler` injects it.
|
|
325
|
+
|
|
326
|
+
In the embedded case a `Handler` must start the workflow through this bean and **never** by injecting `WorkflowClient` into the slice — that would drag `io.temporal` past the §3 quarantine:
|
|
327
|
+
|
|
328
|
+
```java
|
|
329
|
+
// common/temporal
|
|
330
|
+
@ApplicationScoped
|
|
331
|
+
public class PartyOnboardingStarter {
|
|
332
|
+
|
|
333
|
+
@Inject WorkflowClient client; // produced by the extension
|
|
334
|
+
|
|
335
|
+
public void start(String partyId, PartyOnboardingRequestDto request) {
|
|
336
|
+
PartyOnboardingWorkflow workflow = client.newWorkflowStub(
|
|
337
|
+
PartyOnboardingWorkflow.class,
|
|
338
|
+
WorkflowOptions.newBuilder()
|
|
339
|
+
.setTaskQueue(PartyOnboardingWorkflow.TASK_QUEUE) // the interface constant, §2
|
|
340
|
+
.setWorkflowId("party-onboarding-" + partyId) // deterministic = deduplicated
|
|
341
|
+
.setWorkflowIdReusePolicy(
|
|
342
|
+
WorkflowIdReusePolicy.WORKFLOW_ID_REUSE_POLICY_ALLOW_DUPLICATE_FAILED_ONLY)
|
|
343
|
+
.build());
|
|
344
|
+
WorkflowClient.start(workflow::onboard, request); // async — never block the Handler
|
|
345
|
+
}
|
|
346
|
+
}
|
|
347
|
+
```
|
|
348
|
+
|
|
349
|
+
**Starting a workflow inside a database transaction is a dual write.** If `process()` rolls back after the start call succeeded, the workflow orchestrates an entity that does not exist. This is the same hazard the outbox solves for Kafka (see quarkus-kafka-messaging skill), and it gets the same two answers:
|
|
350
|
+
|
|
351
|
+
1. **Deterministic workflow id derived from the business entity** (above) — a retry of the whole operation cannot start a second workflow, so the failure mode is recoverable rather than duplicated. This is mandatory, not optional. Use `ALLOW_DUPLICATE_FAILED_ONLY`, **not** `REJECT_DUPLICATE`: the latter blocks reuse even after the run closed, so a legitimate re-run of that business id becomes permanently impossible. The starter must also catch `WorkflowExecutionAlreadyStarted` and treat it as success — that exception *is* the deduplication working, not an error to propagate.
|
|
352
|
+
2. **Start after commit**, the way `OutboxDispatcher.dispatchAfterCommit(eventId)` already does it — or record an outbox row and let the relay start the workflow. Never call `WorkflowClient.start` in the middle of `execution()` and hope the transaction commits.
|
|
353
|
+
|
|
354
|
+
## 7. Transport: the invocations are gRPC
|
|
355
|
+
|
|
356
|
+
Orchestration is internal by nature — a saga exists because a process crosses services — so **gRPC is the predominant transport at both ends of a workflow**, per the house rule that internal synchronous calls are always gRPC and never internal REST (see quarkus-grpc-services and quarkus-hexagonal-core skills). North-bound REST triggers exist (a customer action that kicks off onboarding) but are the minority; the default assumption for a Temporal capability is gRPC in, gRPC out.
|
|
357
|
+
|
|
358
|
+
**A. Worker-only project** — the calling service starts the workflow; the worker only runs it and calls back out over gRPC:
|
|
359
|
+
|
|
360
|
+
```
|
|
361
|
+
calling service (-ms) worker project (-worker) other services
|
|
362
|
+
┌───────────────────────┐ ┌────────────────────────┐
|
|
363
|
+
│ GrpcService → Handler │ │ PartyOnboardingWorkflowImpl │ (no I/O here)
|
|
364
|
+
│ ↓ │ │ ↓ │
|
|
365
|
+
│ <Capability>Starter ─┼── Temporal ──┼→ PartyOnboardingActivitiesImpl │
|
|
366
|
+
└───────────────────────┘ task queue │ ↓ │
|
|
367
|
+
│ common/client bean ───┼─ gRPC ─→ party-ms, wallet-ms
|
|
368
|
+
└────────────────────────┘
|
|
369
|
+
```
|
|
370
|
+
|
|
371
|
+
**B. Embedded worker** — everything in one deployable, the activity reaching a local `Handler`:
|
|
372
|
+
|
|
373
|
+
```
|
|
374
|
+
<Proto>GrpcService → Handler.process() → <Capability>Starter → Temporal
|
|
375
|
+
(slice folder) (slice folder) (common/temporal)
|
|
376
|
+
↓
|
|
377
|
+
PartyOnboardingWorkflowImpl (no I/O — orchestration only)
|
|
378
|
+
↓
|
|
379
|
+
PartyOnboardingActivitiesImpl → slice Handler
|
|
380
|
+
```
|
|
381
|
+
|
|
382
|
+
Nothing about the inbound chain is special-cased for Temporal: the `*GrpcService` delegates to a `Handler`, and the `Handler` injects the `*Starter`. **A `*GrpcService` never injects `WorkflowClient` directly** — that would put `io.temporal` inside a slice folder and break the §3 quarantine. The one exception is the trigger RPC of a worker-only project, which has no `Handler` to delegate to and calls the `*Starter` itself (§3).
|
|
383
|
+
|
|
384
|
+
### Inbound: the gRPC call starts the workflow, it does not wait for it
|
|
385
|
+
|
|
386
|
+
A gRPC client carries a deadline measured in seconds (`quarkus.grpc.clients.<name>.deadline=2s`); a saga runs for minutes, days or months. So the RPC that triggers a workflow **returns as soon as the workflow is started** — the workflow id is the response — and never blocks on completion:
|
|
387
|
+
|
|
388
|
+
- Start asynchronously (`WorkflowClient.start(...)`, as in §6), return the workflow id to the caller.
|
|
389
|
+
- The caller learns the outcome by polling a `@QueryMethod`, by consuming the domain event the workflow's final activity emits (outbox → Kafka, per the messaging skill), or by its own callback RPC.
|
|
390
|
+
- Blocking the RPC on `workflow.onboard(request)` ties a durable execution to a 2-second deadline: the deadline fires, the caller retries, and the only thing the deterministic workflow id saved you from was a second workflow. The first one is still running, unobserved.
|
|
391
|
+
|
|
392
|
+
### Outbound: activities call other services through `common/client`
|
|
393
|
+
|
|
394
|
+
An activity that reaches another service uses a capability-named bean from `common/client` — the same bean a `Handler` would use, never a raw `@GrpcClient` stub and never an internal REST client:
|
|
395
|
+
|
|
396
|
+
```java
|
|
397
|
+
@ApplicationScoped
|
|
398
|
+
public class PartyOnboardingActivitiesImpl implements PartyOnboardingActivities {
|
|
399
|
+
|
|
400
|
+
@Inject CredentialsValidator credentialsValidator; // common/client — awaits internally
|
|
401
|
+
|
|
402
|
+
@Override
|
|
403
|
+
public boolean validateCredentials(String username, String password) {
|
|
404
|
+
return credentialsValidator.validate(username, password); // plain types in, plain types out
|
|
405
|
+
}
|
|
406
|
+
}
|
|
407
|
+
```
|
|
408
|
+
|
|
409
|
+
Because those beans `await()` internally and expose plain types, the `reactiveIsQuarantined` ArchUnit rule keeps passing unchanged — no Mutiny ever reaches `orchestration/`, exactly as none reaches a `Handler`.
|
|
410
|
+
|
|
411
|
+
**Only activities may call out. A workflow never holds a client**, gRPC or otherwise (§4).
|
|
412
|
+
|
|
413
|
+
### Two timeout layers, and the order between them is not optional
|
|
414
|
+
|
|
415
|
+
```
|
|
416
|
+
grpc deadline < activity startToCloseTimeout < activity scheduleToCloseTimeout
|
|
417
|
+
2s 30s 10m
|
|
418
|
+
```
|
|
419
|
+
|
|
420
|
+
Get this backwards and the activity times out while the gRPC call is still in flight: Temporal schedules a retry, and now two identical calls are hitting a downstream that is already struggling. The gRPC deadline must be the *inner* bound so the call is dead before Temporal reacts to it.
|
|
421
|
+
|
|
422
|
+
**Do not stack retries.** Temporal's activity retry policy and the gRPC/mesh retry policy multiply: 5 activity attempts × 3 mesh retries = 15 calls, sent to the exact service that is failing. Temporal owns the retry — its attempts are durable, backed off, and visible in the workflow history, where a mesh retry is invisible. Keep the client/mesh layer at a single retry for connection-level `UNAVAILABLE` blips at most, and let the activity policy do the rest.
|
|
423
|
+
|
|
424
|
+
### Mapping gRPC failures onto the retry policy
|
|
425
|
+
|
|
426
|
+
A `StatusRuntimeException` from a downstream is classified at the activity boundary. **§9 holds the single authoritative list of which codes are terminal — do not restate it here or anywhere else**, because two copies of a retry-classification table drift and the drift is silent: a code that is terminal in one list and retryable in the other produces either an infinite retry or a saga that compensates on a blip.
|
|
427
|
+
|
|
428
|
+
The downstream already maps its `BusinessException` codes onto gRPC statuses on the way out (see the grpc skill's error-mapping table), so §9's table is that one read in reverse — keep the two consistent when either changes.
|
|
429
|
+
|
|
430
|
+
### The SDK's own transport
|
|
431
|
+
|
|
432
|
+
The worker and the client talk to the Temporal service over gRPC as well, on shaded Netty. That is not trivia: it is why the native build depends on the GraalVM substitutions `quarkus-temporal` ships rather than on any initialization flag (§8), and why a Temporal Cloud connection is mTLS with a client certificate handled as a secret (§2). The connection's own RPC timeouts and keepalive are separate from every timeout above and are configured through `quarkus.temporal.connection.*`, not by hand-building stubs.
|
|
433
|
+
|
|
434
|
+
## 8. Payloads and native image
|
|
435
|
+
|
|
436
|
+
**Payloads are Java records, always annotated `@RegisterForReflection`.** Records are immutable, which is what replay safety wants, and the SDK's default `DataConverter` serializes them with Jackson — which needs reflection metadata that GraalVM will not infer.
|
|
437
|
+
|
|
438
|
+
```java
|
|
439
|
+
@RegisterForReflection
|
|
440
|
+
public record PartyOnboardingRequestDto(
|
|
441
|
+
String tenantId,
|
|
442
|
+
String givenName,
|
|
443
|
+
String familyName,
|
|
444
|
+
String currency) {}
|
|
445
|
+
```
|
|
446
|
+
|
|
447
|
+
This is a deliberate local deviation from the slice DTO convention (Lombok `@Data` classes, per hexagonal-core): a Temporal payload is serialized into an immutable event history, so a mutable DTO with setters is the wrong shape. The `*Dto` suffix and the `dto/` package still apply, so the existing `dtosStayInDtoPackage` ArchUnit rule keeps passing.
|
|
448
|
+
|
|
449
|
+
### The native build needs no Temporal-specific flags
|
|
450
|
+
|
|
451
|
+
```properties
|
|
452
|
+
# That is the whole native configuration — no additional-build-args for Temporal
|
|
453
|
+
quarkus.native.container-build=true
|
|
454
|
+
quarkus.native.builder-image=mandrel
|
|
455
|
+
```
|
|
456
|
+
|
|
457
|
+
**Do not add `--initialize-at-run-time=io.grpc.netty.shaded.io.netty` (or any variant of it).** This is the trap that cost a spike: it looks like the obvious fix, it is what the plain-SDK path appears to need, and **it does not work** — the conflict on `io.grpc.netty.shaded.io.netty.buffer.PooledByteBufAllocator` is a build-time/run-time *policy* clash that an initialization flag cannot resolve in either direction. What resolves it is the substitution classes `quarkus-temporal` ships (§1). Adding the flag on top of the extension buys nothing and misleads the next person debugging a native failure.
|
|
458
|
+
|
|
459
|
+
- If the app sets `quarkus.native.additional-build-args` for some *other* reason, remember the property is a comma-separated list — append, never overwrite.
|
|
460
|
+
- The SDK builds workflow and activity stubs with JDK dynamic proxies, which is precisely the pattern hexagonal-core tells you to avoid. The extension's build step registers what it needs; if a proxy is still missing, use Quarkus's `@RegisterForProxy` rather than a raw flag. (`-H:DynamicProxyConfigurationResources` is gone on GraalVM for JDK 21+ — proxies go through reachability metadata now.)
|
|
461
|
+
- **A Temporal app is not "native-ready" until the binary has actually been built and exercised.** The JVM-mode test proves nothing here: every problem in this section appears only at image-build time.
|
|
462
|
+
|
|
463
|
+
## 9. Errors and retries
|
|
464
|
+
|
|
465
|
+
Temporal's retry policy is the automatic remediation this codebase otherwise hand-writes — but only if failures are classified correctly. **The classification lives in the activity, expressed one way only: throw a non-retryable `ApplicationFailure` for what must not be retried, and let everything else propagate unchanged.**
|
|
466
|
+
|
|
467
|
+
```java
|
|
468
|
+
RuntimeException translate(StatusRuntimeException e, String activityName) {
|
|
469
|
+
Status.Code code = e.getStatus().getCode();
|
|
470
|
+
if (TERMINAL_CODES.contains(code)) {
|
|
471
|
+
String description = e.getStatus().getDescription();
|
|
472
|
+
Log.warnf("%s: terminal gRPC failure code=%s, non-retryable: %s", activityName, code, description);
|
|
473
|
+
return ApplicationFailure.newNonRetryableFailure(
|
|
474
|
+
description != null ? description : e.getMessage(), code.name());
|
|
475
|
+
}
|
|
476
|
+
// UNAVAILABLE, DEADLINE_EXCEEDED, INTERNAL, ... rethrown as-is, so the workflow's
|
|
477
|
+
// ActivityOptions retry policy applies with backoff up to scheduleToCloseTimeout.
|
|
478
|
+
Log.warnf("%s: retryable gRPC failure code=%s: %s", activityName, code, e.getMessage());
|
|
479
|
+
return e;
|
|
480
|
+
}
|
|
481
|
+
```
|
|
482
|
+
|
|
483
|
+
**Do not also configure `setDoNotRetry(...)` on the `ActivityOptions`.** It is a second, independent mechanism that matches on the failure *type* string, so the moment that string and the type you pass to `ApplicationFailure` drift apart the declarative rule silently matches nothing — and a rule that looks wired but is inert is worse than no rule. One mechanism, at the activity boundary.
|
|
484
|
+
|
|
485
|
+
Which gRPC codes are terminal:
|
|
486
|
+
|
|
487
|
+
| Code from the downstream | Treatment |
|
|
488
|
+
|---|---|
|
|
489
|
+
| `INVALID_ARGUMENT`, `ALREADY_EXISTS`, `UNAUTHENTICATED`, `PERMISSION_DENIED` | **non-retryable** `ApplicationFailure` — a malformed or duplicate request, or a token that will not become valid inside this run |
|
|
490
|
+
| everything else (`UNAVAILABLE`, `DEADLINE_EXCEEDED`, `INTERNAL`, `RESOURCE_EXHAUSTED`, …) | rethrown as-is → retried with backoff, bounded by `scheduleToCloseTimeout` |
|
|
491
|
+
|
|
492
|
+
**A compensation activity classifies differently, and getting this wrong deadlocks the rollback.** For a delete-style compensation, `NOT_FOUND` means the thing is already gone and `ALREADY_EXISTS` may encode the downstream's own conflict mapping — both mean *the compensation is as done as it can be*, so the activity must **return normally**, not fail. The same `ALREADY_EXISTS` that correctly fails a forward `create` must not fail its matching `delete`:
|
|
493
|
+
|
|
494
|
+
```java
|
|
495
|
+
} catch (StatusRuntimeException e) {
|
|
496
|
+
Status.Code code = e.getStatus().getCode();
|
|
497
|
+
if (code == Status.Code.NOT_FOUND || code == Status.Code.ALREADY_EXISTS) {
|
|
498
|
+
Log.infof("%s: compensation handled (%s)", activityName, code);
|
|
499
|
+
return; // idempotent: already compensated
|
|
500
|
+
}
|
|
501
|
+
throw translate(e, activityName);
|
|
502
|
+
}
|
|
503
|
+
```
|
|
504
|
+
|
|
505
|
+
The house exception taxonomy applies in both topologies — the worker-only layout in §3 ships `common/exception/BusinessException` too — and splits the same way: `BusinessException` → non-retryable; `TransientPersistenceException` / `PersistenceTimeoutException` / `StaleVersionException` → retryable, which is exactly what the retry policy exists for; `PersistenceException` → retryable but capped by `scheduleToCloseTimeout`.
|
|
506
|
+
|
|
507
|
+
The failure a client ultimately sees is still a house error code: a `WorkflowFailedException` surfacing at the REST edge is translated by `GlobalExceptionHandler` into the catalogued `<MOD>-<HTTP>-<seq>`, localized as usual. A raw Temporal stack trace never reaches an API consumer.
|
|
508
|
+
|
|
509
|
+
## 10. Observability
|
|
510
|
+
|
|
511
|
+
- Propagate trace context into workflows and activities with `OpenTracingClientInterceptor` / `OpenTracingWorkerInterceptor`, registered through the extension's interceptor configuration. They need two dependencies §1 does not list because they are tracing-only — `io.temporal:temporal-opentracing` and `io.opentelemetry:opentelemetry-opentracing-shim` to bridge onto the OTel SDK this platform uses; there is no OTel-native interceptor module. Without them a workflow is a hole in the trace: the REST span ends, and the activity spans belong to no one.
|
|
512
|
+
- Feed SDK metrics into the mandatory Micrometer registry (see quarkus-observability-otel skill). The reporter is `io.temporal.common.reporter.MicrometerClientStatsReporter`, wrapped in a tally `Scope` — but **the stubs are the extension's, so do not hand-build `WorkflowServiceStubsOptions` to attach it**; supply the scope through whatever customization hook `quarkus-temporal` exposes for the connection. Verify the metrics actually appear in `/q/metrics` before considering this done. Workflow task latency and activity failure rates are the client-side signals that a worker fleet is unhealthy. **Task-queue backlog is not among them** — it is server-side state, read from `DescribeTaskQueue` or the Temporal server's own metrics, so wire the autoscaling signal there rather than expecting it in `/q/metrics`.
|
|
513
|
+
- Log inside workflow code only through `Workflow.getLogger(...)`; a plain logger re-emits every line on every replay and makes the log unreadable exactly when you need it.
|
|
514
|
+
- Workflow ids and task queues are bounded, business-meaningful values — good span attributes. The workflow *arguments* are not: they carry the same PII prohibition as everything else.
|
|
515
|
+
|
|
516
|
+
## 11. Testing: `TestWorkflowEnvironment`, in memory
|
|
517
|
+
|
|
518
|
+
Saga behaviour is tested with Temporal's in-memory environment — plain JUnit 5 + Mockito, **no `@QuarkusTest`**, consistent with how `Handler`s are tested (see hexagonal-core testing table). The environment skips time, so a `Workflow.sleep(Duration.ofDays(30))` resolves instantly.
|
|
519
|
+
|
|
520
|
+
```java
|
|
521
|
+
class PartyOnboardingWorkflowTest {
|
|
522
|
+
|
|
523
|
+
private static final String TASK_QUEUE = "test-queue";
|
|
524
|
+
|
|
525
|
+
private TestWorkflowEnvironment testEnv;
|
|
526
|
+
private Worker worker;
|
|
527
|
+
private PartyOnboardingActivities activities;
|
|
528
|
+
|
|
529
|
+
@BeforeEach
|
|
530
|
+
void setUp() {
|
|
531
|
+
testEnv = TestWorkflowEnvironment.newInstance();
|
|
532
|
+
worker = testEnv.newWorker(TASK_QUEUE);
|
|
533
|
+
worker.registerWorkflowImplementationTypes(PartyOnboardingWorkflowImpl.class);
|
|
534
|
+
activities = mock(PartyOnboardingActivities.class);
|
|
535
|
+
worker.registerActivitiesImplementations(activities);
|
|
536
|
+
}
|
|
537
|
+
|
|
538
|
+
@AfterEach
|
|
539
|
+
void tearDown() {
|
|
540
|
+
testEnv.close();
|
|
541
|
+
}
|
|
542
|
+
|
|
543
|
+
@Test
|
|
544
|
+
void shouldCompensateCreatedPartyWhenWalletFails() {
|
|
545
|
+
when(activities.createParty(any())).thenReturn("party-1");
|
|
546
|
+
when(activities.openWallet(any(), any())).thenThrow(new IllegalStateException("wallet down"));
|
|
547
|
+
testEnv.start();
|
|
548
|
+
|
|
549
|
+
PartyOnboardingWorkflow workflow = testEnv.getWorkflowClient().newWorkflowStub(
|
|
550
|
+
PartyOnboardingWorkflow.class,
|
|
551
|
+
WorkflowOptions.newBuilder().setTaskQueue(TASK_QUEUE).build());
|
|
552
|
+
|
|
553
|
+
assertThatThrownBy(() -> workflow.onboard(request()))
|
|
554
|
+
.isInstanceOf(WorkflowException.class);
|
|
555
|
+
|
|
556
|
+
verify(activities).deleteParty("party-1"); // the compensation ran
|
|
557
|
+
verify(activities, never()).closeWallet(any()); // the wallet was never created
|
|
558
|
+
}
|
|
559
|
+
}
|
|
560
|
+
```
|
|
561
|
+
|
|
562
|
+
Mandatory scenarios per workflow — the compensation paths are the point of the whole exercise:
|
|
563
|
+
|
|
564
|
+
| # | Scenario | Assert |
|
|
565
|
+
|---|---|---|
|
|
566
|
+
| 1 | Happy path | result DTO correct; activities invoked in order; **no** compensation invoked |
|
|
567
|
+
| 2 | Failure at step *n* | compensations for steps `1..n-1` ran, in reverse order; none for `n..end`; the workflow still fails |
|
|
568
|
+
| 3 | Non-retryable failure | the activity ran **once** (`verify(activities, times(1))`) — proves the activity threw a non-retryable `ApplicationFailure` instead of being retried |
|
|
569
|
+
| 4 | Compensation is idempotent | a compensation invoked twice leaves the same end state |
|
|
570
|
+
| 5 | Signal / timer paths, where present | time-skipping drives the timer; assert the state the signal produced |
|
|
571
|
+
|
|
572
|
+
- Mockito stubs the **activity interface**, so a workflow test never touches a database or a slice `Handler`. Slice behaviour is already covered by that slice's own `HandlerTest`.
|
|
573
|
+
- Give mocked failing activities a bounded retry policy (or a non-retryable failure), or the test exercises the full retry schedule before failing.
|
|
574
|
+
- Add a `WorkflowReplayer` test against a committed history JSON for any workflow already running in production — it fails the build on a change that would break in-flight executions (§4), which is the one bug integration tests cannot find.
|
|
575
|
+
|
|
576
|
+
## Checklist for a new workflow
|
|
577
|
+
|
|
578
|
+
1. The operation genuinely crosses a slice, service or time boundary — a single `Handler` + one transaction cannot do it.
|
|
579
|
+
2. Topology chosen and written in the app README: **worker-only** (default — deployable named `{module}-{capability}-worker`, no REST, no slices, no datasource, activities call remote services) or **embedded** (stays `-ms`). In a worker-only project, none of the REST/slice/SQL standards were scaffolded and the ArchUnit suite ships the documented subset.
|
|
580
|
+
3. `io.quarkiverse.temporal:quarkus-temporal` on the classpath; `temporal-sdk` **not** declared directly; `temporal-testing` test scope at the version the extension pins; `quarkus-smallrye-health` present.
|
|
581
|
+
4. No bootstrap bean anywhere — lifecycle and discovery are the extension's. `quarkus.temporal.connection.target`, `worker.task-queue` (matching the interface constant) and `termination-timeout` set in `application.properties`.
|
|
582
|
+
5. Endpoint, namespace and certificates come from `${VAR}` / `.env` / vault — never literals (security skill §1).
|
|
583
|
+
6. Workflow lives in `orchestration/<capability>/`; in an embedded app, no `io.temporal` import inside any slice; ArchUnit quarantine + `*Impl` carve-out in place.
|
|
584
|
+
7. `*WorkflowImpl` is deterministic: no clock, no random, no threads, no I/O, no `@Inject`, no unordered-collection iteration.
|
|
585
|
+
8. Saga compensations registered **after** each forward step, LIFO, idempotent; irreversible steps ordered last; `compensate()` then rethrow.
|
|
586
|
+
9. Every activity: idempotent, `scheduleToCloseTimeout` set, terminal failures thrown as non-retryable `ApplicationFailure` (and **no** `setDoNotRetry` alongside it), delete-style compensations treating `NOT_FOUND`/`ALREADY_EXISTS` as success.
|
|
587
|
+
10. Workflow started with a deterministic workflow id, after commit — never mid-transaction.
|
|
588
|
+
11. Transport is gRPC in and out: the trigger reaches the `*Starter` through a `Handler` (embedded) or directly from the worker's own `*GrpcService` (worker-only), it **returns the workflow id without waiting**, and every outbound call goes through a `common/client` bean.
|
|
589
|
+
12. Timeouts ordered `grpc deadline < startToCloseTimeout < scheduleToCloseTimeout`, and retries are not stacked — Temporal owns the retry, the mesh keeps at most one.
|
|
590
|
+
13. Payload records annotated `@RegisterForReflection`; **no** `--initialize-at-run-time` flag for Netty/gRPC (the extension's substitutions handle it); the native binary actually built and exercised, not just the JVM tests.
|
|
591
|
+
14. `TestWorkflowEnvironment` tests cover happy path **and** every compensation path; tracing interceptors and Micrometer metrics scope registered.
|
|
592
|
+
15. `orchestration/<capability>/README.md` documents the steps, their compensations, the timeouts and the task queue.
|
|
@@ -18,17 +18,18 @@ For each story, invoke exactly **one** skill:
|
|
|
18
18
|
bmad-quarkus-build → "Marcus, implement story <story-key>"
|
|
19
19
|
```
|
|
20
20
|
|
|
21
|
-
That is the whole per-story cycle. **Do NOT run `bmad-create-
|
|
21
|
+
That is the whole per-story cycle. **Do NOT run `bmad-create-epics-and-stories`, `bmad-build`,
|
|
22
22
|
or `bmad-code-review` as separate steps** — Marcus owns the full loop (story
|
|
23
23
|
preparation → red-green-refactor implementation → self-review against the ACs) and
|
|
24
|
-
routes into `bmad-build` plus whichever of the
|
|
24
|
+
routes into `bmad-build` plus whichever of the 9 Quarkus domain standards the story
|
|
25
25
|
touches (`quarkus-hexagonal-core`, `quarkus-sql-jdbc-agroal`,
|
|
26
26
|
`quarkus-error-handling-i18n`, `quarkus-openapi-tmforum`, `quarkus-grpc-services`,
|
|
27
|
-
`quarkus-kafka-messaging`, `quarkus-observability-otel`
|
|
27
|
+
`quarkus-kafka-messaging`, `quarkus-observability-otel`, `quarkus-security-standards`,
|
|
28
|
+
and `quarkus-temporal-workflows` on Temporal projects). Splitting the cycle back into
|
|
28
29
|
three commands loses that routing and the native-image/hexagonal review that comes with it.
|
|
29
30
|
|
|
30
31
|
Pass the story key in the invocation so Marcus dispatches directly instead of rendering
|
|
31
|
-
his menu (his activation Step
|
|
32
|
+
his menu (his activation Step 9 skips the menu when the intent is already named).
|
|
32
33
|
|
|
33
34
|
Marcus stays active across stories once activated — do not re-run his activation
|
|
34
35
|
greeting for every story, just hand him the next story key.
|
|
@@ -139,7 +140,7 @@ When every eligible story is `done` or held, stop and report:
|
|
|
139
140
|
## Guardrails
|
|
140
141
|
|
|
141
142
|
- **One `bmad-quarkus-build` invocation per story.** Never decompose it back into
|
|
142
|
-
`bmad-create-
|
|
143
|
+
`bmad-create-epics-and-stories` → `bmad-build` → `bmad-code-review`.
|
|
143
144
|
- Never mark a story `done` unless its tests actually pass and every AC is met. ACs in these
|
|
144
145
|
stories are verbatim Gherkin — do not edit, reword, or renumber them.
|
|
145
146
|
- Never start a story in a `[BUILD-PROHIBITED]`, `[BUILD-PROHIBITED in part]`,
|