bmad-method-quarkus 1.0.5 → 1.0.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,592 @@
1
+ ---
2
+ name: quarkus-temporal-workflows
3
+ description: Standard for durable orchestration and distributed sagas in Quarkus native services using Temporal via the `io.quarkiverse.temporal:quarkus-temporal` extension — the plain `io.temporal:temporal-sdk` with manual wiring cannot produce a working native image. Covers `quarkus.temporal.*` config and auto-discovery (no bootstrap bean), the `io.temporal.workflow.Saga` compensation pattern, `@RegisterForReflection` record payloads, gRPC at both ends, non-retryable `ApplicationFailure`, and `TestWorkflowEnvironment` testing. Defines two topologies — a worker-only deployable (default, `-worker` suffix, no REST, no slices, no datasource) and an embedded worker in a slice app. Use when the user mentions Temporal, quarkus-temporal, workflow orchestration, durable execution, saga, compensation, activities, WorkflowClient, WorkerFactory, task queue, workflow id, `@WorkflowInterface`/`@ActivityInterface`, TestWorkflowEnvironment, a worker-only or orchestrator project, or a long-running process spanning services.
4
+ ---
5
+
6
+ # Temporal Durable Orchestration Standard (Quarkus)
7
+
8
+ Applies to any Quarkus backend project **that uses Temporal**; project directives (CLAUDE.md, ADRs, explicit instructions) override these defaults where they conflict.
9
+
10
+ Temporal is for the work a single `@Transactional Handler.process()` cannot make atomic: a business process that spans several slices, several services, or a span of time (minutes to months), where a failure halfway through must **undo** what already committed elsewhere. A JTA transaction already gives you all-or-nothing inside one database — reach for Temporal only when the operation crosses that boundary. Wrapping a one-slice insert in a workflow buys nothing and costs a worker, a task queue and a replay contract.
11
+
12
+ **Two topologies, and §3 is the first decision you make.** The default is a **worker-only project**: an independent deployable that hosts workflows and activities and has no REST layer, no slices and no database — the shape to reach for when the process spans services. The alternative is an **embedded worker** inside a normal slice-architecture app. §3 gives the layout for each and, for worker-only projects, states which house standards go quiet: no `Resource`, no `<Slice>Handler`/`<Slice>Sql`, no OpenAPI, usually no datasource.
13
+
14
+ **Scope inside a slice architecture (embedded only):** Temporal is quarantined to `orchestration/` and `common/temporal`, the same way Mutiny is quarantined to `*GrpcService` and `common/client` (see quarkus-hexagonal-core skill). No `io.temporal.*` type ever appears in a slice folder, a `Handler`, a `Sql` class or a slice DTO. A slice stays fully testable and deployable with Temporal absent from the classpath.
15
+
16
+ ## 1. Dependencies — the Quarkiverse extension, not the plain SDK
17
+
18
+ ```xml
19
+ <properties>
20
+ <quarkus.platform.version>3.33.3</quarkus.platform.version>
21
+ <quarkus-temporal.version>0.6.0</quarkus-temporal.version>
22
+ <!-- Must track the version quarkus-temporal itself pins (0.6.0 -> 1.38.0):
23
+ temporal-testing is the plain SDK's own artifact and has to match the
24
+ SDK the extension pulls in transitively. -->
25
+ <temporal.version>1.38.0</temporal.version>
26
+ </properties>
27
+
28
+ <dependencies>
29
+ <dependency>
30
+ <groupId>io.quarkiverse.temporal</groupId>
31
+ <artifactId>quarkus-temporal</artifactId>
32
+ <version>${quarkus-temporal.version}</version>
33
+ </dependency>
34
+ <dependency>
35
+ <groupId>io.temporal</groupId>
36
+ <artifactId>temporal-testing</artifactId>
37
+ <version>${temporal.version}</version>
38
+ <scope>test</scope>
39
+ </dependency>
40
+ </dependencies>
41
+ ```
42
+
43
+ **Use `io.quarkiverse.temporal:quarkus-temporal`. Do not wire `io.temporal:temporal-sdk` by hand — it cannot produce a working native image.** This reverses earlier guidance, and the reason is specific enough to be worth recording so nobody re-litigates it: the manual channel-building path hits a hard GraalVM `BUILD_TIME`/`RUN_TIME` policy conflict on `io.grpc.netty.shaded.io.netty.buffer.PooledByteBufAllocator` that **no `--initialize-at-run-time` / `--initialize-at-build-time` combination can override**. The extension ships GraalVM substitution classes for the shaded gRPC/Netty classes, which is what makes the binary build at all. Since the native binary is the delivery artifact (see quarkus-hexagonal-core skill), a JVM-only approach is not an option to fall back on.
44
+
45
+ - The extension pulls `temporal-sdk` transitively — do not declare it yourself, or the two versions will drift.
46
+ - `temporal-testing` stays `test` scope and must match the SDK version the extension pins. Bump both together.
47
+ - `quarkus-smallrye-health` is still required even though the worker serves no REST (see quarkus-hexagonal-core skill, README §11) — liveness/readiness are not optional because the app has no HTTP surface of its own.
48
+ - Everything you write is still plain Temporal SDK code following the official Java samples — `@WorkflowInterface`/`@ActivityInterface` pairs, activity stubs built inside the workflow, `Saga` for compensation. The extension owns the *lifecycle*, not your workflow code. Do not build a house abstraction over `Workflow.newActivityStub`; it hides the options that make a workflow correct.
49
+
50
+ ## 2. Configuration and lifecycle — owned by the extension
51
+
52
+ There is **no composition root to write**. The extension owns the `WorkflowServiceStubs` / `WorkflowClient` / `WorkerFactory` lifecycle and **auto-discovers** your workflow and activity implementations from their `@WorkflowInterface` / `@ActivityInterface` contracts. Everything is configuration:
53
+
54
+ ```properties
55
+ quarkus.temporal.connection.target=${TEMPORAL_SERVICE_TARGET:127.0.0.1:7233}
56
+ quarkus.temporal.worker.task-queue=alva.customer.party-onboarding.v1
57
+ # Blocks shutdown until in-flight activity tasks finish, or this timeout hits
58
+ quarkus.temporal.termination-timeout=20s
59
+ ```
60
+
61
+ - **Do not hand-write a `TemporalBootstrap`/`TemporalLifecycle` bean, a `@Produces WorkflowClient`, or a `WorkerFactory.start()` in a `StartupEvent` observer.** The extension already did it; a second registration on the same task queue is a duplicate poller. If you are migrating from the old manual wiring, delete the bean and the `@ConfigMapping` that fed it.
62
+ - `quarkus.temporal.termination-timeout` replaces the manual `factory.shutdown()` + `awaitTermination` drain. Budget it inside the pod's grace period alongside the telemetry flush (see quarkus-observability-otel skill, "The shutdown budget").
63
+ - The task queue is configured **and** referenced from code, so pin it once and keep the two in sync — a constant on the workflow interface, mirrored by the property:
64
+
65
+ ```java
66
+ @WorkflowInterface
67
+ public interface PartyOnboardingWorkflow {
68
+
69
+ /** Must match quarkus.temporal.worker.task-queue in application.properties. */
70
+ String TASK_QUEUE = "alva.customer.party-onboarding.v1";
71
+
72
+ @WorkflowMethod
73
+ PartyOnboardingResultDto onboard(PartyOnboardingRequestDto request);
74
+ }
75
+ ```
76
+
77
+ - Inject `WorkflowClient` wherever you start a workflow (§6); the extension produces it.
78
+ - The endpoint, namespace and any mTLS certificate path are environment-specific — `.env` locally, platform vault in deployed environments, `${VAR}` in properties (see quarkus-security-standards skill §1–2). A Temporal Cloud client certificate is a secret and is never committed.
79
+
80
+ **What has not changed:** a workflow implementation is still instantiated fresh per execution and replayed, so **CDI cannot inject into it** — its only collaborators are the activity stubs it builds itself (§4). Activity implementations, by contrast, are ordinary `@ApplicationScoped` beans and are injected normally.
81
+
82
+ ## 3. Topology: a worker-only project (default), or an embedded worker
83
+
84
+ Decide this **before** the first file, and record it in the app's `README.md` — the two shapes have different layouts and suspend different house rules.
85
+
86
+ ### A. Worker-only project — the default
87
+
88
+ A deployable that hosts workflows and activities **and nothing else**: no REST, no slices, and normally no database. Every side effect is a remote call to the service that owns that data (§7). This is the right shape whenever the process spans services, for two reasons that matter in production:
89
+
90
+ - Workflow code is **replay-sensitive** (§4). Keeping it in its own deployable means a workflow change is released on its own cadence instead of riding along with an unrelated business-service deploy — and a business service can be redeployed freely without touching in-flight executions.
91
+ - Workers scale on **task-queue backlog**, not on request rate. A pod sized for orchestration has nothing in common with one sized for an API.
92
+
93
+ ```
94
+ apps/customer-onboarding-worker/ # worker-only deployable — note the -worker suffix
95
+ ├── src/main/java/com/alva/customer/onboarding/
96
+ │ ├── common/
97
+ │ │ │ # no TemporalConfig, no bootstrap bean — §2
98
+ │ │ ├── client/PartyRegistrar.java # gRPC beans — how activities reach other services
99
+ │ │ ├── client/WalletOpener.java
100
+ │ │ └── exception/BusinessException.java # the house error taxonomy still applies (§9)
101
+ │ └── orchestration/
102
+ │ └── party_onboarding/ # snake_case, capability-named
103
+ │ ├── dto/
104
+ │ │ ├── PartyOnboardingRequestDto.java # record + @RegisterForReflection
105
+ │ │ └── PartyOnboardingResultDto.java
106
+ │ ├── PartyOnboardingWorkflow.java # @WorkflowInterface — the contract
107
+ │ ├── PartyOnboardingWorkflowImpl.java # the saga; deterministic, zero I/O
108
+ │ ├── PartyOnboardingActivities.java # @ActivityInterface
109
+ │ ├── PartyOnboardingActivitiesImpl.java # injects common/client beans — no local Handler
110
+ │ └── README.md # steps, compensations, timeouts, task queue
111
+ ├── src/main/proto/ # ONLY if this app owns the trigger RPC (below)
112
+ ├── Dockerfile · service.yaml · README.md
113
+ ```
114
+
115
+ There is no `Resource`, no `@Path`, no slice folder, no `Sql` class and no `db/` migration. **Do not scaffold them "for later"** — an unused datasource is a connection pool, a credential and a readiness dependency bought for nothing.
116
+
117
+ **The deployable carries the `-worker` suffix, never `-ms`** — `{module}-{capability}-worker`, e.g. `customer-onboarding-worker` (see the suffix table in the quarkus-hexagonal-core skill). The suffix is a deploy-time contract, and this kind of app contradicts everything `-ms` implies: no ingress, no HTTP route to probe, no API to publish, and scaling driven by task-queue backlog instead of request rate. Naming it `-ms` would hand the platform a service it would try to route to and scale on RPS it never receives. Its `service.yaml` matches: no provided APIs, task queues listed instead.
118
+
119
+ **Which monorepo owns it:** the domain that owns the business **process**, even when most steps call other domains. Cross-domain calls go through contracts published in `contracts/`, exactly like any other inter-domain call.
120
+
121
+ #### What does *not* apply in a worker-only project
122
+
123
+ Marcus reads eight universal standards; several of them have nothing to govern here. Applying them anyway produces a REST layer nobody calls and a database nobody reads.
124
+
125
+ | Standard | Status in a worker-only project |
126
+ |---|---|
127
+ | `quarkus-openapi-tmforum` | **Does not apply.** No `Resource`, no OpenAPI document, no Swagger. The only HTTP surface is `/q/health` and `/q/metrics`. |
128
+ | Vertical slices (`quarkus-hexagonal-core`) | **Does not apply.** No slice folders, no `<Slice>Handler`, no `<Slice>Sql`. `orchestration/` + `common/` *is* the structure. |
129
+ | `quarkus-sql-jdbc-agroal` | **Usually does not apply** — Temporal holds the process state, so there is no datasource. If the app genuinely owns a table (an idempotency/dedup ledger), its SQL follows that skill, with one adjustment: there is no `Handler` below the activity, so **the activity itself plays the Handler role** — it carries the `@Transactional`, opens the single `Connection` and passes it to every `Sql` call. That is the only case where an activity is transactional. |
130
+ | `quarkus-grpc-services` | **Applies fully** — it is how activities reach every other service (§7), and how the trigger RPC is exposed when this app owns one. |
131
+ | `quarkus-error-handling-i18n` | Error codes and the `BusinessException` taxonomy apply (§9). The REST edge mapper is absent; the gRPC interceptor applies only if the app exposes an RPC. |
132
+ | `quarkus-observability-otel` | **Applies fully, and matters more** — there is no request log to fall back on when a workflow misbehaves. |
133
+ | `quarkus-security-standards` | **Applies fully** — Temporal endpoint, namespace and mTLS certificates are secrets (§2). |
134
+ | `quarkus-kafka-messaging` | Applies when a workflow is triggered by, or emits, domain events. |
135
+ | Rest of `quarkus-hexagonal-core` | Naming, `apps/` layout, Java 25, native build + Dockerfile, `service.yaml` + README: **all unchanged**. |
136
+
137
+ ArchUnit ships a **subset**, and the two lists below partition hexagonal-core's canonical suite — every rule it defines is either kept or explicitly dropped, so nothing is left ambiguous.
138
+
139
+ **Keep:** `bannedSuffixes` (with the `..orchestration..` `*Impl` carve-out below), `serviceSuffixIsReserved` and `grpcServices` (they bite the moment the worker owns its trigger RPC), `reactiveIsQuarantined`, `noOrm`, `businessExceptionIsTheOnlyOne`, `dtosStayInDtoPackage`, `dtosHaveNoLogic`, plus this skill's own `workflowsDoNoIo`.
140
+
141
+ **Drop:** `slicesAreIndependent`, `onlyHandlersTouchSql`, `transactionalOnlyOnProcess`, `handlersExposeOnlyProcess`, `handlerKnowsNoTransport`, `crossCuttingWritesAreCentralised`, `restResources` and `transportHasNoJdbc` — with no slices, no `Handler` and no `Sql` to match, they can never fail, and a rule that cannot fail is noise that makes the suite look stronger than it is.
142
+
143
+ `temporalIsQuarantined` is **dropped too, and only here**: in a worker-only app the whole codebase is Temporal, so the rule would forbid what the app exists to do. It is mandatory in the embedded topology (B).
144
+
145
+ #### Who starts the workflow
146
+
147
+ - **Default: the calling service does.** The worker exposes **no inbound API at all** — it polls its task queue and nothing else. The service that owns the triggering business event holds its own `*Starter` bean (§6); the workflow id and the payload record are the contract between them. Nothing to route, nothing to secure at the edge.
148
+ - **Alternative: the worker owns the trigger RPC.** Then it has `src/main/proto/` and one `*GrpcService`, and that adapter calls the `*Starter` **directly — with no `Handler` in between.** This is a deliberate deviation from "a transport adapter always delegates to a `Handler`": a `Handler` exists to hold business logic and a transaction, and a trigger RPC has neither. An empty pass-through `Handler` would be pure ceremony. The adapter still does only proto↔DTO translation plus the start call.
149
+
150
+ Readiness reflects the **worker**, not a route: can it reach the Temporal frontend, and is the factory polling. Scale on task-queue backlog (§10), never on RPS.
151
+
152
+ ### B. Embedded worker — the exception
153
+
154
+ A normal `-ms` app that also hosts a worker — and it **stays `-ms`**, because it still serves an API; the `-worker` suffix is reserved for deployables that serve none. Use it only when every step of the saga is a capability **this same app already owns**, and Temporal is there for durability, retries or long waits rather than for crossing a service boundary. The cost is that the app now has two runtime roles and a replay-sensitive workflow tied to the service's release cadence.
155
+
156
+ ```
157
+ src/main/java/com/<company>/<module>/
158
+ ├── common/temporal/PartyOnboardingStarter.java # §6 — the only Temporal-aware class outside orchestration/
159
+ ├── orchestration/
160
+ │ └── party_onboarding/… # same contents; activities call local Handlers
161
+ └── create_party_individual/ # slices — untouched, no Temporal import
162
+ ```
163
+
164
+ Here the dependency direction is deliberately **one-way**: `orchestration/` may depend on slice `Handler`s and slice DTOs; a slice may never depend on `orchestration/`. A slice that imports a workflow has been made un-deployable without a Temporal cluster. Extend `slicesAreIndependent` with the single asymmetric exemption shown below, and keep every other slice rule intact — in this topology they all still bite.
165
+
166
+ ### Naming
167
+
168
+ | Artifact | Convention | Example |
169
+ |---|---|---|
170
+ | Workflow interface | `<Capability>Workflow`, `@WorkflowInterface` | `PartyOnboardingWorkflow` |
171
+ | Workflow implementation | `<Capability>WorkflowImpl` | `PartyOnboardingWorkflowImpl` |
172
+ | Activity interface | `<Capability>Activities`, `@ActivityInterface` | `PartyOnboardingActivities` |
173
+ | Activity implementation | `<Capability>ActivitiesImpl` | `PartyOnboardingActivitiesImpl` |
174
+ | Workflow method | imperative verb, `@WorkflowMethod` | `onboard` |
175
+ | Signal / Query | `<verb>` / `get<Noun>` | `approve` / `getStatus` |
176
+ | Starter bean | capability noun in `common/temporal` | `PartyOnboardingStarter` |
177
+ | Payload DTO | `*Dto` **record** in `dto/` | `PartyOnboardingRequestDto` |
178
+ | Task queue | `<org>.<module>.<capability-kebab>.v<major>` | `alva.customer.party-onboarding.v1` |
179
+ | Workflow id | `<capability-kebab>-<business-id>` — deterministic | `party-onboarding-9f3e…` |
180
+
181
+ **`*Impl` is banned everywhere else in this codebase and sanctioned here.** The SDK builds its client stubs from the annotated interface, so the interface/implementation pair is structural, not stylistic, and the official samples name it `*Impl`. Carve it out of the ArchUnit `bannedSuffixes` rule **by package**, exactly as `common/util/StringUtils` and the `Jdbc` helper are carved out — never by weakening the rule:
182
+
183
+ ```java
184
+ @ArchTest
185
+ static final ArchRule bannedSuffixesOutsideOrchestration = noClasses()
186
+ .that().resideOutsideOfPackage("..orchestration..")
187
+ .should().haveSimpleNameEndingWith("Impl"); // …plus the existing suffixes
188
+
189
+ // Temporal stays quarantined
190
+ @ArchTest
191
+ static final ArchRule temporalIsQuarantined = noClasses()
192
+ .that().resideOutsideOfPackage("..orchestration..")
193
+ .and().resideOutsideOfPackage("..common.temporal..")
194
+ .should().dependOnClassesThat().resideInAPackage("io.temporal..");
195
+
196
+ // A workflow implementation performs no I/O of its own (§4)
197
+ @ArchTest
198
+ static final ArchRule workflowsDoNoIo = noClasses()
199
+ .that().haveSimpleNameEndingWith("WorkflowImpl")
200
+ .should().dependOnClassesThat().resideInAnyPackage(
201
+ "java.sql..", "javax.sql..", "jakarta.ws.rs..", "io.quarkus.grpc..");
202
+ ```
203
+
204
+ Extend the existing `slicesAreIndependent` rule with **one** asymmetric exemption — orchestration may reach into slices, never the reverse:
205
+
206
+ ```java
207
+ .ignoreDependency(resideInAPackage("..orchestration.."), alwaysTrue())
208
+ ```
209
+
210
+ ## 4. The workflow implementation: deterministic, and that is a hard constraint
211
+
212
+ Temporal replays workflow code from its event history after every worker restart. Replay must produce the identical sequence of commands, so a workflow implementation is **pure orchestration**: it calls activity stubs and does nothing else.
213
+
214
+ Banned inside a `*WorkflowImpl`, with the replacement the SDK provides:
215
+
216
+ | Never | Always |
217
+ |---|---|
218
+ | `System.currentTimeMillis()`, `LocalDateTime.now()` | `Workflow.currentTimeMillis()` |
219
+ | `UUID.randomUUID()`, `Math.random()` | `Workflow.randomUUID()`, `Workflow.newRandom()` |
220
+ | `Thread.sleep(…)` | `Workflow.sleep(Duration…)` |
221
+ | `new Thread(…)`, executor services | `Async.function(…)` / `Async.procedure(…)`, scoped with `Workflow.newCancellationScope` |
222
+ | JDBC, HTTP, gRPC, Kafka, file I/O | an **activity** |
223
+ | `Logger.getLogger(…)` | `Workflow.getLogger(…)` — suppresses duplicate lines on replay |
224
+ | iteration over `HashMap`/`HashSet` | `LinkedHashMap`/`TreeMap`, or sort explicitly |
225
+
226
+ `@Inject` does not work here either (§2) — the workflow's only collaborators are its activity stubs. A workflow that needs configuration receives it as a **workflow argument**, never read from `ConfigProvider` at replay time.
227
+
228
+ **Versioning is the production footgun.** Editing a workflow implementation while executions of the old code are still running causes a non-deterministic replay and a stuck workflow. Any change to the *sequence* of activity calls requires `Workflow.getVersion("addWallet", DEFAULT_VERSION, 1)` branching, or a new task queue (`…v2`) drained in parallel. Changing an activity's *body* is always safe; changing the workflow's call order never is.
229
+
230
+ ## 5. The saga: `io.temporal.workflow.Saga` with explicit compensations
231
+
232
+ ```java
233
+ public class PartyOnboardingWorkflowImpl implements PartyOnboardingWorkflow {
234
+
235
+ private static final Logger log = Workflow.getLogger(PartyOnboardingWorkflowImpl.class);
236
+
237
+ private final PartyOnboardingActivities activities = Workflow.newActivityStub(
238
+ PartyOnboardingActivities.class,
239
+ ActivityOptions.newBuilder()
240
+ .setStartToCloseTimeout(Duration.ofSeconds(30))
241
+ .setScheduleToCloseTimeout(Duration.ofMinutes(10)) // the outer bound — never omit
242
+ .setRetryOptions(RetryOptions.newBuilder()
243
+ .setInitialInterval(Duration.ofSeconds(1))
244
+ .setMaximumInterval(Duration.ofSeconds(30))
245
+ .setBackoffCoefficient(2.0)
246
+ .build()) // no setDoNotRetry — see §9
247
+ .build());
248
+
249
+ @Override
250
+ public PartyOnboardingResultDto onboard(PartyOnboardingRequestDto request) {
251
+ Saga saga = new Saga(new Saga.Options.Builder()
252
+ .setParallelCompensation(false) // compensate in reverse order, one at a time
253
+ .setContinueWithError(true) // attempt every compensation — see below
254
+ .build());
255
+ try {
256
+ String partyId = activities.createParty(request);
257
+ saga.addCompensation(activities::deleteParty, partyId);
258
+
259
+ String walletId = activities.openWallet(partyId, request.currency());
260
+ saga.addCompensation(activities::closeWallet, walletId);
261
+
262
+ activities.sendWelcomeNotification(partyId);
263
+ return new PartyOnboardingResultDto(partyId, walletId);
264
+
265
+ } catch (RuntimeException e) { // deliberately broad — see below
266
+ log.warn("Onboarding failed, compensating: {}", e.getMessage());
267
+ saga.compensate();
268
+ throw e; // the workflow still fails — do not swallow
269
+ }
270
+ }
271
+ }
272
+ ```
273
+
274
+ Rules:
275
+ - **Catch `RuntimeException`, not `ActivityFailure`.** `ActivityFailure` misses `ChildWorkflowFailure`, `CanceledFailure` and every other `TemporalFailure` subtype, and a miss means the compensation chain silently never runs — the exact failure the saga exists to prevent. `RuntimeException` is deliberately broader than even `TemporalFailure`: anything that escapes the forward path should compensate.
276
+ - **`setContinueWithError(true)` is the default here.** Out of the box it is `false`, which makes the *first* failing compensation abort the rest of the chain and leave earlier steps uncompensated — almost never what you want, since the compensations are usually independent deletes. Leave it `false` only when a failed compensation genuinely makes the remaining ones unsafe.
277
+ - **A step whose effect is an idempotent, resubmittable correction needs no compensation at all.** Register one only for steps that created something. Say so in a comment where you skip it, or the next reader will assume it was forgotten.
278
+ - **Register the compensation *after* the forward activity returns**, using its result. Registering before means compensating something that never happened.
279
+ - Compensations run in **reverse registration order** (LIFO). Keep `setParallelCompensation(false)` unless the steps are provably independent — parallel compensation of dependent resources reintroduces the ordering bug the saga exists to avoid.
280
+ - **`saga.compensate()` then rethrow.** A saga that compensates and returns normally reports success for a business process that did not happen.
281
+ - A compensation is itself an activity, so it is retried like any other — and it must be **idempotent**: `deleteParty` on an already-deleted party succeeds quietly rather than failing the compensation chain.
282
+ - No compensation is possible for some steps (an email that went out). Order the workflow so irreversible steps come **last**, after everything reversible has succeeded.
283
+
284
+ ## 6. Activities: the only side-effecting layer
285
+
286
+ An activity implementation is thin: extract the arguments, perform **exactly one** side effect, return a plain value. **No business logic, no orchestration, no `@Transactional` of its own** — the transaction always lives one level down. What it delegates *to* is the one thing the topology (§3) changes:
287
+
288
+ ```java
289
+ // A. Worker-only project (default): the side effect is a remote call
290
+ @ApplicationScoped
291
+ public class PartyOnboardingActivitiesImpl implements PartyOnboardingActivities {
292
+
293
+ @Inject PartyRegistrar partyRegistrar; // common/client — gRPC, awaits internally (§7)
294
+ @Inject WalletOpener walletOpener;
295
+
296
+ @Override
297
+ public String createParty(PartyOnboardingRequestDto request) {
298
+ return partyRegistrar.register(request.tenantId(), request.givenName(), request.familyName());
299
+ }
300
+ }
301
+ ```
302
+
303
+ ```java
304
+ // B. Embedded worker: the side effect is a Handler in this same app
305
+ @Inject CreatePartyIndividualHandler createPartyHandler;
306
+
307
+ @Override
308
+ public String createParty(PartyOnboardingRequestDto request) {
309
+ return createPartyHandler.process(toSliceRequest(request)).getPartyId();
310
+ }
311
+ ```
312
+
313
+ Either way the activity is a transport adapter in the sense the hexagonal-core skill defines, and the `@Transactional` boundary is the remote service's `Handler` (A) or the local one (B) — never here. The single exception is a worker-only app that owns a table of its own (§3): with no `Handler` beneath it, that activity plays the `Handler` role and carries the transaction itself.
314
+
315
+ - **Activities are at-least-once.** Temporal retries on worker crash, timeout and failure, so every activity must be idempotent. Where the underlying operation is not naturally idempotent, dedupe on a business key or on `Activity.getExecutionContext().getInfo().getWorkflowId()`.
316
+ - Every activity needs a `setScheduleToCloseTimeout` (or a `setMaximumAttempts`). An activity with unlimited retries and no outer bound does not fail — it hangs forever, in production and in your tests.
317
+ - Long-running activities heartbeat (`Activity.getExecutionContext().heartbeat(progress)`) so a dead worker is detected in seconds rather than at the start-to-close timeout.
318
+
319
+ ### Starting a workflow
320
+
321
+ The `*Starter` bean below is the same in both topologies — what differs is **which deployable owns it**:
322
+
323
+ - **Worker-only (A):** the `*Starter` lives in the *calling* service, which therefore carries `quarkus-temporal` of its own — configured as a **client without a worker**, since it starts workflows rather than running them. That falls out of auto-discovery on its own: the calling service contains no `@WorkflowInterface`/`@ActivityInterface` implementations, so the extension has nothing to register and starts no poller — just set `quarkus.temporal.connection.target` and leave the worker's task queue unset. Never register a worker on the orchestrator's task queue from a service that cannot execute the workflow. When the worker app owns its own trigger RPC (§3), its `*GrpcService` injects `WorkflowClient` and starts the workflow inline — no separate bean.
324
+ - **Embedded (B):** it lives in this app, and a slice `Handler` injects it.
325
+
326
+ In the embedded case a `Handler` must start the workflow through this bean and **never** by injecting `WorkflowClient` into the slice — that would drag `io.temporal` past the §3 quarantine:
327
+
328
+ ```java
329
+ // common/temporal
330
+ @ApplicationScoped
331
+ public class PartyOnboardingStarter {
332
+
333
+ @Inject WorkflowClient client; // produced by the extension
334
+
335
+ public void start(String partyId, PartyOnboardingRequestDto request) {
336
+ PartyOnboardingWorkflow workflow = client.newWorkflowStub(
337
+ PartyOnboardingWorkflow.class,
338
+ WorkflowOptions.newBuilder()
339
+ .setTaskQueue(PartyOnboardingWorkflow.TASK_QUEUE) // the interface constant, §2
340
+ .setWorkflowId("party-onboarding-" + partyId) // deterministic = deduplicated
341
+ .setWorkflowIdReusePolicy(
342
+ WorkflowIdReusePolicy.WORKFLOW_ID_REUSE_POLICY_ALLOW_DUPLICATE_FAILED_ONLY)
343
+ .build());
344
+ WorkflowClient.start(workflow::onboard, request); // async — never block the Handler
345
+ }
346
+ }
347
+ ```
348
+
349
+ **Starting a workflow inside a database transaction is a dual write.** If `process()` rolls back after the start call succeeded, the workflow orchestrates an entity that does not exist. This is the same hazard the outbox solves for Kafka (see quarkus-kafka-messaging skill), and it gets the same two answers:
350
+
351
+ 1. **Deterministic workflow id derived from the business entity** (above) — a retry of the whole operation cannot start a second workflow, so the failure mode is recoverable rather than duplicated. This is mandatory, not optional. Use `ALLOW_DUPLICATE_FAILED_ONLY`, **not** `REJECT_DUPLICATE`: the latter blocks reuse even after the run closed, so a legitimate re-run of that business id becomes permanently impossible. The starter must also catch `WorkflowExecutionAlreadyStarted` and treat it as success — that exception *is* the deduplication working, not an error to propagate.
352
+ 2. **Start after commit**, the way `OutboxDispatcher.dispatchAfterCommit(eventId)` already does it — or record an outbox row and let the relay start the workflow. Never call `WorkflowClient.start` in the middle of `execution()` and hope the transaction commits.
353
+
354
+ ## 7. Transport: the invocations are gRPC
355
+
356
+ Orchestration is internal by nature — a saga exists because a process crosses services — so **gRPC is the predominant transport at both ends of a workflow**, per the house rule that internal synchronous calls are always gRPC and never internal REST (see quarkus-grpc-services and quarkus-hexagonal-core skills). North-bound REST triggers exist (a customer action that kicks off onboarding) but are the minority; the default assumption for a Temporal capability is gRPC in, gRPC out.
357
+
358
+ **A. Worker-only project** — the calling service starts the workflow; the worker only runs it and calls back out over gRPC:
359
+
360
+ ```
361
+ calling service (-ms) worker project (-worker) other services
362
+ ┌───────────────────────┐ ┌────────────────────────┐
363
+ │ GrpcService → Handler │ │ PartyOnboardingWorkflowImpl │ (no I/O here)
364
+ │ ↓ │ │ ↓ │
365
+ │ <Capability>Starter ─┼── Temporal ──┼→ PartyOnboardingActivitiesImpl │
366
+ └───────────────────────┘ task queue │ ↓ │
367
+ │ common/client bean ───┼─ gRPC ─→ party-ms, wallet-ms
368
+ └────────────────────────┘
369
+ ```
370
+
371
+ **B. Embedded worker** — everything in one deployable, the activity reaching a local `Handler`:
372
+
373
+ ```
374
+ <Proto>GrpcService → Handler.process() → <Capability>Starter → Temporal
375
+ (slice folder) (slice folder) (common/temporal)
376
+ ↓
377
+ PartyOnboardingWorkflowImpl (no I/O — orchestration only)
378
+ ↓
379
+ PartyOnboardingActivitiesImpl → slice Handler
380
+ ```
381
+
382
+ Nothing about the inbound chain is special-cased for Temporal: the `*GrpcService` delegates to a `Handler`, and the `Handler` injects the `*Starter`. **A `*GrpcService` never injects `WorkflowClient` directly** — that would put `io.temporal` inside a slice folder and break the §3 quarantine. The one exception is the trigger RPC of a worker-only project, which has no `Handler` to delegate to and calls the `*Starter` itself (§3).
383
+
384
+ ### Inbound: the gRPC call starts the workflow, it does not wait for it
385
+
386
+ A gRPC client carries a deadline measured in seconds (`quarkus.grpc.clients.<name>.deadline=2s`); a saga runs for minutes, days or months. So the RPC that triggers a workflow **returns as soon as the workflow is started** — the workflow id is the response — and never blocks on completion:
387
+
388
+ - Start asynchronously (`WorkflowClient.start(...)`, as in §6), return the workflow id to the caller.
389
+ - The caller learns the outcome by polling a `@QueryMethod`, by consuming the domain event the workflow's final activity emits (outbox → Kafka, per the messaging skill), or by its own callback RPC.
390
+ - Blocking the RPC on `workflow.onboard(request)` ties a durable execution to a 2-second deadline: the deadline fires, the caller retries, and the only thing the deterministic workflow id saved you from was a second workflow. The first one is still running, unobserved.
391
+
392
+ ### Outbound: activities call other services through `common/client`
393
+
394
+ An activity that reaches another service uses a capability-named bean from `common/client` — the same bean a `Handler` would use, never a raw `@GrpcClient` stub and never an internal REST client:
395
+
396
+ ```java
397
+ @ApplicationScoped
398
+ public class PartyOnboardingActivitiesImpl implements PartyOnboardingActivities {
399
+
400
+ @Inject CredentialsValidator credentialsValidator; // common/client — awaits internally
401
+
402
+ @Override
403
+ public boolean validateCredentials(String username, String password) {
404
+ return credentialsValidator.validate(username, password); // plain types in, plain types out
405
+ }
406
+ }
407
+ ```
408
+
409
+ Because those beans `await()` internally and expose plain types, the `reactiveIsQuarantined` ArchUnit rule keeps passing unchanged — no Mutiny ever reaches `orchestration/`, exactly as none reaches a `Handler`.
410
+
411
+ **Only activities may call out. A workflow never holds a client**, gRPC or otherwise (§4).
412
+
413
+ ### Two timeout layers, and the order between them is not optional
414
+
415
+ ```
416
+ grpc deadline < activity startToCloseTimeout < activity scheduleToCloseTimeout
417
+ 2s 30s 10m
418
+ ```
419
+
420
+ Get this backwards and the activity times out while the gRPC call is still in flight: Temporal schedules a retry, and now two identical calls are hitting a downstream that is already struggling. The gRPC deadline must be the *inner* bound so the call is dead before Temporal reacts to it.
421
+
422
+ **Do not stack retries.** Temporal's activity retry policy and the gRPC/mesh retry policy multiply: 5 activity attempts × 3 mesh retries = 15 calls, sent to the exact service that is failing. Temporal owns the retry — its attempts are durable, backed off, and visible in the workflow history, where a mesh retry is invisible. Keep the client/mesh layer at a single retry for connection-level `UNAVAILABLE` blips at most, and let the activity policy do the rest.
423
+
424
+ ### Mapping gRPC failures onto the retry policy
425
+
426
+ A `StatusRuntimeException` from a downstream is classified at the activity boundary. **§9 holds the single authoritative list of which codes are terminal — do not restate it here or anywhere else**, because two copies of a retry-classification table drift and the drift is silent: a code that is terminal in one list and retryable in the other produces either an infinite retry or a saga that compensates on a blip.
427
+
428
+ The downstream already maps its `BusinessException` codes onto gRPC statuses on the way out (see the grpc skill's error-mapping table), so §9's table is that one read in reverse — keep the two consistent when either changes.
429
+
430
+ ### The SDK's own transport
431
+
432
+ The worker and the client talk to the Temporal service over gRPC as well, on shaded Netty. That is not trivia: it is why the native build depends on the GraalVM substitutions `quarkus-temporal` ships rather than on any initialization flag (§8), and why a Temporal Cloud connection is mTLS with a client certificate handled as a secret (§2). The connection's own RPC timeouts and keepalive are separate from every timeout above and are configured through `quarkus.temporal.connection.*`, not by hand-building stubs.
433
+
434
+ ## 8. Payloads and native image
435
+
436
+ **Payloads are Java records, always annotated `@RegisterForReflection`.** Records are immutable, which is what replay safety wants, and the SDK's default `DataConverter` serializes them with Jackson — which needs reflection metadata that GraalVM will not infer.
437
+
438
+ ```java
439
+ @RegisterForReflection
440
+ public record PartyOnboardingRequestDto(
441
+ String tenantId,
442
+ String givenName,
443
+ String familyName,
444
+ String currency) {}
445
+ ```
446
+
447
+ This is a deliberate local deviation from the slice DTO convention (Lombok `@Data` classes, per hexagonal-core): a Temporal payload is serialized into an immutable event history, so a mutable DTO with setters is the wrong shape. The `*Dto` suffix and the `dto/` package still apply, so the existing `dtosStayInDtoPackage` ArchUnit rule keeps passing.
448
+
449
+ ### The native build needs no Temporal-specific flags
450
+
451
+ ```properties
452
+ # That is the whole native configuration — no additional-build-args for Temporal
453
+ quarkus.native.container-build=true
454
+ quarkus.native.builder-image=mandrel
455
+ ```
456
+
457
+ **Do not add `--initialize-at-run-time=io.grpc.netty.shaded.io.netty` (or any variant of it).** This is the trap that cost a spike: it looks like the obvious fix, it is what the plain-SDK path appears to need, and **it does not work** — the conflict on `io.grpc.netty.shaded.io.netty.buffer.PooledByteBufAllocator` is a build-time/run-time *policy* clash that an initialization flag cannot resolve in either direction. What resolves it is the substitution classes `quarkus-temporal` ships (§1). Adding the flag on top of the extension buys nothing and misleads the next person debugging a native failure.
458
+
459
+ - If the app sets `quarkus.native.additional-build-args` for some *other* reason, remember the property is a comma-separated list — append, never overwrite.
460
+ - The SDK builds workflow and activity stubs with JDK dynamic proxies, which is precisely the pattern hexagonal-core tells you to avoid. The extension's build step registers what it needs; if a proxy is still missing, use Quarkus's `@RegisterForProxy` rather than a raw flag. (`-H:DynamicProxyConfigurationResources` is gone on GraalVM for JDK 21+ — proxies go through reachability metadata now.)
461
+ - **A Temporal app is not "native-ready" until the binary has actually been built and exercised.** The JVM-mode test proves nothing here: every problem in this section appears only at image-build time.
462
+
463
+ ## 9. Errors and retries
464
+
465
+ Temporal's retry policy is the automatic remediation this codebase otherwise hand-writes — but only if failures are classified correctly. **The classification lives in the activity, expressed one way only: throw a non-retryable `ApplicationFailure` for what must not be retried, and let everything else propagate unchanged.**
466
+
467
+ ```java
468
+ RuntimeException translate(StatusRuntimeException e, String activityName) {
469
+ Status.Code code = e.getStatus().getCode();
470
+ if (TERMINAL_CODES.contains(code)) {
471
+ String description = e.getStatus().getDescription();
472
+ Log.warnf("%s: terminal gRPC failure code=%s, non-retryable: %s", activityName, code, description);
473
+ return ApplicationFailure.newNonRetryableFailure(
474
+ description != null ? description : e.getMessage(), code.name());
475
+ }
476
+ // UNAVAILABLE, DEADLINE_EXCEEDED, INTERNAL, ... rethrown as-is, so the workflow's
477
+ // ActivityOptions retry policy applies with backoff up to scheduleToCloseTimeout.
478
+ Log.warnf("%s: retryable gRPC failure code=%s: %s", activityName, code, e.getMessage());
479
+ return e;
480
+ }
481
+ ```
482
+
483
+ **Do not also configure `setDoNotRetry(...)` on the `ActivityOptions`.** It is a second, independent mechanism that matches on the failure *type* string, so the moment that string and the type you pass to `ApplicationFailure` drift apart the declarative rule silently matches nothing — and a rule that looks wired but is inert is worse than no rule. One mechanism, at the activity boundary.
484
+
485
+ Which gRPC codes are terminal:
486
+
487
+ | Code from the downstream | Treatment |
488
+ |---|---|
489
+ | `INVALID_ARGUMENT`, `ALREADY_EXISTS`, `UNAUTHENTICATED`, `PERMISSION_DENIED` | **non-retryable** `ApplicationFailure` — a malformed or duplicate request, or a token that will not become valid inside this run |
490
+ | everything else (`UNAVAILABLE`, `DEADLINE_EXCEEDED`, `INTERNAL`, `RESOURCE_EXHAUSTED`, …) | rethrown as-is → retried with backoff, bounded by `scheduleToCloseTimeout` |
491
+
492
+ **A compensation activity classifies differently, and getting this wrong deadlocks the rollback.** For a delete-style compensation, `NOT_FOUND` means the thing is already gone and `ALREADY_EXISTS` may encode the downstream's own conflict mapping — both mean *the compensation is as done as it can be*, so the activity must **return normally**, not fail. The same `ALREADY_EXISTS` that correctly fails a forward `create` must not fail its matching `delete`:
493
+
494
+ ```java
495
+ } catch (StatusRuntimeException e) {
496
+ Status.Code code = e.getStatus().getCode();
497
+ if (code == Status.Code.NOT_FOUND || code == Status.Code.ALREADY_EXISTS) {
498
+ Log.infof("%s: compensation handled (%s)", activityName, code);
499
+ return; // idempotent: already compensated
500
+ }
501
+ throw translate(e, activityName);
502
+ }
503
+ ```
504
+
505
+ The house exception taxonomy applies in both topologies — the worker-only layout in §3 ships `common/exception/BusinessException` too — and splits the same way: `BusinessException` → non-retryable; `TransientPersistenceException` / `PersistenceTimeoutException` / `StaleVersionException` → retryable, which is exactly what the retry policy exists for; `PersistenceException` → retryable but capped by `scheduleToCloseTimeout`.
506
+
507
+ The failure a client ultimately sees is still a house error code: a `WorkflowFailedException` surfacing at the REST edge is translated by `GlobalExceptionHandler` into the catalogued `<MOD>-<HTTP>-<seq>`, localized as usual. A raw Temporal stack trace never reaches an API consumer.
508
+
509
+ ## 10. Observability
510
+
511
+ - Propagate trace context into workflows and activities with `OpenTracingClientInterceptor` / `OpenTracingWorkerInterceptor`, registered through the extension's interceptor configuration. They need two dependencies §1 does not list because they are tracing-only — `io.temporal:temporal-opentracing` and `io.opentelemetry:opentelemetry-opentracing-shim` to bridge onto the OTel SDK this platform uses; there is no OTel-native interceptor module. Without them a workflow is a hole in the trace: the REST span ends, and the activity spans belong to no one.
512
+ - Feed SDK metrics into the mandatory Micrometer registry (see quarkus-observability-otel skill). The reporter is `io.temporal.common.reporter.MicrometerClientStatsReporter`, wrapped in a tally `Scope` — but **the stubs are the extension's, so do not hand-build `WorkflowServiceStubsOptions` to attach it**; supply the scope through whatever customization hook `quarkus-temporal` exposes for the connection. Verify the metrics actually appear in `/q/metrics` before considering this done. Workflow task latency and activity failure rates are the client-side signals that a worker fleet is unhealthy. **Task-queue backlog is not among them** — it is server-side state, read from `DescribeTaskQueue` or the Temporal server's own metrics, so wire the autoscaling signal there rather than expecting it in `/q/metrics`.
513
+ - Log inside workflow code only through `Workflow.getLogger(...)`; a plain logger re-emits every line on every replay and makes the log unreadable exactly when you need it.
514
+ - Workflow ids and task queues are bounded, business-meaningful values — good span attributes. The workflow *arguments* are not: they carry the same PII prohibition as everything else.
515
+
516
+ ## 11. Testing: `TestWorkflowEnvironment`, in memory
517
+
518
+ Saga behaviour is tested with Temporal's in-memory environment — plain JUnit 5 + Mockito, **no `@QuarkusTest`**, consistent with how `Handler`s are tested (see hexagonal-core testing table). The environment skips time, so a `Workflow.sleep(Duration.ofDays(30))` resolves instantly.
519
+
520
+ ```java
521
+ class PartyOnboardingWorkflowTest {
522
+
523
+ private static final String TASK_QUEUE = "test-queue";
524
+
525
+ private TestWorkflowEnvironment testEnv;
526
+ private Worker worker;
527
+ private PartyOnboardingActivities activities;
528
+
529
+ @BeforeEach
530
+ void setUp() {
531
+ testEnv = TestWorkflowEnvironment.newInstance();
532
+ worker = testEnv.newWorker(TASK_QUEUE);
533
+ worker.registerWorkflowImplementationTypes(PartyOnboardingWorkflowImpl.class);
534
+ activities = mock(PartyOnboardingActivities.class);
535
+ worker.registerActivitiesImplementations(activities);
536
+ }
537
+
538
+ @AfterEach
539
+ void tearDown() {
540
+ testEnv.close();
541
+ }
542
+
543
+ @Test
544
+ void shouldCompensateCreatedPartyWhenWalletFails() {
545
+ when(activities.createParty(any())).thenReturn("party-1");
546
+ when(activities.openWallet(any(), any())).thenThrow(new IllegalStateException("wallet down"));
547
+ testEnv.start();
548
+
549
+ PartyOnboardingWorkflow workflow = testEnv.getWorkflowClient().newWorkflowStub(
550
+ PartyOnboardingWorkflow.class,
551
+ WorkflowOptions.newBuilder().setTaskQueue(TASK_QUEUE).build());
552
+
553
+ assertThatThrownBy(() -> workflow.onboard(request()))
554
+ .isInstanceOf(WorkflowException.class);
555
+
556
+ verify(activities).deleteParty("party-1"); // the compensation ran
557
+ verify(activities, never()).closeWallet(any()); // the wallet was never created
558
+ }
559
+ }
560
+ ```
561
+
562
+ Mandatory scenarios per workflow — the compensation paths are the point of the whole exercise:
563
+
564
+ | # | Scenario | Assert |
565
+ |---|---|---|
566
+ | 1 | Happy path | result DTO correct; activities invoked in order; **no** compensation invoked |
567
+ | 2 | Failure at step *n* | compensations for steps `1..n-1` ran, in reverse order; none for `n..end`; the workflow still fails |
568
+ | 3 | Non-retryable failure | the activity ran **once** (`verify(activities, times(1))`) — proves the activity threw a non-retryable `ApplicationFailure` instead of being retried |
569
+ | 4 | Compensation is idempotent | a compensation invoked twice leaves the same end state |
570
+ | 5 | Signal / timer paths, where present | time-skipping drives the timer; assert the state the signal produced |
571
+
572
+ - Mockito stubs the **activity interface**, so a workflow test never touches a database or a slice `Handler`. Slice behaviour is already covered by that slice's own `HandlerTest`.
573
+ - Give mocked failing activities a bounded retry policy (or a non-retryable failure), or the test exercises the full retry schedule before failing.
574
+ - Add a `WorkflowReplayer` test against a committed history JSON for any workflow already running in production — it fails the build on a change that would break in-flight executions (§4), which is the one bug integration tests cannot find.
575
+
576
+ ## Checklist for a new workflow
577
+
578
+ 1. The operation genuinely crosses a slice, service or time boundary — a single `Handler` + one transaction cannot do it.
579
+ 2. Topology chosen and written in the app README: **worker-only** (default — deployable named `{module}-{capability}-worker`, no REST, no slices, no datasource, activities call remote services) or **embedded** (stays `-ms`). In a worker-only project, none of the REST/slice/SQL standards were scaffolded and the ArchUnit suite ships the documented subset.
580
+ 3. `io.quarkiverse.temporal:quarkus-temporal` on the classpath; `temporal-sdk` **not** declared directly; `temporal-testing` test scope at the version the extension pins; `quarkus-smallrye-health` present.
581
+ 4. No bootstrap bean anywhere — lifecycle and discovery are the extension's. `quarkus.temporal.connection.target`, `worker.task-queue` (matching the interface constant) and `termination-timeout` set in `application.properties`.
582
+ 5. Endpoint, namespace and certificates come from `${VAR}` / `.env` / vault — never literals (security skill §1).
583
+ 6. Workflow lives in `orchestration/<capability>/`; in an embedded app, no `io.temporal` import inside any slice; ArchUnit quarantine + `*Impl` carve-out in place.
584
+ 7. `*WorkflowImpl` is deterministic: no clock, no random, no threads, no I/O, no `@Inject`, no unordered-collection iteration.
585
+ 8. Saga compensations registered **after** each forward step, LIFO, idempotent; irreversible steps ordered last; `compensate()` then rethrow.
586
+ 9. Every activity: idempotent, `scheduleToCloseTimeout` set, terminal failures thrown as non-retryable `ApplicationFailure` (and **no** `setDoNotRetry` alongside it), delete-style compensations treating `NOT_FOUND`/`ALREADY_EXISTS` as success.
587
+ 10. Workflow started with a deterministic workflow id, after commit — never mid-transaction.
588
+ 11. Transport is gRPC in and out: the trigger reaches the `*Starter` through a `Handler` (embedded) or directly from the worker's own `*GrpcService` (worker-only), it **returns the workflow id without waiting**, and every outbound call goes through a `common/client` bean.
589
+ 12. Timeouts ordered `grpc deadline < startToCloseTimeout < scheduleToCloseTimeout`, and retries are not stacked — Temporal owns the retry, the mesh keeps at most one.
590
+ 13. Payload records annotated `@RegisterForReflection`; **no** `--initialize-at-run-time` flag for Netty/gRPC (the extension's substitutions handle it); the native binary actually built and exercised, not just the JVM tests.
591
+ 14. `TestWorkflowEnvironment` tests cover happy path **and** every compensation path; tracing interceptors and Micrometer metrics scope registered.
592
+ 15. `orchestration/<capability>/README.md` documents the steps, their compensations, the timeouts and the task queue.
@@ -18,17 +18,18 @@ For each story, invoke exactly **one** skill:
18
18
  bmad-quarkus-build → "Marcus, implement story <story-key>"
19
19
  ```
20
20
 
21
- That is the whole per-story cycle. **Do NOT run `bmad-create-story`, `bmad-dev-story`,
21
+ That is the whole per-story cycle. **Do NOT run `bmad-create-epics-and-stories`, `bmad-build`,
22
22
  or `bmad-code-review` as separate steps** — Marcus owns the full loop (story
23
23
  preparation → red-green-refactor implementation → self-review against the ACs) and
24
- routes into `bmad-build` plus whichever of the 7 Quarkus domain standards the story
24
+ routes into `bmad-build` plus whichever of the 9 Quarkus domain standards the story
25
25
  touches (`quarkus-hexagonal-core`, `quarkus-sql-jdbc-agroal`,
26
26
  `quarkus-error-handling-i18n`, `quarkus-openapi-tmforum`, `quarkus-grpc-services`,
27
- `quarkus-kafka-messaging`, `quarkus-observability-otel`). Splitting the cycle back into
27
+ `quarkus-kafka-messaging`, `quarkus-observability-otel`, `quarkus-security-standards`,
28
+ and `quarkus-temporal-workflows` on Temporal projects). Splitting the cycle back into
28
29
  three commands loses that routing and the native-image/hexagonal review that comes with it.
29
30
 
30
31
  Pass the story key in the invocation so Marcus dispatches directly instead of rendering
31
- his menu (his activation Step 8 skips the menu when the intent is already named).
32
+ his menu (his activation Step 9 skips the menu when the intent is already named).
32
33
 
33
34
  Marcus stays active across stories once activated — do not re-run his activation
34
35
  greeting for every story, just hand him the next story key.
@@ -139,7 +140,7 @@ When every eligible story is `done` or held, stop and report:
139
140
  ## Guardrails
140
141
 
141
142
  - **One `bmad-quarkus-build` invocation per story.** Never decompose it back into
142
- `bmad-create-story` → `bmad-dev-story` → `bmad-code-review`.
143
+ `bmad-create-epics-and-stories` → `bmad-build` → `bmad-code-review`.
143
144
  - Never mark a story `done` unless its tests actually pass and every AC is met. ACs in these
144
145
  stories are verbatim Gherkin — do not edit, reword, or renumber them.
145
146
  - Never start a story in a `[BUILD-PROHIBITED]`, `[BUILD-PROHIBITED in part]`,