@aws-cdk/aws-glue-alpha 2.268.0-alpha.0 → 2.270.0-alpha.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.jsii +1820 -1612
- package/.jsii.tabl.json.gz +0 -0
- package/.warnings.jsii.js +1649 -10
- package/README.md +105 -35
- package/adr/job-arguments.md +212 -0
- package/lib/catalog.js +88 -4
- package/lib/code.js +59 -4
- package/lib/connection.d.ts +44 -29
- package/lib/connection.js +147 -31
- package/lib/constants.d.ts +17 -0
- package/lib/constants.js +20 -2
- package/lib/data-format.js +43 -6
- package/lib/data-quality-ruleset.js +60 -4
- package/lib/database.js +48 -2
- package/lib/external-table.js +48 -2
- package/lib/jobs/job.d.ts +90 -9
- package/lib/jobs/job.js +219 -38
- package/lib/jobs/pyspark-etl-job.d.ts +1 -3
- package/lib/jobs/pyspark-etl-job.js +28 -14
- package/lib/jobs/pyspark-flex-etl-job.d.ts +1 -3
- package/lib/jobs/pyspark-flex-etl-job.js +28 -14
- package/lib/jobs/pyspark-streaming-job.d.ts +1 -3
- package/lib/jobs/pyspark-streaming-job.js +28 -14
- package/lib/jobs/python-shell-job.d.ts +20 -7
- package/lib/jobs/python-shell-job.js +49 -33
- package/lib/jobs/ray-job.js +8 -13
- package/lib/jobs/scala-spark-etl-job.d.ts +1 -3
- package/lib/jobs/scala-spark-etl-job.js +29 -14
- package/lib/jobs/scala-spark-flex-etl-job.d.ts +1 -3
- package/lib/jobs/scala-spark-flex-etl-job.js +29 -14
- package/lib/jobs/scala-spark-streaming-job.d.ts +1 -3
- package/lib/jobs/scala-spark-streaming-job.js +29 -14
- package/lib/jobs/spark-job.d.ts +9 -7
- package/lib/jobs/spark-job.js +26 -32
- package/lib/partition-projection.d.ts +36 -47
- package/lib/partition-projection.js +34 -40
- package/lib/s3-table.js +106 -5
- package/lib/schema.js +50 -3
- package/lib/security-configuration.js +60 -5
- package/lib/storage-parameter.js +91 -2
- package/lib/table-base.js +32 -2
- package/lib/triggers/trigger-options.d.ts +102 -42
- package/lib/triggers/trigger-options.js +242 -3
- package/lib/triggers/workflow.d.ts +31 -47
- package/lib/triggers/workflow.js +87 -142
- package/package.json +10 -10
package/README.md
CHANGED
|
@@ -3,18 +3,16 @@
|
|
|
3
3
|
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-

|
|
7
7
|
|
|
8
|
-
>
|
|
9
|
-
> They are subject to non-backward compatible changes or removal in any future version. These are
|
|
10
|
-
> not subject to the [Semantic Versioning](https://semver.org/) model and breaking changes will be
|
|
11
|
-
> announced in the release notes. This means that while you may use them, you may need to update
|
|
12
|
-
> your source code when upgrading to a newer version of this package.
|
|
8
|
+
> This API may emit warnings. Backward compatibility is not guaranteed.
|
|
13
9
|
|
|
14
10
|
---
|
|
15
11
|
|
|
16
12
|
<!--END STABILITY BANNER-->
|
|
17
13
|
|
|
14
|
+
All constructs moved to aws-cdk-lib/aws-glue.
|
|
15
|
+
|
|
18
16
|
This module is part of the [AWS Cloud Development Kit](https://github.com/aws/aws-cdk) project.
|
|
19
17
|
|
|
20
18
|
## README
|
|
@@ -81,6 +79,17 @@ The Spark UI (`—enable-spark-ui`) is off by default; enable it by setting the
|
|
|
81
79
|
You can find more details about version, worker type and other features in
|
|
82
80
|
[Glue's public documentation](https://docs.aws.amazon.com/glue/latest/dg/aws-glue-api-jobs-job.html).
|
|
83
81
|
|
|
82
|
+
> **Note on continuous logging and encryption:** Because continuous logging is
|
|
83
|
+
> enabled by default, job driver and executor stdout/stderr are streamed to
|
|
84
|
+
> CloudWatch. Unless you attach a [`SecurityConfiguration`](#securityconfiguration)
|
|
85
|
+
> with `cloudWatchEncryption`, these logs are written to the account-shared,
|
|
86
|
+
> default Glue log group (`/aws-glue/jobs/logs-v2/`), which is **not** encrypted
|
|
87
|
+
> with a customer-managed key. Since job logs can contain sensitive runtime data
|
|
88
|
+
> (SQL statements, row values, error stack traces), attach a `SecurityConfiguration`
|
|
89
|
+
> with `cloudWatchEncryption` for regulated workloads. The construct emits a
|
|
90
|
+
> synthesis-time warning when continuous logging is on and no `SecurityConfiguration`
|
|
91
|
+
> is attached.
|
|
92
|
+
|
|
84
93
|
Reference the pyspark-etl-jobs.test.ts and scalaspark-etl-jobs.test.ts unit tests
|
|
85
94
|
for examples of required-only and optional job parameters when creating these
|
|
86
95
|
types of jobs.
|
|
@@ -256,8 +265,9 @@ Python shell jobs support a Python version that depends on the AWS Glue
|
|
|
256
265
|
version you use. These can be used to schedule and run tasks that don't
|
|
257
266
|
require an Apache Spark environment. Python shell jobs default to
|
|
258
267
|
Python 3.9 and a MaxCapacity of `0.0625`. Python 3.9 supports pre-loaded
|
|
259
|
-
analytics libraries
|
|
260
|
-
|
|
268
|
+
analytics libraries, enabled by default (`librarySet: glue.LibrarySet.ANALYTICS`).
|
|
269
|
+
Set `librarySet: glue.LibrarySet.NONE` when your libraries are custom or
|
|
270
|
+
conflict with the pre-installed ones.
|
|
261
271
|
|
|
262
272
|
Reference the pyspark-shell-job.test.ts unit tests for examples of
|
|
263
273
|
required-only and optional job parameters when creating these types of jobs.
|
|
@@ -342,6 +352,51 @@ new glue.PySparkEtlJob(stack, 'SelectiveJob', {
|
|
|
342
352
|
|
|
343
353
|
This feature is available for all Spark job types (ETL, Streaming, Flex).
|
|
344
354
|
|
|
355
|
+
### Job Arguments
|
|
356
|
+
|
|
357
|
+
Glue jobs are configured through a map of name-value arguments (`DefaultArguments`). This construct
|
|
358
|
+
manages several of these arguments on your behalf and exposes each one through a dedicated,
|
|
359
|
+
strongly-typed prop:
|
|
360
|
+
|
|
361
|
+
| Managed argument(s) | Prop |
|
|
362
|
+
|--------------------------------------------------------------------------|-----------------------------------------------------------------|
|
|
363
|
+
| `--enable-continuous-cloudwatch-log`, `--continuous-log-*` | `continuousLogging` |
|
|
364
|
+
| `--enable-metrics` | `enableMetrics` |
|
|
365
|
+
| `--enable-observability-metrics` | `enableObservabilityMetrics` |
|
|
366
|
+
| `--enable-spark-ui`, `--spark-event-logs-path` | `sparkUI` |
|
|
367
|
+
| `--job-language`, `--class` | job class / `className` |
|
|
368
|
+
| `--extra-jars`, `--user-jars-first`, `--extra-py-files`, `--extra-files` | `extraJars`, `extraJarsFirst`, `extraPythonFiles`, `extraFiles` |
|
|
369
|
+
| `library-set` | `librarySet` (Python Shell) |
|
|
370
|
+
|
|
371
|
+
The `defaultArguments` prop is the escape hatch for arguments this construct does **not** model.
|
|
372
|
+
Use it for any argument without a dedicated prop:
|
|
373
|
+
|
|
374
|
+
```ts
|
|
375
|
+
import * as cdk from 'aws-cdk-lib';
|
|
376
|
+
import * as iam from 'aws-cdk-lib/aws-iam';
|
|
377
|
+
declare const stack: cdk.Stack;
|
|
378
|
+
declare const role: iam.IRole;
|
|
379
|
+
declare const script: glue.Code;
|
|
380
|
+
|
|
381
|
+
new glue.PySparkEtlJob(stack, 'PySparkETLJob', {
|
|
382
|
+
role,
|
|
383
|
+
script,
|
|
384
|
+
defaultArguments: {
|
|
385
|
+
// an argument this construct does not manage
|
|
386
|
+
'--enable-glue-datacatalog': 'true',
|
|
387
|
+
},
|
|
388
|
+
});
|
|
389
|
+
```
|
|
390
|
+
|
|
391
|
+
To keep a single, unambiguous way to express each intent, setting a **construct-managed** argument
|
|
392
|
+
(any argument in the table above) or a **Glue-reserved** argument (`--debug`, `--mode`,
|
|
393
|
+
`--JOB_NAME`, `--endpoint`) through `defaultArguments` throws at synthesis time. This holds even
|
|
394
|
+
when the feature is turned off — for example, `enableMetrics: false` combined with
|
|
395
|
+
`defaultArguments: { '--enable-metrics': '' }` throws rather than silently re-enabling metrics.
|
|
396
|
+
Configure managed arguments through their dedicated prop instead — for example, use
|
|
397
|
+
`continuousLogging: { enabled: false }` rather than
|
|
398
|
+
`defaultArguments: { '--enable-continuous-cloudwatch-log': 'false' }`.
|
|
399
|
+
|
|
345
400
|
### Enable Job Run Queuing
|
|
346
401
|
|
|
347
402
|
AWS Glue job queuing monitors your account level quotas and limits. If quotas or limits are insufficient to start a Glue job run, AWS Glue will automatically queue the job and wait for limits to free up. Once limits become available, AWS Glue will retry the job run. Glue jobs will queue for limits like max concurrent job runs per account, max concurrent Data Processing Units (DPU), and resource unavailable due to IP address exhaustion in Amazon Virtual Private Cloud (Amazon VPC).
|
|
@@ -405,7 +460,7 @@ const job = new glue.PySparkEtlJob(stack, 'Job', { role, script });
|
|
|
405
460
|
// Create a workflow and add a trigger that runs the job
|
|
406
461
|
const workflow = new glue.Workflow(stack, 'Workflow');
|
|
407
462
|
workflow.addOnDemandTrigger('OnDemandTrigger', {
|
|
408
|
-
actions: [
|
|
463
|
+
actions: [glue.Action.job(job)],
|
|
409
464
|
});
|
|
410
465
|
```
|
|
411
466
|
|
|
@@ -418,21 +473,35 @@ actions list using the job or crawler objects using conditional types.
|
|
|
418
473
|
|
|
419
474
|
#### **2. Scheduled Triggers**
|
|
420
475
|
|
|
421
|
-
|
|
422
|
-
|
|
423
|
-
|
|
424
|
-
|
|
425
|
-
without
|
|
426
|
-
|
|
427
|
-
takes an optional description and a list of jobs or crawlers as actions.
|
|
476
|
+
Use `addScheduledTrigger` with a `TriggerSchedule` to fire on a cron schedule.
|
|
477
|
+
`TriggerSchedule.daily()` and `TriggerSchedule.weekly()` are convenience
|
|
478
|
+
factories; `TriggerSchedule.cron(...)` lets you build any schedule from the
|
|
479
|
+
[existing event Schedule class](https://docs.aws.amazon.com/cdk/api/v2/docs/aws-cdk-lib.aws_events.Schedule.html)
|
|
480
|
+
without writing raw cron expressions. The L2 extracts the expression that Glue
|
|
481
|
+
requires from the `TriggerSchedule`.
|
|
428
482
|
|
|
429
|
-
|
|
483
|
+
```ts
|
|
484
|
+
import * as cdk from 'aws-cdk-lib';
|
|
485
|
+
import * as iam from 'aws-cdk-lib/aws-iam';
|
|
486
|
+
declare const stack: cdk.Stack;
|
|
487
|
+
declare const role: iam.IRole;
|
|
488
|
+
declare const script: glue.Code;
|
|
489
|
+
const job = new glue.PySparkEtlJob(stack, 'Job', { role, script });
|
|
490
|
+
const workflow = new glue.Workflow(stack, 'Workflow');
|
|
491
|
+
|
|
492
|
+
workflow.addScheduledTrigger('WeeklyTrigger', {
|
|
493
|
+
actions: [glue.Action.job(job)],
|
|
494
|
+
schedule: glue.TriggerSchedule.weekly(),
|
|
495
|
+
});
|
|
496
|
+
```
|
|
430
497
|
|
|
431
|
-
|
|
432
|
-
|
|
433
|
-
|
|
434
|
-
|
|
435
|
-
|
|
498
|
+
#### **3. Event Triggers**
|
|
499
|
+
|
|
500
|
+
Use `addEventTrigger` for EventBridge event-based triggers. There are two types:
|
|
501
|
+
batching and non-batching. For batching triggers, you must specify `batchSize`.
|
|
502
|
+
For non-batching triggers, `batchSize` defaults to 1. For both, `batchWindow`
|
|
503
|
+
defaults to 900 seconds, but you can override the window to align with your
|
|
504
|
+
workload's requirements.
|
|
436
505
|
|
|
437
506
|
#### **4. Conditional Triggers**
|
|
438
507
|
|
|
@@ -451,13 +520,14 @@ certain types of data stores.
|
|
|
451
520
|
|
|
452
521
|
* **Networking - the CDK determines the best fit subnet for Glue connection
|
|
453
522
|
configuration**
|
|
454
|
-
|
|
455
|
-
|
|
456
|
-
`vpcSubnets`
|
|
523
|
+
Configure VPC placement through the `network` property, built with
|
|
524
|
+
`ConnectionNetwork.subnet(subnet)` to pin a specific subnet, or
|
|
525
|
+
`ConnectionNetwork.vpc(vpc, vpcSubnets?)` to let the L2 select one via the
|
|
526
|
+
existing
|
|
457
527
|
[EC2 Subnet Selection](https://docs.aws.amazon.com/cdk/api/v2/python/aws_cdk.aws_ec2/SubnetSelection.html)
|
|
458
|
-
library
|
|
459
|
-
|
|
460
|
-
|
|
528
|
+
library. A Glue connection targets a single subnet, so the first subnet of
|
|
529
|
+
the selection is used. The two factories are mutually exclusive, so a subnet
|
|
530
|
+
and a VPC can never be combined.
|
|
461
531
|
|
|
462
532
|
Pin the connection to a specific subnet:
|
|
463
533
|
|
|
@@ -469,7 +539,7 @@ new glue.Connection(this, 'MyConnection', {
|
|
|
469
539
|
// The security groups granting AWS Glue inbound access to the data source within the VPC
|
|
470
540
|
securityGroups: [securityGroup],
|
|
471
541
|
// The VPC subnet which contains the data source
|
|
472
|
-
subnet,
|
|
542
|
+
network: glue.ConnectionNetwork.subnet(subnet),
|
|
473
543
|
});
|
|
474
544
|
```
|
|
475
545
|
|
|
@@ -481,9 +551,8 @@ declare const vpc: ec2.Vpc;
|
|
|
481
551
|
new glue.Connection(this, 'MyConnection', {
|
|
482
552
|
type: glue.ConnectionType.NETWORK,
|
|
483
553
|
securityGroups: [securityGroup],
|
|
484
|
-
|
|
485
|
-
|
|
486
|
-
vpcSubnets: { subnetType: ec2.SubnetType.PRIVATE_WITH_EGRESS },
|
|
554
|
+
// vpcSubnets is optional - defaults to private subnets
|
|
555
|
+
network: glue.ConnectionNetwork.vpc(vpc, { subnetType: ec2.SubnetType.PRIVATE_WITH_EGRESS }),
|
|
487
556
|
});
|
|
488
557
|
```
|
|
489
558
|
|
|
@@ -496,7 +565,7 @@ declare const db: rds.DatabaseCluster;
|
|
|
496
565
|
new glue.Connection(this, "RdsConnection", {
|
|
497
566
|
type: glue.ConnectionType.JDBC,
|
|
498
567
|
securityGroups: [securityGroup],
|
|
499
|
-
subnet,
|
|
568
|
+
network: glue.ConnectionNetwork.subnet(subnet),
|
|
500
569
|
secret: db.secret,
|
|
501
570
|
properties: {
|
|
502
571
|
JDBC_CONNECTION_URL: `jdbc:mysql://${db.clusterEndpoint.socketAddress}/databasename`,
|
|
@@ -945,8 +1014,9 @@ new glue.S3Table(this, 'MyTable', {
|
|
|
945
1014
|
min: '2020-01-01',
|
|
946
1015
|
max: '2023-12-31',
|
|
947
1016
|
format: 'yyyy-MM-dd',
|
|
948
|
-
interval
|
|
949
|
-
|
|
1017
|
+
// `step` bundles interval + unit (supply both or neither). Optional at day
|
|
1018
|
+
// precision or coarser; required when the format is sub-day (e.g. hours).
|
|
1019
|
+
step: { interval: 1, intervalUnit: glue.DateIntervalUnit.DAYS },
|
|
950
1020
|
}),
|
|
951
1021
|
},
|
|
952
1022
|
});
|
|
@@ -0,0 +1,212 @@
|
|
|
1
|
+
# Glue Job Arguments
|
|
2
|
+
|
|
3
|
+
## Status
|
|
4
|
+
|
|
5
|
+
accepted
|
|
6
|
+
|
|
7
|
+
## Context
|
|
8
|
+
|
|
9
|
+
Every Glue job resource (`AWS::Glue::Job`) accepts a `DefaultArguments` map — a
|
|
10
|
+
flat `string → string` dictionary of `--flag`/value pairs that Glue passes to the
|
|
11
|
+
job script on every run. Some of these arguments are ordinary user configuration
|
|
12
|
+
(`--additional-python-modules`, `--enable-glue-datacatalog`, `--TempDir`, …), but
|
|
13
|
+
others are the wire form of features the CDK L2 models with strongly-typed props.
|
|
14
|
+
|
|
15
|
+
The L2 job constructs therefore populate `DefaultArguments` from two sources:
|
|
16
|
+
|
|
17
|
+
1. **Construct-managed arguments** — derived by the construct from typed props or
|
|
18
|
+
from the job class itself. Examples:
|
|
19
|
+
- `continuousLogging` → `--enable-continuous-cloudwatch-log`,
|
|
20
|
+
`--continuous-log-logGroup`, `--continuous-log-logStreamPrefix`,
|
|
21
|
+
`--continuous-log-conversionPattern`, `--enable-continuous-log-filter`
|
|
22
|
+
- `enableMetrics` → `--enable-metrics`
|
|
23
|
+
- `enableObservabilityMetrics` → `--enable-observability-metrics`
|
|
24
|
+
- `sparkUI` → `--enable-spark-ui`, `--spark-event-logs-path`
|
|
25
|
+
- `extraJars` / `extraJarsFirst` / `extraPythonFiles` / `extraFiles` →
|
|
26
|
+
`--extra-jars`, `--user-jars-first`, `--extra-py-files`, `--extra-files`
|
|
27
|
+
- `className` → `--class` (Scala only)
|
|
28
|
+
- the job language itself → `--job-language`
|
|
29
|
+
- `librarySet` → `library-set` (Python Shell only)
|
|
30
|
+
2. **`defaultArguments`** — the untyped escape-hatch map the user supplies directly,
|
|
31
|
+
for arguments the L2 does *not* model.
|
|
32
|
+
|
|
33
|
+
The two sources can collide. Before this decision, the collision was resolved
|
|
34
|
+
silently and inconsistently across job types:
|
|
35
|
+
|
|
36
|
+
- `SparkJob` / `PythonShellJob` merged as `{ ...managed, ...userDefaultArguments }`,
|
|
37
|
+
so the user value won — a user could pass
|
|
38
|
+
`defaultArguments: { '--enable-continuous-cloudwatch-log': 'false' }` and silently
|
|
39
|
+
turn off a secure default.
|
|
40
|
+
- `RayJob` merged the other way, so the construct value won — a user's
|
|
41
|
+
`defaultArguments` entry for a managed key was silently dropped.
|
|
42
|
+
|
|
43
|
+
Both behaviors are footguns: one weakens the construct's secure/observable defaults
|
|
44
|
+
without warning, the other ignores explicit user input without warning. Because Glue
|
|
45
|
+
enables continuous CloudWatch logging by default and that data can contain sensitive
|
|
46
|
+
runtime values (SQL, row data, stack traces), the "user silently wins" case is also a
|
|
47
|
+
security concern.
|
|
48
|
+
|
|
49
|
+
Separately, Glue itself reserves a handful of argument keys for its own internal use
|
|
50
|
+
(`--debug`, `--mode`, `--JOB_NAME`, `--endpoint`). These are never valid user input on
|
|
51
|
+
any job type.
|
|
52
|
+
|
|
53
|
+
## Constraints
|
|
54
|
+
|
|
55
|
+
- `DefaultArguments` is a single flat map on the L1; there is no separate channel to
|
|
56
|
+
distinguish "managed" from "user" keys once they are merged. Whatever the L2 does,
|
|
57
|
+
it must produce one merged map.
|
|
58
|
+
- Which keys are managed varies by job type: `--class` exists only for Scala jobs,
|
|
59
|
+
`library-set` only for Python Shell, `--enable-spark-ui` only for Spark, and so on.
|
|
60
|
+
A single global list would either over-reject (block a key that is a legitimate
|
|
61
|
+
escape hatch for a job type that doesn't manage it — e.g. `--extra-py-files` on a
|
|
62
|
+
job with no `extraPythonFiles` prop) or under-reject.
|
|
63
|
+
- Argument keys can be tokens (e.g. produced by `CfnJson`) that only resolve at
|
|
64
|
+
deploy time. String comparison cannot see through them at synthesis.
|
|
65
|
+
- The set of managed keys must not be tied to the set the construct *happens to emit*
|
|
66
|
+
for a given configuration: a prop that turns a feature off (`enableMetrics:
|
|
67
|
+
false`) emits nothing, but the key is still construct-managed and must stay reserved.
|
|
68
|
+
|
|
69
|
+
## Decision
|
|
70
|
+
|
|
71
|
+
**A managed argument has exactly one way to be configured: its typed prop.** Passing a
|
|
72
|
+
construct-managed or Glue-reserved key through `defaultArguments` throws a
|
|
73
|
+
`ValidationError` at synthesis time rather than silently winning or being dropped.
|
|
74
|
+
`defaultArguments` remains the escape hatch for every argument the L2 does not model.
|
|
75
|
+
|
|
76
|
+
### Data flow
|
|
77
|
+
|
|
78
|
+
There is a single sink for every construct-managed argument — the base-class method:
|
|
79
|
+
|
|
80
|
+
```ts
|
|
81
|
+
protected setManagedArgument(key: string, value?: string): void
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
It records `key` in the reserved set and, when `value !== undefined`, emits it. A
|
|
85
|
+
subclass calls it once per managed key, passing `undefined` when the feature is off or
|
|
86
|
+
unset — the key is reserved either way. Subclasses do not build local argument maps, so
|
|
87
|
+
this is the *only* way to emit a managed argument: declaration and emission happen in the
|
|
88
|
+
same call, and the reserved set therefore cannot drift from what is emitted.
|
|
89
|
+
|
|
90
|
+
Each job subclass, in its constructor:
|
|
91
|
+
|
|
92
|
+
1. Registers its managed arguments through `setManagedArgument` — directly, or through
|
|
93
|
+
the shared helpers `setupContinuousLogging` (all job types),
|
|
94
|
+
`nonExecutableCommonArguments` and `setupExtraCodeArguments` (Spark), and its own
|
|
95
|
+
`executableArguments` (`--job-language`; plus `--class` for Scala, `library-set` for
|
|
96
|
+
Python Shell).
|
|
97
|
+
2. Calls the base-class method:
|
|
98
|
+
|
|
99
|
+
```ts
|
|
100
|
+
protected mergeDefaultArguments(
|
|
101
|
+
defaultArguments?: { [key: string]: string },
|
|
102
|
+
): { [key: string]: string }
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
which validates the user-supplied `defaultArguments` against the accumulated reserved
|
|
106
|
+
set and returns the merged map, passed as `DefaultArguments` on the `CfnJob`.
|
|
107
|
+
|
|
108
|
+
`setManagedArgument` declares managed keys; `mergeDefaultArguments` validates and merges
|
|
109
|
+
them. Each is the sole choke point for its job.
|
|
110
|
+
|
|
111
|
+
Which keys a job type reserves falls out of which `setManagedArgument` calls its
|
|
112
|
+
constructor makes: only Scala jobs register `--class`, only Python Shell registers
|
|
113
|
+
`library-set`, only Spark registers `--enable-spark-ui`, and so on. Per-job-type scoping
|
|
114
|
+
is automatic — there is no separate list to maintain per type.
|
|
115
|
+
|
|
116
|
+
### The reserved set
|
|
117
|
+
|
|
118
|
+
The keys a user may not set through `defaultArguments` are the union of:
|
|
119
|
+
|
|
120
|
+
- **`GLUE_RESERVED_ARGUMENTS`** — `--debug`, `--mode`, `--JOB_NAME`, `--endpoint`. Owned
|
|
121
|
+
by the Glue service, reserved on every job type. This is the one static list, because
|
|
122
|
+
it is external to the constructs — nothing derives it from a prop.
|
|
123
|
+
- **`_managedArgumentKeys`** — every key that any `setManagedArgument` call registered on
|
|
124
|
+
this instance, whether or not a value was emitted for it.
|
|
125
|
+
|
|
126
|
+
### Validation rules (per user-supplied key)
|
|
127
|
+
|
|
128
|
+
For each key in `defaultArguments`:
|
|
129
|
+
|
|
130
|
+
1. If the key is an unresolved token, the conflict check is skipped (equality is
|
|
131
|
+
unknowable at synth time) and a warning
|
|
132
|
+
(`@aws-cdk/aws-glue-alpha:tokenJobArgumentKey`) is emitted. If it resolves to a
|
|
133
|
+
managed key at deploy time, the construct-managed value wins (see merge order).
|
|
134
|
+
2. If the key is in **`GLUE_RESERVED_ARGUMENTS`** → throw. Glue-reserved keys are
|
|
135
|
+
never emitted by the construct, so there is no construct value to reconcile against.
|
|
136
|
+
3. If the key is in the reserved set (`_managedArgumentKeys`):
|
|
137
|
+
- If the construct actually emitted a value for that key
|
|
138
|
+
(`Object.hasOwn(_managedArguments, key)`) and the supplied value is identical →
|
|
139
|
+
allowed. Passing the same value the construct would produce is not
|
|
140
|
+
contradictory; autocorrecting config is preferred over an error.
|
|
141
|
+
- Otherwise (different value, or the construct emitted nothing because the feature
|
|
142
|
+
is off) → throw.
|
|
143
|
+
4. Otherwise, the key is genuinely custom → allowed, flows through untouched.
|
|
144
|
+
|
|
145
|
+
`Object.hasOwn` is used deliberately instead of the `in` operator so that inherited
|
|
146
|
+
`Object.prototype` members (`toString`, `constructor`, `hasOwnProperty`, …) supplied as
|
|
147
|
+
argument keys are treated as ordinary custom keys rather than falsely matching a
|
|
148
|
+
managed key.
|
|
149
|
+
|
|
150
|
+
### Merge order
|
|
151
|
+
|
|
152
|
+
After validation, the result is:
|
|
153
|
+
|
|
154
|
+
```ts
|
|
155
|
+
return { ...defaultArguments, ...this._managedArguments };
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
Managed arguments are spread last, so they win on any residual overlap. By this point
|
|
159
|
+
the only overlaps that can remain are (a) exact-value matches allowed by rule 3, which
|
|
160
|
+
are indistinguishable either way, and (b) token keys from rule 1, for which
|
|
161
|
+
managed-wins is the documented and warned-about behavior.
|
|
162
|
+
|
|
163
|
+
### Related synthesis-time warnings
|
|
164
|
+
|
|
165
|
+
Two other warnings live in the same flow:
|
|
166
|
+
|
|
167
|
+
- **`@aws-cdk/aws-glue-alpha:unencryptedContinuousLogging`** — continuous logging is on
|
|
168
|
+
(explicitly or by default) but no `SecurityConfiguration` is attached, so driver /
|
|
169
|
+
executor logs land in an unencrypted, account-shared CloudWatch log group. We only
|
|
170
|
+
warn when *no* security configuration is attached at all, because
|
|
171
|
+
`ISecurityConfiguration` exposes only the name and we cannot introspect whether it
|
|
172
|
+
actually configures `cloudWatchEncryption` (avoiding false positives).
|
|
173
|
+
- **`@aws-cdk/aws-glue-alpha:plaintextJobArgumentSecret`** — a `defaultArguments` key
|
|
174
|
+
looks like a credential and holds a plaintext literal. `DefaultArguments` is emitted
|
|
175
|
+
verbatim into the template; secrets belong in AWS Secrets Manager.
|
|
176
|
+
|
|
177
|
+
## Alternatives
|
|
178
|
+
|
|
179
|
+
### Invert precedence so the construct always wins
|
|
180
|
+
|
|
181
|
+
Merge as `{ ...userDefaultArguments, ...managed }` everywhere (which is what `RayJob`
|
|
182
|
+
already did). This is more secure than "user wins" but still silent: a user who
|
|
183
|
+
deliberately sets a managed key via `defaultArguments` has it dropped with no
|
|
184
|
+
indication. It also still leaves two channels for one setting. Rejected in favor of a
|
|
185
|
+
single, explicit way to express each intent.
|
|
186
|
+
|
|
187
|
+
### Derive the reserved set from the emitted arguments
|
|
188
|
+
|
|
189
|
+
Compute conflicts from the keys the construct actually emits. Attractive: adding a typed
|
|
190
|
+
prop reserves its key automatically, with no separate list to maintain. Rejected: it ties
|
|
191
|
+
the reserved set to the current configuration. A prop that turns a feature off emits no
|
|
192
|
+
key, so `defaultArguments` could silently re-enable it (`enableMetrics: false` +
|
|
193
|
+
`defaultArguments: { '--enable-metrics': '' }`). The reserved set must include keys the
|
|
194
|
+
construct manages even when it emits nothing for them. That is why `setManagedArgument`
|
|
195
|
+
records the key regardless of value.
|
|
196
|
+
|
|
197
|
+
### A per-job-type list of managed keys
|
|
198
|
+
|
|
199
|
+
Give each job type an explicit `string[]` of the keys it manages, passed to the
|
|
200
|
+
validation step alongside the emitted map. Correct, but it names every managed key in
|
|
201
|
+
two places — the emission site (inside an `enabled ? {...} : {}` expression) and the
|
|
202
|
+
list — which drift: adding a typed prop requires updating the list too, and a forgotten
|
|
203
|
+
entry silently reopens the re-enable bypass with no compile-time signal. The
|
|
204
|
+
`setManagedArgument` accumulator collapses the two into one call, so the key is named
|
|
205
|
+
once.
|
|
206
|
+
|
|
207
|
+
### One global reserved list on the base class
|
|
208
|
+
|
|
209
|
+
A single list of every managed key across all job types. Rejected because it
|
|
210
|
+
over-rejects: it would block, for example, `--extra-py-files` on a job type that has no
|
|
211
|
+
`extraPythonFiles` prop, removing a legitimate escape hatch with no typed replacement.
|
|
212
|
+
Managed keys must be scoped per job type.
|