singer-tap-kit 0.1.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,229 @@
1
+ Metadata-Version: 2.3
2
+ Name: singer-tap-kit
3
+ Version: 0.1.0
4
+ Summary: Singer.io tap for extracting data from the Kit API
5
+ Author: Andrew Jones
6
+ Author-email: Andrew Jones <andrew@andrew-jones.com>
7
+ License: MIT License
8
+
9
+ Copyright (c) 2025 Andrew Jones
10
+
11
+ Permission is hereby granted, free of charge, to any person obtaining a copy
12
+ of this software and associated documentation files (the "Software"), to deal
13
+ in the Software without restriction, including without limitation the rights
14
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
15
+ copies of the Software, and to permit persons to whom the Software is
16
+ furnished to do so, subject to the following conditions:
17
+
18
+ The above copyright notice and this permission notice shall be included in all
19
+ copies or substantial portions of the Software.
20
+
21
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
22
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
23
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
24
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
25
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
26
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
27
+ SOFTWARE.
28
+ Classifier: License :: OSI Approved :: Apache Software License
29
+ Classifier: Programming Language :: Python :: 3
30
+ Classifier: Programming Language :: Python :: 3.9
31
+ Requires-Dist: singer-python>=5.13.0
32
+ Requires-Dist: requests>=2.31.0
33
+ Requires-Dist: python-dateutil>=2.8.2
34
+ Requires-Dist: pytest>=7.0.0 ; extra == 'dev'
35
+ Requires-Dist: pytest-mock>=3.10.0 ; extra == 'dev'
36
+ Requires-Python: >=3.9
37
+ Project-URL: Homepage, https://github.com/andrewjones/tap-kit
38
+ Provides-Extra: dev
39
+ Description-Content-Type: text/markdown
40
+
41
+ # tap-kit
42
+
43
+ A [Singer](https://singer.io) tap for extracting data from the [Kit API](https://developers.kit.com/).
44
+
45
+ ## Installation
46
+
47
+ ```bash
48
+ pip install singer-tap-kit
49
+ ```
50
+
51
+ Or run directly with `uvx`.
52
+
53
+ ```bash
54
+ uvx --from singer-tap-kit tap-kit --help
55
+
56
+ # Example with CSV target
57
+ uvx --from singer-tap-kit tap-kit --config config.json | uvx --with setuptools target-csv
58
+ ```
59
+
60
+ ## Configuration
61
+
62
+ Create a `config.json` file with your Kit v4 API key:
63
+
64
+ ```json
65
+ {
66
+ "api_key": "your_kit_v4_api_key_here",
67
+ "start_date": "2024-01-01T00:00:00Z"
68
+ }
69
+ ```
70
+
71
+ - `api_key` (required): Your Kit v4 API key, sent as the `X-Kit-Api-Key` header.
72
+ - `start_date` (optional): ISO 8601 timestamp used as the initial incremental bookmark.
73
+
74
+ You can find your API key in your [Kit API settings](https://app.kit.com/account_settings/developer_settings).
75
+
76
+ ## Usage
77
+
78
+ ### Discover available streams
79
+
80
+ ```bash
81
+ tap-kit --config config.json --discover
82
+ ```
83
+
84
+ This will output a catalog of available streams in JSON format.
85
+
86
+ ### Run the tap
87
+
88
+ ```bash
89
+ tap-kit --config config.json --catalog catalog.json
90
+ ```
91
+
92
+ Or to use the default catalog:
93
+
94
+ ```bash
95
+ tap-kit --config config.json
96
+ ```
97
+
98
+ ### State management
99
+
100
+ To maintain state between runs (for incremental syncs):
101
+
102
+ ```bash
103
+ tap-kit --config config.json --state state.json
104
+ ```
105
+
106
+ ## Streams
107
+
108
+ Stats are modelled as their own snapshot streams (`broadcast_stats`,
109
+ `subscriber_stats`) rather than merged into the entity records. Engagement metrics
110
+ change over time, so each sync stamps every stat row with a `synced_at` timestamp
111
+ and uses a composite primary key of `[id, synced_at]`. This captures a point-in-time
112
+ snapshot per run: on append-style targets every run adds rows; on merge/upsert
113
+ targets (e.g. BigQuery MERGE on key properties) the new `synced_at` inserts rather
114
+ than overwriting. Partition the snapshot tables on `synced_at` in your warehouse.
115
+
116
+ ### Broadcasts
117
+
118
+ The `broadcasts` stream extracts broadcast entity data from your Kit account.
119
+
120
+ **Schema:**
121
+ - `id` (integer): Unique broadcast ID
122
+ - `publication_id` (integer, nullable): Publication ID
123
+ - `created_at` (datetime): When the broadcast was created
124
+ - `subject` (string, nullable): Broadcast subject line
125
+ - `preview_text` (string, nullable): Inbox preview text
126
+ - `description` (string, nullable): Broadcast description
127
+ - `content` (string, nullable): Broadcast content
128
+ - `public` (boolean, nullable): Whether the broadcast is public
129
+ - `published_at` (datetime, nullable): When the broadcast was published
130
+ - `send_at` (datetime, nullable): When the broadcast was/will be sent
131
+ - `thumbnail_alt` (string, nullable): Thumbnail alt text
132
+ - `thumbnail_url` (string, nullable): Thumbnail URL
133
+ - `public_url` (string, nullable): Public URL for the broadcast
134
+ - `email_address` (string, nullable): Sender email address
135
+ - `email_template` (object, nullable): Email template (id, name)
136
+ - `subscriber_filter` (array/object, nullable): Subscriber targeting filter
137
+ - `status` (string, nullable): Broadcast status
138
+
139
+ **Replication Method:** INCREMENTAL
140
+ **Replication Key:** `created_at`
141
+ **Primary Key:** `id`
142
+
143
+ ### Broadcast Stats
144
+
145
+ The `broadcast_stats` stream snapshots delivery/engagement stats for every
146
+ broadcast, fetched in bulk from the `/broadcasts/stats` batch endpoint.
147
+
148
+ **Schema:**
149
+ - `id` (integer): Broadcast ID
150
+ - `synced_at` (datetime): Snapshot timestamp for this sync run
151
+ - `recipients` (integer, nullable): Number of recipients
152
+ - `open_rate` (number, nullable): Email open rate
153
+ - `emails_opened` (integer, nullable): Number of emails opened
154
+ - `click_rate` (number, nullable): Email click rate
155
+ - `unsubscribe_rate` (number, nullable): Unsubscribe rate
156
+ - `unsubscribes` (integer, nullable): Number of unsubscribes
157
+ - `total_clicks` (integer, nullable): Total number of clicks
158
+ - `show_total_clicks` (boolean, nullable): Whether to show total clicks
159
+ - `status` (string, nullable): Broadcast status
160
+ - `progress` (number, nullable): Send progress percentage
161
+ - `open_tracking_disabled` (boolean, nullable): Whether open tracking is off
162
+ - `click_tracking_disabled` (boolean, nullable): Whether click tracking is off
163
+
164
+ **Replication Method:** FULL_TABLE
165
+ **Primary Key:** `id`, `synced_at`
166
+
167
+ ### Subscribers
168
+
169
+ The `subscribers` stream extracts subscriber data from your Kit account.
170
+
171
+ **Schema:**
172
+ - `id` (integer): Unique subscriber ID
173
+ - `first_name` (string, nullable): Subscriber first name
174
+ - `email_address` (string): Subscriber email address
175
+ - `state` (string, nullable): Subscriber state (e.g., "active", "inactive")
176
+ - `created_at` (datetime): When the subscriber was created
177
+ - `fields` (object, nullable): Custom fields
178
+
179
+ **Replication Method:** INCREMENTAL
180
+ **Replication Key:** `created_at`
181
+
182
+ ### Subscriber Stats
183
+
184
+ The `subscriber_stats` stream snapshots per-subscriber engagement stats.
185
+
186
+ **Schema:**
187
+ - `id` (integer): Subscriber ID
188
+ - `synced_at` (datetime): Snapshot timestamp for this sync run
189
+ - `sent` (integer, nullable): Emails sent
190
+ - `opened` (integer, nullable): Emails opened
191
+ - `clicked` (integer, nullable): Emails clicked
192
+ - `bounced` (integer, nullable): Emails bounced
193
+ - `open_rate` (number, nullable): Open rate
194
+ - `click_rate` (number, nullable): Click rate
195
+ - `last_sent` (datetime, nullable): Last send timestamp
196
+ - `last_opened` (datetime, nullable): Last open timestamp
197
+ - `last_clicked` (datetime, nullable): Last click timestamp
198
+ - `sends_since_last_open` (integer, nullable): Sends since last open
199
+ - `sends_since_last_click` (integer, nullable): Sends since last click
200
+
201
+ **Replication Method:** FULL_TABLE
202
+ **Primary Key:** `id`, `synced_at`
203
+
204
+ **Note:** This stream makes one API call per subscriber. Kit rate-limits to 120
205
+ requests per rolling 60 seconds, so large subscriber lists will sync slowly; the
206
+ tap automatically retries on rate-limit (HTTP 429) responses.
207
+
208
+ ## Development
209
+
210
+ This project is managed with `uv`.
211
+
212
+ ### Setup
213
+
214
+ 1. Clone the repository
215
+ 1. Create a `config.json` file with your API key
216
+
217
+ ### Testing
218
+
219
+ ```bash
220
+ # Discover streams
221
+ uv run tap-kit --config config.json --discover
222
+
223
+ # Run sync
224
+ uv run tap-kit --config config.json
225
+ ```
226
+
227
+ ## License
228
+
229
+ This project is licensed under the MIT License.
@@ -0,0 +1,5 @@
1
+ tap_kit/__init__.py,sha256=8c4e763135b88dbdc5398a83fe6f1904fdf98af3d6080e5c0b2a01dd53cf0225,14616
2
+ singer_tap_kit-0.1.0.dist-info/WHEEL,sha256=b6dc288e80aa2d1b1518ddb3502fd5b53e8fd6cb507ed2a4f932e9e6088b264a,78
3
+ singer_tap_kit-0.1.0.dist-info/entry_points.txt,sha256=de8aef1cb693b14f086a3001936773ebc77238d5bfa4ed061b3443f37e25fc27,42
4
+ singer_tap_kit-0.1.0.dist-info/METADATA,sha256=19ef178f71bd5770f74906714aa1aa23a8721146912e22484031a845cbc52a00,8027
5
+ singer_tap_kit-0.1.0.dist-info/RECORD,,
@@ -0,0 +1,4 @@
1
+ Wheel-Version: 1.0
2
+ Generator: uv 0.8.4
3
+ Root-Is-Purelib: true
4
+ Tag: py3-none-any
@@ -0,0 +1,3 @@
1
+ [console_scripts]
2
+ tap-kit = tap_kit:main
3
+
tap_kit/__init__.py ADDED
@@ -0,0 +1,446 @@
1
+ import json
2
+ import sys
3
+ import time
4
+ from datetime import datetime, timezone
5
+ from typing import Dict, Any, List, Optional
6
+ import requests
7
+ from singer import get_bookmark, write_record, write_schema, get_logger, parse_args
8
+ from singer.catalog import Catalog, CatalogEntry, Schema
9
+
10
+
11
+ REQUIRED_CONFIG_KEYS = ["api_key"]
12
+ CONFIG = {}
13
+ STATE = {}
14
+ LOGGER = get_logger()
15
+
16
+ DEFAULT_PER_PAGE = 500
17
+ MAX_RATE_LIMIT_RETRIES = 5
18
+
19
+ STREAM_CONFIG = {
20
+ "broadcasts": {
21
+ "replication_method": "INCREMENTAL",
22
+ "replication_key": "created_at",
23
+ "key_properties": ["id"],
24
+ },
25
+ "broadcast_stats": {
26
+ "replication_method": "FULL_TABLE",
27
+ "replication_key": None,
28
+ "key_properties": ["id", "synced_at"],
29
+ },
30
+ "subscribers": {
31
+ "replication_method": "INCREMENTAL",
32
+ "replication_key": "created_at",
33
+ "key_properties": ["id"],
34
+ },
35
+ "subscriber_stats": {
36
+ "replication_method": "FULL_TABLE",
37
+ "replication_key": None,
38
+ "key_properties": ["id", "synced_at"],
39
+ },
40
+ }
41
+
42
+
43
+ def utc_now() -> str:
44
+ """Current UTC time as an ISO 8601 string, used to stamp stat snapshots."""
45
+ return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
46
+
47
+
48
+ class KitAPI:
49
+ """Client for the Kit v4 API"""
50
+
51
+ def __init__(self, api_key: str):
52
+ self.api_key = api_key
53
+ self.base_url = "https://api.kit.com/v4"
54
+ self.headers = {
55
+ "X-Kit-Api-Key": api_key,
56
+ "Accept": "application/json",
57
+ }
58
+
59
+ def _get(
60
+ self, path: str, params: Optional[Dict[str, Any]] = None
61
+ ) -> Dict[str, Any]:
62
+ """GET a path, retrying on rate limit (HTTP 429)."""
63
+ url = f"{self.base_url}{path}"
64
+
65
+ for attempt in range(MAX_RATE_LIMIT_RETRIES + 1):
66
+ response = requests.get(url, headers=self.headers, params=params)
67
+
68
+ if response.status_code == 429 and attempt < MAX_RATE_LIMIT_RETRIES:
69
+ retry_after = response.headers.get("Retry-After")
70
+ wait = float(retry_after) if retry_after else 2**attempt
71
+ LOGGER.warning(
72
+ f"Rate limited on {path}, retrying in {wait}s "
73
+ f"(attempt {attempt + 1}/{MAX_RATE_LIMIT_RETRIES})"
74
+ )
75
+ time.sleep(wait)
76
+ continue
77
+
78
+ response.raise_for_status()
79
+ return response.json()
80
+
81
+ response.raise_for_status()
82
+ return response.json()
83
+
84
+ def list_broadcasts(
85
+ self, after: Optional[str] = None, per_page: int = DEFAULT_PER_PAGE
86
+ ) -> Dict[str, Any]:
87
+ """Fetch a page of broadcasts."""
88
+ params = {"per_page": per_page}
89
+ if after:
90
+ params["after"] = after
91
+ return self._get("/broadcasts", params)
92
+
93
+ def list_broadcast_stats(
94
+ self, after: Optional[str] = None, per_page: int = DEFAULT_PER_PAGE
95
+ ) -> Dict[str, Any]:
96
+ """Fetch a page of broadcast stats (batch endpoint)."""
97
+ params = {"per_page": per_page}
98
+ if after:
99
+ params["after"] = after
100
+ return self._get("/broadcasts/stats", params)
101
+
102
+ def list_subscribers(
103
+ self, after: Optional[str] = None, per_page: int = DEFAULT_PER_PAGE
104
+ ) -> Dict[str, Any]:
105
+ """Fetch a page of subscribers."""
106
+ params = {"per_page": per_page}
107
+ if after:
108
+ params["after"] = after
109
+ return self._get("/subscribers", params)
110
+
111
+ def get_subscriber_stats(self, subscriber_id: int) -> Dict[str, Any]:
112
+ """Fetch stats for a single subscriber."""
113
+ return self._get(f"/subscribers/{subscriber_id}/stats")
114
+
115
+
116
+ def load_schemas() -> Dict[str, Schema]:
117
+ """Define schemas for each stream."""
118
+ schemas = {}
119
+
120
+ broadcast_schema = {
121
+ "type": "object",
122
+ "properties": {
123
+ "id": {"type": "integer"},
124
+ "publication_id": {"type": ["null", "integer"]},
125
+ "created_at": {"type": "string", "format": "date-time"},
126
+ "subject": {"type": ["null", "string"]},
127
+ "preview_text": {"type": ["null", "string"]},
128
+ "description": {"type": ["null", "string"]},
129
+ "content": {"type": ["null", "string"]},
130
+ "public": {"type": ["null", "boolean"]},
131
+ "published_at": {"type": ["null", "string"], "format": "date-time"},
132
+ "send_at": {"type": ["null", "string"], "format": "date-time"},
133
+ "thumbnail_alt": {"type": ["null", "string"]},
134
+ "thumbnail_url": {"type": ["null", "string"]},
135
+ "public_url": {"type": ["null", "string"]},
136
+ "email_address": {"type": ["null", "string"]},
137
+ "email_template": {"type": ["null", "object"]},
138
+ "subscriber_filter": {"type": ["null", "array", "object"]},
139
+ "status": {"type": ["null", "string"]},
140
+ },
141
+ }
142
+
143
+ broadcast_stats_schema = {
144
+ "type": "object",
145
+ "properties": {
146
+ "id": {"type": "integer"},
147
+ "synced_at": {"type": "string", "format": "date-time"},
148
+ "recipients": {"type": ["null", "integer"]},
149
+ "open_rate": {"type": ["null", "number"]},
150
+ "emails_opened": {"type": ["null", "integer"]},
151
+ "click_rate": {"type": ["null", "number"]},
152
+ "unsubscribe_rate": {"type": ["null", "number"]},
153
+ "unsubscribes": {"type": ["null", "integer"]},
154
+ "total_clicks": {"type": ["null", "integer"]},
155
+ "show_total_clicks": {"type": ["null", "boolean"]},
156
+ "status": {"type": ["null", "string"]},
157
+ "progress": {"type": ["null", "number"]},
158
+ "open_tracking_disabled": {"type": ["null", "boolean"]},
159
+ "click_tracking_disabled": {"type": ["null", "boolean"]},
160
+ },
161
+ }
162
+
163
+ subscriber_schema = {
164
+ "type": "object",
165
+ "properties": {
166
+ "id": {"type": "integer"},
167
+ "first_name": {"type": ["null", "string"]},
168
+ "email_address": {"type": "string"},
169
+ "state": {"type": ["null", "string"]},
170
+ "created_at": {"type": "string", "format": "date-time"},
171
+ "fields": {"type": ["null", "object"]},
172
+ },
173
+ }
174
+
175
+ subscriber_stats_schema = {
176
+ "type": "object",
177
+ "properties": {
178
+ "id": {"type": "integer"},
179
+ "synced_at": {"type": "string", "format": "date-time"},
180
+ "sent": {"type": ["null", "integer"]},
181
+ "opened": {"type": ["null", "integer"]},
182
+ "clicked": {"type": ["null", "integer"]},
183
+ "bounced": {"type": ["null", "integer"]},
184
+ "open_rate": {"type": ["null", "number"]},
185
+ "click_rate": {"type": ["null", "number"]},
186
+ "last_sent": {"type": ["null", "string"], "format": "date-time"},
187
+ "last_opened": {"type": ["null", "string"], "format": "date-time"},
188
+ "last_clicked": {"type": ["null", "string"], "format": "date-time"},
189
+ "sends_since_last_open": {"type": ["null", "integer"]},
190
+ "sends_since_last_click": {"type": ["null", "integer"]},
191
+ },
192
+ }
193
+
194
+ schemas["broadcasts"] = Schema.from_dict(broadcast_schema)
195
+ schemas["broadcast_stats"] = Schema.from_dict(broadcast_stats_schema)
196
+ schemas["subscribers"] = Schema.from_dict(subscriber_schema)
197
+ schemas["subscriber_stats"] = Schema.from_dict(subscriber_stats_schema)
198
+
199
+ return schemas
200
+
201
+
202
+ def discover() -> Catalog:
203
+ """Discover available streams and schemas."""
204
+ schemas = load_schemas()
205
+ streams = []
206
+
207
+ for stream_id, schema in schemas.items():
208
+ stream_settings = STREAM_CONFIG[stream_id]
209
+ replication_method = stream_settings["replication_method"]
210
+ replication_key = stream_settings["replication_key"]
211
+ key_properties = stream_settings["key_properties"]
212
+
213
+ metadata_entry = {
214
+ "inclusion": "available",
215
+ "table-key-properties": key_properties,
216
+ "schema-name": stream_id,
217
+ }
218
+ if replication_key:
219
+ metadata_entry["valid-replication-keys"] = [replication_key]
220
+
221
+ stream_metadata = [{"metadata": metadata_entry, "breadcrumb": []}]
222
+
223
+ catalog_entry = CatalogEntry(
224
+ stream=stream_id,
225
+ tap_stream_id=stream_id,
226
+ schema=schema,
227
+ key_properties=key_properties,
228
+ replication_key=replication_key,
229
+ metadata=stream_metadata,
230
+ replication_method=replication_method,
231
+ )
232
+
233
+ streams.append(catalog_entry)
234
+
235
+ return Catalog(streams)
236
+
237
+
238
+ def sync_broadcasts(
239
+ config: Dict[str, Any], state: Dict[str, Any], catalog: Catalog
240
+ ) -> None:
241
+ """Sync broadcasts stream (entity records, no stats)."""
242
+ api = KitAPI(config["api_key"])
243
+
244
+ broadcasts_catalog = catalog.get_stream("broadcasts")
245
+ if not broadcasts_catalog:
246
+ LOGGER.error("Broadcasts stream not found in catalog")
247
+ return
248
+
249
+ get_bookmark(
250
+ state,
251
+ "broadcasts",
252
+ broadcasts_catalog.replication_key,
253
+ config.get("start_date"),
254
+ )
255
+
256
+ write_schema(
257
+ "broadcasts",
258
+ broadcasts_catalog.schema.to_dict(),
259
+ key_properties=broadcasts_catalog.key_properties,
260
+ )
261
+
262
+ after = None
263
+ total_records = 0
264
+
265
+ while True:
266
+ response = api.list_broadcasts(after=after)
267
+ broadcasts = response.get("broadcasts") or []
268
+
269
+ for broadcast in broadcasts:
270
+ write_record("broadcasts", broadcast)
271
+ total_records += 1
272
+
273
+ pagination = response.get("pagination") or {}
274
+ if not pagination.get("has_next_page"):
275
+ break
276
+ after = pagination.get("end_cursor")
277
+ if not after:
278
+ break
279
+
280
+ LOGGER.info(f"Synced {total_records} broadcast records")
281
+
282
+
283
+ def sync_broadcast_stats(
284
+ config: Dict[str, Any], catalog: Catalog, synced_at: str
285
+ ) -> None:
286
+ """Sync broadcast_stats snapshot stream via the batch stats endpoint."""
287
+ api = KitAPI(config["api_key"])
288
+
289
+ stats_catalog = catalog.get_stream("broadcast_stats")
290
+ if not stats_catalog:
291
+ LOGGER.error("Broadcast_stats stream not found in catalog")
292
+ return
293
+
294
+ write_schema(
295
+ "broadcast_stats",
296
+ stats_catalog.schema.to_dict(),
297
+ key_properties=stats_catalog.key_properties,
298
+ )
299
+
300
+ after = None
301
+ total_records = 0
302
+
303
+ while True:
304
+ response = api.list_broadcast_stats(after=after)
305
+ broadcasts = response.get("broadcasts") or []
306
+
307
+ for broadcast in broadcasts:
308
+ stats = broadcast.get("stats") or {}
309
+ record = {"id": broadcast["id"], "synced_at": synced_at, **stats}
310
+ write_record("broadcast_stats", record)
311
+ total_records += 1
312
+
313
+ pagination = response.get("pagination") or {}
314
+ if not pagination.get("has_next_page"):
315
+ break
316
+ after = pagination.get("end_cursor")
317
+ if not after:
318
+ break
319
+
320
+ LOGGER.info(f"Synced {total_records} broadcast_stats records")
321
+
322
+
323
+ def sync_subscribers(
324
+ config: Dict[str, Any], state: Dict[str, Any], catalog: Catalog
325
+ ) -> List[int]:
326
+ """Sync subscribers stream. Returns the list of subscriber ids seen."""
327
+ api = KitAPI(config["api_key"])
328
+
329
+ subscribers_catalog = catalog.get_stream("subscribers")
330
+ if not subscribers_catalog:
331
+ LOGGER.error("Subscribers stream not found in catalog")
332
+ return []
333
+
334
+ get_bookmark(
335
+ state,
336
+ "subscribers",
337
+ subscribers_catalog.replication_key,
338
+ config.get("start_date"),
339
+ )
340
+
341
+ write_schema(
342
+ "subscribers",
343
+ subscribers_catalog.schema.to_dict(),
344
+ key_properties=subscribers_catalog.key_properties,
345
+ )
346
+
347
+ after = None
348
+ total_records = 0
349
+ subscriber_ids: List[int] = []
350
+
351
+ while True:
352
+ response = api.list_subscribers(after=after)
353
+ subscribers = response.get("subscribers") or []
354
+
355
+ for subscriber in subscribers:
356
+ write_record("subscribers", subscriber)
357
+ subscriber_ids.append(subscriber["id"])
358
+ total_records += 1
359
+
360
+ pagination = response.get("pagination") or {}
361
+ if not pagination.get("has_next_page"):
362
+ break
363
+ after = pagination.get("end_cursor")
364
+ if not after:
365
+ break
366
+
367
+ LOGGER.info(f"Synced {total_records} subscriber records")
368
+ return subscriber_ids
369
+
370
+
371
+ def sync_subscriber_stats(
372
+ config: Dict[str, Any],
373
+ catalog: Catalog,
374
+ subscriber_ids: List[int],
375
+ synced_at: str,
376
+ ) -> None:
377
+ """Sync subscriber_stats snapshot stream (one API call per subscriber)."""
378
+ api = KitAPI(config["api_key"])
379
+
380
+ stats_catalog = catalog.get_stream("subscriber_stats")
381
+ if not stats_catalog:
382
+ LOGGER.error("Subscriber_stats stream not found in catalog")
383
+ return
384
+
385
+ write_schema(
386
+ "subscriber_stats",
387
+ stats_catalog.schema.to_dict(),
388
+ key_properties=stats_catalog.key_properties,
389
+ )
390
+
391
+ total_records = 0
392
+
393
+ for subscriber_id in subscriber_ids:
394
+ try:
395
+ response = api.get_subscriber_stats(subscriber_id)
396
+ except requests.exceptions.RequestException as e:
397
+ LOGGER.warning(f"Error fetching stats for subscriber {subscriber_id}: {e}")
398
+ continue
399
+
400
+ subscriber = response.get("subscriber") or {}
401
+ stats = subscriber.get("stats") or {}
402
+ record = {
403
+ "id": subscriber.get("id", subscriber_id),
404
+ "synced_at": synced_at,
405
+ **stats,
406
+ }
407
+ write_record("subscriber_stats", record)
408
+ total_records += 1
409
+
410
+ LOGGER.info(f"Synced {total_records} subscriber_stats records")
411
+
412
+
413
+ def sync(config: Dict[str, Any], state: Dict[str, Any], catalog: Catalog) -> None:
414
+ """Sync all selected streams."""
415
+ stream_ids = {stream.tap_stream_id for stream in catalog.streams}
416
+ synced_at = utc_now()
417
+
418
+ if "broadcasts" in stream_ids:
419
+ sync_broadcasts(config, state, catalog)
420
+
421
+ if "broadcast_stats" in stream_ids:
422
+ sync_broadcast_stats(config, catalog, synced_at)
423
+
424
+ if "subscribers" in stream_ids or "subscriber_stats" in stream_ids:
425
+ subscriber_ids = sync_subscribers(config, state, catalog)
426
+ if "subscriber_stats" in stream_ids:
427
+ sync_subscriber_stats(config, catalog, subscriber_ids, synced_at)
428
+
429
+
430
+ def main() -> None:
431
+ """Main entry point."""
432
+ args = parse_args(REQUIRED_CONFIG_KEYS)
433
+
434
+ config = args.config
435
+ state = args.state
436
+ catalog = args.catalog if args.catalog else discover()
437
+
438
+ if args.discover:
439
+ catalog = discover()
440
+ json.dump(catalog.to_dict(), sys.stdout, indent=2)
441
+ else:
442
+ sync(config, state, catalog)
443
+
444
+
445
+ if __name__ == "__main__":
446
+ main()