zen-fs-config 0.5.28 → 0.5.29

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/DESIGN.md CHANGED
@@ -1,1236 +1,1236 @@
1
- # zen-fs-config — Design Document
2
-
3
- > **Architecture direction (REQUIREMENTS §9 D2)**: the core uses a **GenericSyncGroup base with two type implementations** — `config-sync` (versioning/tombstone/conflict + hosts `.meta/app-data-groups/`) and `data-sync` (plain file sync, always governed as an app-data-group under config-sync via `ConfigRepo.createAppDataGroup`); **backend-type management and data-backend info persistence stay in the core `zen-fs-config` (UI is presentation-only)**. This document describes the **current implementation (two classes)**; the target architecture follows D2.
4
-
5
- ## 1. Overview
6
-
7
- zen-fs-config is a distributed configuration management library built on top of:
8
- - **ZenFS** (`@zenfs/core`) — Virtual file system with pluggable backends
9
- - **zen-fs-cache** — Caching layer with ETag/304 revalidation
10
- - **zen-fs-sync** — Sync engine for mirroring configs across backends
11
-
12
- It allows multiple application instances (programs) running on different nodes to share configuration through a network of ZenFS backends, with per-app isolation, shared config spaces, node-local config, and conflict safety.
13
-
14
- zen-fs-config supports two types of **sync groups**, each backed by a multi-backend sync network but serving different purposes:
15
-
16
- - **Config-sync group** — Synchronizes the configuration repository itself: backend topology, app configs, shared configs, node-local configs. This is the "meta layer."
17
- - **Data-sync group** — Synchronizes application data only. A data-sync group is always referenced and managed by a config-sync group as an app's data storage layer (via `ConfigRepo.createAppDataGroup`).
18
-
19
- A config-sync group can reference one or more data-sync groups per app, allowing apps to store bulk data on separate backends (e.g., a different repo/branch under the same account) while keeping configuration management unified.
20
-
21
- ## 2. Architecture
22
-
23
- ### 2.1 Three-Layer Stack
24
-
25
- ```
26
- Application code
27
- ↓ (reads/writes via standard node:fs API)
28
- ConfigRepo (this library)
29
- ├─ zen-fs-cache → CachedFileSystem (ETag/TTL read cache)
30
- ├─ ZenFS VFS → Context-isolated fs per app (chroot)
31
- └─ zen-fs-sync → Change detection + conflict resolution
32
- ├─ Backend X (replica)
33
- ├─ Backend Y (replica)
34
- └─ Backend Z (replica)
35
- ```
36
-
37
- ### 2.2 IndexedDB as Local Primary (Offline-First)
38
-
39
- Every program instance uses **IndexedDB** as its local primary backend. All config reads and writes target IndexedDB directly, ensuring offline availability and fast local access.
40
-
41
- User-provided backends (Gitee, S3, RemoteStorage, etc.) are added as **replicas** — they receive bi-directional sync with the local IndexedDB but are never the direct target of config operations.
42
-
43
- ```
44
- Program A → Primary = IndexedDB (local), Replicas = [Gitee, S3]
45
- Program B → Primary = IndexedDB (local), Replicas = [Gitee, S3]
46
- ```
47
-
48
- This means:
49
- - Config is always available offline (IndexedDB persists in the browser)
50
- - Remote backends accelerate multi-device sync, not local access
51
- - Re-opening the app requires zero backend parameters — IndexedDB + `.meta/backends/` contain everything needed
52
-
53
- ### 2.3 Self-Describing Configuration
54
-
55
- Backend topology and sync rules are stored **inside** the configuration repository (in `.meta/`), not passed as external parameters. This means any node that can read the config repo can bootstrap the entire sync network.
56
-
57
- External input at startup is limited to: **which backend to connect to** and optionally **bootstrap data** (if the repo doesn't exist yet).
58
-
59
- ### 2.4 Sync Group Types
60
-
61
- A **sync group** is a set of backends that synchronize with each other. There are two types:
62
-
63
- ```
64
- ┌─────────────────────────────────────────────────────────────────┐
65
- │ Config-Sync Group │
66
- │ ┌──────────────────────────────────────────────────────┐ │
67
- │ │ .meta/backends/ ← config-sync backends │ │
68
- │ │ .meta/app-data-groups/ ← references to data-sync │ │
69
- │ │ /{appId}/ ← app config data │ │
70
- │ │ /shared/ ← shared config │ │
71
- │ │ /nodes/ ← node-local config │ │
72
- │ └──────────────────────────────────────────────────────┘ │
73
- │ │ references │
74
- │ ▼ │
75
- │ ┌──────────────────────────────────────────────────────┐ │
76
- │ │ Data-Sync Group (per-app) │ │
77
- │ │ .meta/backends/ ← data-sync backends │ │
78
- │ │ / ← app data files │ │
79
- │ └──────────────────────────────────────────────────────┘ │
80
- └─────────────────────────────────────────────────────────────────┘
81
- ```
82
-
83
- **Config-sync group**:
84
- - Synced content: `.meta/` (topology), `/{appId}/` (app config), `/shared/`, `/nodes/`
85
- - Backends: full credentials + storage location (e.g., Gitee: token + owner + repo + branch)
86
- - Can reference data-sync groups via `.meta/app-data-groups/{appId}/`
87
-
88
- **Data-sync group**:
89
- - Synced content: application data files only (no config meta layer)
90
- - Backends: full credentials + storage location, but typically reusing the same account as a config-sync backend with a different storage target (e.g., same token/owner, different repo/branch)
91
- - Can be referenced by a config-sync group as an app's data storage (via `ConfigRepo.createAppDataGroup`)
92
-
93
- **Group type detection**: When connecting to a backend, the library reads `.meta/group-type` to determine the group type. If the file is absent, the backend is treated as a new empty group and the caller decides which type to create.
94
-
95
- | `.meta/group-type` | Behavior |
96
- |---|---|
97
- | `config-sync` | Full system: IndexedDB primary + config replicas + optional data-sync groups |
98
- | `data-sync` | Lightweight: data backends only, direct read/write, no config meta layer |
99
- | absent | New empty backend; caller decides group type at creation time |
100
-
101
- ## 3. File System Structure
102
-
103
- ### 3.1 Config-Sync Group
104
-
105
- ```
106
- /
107
- ├─ .meta/ [synced to replicas]
108
- │ ├─ group-type Group type marker: "config-sync"
109
- │ ├─ backends/ Backend topology (one file per backend)
110
- │ │ ├─ local-idb.json { id, type, options, description }
111
- │ │ ├─ gitee-prod.json
112
- │ │ └─ ...
113
- │ ├─ app-data-groups/ References to data-sync groups (per app)
114
- │ │ └─ {appId}/
115
- │ │ └─ {dataGroupId}.json { groupType: "data-sync", backends: [...] }
116
- │ ├─ .deleted/ Tombstones for deletion propagation
117
- │ │ └─ {encoded-path}.json
118
- │ └─ .conflicts/ Conflict archives (safekeeping)
119
- │ └─ {timestamp}_{path}/
120
- │ ├─ meta.json
121
- │ ├─ source
122
- │ └─ target
123
- │
124
- ├─ {appId}/ [synced: owner → replicas]
125
- │ ├─ db.json
126
- │ ├─ cache.json
127
- │ └─ .db.json.version Sidecar version file
128
- │
129
- ├─ shared/ [synced: bi-directional]
130
- │ ├─ feature-flags.json
131
- │ ├─ api-version.json
132
- │ └─ .feature-flags.json.version
133
- │
134
- └─ nodes/ [synced — bidirectional, not node-scoped]
135
- ├─ {nodeId}/
136
- │ ├─ local.json Node-local config
137
- │ └─ env.json
138
- └─ .node-id Current node's ID (auto-generated)
139
- ```
140
-
141
- ### 3.2 Data-Sync Group
142
-
143
- ```
144
- /
145
- ├─ .meta/ [synced to replicas]
146
- │ ├─ group-type Group type marker: "data-sync"
147
- │ └─ backends/ Data backend topology (one file per backend)
148
- │ ├─ gitee-data-1.json { id, type, options, description }
149
- │ └─ gitee-data-2.json
150
- │
151
- └─ (application data files) [synced: bi-directional]
152
- ├─ documents/
153
- │ ├─ note-1.json
154
- │ └─ note-2.json
155
- └─ media/
156
- └─ config.json
157
- ```
158
-
159
- A data-sync group has a much simpler structure: no version sidecars, no tombstones, no conflict archives — just raw data files and a minimal `.meta/` for backend topology and group type identification.
160
-
161
- ### 3.3 Directory Semantics (Config-Sync Group)
162
-
163
- | Directory | Sync Direction | Conflict Risk | Purpose |
164
- |---|---|---|---|
165
- | `/{appId}/` | Bi-directional (primary ↔ replicas) | Low (single device) | Per-app private config |
166
- | `/shared/` | Bi-directional | Possible (multiple writers) | Cross-app shared config |
167
- | `/nodes/` | None (by default) | None | Per-node local config |
168
- | `/.meta/` | Bi-directional | None (topology files) | Backend topology, tombstones, conflict archives |
169
-
170
- ### 3.4 Config-to-File Mapping
171
-
172
- Each config key maps to one file. The mapping is straightforward:
173
-
174
- - `setConfig('/db/host', { hostname: 'localhost' })` → writes file `/app-a/db/host.json` with content `{"hostname":"localhost"}`
175
- - `getConfig('/db/host')` → reads file `/app-a/db/host.json`, parses based on extension
176
- - Path is relative to the app's root (`/{appId}/`), with `.json` extension appended automatically
177
- - If path already has an extension (e.g., `/readme.md`), the extension is preserved
178
-
179
- ### 3.5 Serialization
180
-
181
- The serializer is determined by file extension:
182
-
183
- | Extension | Serialize | Deserialize |
184
- |---|---|---|
185
- | `.json` (default) | `JSON.stringify` | `JSON.parse` |
186
- | `.yaml` | YAML dump | YAML parse |
187
- | `.toml` | TOML dump | TOML parse |
188
- | `.txt` / no struct extension | `String(data)` | Return as string |
189
-
190
- Users can inject a custom `ConfigSerializer` for other formats.
191
-
192
- ## 4. Backend Topology (`.meta/backends/*.json`)
193
-
194
- Each backend is stored as an individual JSON file in `.meta/backends/`. This allows atomic add/remove operations without rewriting the entire topology.
195
-
196
- **`.meta/backends/local-idb.json`** (always present):
197
- ```json
198
- {
199
- "id": "local-idb",
200
- "type": "IndexedDB",
201
- "options": { "storeName": "zen-fs-config-my-app" },
202
- "description": "Local IndexedDB primary backend"
203
- }
204
- ```
205
-
206
- **`.meta/backends/gitee-prod.json`** (user-added replica):
207
- ```json
208
- {
209
- "id": "gitee-prod",
210
- "type": "Gitee",
211
- "options": { "token": "...", "owner": "...", "repo": "...", "branch": "main" },
212
- "description": "Production Gitee config repo"
213
- }
214
- ```
215
-
216
- The local IndexedDB backend (`local-idb`) is always the primary — all config operations target it directly. All other backends are replicas with bi-directional sync.
217
-
218
- **Migration**: If a legacy `.meta/backends.json` file exists (pre-0.4.0), it is automatically migrated to individual files on startup.
219
-
220
- ## 4.1 App Data Groups (`.meta/app-data-groups/{appId}/`)
221
-
222
- A config-sync group can reference data-sync groups on behalf of specific apps. Each reference is stored as a JSON file under `.meta/app-data-groups/{appId}/`.
223
-
224
- **`.meta/app-data-groups/my-app/data-store-1.json`**:
225
- ```json
226
- {
227
- "id": "data-store-1",
228
- "groupType": "data-sync",
229
- "backends": [
230
- {
231
- "id": "gitee-data",
232
- "type": "Gitee",
233
- "options": { "token": "...", "owner": "...", "repo": "my-app-data", "branch": "main" },
234
- "accountBackendId": "gitee-prod",
235
- "description": "App data on Gitee (reuses gitee-prod account)"
236
- }
237
- ]
238
- }
239
- ```
240
-
241
- ### Account Reuse
242
-
243
- The `accountBackendId` field (optional) references a config-sync backend whose account fields (e.g., `token`, `owner`, `baseUrl` for Gitee/GitHub) are reused. When present, the data-sync backend's `options` only need to specify the **storage location** fields (e.g., `repo`, `branch`); account fields are merged in from the referenced config-sync backend at creation time.
244
-
245
- When `accountBackendId` is absent or null, the data-sync backend must provide a complete set of options (including credentials).
246
-
247
- The `accountFields` metadata for each backend type (registered via `BackendMetadata.accountFields`) determines which fields are "account" vs "storage location":
248
-
249
- | Backend Type | Account Fields | Storage Location Fields |
250
- |---|---|---|
251
- | Gitee | `token`, `owner`, `baseUrl` | `repo`, `branch` |
252
- | GitHub | `token`, `owner`, `baseUrl` | `repo`, `branch` |
253
- | WebDAV | `url`, `username`, `password` | `rootPath` |
254
- | RemoteStorage | `userAddress`, `token` | (none) |
255
-
256
- ## 5. Sync Rules (`.meta/sync-rules.json`)
257
-
258
- ```json
259
- {
260
- "version": 1,
261
- "rules": [
262
- {
263
- "prefix": "/app-a/",
264
- "direction": "one-way",
265
- "conflictStrategy": "source-wins",
266
- "replicas": ["local-idb", "remote-s3"]
267
- },
268
- {
269
- "prefix": "/app-b/",
270
- "direction": "one-way",
271
- "conflictStrategy": "source-wins",
272
- "replicas": ["local-idb", "remote-s3"]
273
- },
274
- {
275
- "prefix": "/shared/",
276
- "direction": "bi-directional",
277
- "conflictStrategy": "merge",
278
- "replicas": ["local-idb", "remote-s3"]
279
- },
280
- {
281
- "prefix": "/nodes/",
282
- "direction": "none"
283
- },
284
- {
285
- "prefix": "/.meta/",
286
- "direction": "none"
287
- }
288
- ]
289
- }
290
- ```
291
-
292
- - Private app directories (`/{appId}/`): one-way push, no conflict possible
293
- - Shared directory (`/shared/`): bi-directional, conflict possible, merge strategy
294
- - Nodes directory (`/nodes/{nodeId}/`): bi-directional via the main sync pair (not node-scoped — every node's config syncs to every replica)
295
- - Meta directory (`/.meta/`): bi-directional (topology, tombstones, conflict archives)
296
-
297
- ## 6. Versioning & Change Detection
298
-
299
- ### 6.1 Sidecar Version Files
300
-
301
- Each config file has a companion version file:
302
-
303
- ```
304
- /app-a/db.json → Config content
305
- /app-a/.db.json.version → Version metadata
306
- ```
307
-
308
- Version file content:
309
- ```json
310
- {
311
- "version": 5,
312
- "hash": "sha256:a1b2c3d4...",
313
- "author": "app-a",
314
- "timestamp": 1689686400000
315
- }
316
- ```
317
-
318
- ### 6.2 Comparison Logic (extends zen-fs-sync's FileSnapshot)
319
-
320
- | Condition | Action |
321
- |---|---|
322
- | hash same | Skip (content unchanged) |
323
- | hash different, version different | Higher version wins |
324
- | hash different, version same | **Conflict** → conflict safety mechanism |
325
- | version/hash missing | Fall back to mtime+size comparison (backward compat) |
326
-
327
- ### 6.3 Version Increment
328
-
329
- On each write:
330
- 1. Read current version file (if exists)
331
- 2. Increment version by 1
332
- 3. Compute SHA-256 hash of new content
333
- 4. Set author to current instance's `{appId}/{nodeId}`
334
- 5. Write config file first, then version file
335
-
336
- Crash recovery: on startup, if hash in version file doesn't match actual file content, auto-increment version and update hash.
337
-
338
- ## 7. Conflict Safety Mechanism
339
-
340
- When a conflict is detected (same version, different hash on `/shared/` files):
341
-
342
- ### 7.1 Archive Both Versions
343
-
344
- Both conflicting versions are saved to `.meta/.conflicts/` before any resolution:
345
-
346
- ```
347
- .meta/.conflicts/1689686400000_shared-feature-flags.from-app-a.to-app-b.json
348
- ```
349
-
350
- Archive file content:
351
- ```json
352
- {
353
- "conflictPath": "/shared/feature-flags.json",
354
- "timestamp": 1689686400000,
355
- "sourceAuthor": "app-a/server-1",
356
- "targetAuthor": "app-b/server-2",
357
- "sourceContent": { "darkMode": true, "newFeature": true },
358
- "targetContent": { "darkMode": false, "newFeature": false },
359
- "sourceVersion": 3,
360
- "targetVersion": 3
361
- }
362
- ```
363
-
364
- ### 7.2 Resolution Strategies
365
-
366
- After archiving, resolve according to the configured strategy:
367
-
368
- | Strategy | Behavior |
369
- |---|---|
370
- | `source-wins` | Source content overwrites target. Target content archived. |
371
- | `target-wins` | Target content preserved. Source content archived. |
372
- | `merge` | JSON deep merge. Both originals archived. Non-JSON falls back to source-wins. |
373
-
374
- ### 7.3 Event Notification
375
-
376
- zen-fs-sync emits a `conflict` event with full conflict details. Application can:
377
- - Accept the auto-resolved result
378
- - Read `.meta/.conflicts/` archives to manually merge
379
- - Call `configRepo.resolveConflict(conflictId, mergedContent)` to submit a custom merge
380
-
381
- **Guarantee**: Neither side's content is ever lost. Recovery is always possible from `.meta/.conflicts/`.
382
-
383
- ## 8. Node-Local Configuration
384
-
385
- Some configs are specific to a single node and live under `/nodes/{nodeId}/`. They are synced to backends by the main bidirectional sync pair (there is no `direction: "none"` exclusion).
386
-
387
- ### 8.1 Storage
388
-
389
- Node-local configs live under `/nodes/{nodeId}/`. There is no `direction: "none"` exclusion for `/nodes/` — the main bidirectional sync pair (`fullFS` ↔ replica, root `/`, no filter) replicates it to every replica exactly like `/{appId}/` or `/shared/` (sync is not node-scoped).
390
-
391
- ```
392
- /nodes/server-1/
393
- ├─ local.json → { "ip": "10.0.0.1", "cpuCount": 8 }
394
- └─ env.json → { "NODE_ENV": "production" }
395
- ```
396
-
397
- ### 8.2 Node ID Source
398
-
399
- Priority order:
400
- 1. Explicit parameter: `createConfigRepo('app-a', { nodeId: 'server-1', ... })`
401
- 2. Environment variable: `process.env.NODE_ID`
402
- 3. Auto-generated: random ID written to `/nodes/.node-id` on first startup
403
-
404
- ### 8.3 API
405
-
406
- ```typescript
407
- // Write node-local config (writes local primary; synced to replicas on next poll/flush)
408
- repo.setNodeConfig('server-1', '/local.json', { ip: '10.0.0.1' });
409
-
410
- // Read node-local config
411
- const config = repo.getNodeConfig<{ ip: string }>('server-1', '/local.json');
412
-
413
- // Publish node config to sync backends (one-time, for debugging)
414
- const result = await repo.publishNodeConfig('server-1');
415
- // or publish specific files only:
416
- const result = await repo.publishNodeConfig('server-1', { paths: ['/local.json'] });
417
-
418
- // Peek at other nodes' published configs (read-only)
419
- const otherConfig = repo.peekNodeConfig<{ ip: string }>('server-2', '/local.json');
420
- ```
421
-
422
- | API | Write Target | Persisted | Synced | Purpose |
423
- |---|---|---|---|---|
424
- | `getConfig` / `setConfig` | CachedFS → auto-sync to replicas | Yes | Yes | Normal config |
425
- | `getNodeConfig` / `setNodeConfig` | Local primary → synced to replicas via main pair (eventual) | Yes | Yes (all `/nodes/*`, not node-scoped) | Node config, namespaced by nodeId |
426
- | `publishNodeConfig` | Explicit one-shot push to replicas | Yes | Yes (one-shot) | Force immediate sync of node files |
427
- | `peekNodeConfig` | Read local (synced-in) copy | N/A | N/A | Read another node's config (after it has synced in) |
428
-
429
- ## 9. ConfigRepo Interface
430
-
431
- ```typescript
432
- interface ConfigRepo {
433
- /** Application ID (e.g., "app-a") */
434
- readonly appId: string;
435
- /** Node ID (e.g., "server-1") */
436
- readonly nodeId: string;
437
- /** ZenFS-compatible fs object (node:fs API), context-isolated to own directories */
438
- readonly fs: typeof import('node:fs');
439
-
440
- /** Load/reload config from raw string (for initial setup) */
441
- load(rawConfig: string): Promise<void>;
442
-
443
- /** Read config value */
444
- getConfig<T>(path: string): T;
445
-
446
- /** Write config value (auto-synced) */
447
- setConfig(path: string, data: any): void;
448
-
449
- /** Read node-local config */
450
- getNodeConfig<T>(nodeId: string, path: string): T;
451
-
452
- /** Write node-local config (writes local primary; synced to replicas via main pair) */
453
- setNodeConfig(nodeId: string, path: string, data: any): void;
454
-
455
- /** Publish node-local config to sync backends (one-time, for debugging) */
456
- publishNodeConfig(nodeId: string, options?: {
457
- paths?: string[];
458
- }): Promise<SyncResult>;
459
-
460
- /** Peek at another node's published config (read-only) */
461
- peekNodeConfig<T>(nodeId: string, path: string): T;
462
-
463
- /** Manually flush all pending sync */
464
- flush(): Promise<SyncResult[]>;
465
-
466
- /** Get sync status for all sync pairs */
467
- getSyncStatuses(): Map<string, SyncPairStatus>;
468
-
469
- /** Resolve a conflict with custom merged content */
470
- resolveConflict(conflictId: string, mergedContent: any): Promise<void>;
471
-
472
- /** List conflict archives */
473
- listConflicts(): Promise<ConflictArchive[]>;
474
-
475
- /** Read backend topology (aggregated from .meta/backends/*.json) */
476
- getBackends(): Promise<BackendsMeta | null>;
477
-
478
- /** Write backend topology (writes each backend as individual file) */
479
- updateBackends(meta: BackendsMeta): Promise<void>;
480
-
481
- /** Dynamically add a replica backend */
482
- addBackend(id: string, type: string, options: Record<string, unknown>, description?: string): Promise<void>;
483
-
484
- /** Dynamically remove a replica backend */
485
- removeBackend(id: string): Promise<void>;
486
-
487
- /** Delete a file with tombstone (propagates deletion to all backends) */
488
- deleteFile(path: string): Promise<void>;
489
-
490
- /** Sync .meta/ files to all replicas */
491
- syncMetaToReplicas(): Promise<void>;
492
-
493
- // --- App Data Storage (data-sync groups) ---
494
-
495
- /**
496
- * Create a data-sync group for this app, referencing a config-sync backend's account.
497
- * The data-sync group gets its own set of backends (typically reusing an account
498
- * from a config-sync backend but with different storage location like repo/branch).
499
- *
500
- * @param id Data group ID (e.g., "data-store-1")
501
- * @param backends Array of backend descriptors for the data-sync group.
502
- * Each can optionally specify `accountBackendId` to reuse credentials.
503
- */
504
- createAppDataGroup(
505
- id: string,
506
- backends: AppDataBackendDescriptor[],
507
- ): Promise<AppDataGroup>;
508
-
509
- /**
510
- * Get an existing data-sync group for this app.
511
- * Returns a handle with its own fs for reading/writing data files.
512
- */
513
- getAppDataGroup(id: string): Promise<AppDataGroup>;
514
-
515
- /** List all data-sync groups registered for this app. */
516
- listAppDataGroups(): Promise<AppDataGroupDescriptor[]>;
517
-
518
- /** Remove a data-sync group (stops sync, removes descriptor). */
519
- removeAppDataGroup(id: string): Promise<void>;
520
-
521
- /** Dispose: stop sync, release resources */
522
- dispose(): Promise<void>;
523
- }
524
-
525
- /**
526
- * A data-sync group handle. Provides direct file system access
527
- * to the app's data storage, independent of the config-sync layer.
528
- */
529
- interface AppDataGroup {
530
- readonly groupId: string;
531
- readonly appId: string;
532
- /** Direct fs for reading/writing data files (chroot to this group's root) */
533
- readonly fs: typeof import('node:fs');
534
- /** Get sync status for this data group's sync pairs */
535
- getSyncStatuses(): Map<string, SyncPairStatus>;
536
- /** Manually flush pending sync */
537
- flush(): Promise<SyncResult[]>;
538
- /** Stop sync and release resources */
539
- dispose(): Promise<void>;
540
- }
541
-
542
- /** Descriptor for a backend within a data-sync group. */
543
- interface AppDataBackendDescriptor {
544
- id: string;
545
- type: string;
546
- options: Record<string, unknown>;
547
- /** Optional: reuse account fields from a config-sync backend */
548
- accountBackendId?: string;
549
- description?: string;
550
- }
551
- ```
552
-
553
- ## 10. Initialization
554
-
555
- The recommended entry point is `connect`, which always anchors on a config-sync repo (the host for data groups) and auto-detects the backend type. The lower-level `createConfigRepo` is also available; data groups are always created via `ConfigRepo.createAppDataGroup` (decision A / T5).
556
-
557
- ### `connect` (recommended — auto-detect)
558
-
559
- ```typescript
560
- import { connect } from 'zen-fs-config';
561
-
562
- // User provides a backend — connect auto-detects config-sync vs data-sync
563
- const result = await connect('my-app', {
564
- backendInfo: {
565
- type: 'Gitee',
566
- options: { token: '...', owner: '...', repo: '...', branch: 'main' },
567
- },
568
- });
569
-
570
- // result.groupType → "config-sync" or "data-sync"
571
- // result.repo → ConfigRepo (always — it hosts the data groups)
572
- // result.dataGroup → config-managed data group (default group on first launch / data-sync connect)
573
- ```
574
-
575
- ### Zero-parameter (offline-first)
576
-
577
- ```typescript
578
- import { createConfigRepo } from 'zen-fs-config';
579
-
580
- // No parameters needed — IndexedDB is always created as primary
581
- const repo = await createConfigRepo('my-app');
582
-
583
- // Config is immediately available from IndexedDB
584
- repo.setConfig('/db/host', { hostname: 'localhost', port: 3306 });
585
- ```
586
-
587
- ### With initial replica backend
588
-
589
- ```typescript
590
- const repo = await createConfigRepo('my-app', {
591
- // Optional: provide a remote backend as initial replica
592
- primaryBackendId: 'gitee-prod',
593
- backendInfo: {
594
- type: 'Gitee',
595
- options: { token: '...', owner: '...', repo: '...', branch: 'main' },
596
- },
597
- // Optional: customize IndexedDB store name
598
- idbStoreName: 'my-app-config',
599
- // Optional: node ID (auto-detected if not provided)
600
- nodeId: 'server-1',
601
- });
602
-
603
- // Later, add more backends dynamically
604
- await repo.addBackend('s3-backup', 'S3Bucket', {
605
- bucket: 'app-config',
606
- region: 'us-east-1',
607
- }, 'S3 backup');
608
-
609
- // Remove a backend
610
- await repo.removeBackend('gitee-prod');
611
-
612
- // Cleanup
613
- await repo.dispose();
614
- ```
615
-
616
- ### Re-opening (zero parameters)
617
-
618
- ```typescript
619
- // On subsequent opens, just pass appId
620
- // IndexedDB + .meta/backends/ contain all state
621
- const repo = await createConfigRepo('my-app');
622
-
623
- // All previously added backends are automatically reconnected
624
- const backends = await repo.getBackends();
625
- // backends.backends = [{ id: 'local-idb', ... }, { id: 's3-backup', ... }]
626
- ```
627
-
628
- ### Data-Sync Group via Config-Sync (recommended)
629
-
630
- > Data groups are always managed by a config-sync repo via `createAppDataGroup` (decision A / T5). The snippet below shows the supported path.
631
-
632
- ```typescript
633
- import { connect } from 'zen-fs-config';
634
-
635
- // connect always returns a config-sync repo; data groups live under it
636
- const result = await connect('my-app', { nodeId: 'node-1' });
637
- const repo = result.repo!;
638
-
639
- // Create a data group and attach a data backend — config-sync stores the
640
- // data group's backend topology in .meta/app-data-groups/{appId}/{id}.json
641
- const dataGroup = await repo.createAppDataGroup('notes', [
642
- { id: 'gitee-data', type: 'Gitee', options: { token: '...', owner: '...', repo: 'my-app-data', branch: 'main' } },
643
- ]);
644
-
645
- // Read/write data files directly
646
- await dataGroup.fs.promises.writeFile('/notes/todo.json', JSON.stringify({ task: 'buy milk' }));
647
- const data = JSON.parse(await dataGroup.fs.promises.readFile('/notes/todo.json', 'utf-8'));
648
-
649
- // Add more data backends later (multi-backend sync)
650
- await dataGroup.addBackend('gitee-backup', 'Gitee', {
651
- token: '...', owner: '...', repo: 'my-app-data-backup', branch: 'main',
652
- });
653
-
654
- // Cleanup
655
- await dataGroup.dispose();
656
- await repo.dispose();
657
- ```
658
-
659
- ### Config-Sync with App Data Group (account reuse)
660
-
661
- ```typescript
662
- const repo = await createConfigRepo('my-app', {
663
- backendInfo: {
664
- type: 'Gitee',
665
- options: { token: '...', owner: '...', repo: 'configs', branch: 'main' },
666
- },
667
- });
668
-
669
- // Create a data-sync group that reuses the config backend's account
670
- // but stores data in a different repo
671
- await repo.createAppDataGroup('data-store-1', [
672
- {
673
- id: 'gitee-data',
674
- type: 'Gitee',
675
- accountBackendId: 'gitee-prod', // reuse token + owner from this config backend
676
- options: { repo: 'my-app-data', branch: 'main' }, // only storage location
677
- },
678
- ]);
679
-
680
- // Get the data group handle for direct file access
681
- const dataGroup = await repo.getAppDataGroup('data-store-1');
682
- await dataGroup.fs.promises.writeFile('/cache.json', '{"key":"value"}');
683
- ```
684
-
685
- ## 11. Initialization Flow
686
-
687
- ### 11.1 Config-Sync Group (`createConfigRepo`)
688
-
689
- ```
690
- createConfigRepo('my-app', options?)
691
- │
692
- ├─ 1. Create IndexedDB backend (always, ID = 'local-idb')
693
- │ storeName = options.idbStoreName || `zen-fs-config-${appId}`
694
- │
695
- ├─ 2. Ensure /.meta/ directory exists
696
- │
697
- ├─ 3. Write /.meta/group-type = "config-sync" (if not exists)
698
- │
699
- ├─ 4. Migrate legacy .meta/backends.json → .meta/backends/*.json (if exists)
700
- │
701
- ├─ 5. If options.backendInfo provided:
702
- │ ├─ Generate replica ID (options.primaryBackendId or auto)
703
- │ ├─ Dedup check: same type + options (stable key) already registered?
704
- │ └─ Write descriptor to .meta/backends/{replicaId}.json (if not duplicate)
705
- │
706
- ├─ 6. Read all backend descriptors from .meta/backends/
707
- │ └─ Dedup: remove duplicates (same type + options, different ID)
708
- │ ├─ Delete duplicate files on ALL replicas directly
709
- │ └─ Create tombstone + delete local file
710
- │
711
- ├─ 7. Determine nodeId (explicit parameter > auto-generated)
712
- │
713
- ├─ 8. Create final ConfigRepo instance (primary = 'local-idb')
714
- │
715
- ├─ 9. setupSync: for each replica backend:
716
- │ ├─ Create backend instance (e.g., Gitee, RemoteStorage)
717
- │ ├─ Create SyncPair(IndexedDB ↔ replica, bi-directional)
718
- │ ├─ Register conflict handler
719
- │ └─ NOTE: Does NOT call watch() yet (see §11.4 for why)
720
- │
721
- ├─ 10. Load config cache from IndexedDB (fast, local-only)
722
- │
723
- ├─ 11. initialSyncAndDedup() — only if replicas exist:
724
- │ ├─ unwatchAll() — safety: clear any stale snapshots
725
- │ ├─ syncAll() — full bidirectional sync (no cached snapshot
726
- │ │ → every file is compared, remote-only files
727
- │ │ are pulled to local)
728
- │ ├─ readAllBackendDescriptors() — dedup duplicates pulled from remote
729
- │ │ ├─ Delete dup files on ALL replicas directly
730
- │ │ └─ Create tombstones for deduped descriptors
731
- │ ├─ processTombstones() — delete deduped files on all replicas
732
- │ └─ watchAll() — start monitoring for future changes
733
- │ (snapshots now reflect the fully synced state)
734
- │
735
- ├─ 12. syncMetaToReplicas() — background push of .meta/ changes
736
- │ (watchers already running, this just speeds up initial propagation)
737
- │
738
- └─ 13. Return ConfigRepo instance
739
- ```
740
-
741
- ### 11.2 Why "Sync Before Watch" (Critical Design Decision)
742
-
743
- The sync engine (`zen-fs-sync`) uses **snapshot-based change detection**. When `watch()` is called on a SyncPair, it triggers `buildInitialSnapshots()` which:
744
-
745
- 1. Builds a snapshot of the source (IndexedDB) — walks all files, records `path`, `size`, `mtimeMs`
746
- 2. Builds a snapshot of the target (remote backend) — same process
747
- 3. Caches **separate** snapshots: `prevSrcSnap` (source) and `prevTgtSnap` (target)
748
-
749
- On the next `syncAll()`, `syncBidirectional()` compares each side's current snapshot with its own previous snapshot independently. If both sides are unchanged → **"unchanged" → skip sync entirely**.
750
-
751
- **The problem**: `buildInitialSnapshots()` only *reads* file metadata — it does NOT copy any files. So if the remote has files that the local doesn't (e.g., duplicate backend descriptors written by another node), the cached snapshots reflect the un-synced state. The subsequent sync sees "snapshots already match this state" and skips — the file is never actually copied to local, and local-only dedup logic never runs.
752
-
753
- **The fix**: Always perform a full `syncAll()` **before** `watch()`. With no cached snapshot, `syncBidirectional()` does a complete comparison and copies all missing files. After sync completes, `watch()` builds snapshots from the now-consistent state.
754
-
755
- This pattern is applied in three places:
756
- - `createConfigRepo()` → `initialSyncAndDedup()` (sync → dedup → watch)
757
- - `addBackend()` → `syncMetaToReplicas()` then `watch()` (sync → watch)
758
- - `AppDataGroupImpl.connect()` → `syncAll()` then `watchAll()` (sync → watch)
759
-
760
- ### 11.3 `flush()` — Manual Sync Trigger
761
-
762
- ```
763
- flush()
764
- │
765
- ├─ 1. processTombstones()
766
- │ For each tombstone in /.meta/.deleted/:
767
- │ ├─ Delete the actual file on primary (in case re-created)
768
- │ ├─ Delete the actual file on ALL replicas
769
- │ └─ Delete version sidecars on all replicas
770
- │
771
- ├─ 2. syncAll()
772
- │ For each SyncPair (IndexedDB ↔ replica):
773
- │ ├─ Build current snapshots of both sides
774
- │ ├─ Compare with cached snapshot (if any)
775
- │ ├─ Detect changes: Created / Modified / Deleted
776
- │ ├─ Resolve conflicts (source-wins strategy)
777
- │ └─ Copy files in both directions as needed
778
- │
779
- ├─ 3. readAllBackendDescriptors() — post-sync dedup
780
- │ Sync may have pulled duplicate backend descriptors from remote.
781
- │ Re-run dedup to catch and remove them.
782
- │ ├─ Delete dup files on ALL replicas directly
783
- │ └─ Create tombstones for deduped descriptors
784
- │
785
- ├─ 4. processTombstones() — process any new tombstones from step 3
786
- │
787
- ├─ 5. updateTombstoneConfirmations()
788
- │ Mark each tombstone as confirmed by all replica backends
789
- │
790
- ├─ 6. gcTombstones()
791
- │ Remove tombstones confirmed by ALL backends in the topology
792
- │
793
- └─ Return SyncResult[] (one per sync pair)
794
- ```
795
-
796
- ### 11.4 Tombstone-Based Deletion Propagation
797
-
798
- When a file is deleted via `deleteFile(path)`:
799
-
800
- ```
801
- deleteFile('/.meta/backends/old-backend.json')
802
- │
803
- ├─ 1. Write tombstone: /.meta/.deleted/++meta__backends__old-backend++json.json
804
- │ { path, deletedAt, deletedBy, confirmedBy: [primaryBackendId] }
805
- │
806
- ├─ 2. Delete the actual file on primary (IndexedDB)
807
- │
808
- └─ 3. Delete version sidecar (.old-backend.json.version) on primary
809
- ```
810
-
811
- On the next `processTombstones()` (called by `flush()` or `initialSyncAndDedup()`):
812
-
813
- ```
814
- For each tombstone:
815
- ├─ Delete file on primary (in case sync re-created it)
816
- ├─ Delete file on ALL replicas
817
- ├─ Delete version sidecar on ALL replicas
818
- └─ Tombstone file itself is synced to replicas via syncAll()
819
- → Late-joining replicas see the tombstone and delete the file
820
- ```
821
-
822
- **Why tombstones?** Without them, bi-directional sync treats a deleted local file as "missing → needs to be copied from remote". The tombstone explicitly signals "this file was intentionally deleted" so all replicas honor the deletion. Tombstones are garbage-collected after all backends confirm receipt.
823
-
824
- ### 11.5 Backend Deduplication
825
-
826
- When `readAllBackendDescriptors()` detects two backends with the same `type` + `options` (using stable key ordering) but different IDs:
827
-
828
- ```
829
- Detected: rs-1 and rs-2 have identical type + options
830
- │
831
- ├─ 1. Keep the one with the earliest mtime (created first)
832
- │
833
- ├─ 2. For each duplicate:
834
- │ ├─ Delete descriptor file on ALL replicas directly
835
- │ │ (prevents sync from pulling it back)
836
- │ ├─ Delete version sidecar on ALL replicas
837
- │ └─ Create tombstone + delete local file
838
- │
839
- └─ 3. Return deduplicated list (duplicates removed)
840
- ```
841
-
842
- The stable key function (`backendDedupKey`) sorts object keys recursively, so `{ token: 'a', owner: 'b' }` and `{ owner: 'b', token: 'a' }` produce the same key and are correctly detected as duplicates.
843
-
844
- ### 11.6 Dynamic Backend Management
845
-
846
- **`addBackend(id, type, options)`**:
847
-
848
- ```
849
- ├─ 1. Dedup check: reject if same type+options already registered
850
- ├─ 2. Create backend instance
851
- ├─ 3. Write descriptor to .meta/backends/{id}.json
852
- ├─ 4. Create SyncPair (IndexedDB ↔ new replica, bi-directional)
853
- ├─ 5. syncMetaToReplicas() — full sync FIRST (pull + push)
854
- └─ 6. watch(pairId) — start monitoring AFTER sync completes
855
- ```
856
-
857
- **`removeBackend(id)`**:
858
-
859
- ```
860
- ├─ 1. Delete descriptor file on the remote backend DIRECTLY
861
- │ (must happen before removing sync pair — otherwise can't reach remote)
862
- ├─ 2. Delete version sidecar on remote
863
- ├─ 3. Create tombstone + delete local descriptor file
864
- ├─ 4. Remove sync pair (stops watching + disposes)
865
- ├─ 5. Remove from replicaBackends map
866
- ├─ 6. Dispose backend instance
867
- ├─ 7. processTombstones() — propagate deletion to remaining replicas
868
- └─ 8. flush() — sync + GC tombstones
869
- ```
870
-
871
- ### 11.7 Watch Mode (Auto-Sync)
872
-
873
- After initialization, each SyncPair runs in **watch mode** with hybrid change detection:
874
-
875
- ```
876
- watch() triggers:
877
- │
878
- ├─ 1. Register onChange callbacks (if backend supports it)
879
- │ Local backends (IndexedDB) push change notifications
880
- │ → triggers debounced sync (default 300ms)
881
- │
882
- ├─ 2. buildInitialSnapshots()
883
- │ ├─ BiDirectional: cache separate source and target snapshots
884
- │ └─ OneWay: cache source snapshot only
885
- │
886
- └─ 3. Start poll timers (if backend supports shouldSync)
887
- ├─ Remote backends poll shouldSync() every pollIntervalMs (default 30min)
888
- └─ Fallback: if no onChange and no shouldSync, poll every interval
889
- ```
890
-
891
- **State guard**: If `unwatch()` is called during `buildInitialSnapshots()` (which is async), the snapshots are discarded — they won't be cached. This prevents stale snapshots from causing sync skips.
892
-
893
- **Snapshot comparison** in `syncBidirectional()`:
894
- 1. Build current snapshots of both sides (via `getSnapshot()`)
895
- 2. Compare each side independently against its own previous snapshot:
896
- - `srcChanged = !snapshotsEqual(prevSrcSnap, currentSrcSnap)`
897
- - `tgtChanged = !snapshotsEqual(prevTgtSnap, currentTgtSnap)`
898
- - If neither changed and both previous snapshots exist → skip sync entirely
899
- 3. Cache current snapshots as `prevSrcSnap` and `prevTgtSnap` for next comparison
900
- 4. If either side changed, proceed with full diff and file operations
901
-
902
- **Key difference from previous design**: The old approach merged source and target snapshots into a single map (`source ∪ target`), which lost information about which filesystem a file belonged to. The new approach keeps them separate, enabling precise per-side change detection and bidirectional deletion propagation (see §11.10).
903
-
904
- ### 11.8 Standalone Data-Sync Entry — REMOVED
905
-
906
- The standalone data-sync entry has been **removed** from the public API (decision A / T5). Data groups are created exclusively via `ConfigRepo.createAppDataGroup` and are always owned by a config-sync repo, whose `.meta/app-data-groups/{appId}/{id}.json` is the authoritative source of the data group's backend topology.
907
-
908
- ### 11.9 Unified Entry Point (`connect`)
909
-
910
- `createConfigRepo` is the lower-level factory for a config-sync repo. The recommended entry point is `connect`, which always anchors on a config-sync repo and dispatches the backend as follows:
911
-
912
- ```
913
- connect('my-app', options?)
914
- │
915
- ├─ 1. Determine config-sync repo (create or reuse) — this always hosts the data groups
916
- │
917
- ├─ 2. If options.backendInfo provided, read /.meta/group-type
918
- │
919
- ├─ "config-sync" → connect the repo to that backend; then load every app
920
- │ data group from .meta/app-data-groups/{appId}/*.json (UC2)
921
- │ return { groupType: "config-sync", repo, appDataGroups }
922
- │
923
- ├─ "data-sync" → keep the config repo local-only; ensure the default data
924
- │ group exists and attach the backend to it (UC3). The backend
925
- │ info is written back into app-data-groups by addBackend.
926
- │ return { groupType: "data-sync", repo, dataGroup }
927
- │
928
- └─ absent (no backendInfo) → local-only config repo + a default data group
929
- └─ return { groupType: "config-sync" (or options.groupType), repo, dataGroup }
930
- ```
931
-
932
- **Usage**:
933
-
934
- ```typescript
935
- import { connect } from 'zen-fs-config';
936
-
937
- // Auto-detect: connects to backend, reads group-type, dispatches accordingly
938
- const result = await connect('my-app', {
939
- backendInfo: {
940
- type: 'Gitee',
941
- options: { token: '...', owner: '...', repo: '...', branch: 'main' },
942
- },
943
- });
944
-
945
- if (result.groupType === 'config-sync') {
946
- // result.repo is a ConfigRepo — full config system
947
- const repo = result.repo;
948
- repo.setConfig('/db/host', { hostname: 'localhost' });
949
- } else {
950
- // result.dataGroup is a config-managed data group (result.repo is its host)
951
- const dataGroup = result.dataGroup;
952
- await dataGroup.fs.promises.writeFile('/data.json', '{"key":"value"}');
953
- }
954
-
955
- // Explicit override (skip detection, force a specific group type)
956
- const result = await connect('my-app', {
957
- backendInfo: { type: 'Gitee', options: {...} },
958
- groupType: 'data-sync', // force data-sync even if backend has no group-type yet
959
- });
960
- ```
961
-
962
- **Return type**:
963
-
964
- ```typescript
965
- interface ConnectResult {
966
- /** Detected or forced group type */
967
- groupType: 'config-sync' | 'data-sync';
968
- /** The config-sync repo — always present; it hosts the data groups */
969
- repo?: ConfigRepo;
970
- /** The default app data group (config-managed); present on first launch / data-sync connect */
971
- dataGroup?: AppDataGroup;
972
- /** All app data groups loaded for this app (may be empty) */
973
- appDataGroups?: AppDataGroup[];
974
- }
975
- ```
976
-
977
- **Offline / zero-parameter mode**: When no `backendInfo` is provided, `connect` defaults to `config-sync`, creates an IndexedDB-only config repo, and also creates a default app data group under it via `createAppDataGroup` (same as `createConfigRepo` with no options, plus the default data group).
978
-
979
- ### 11.10 Snapshot Optimization Design
980
-
981
- This section describes three interrelated optimizations to the sync engine's snapshot mechanism.
982
-
983
- #### 11.10.1 FS-Provided `createSnapshot()`
984
-
985
- The `SyncableFS` interface now includes an optional `createSnapshot()` method:
986
-
987
- ```typescript
988
- interface SyncableFS {
989
- // ... existing methods ...
990
-
991
- /**
992
- * Optional: Build a filesystem snapshot.
993
- * Returns a map of relative path → {size, mtimeMs} for all files under root.
994
- * Returns null if the filesystem is unreachable.
995
- *
996
- * Backends that can provide a more efficient snapshot than the generic
997
- * walkFiles+stat approach should implement this method.
998
- */
999
- createSnapshot?(root: string, filter?: SyncFilter): Promise<Map<string, FileSnapshot> | null>;
1000
- }
1001
- ```
1002
-
1003
- The sync engine's `getSnapshot()` helper dispatches to the FS-provided method when available, falling back to the generic `buildSnapshot()` (walkFiles + stat) otherwise:
1004
-
1005
- ```typescript
1006
- private async getSnapshot(fs: SyncableFS): Promise<Map<string, FileSnapshot> | null> {
1007
- if (fs.createSnapshot) {
1008
- return fs.createSnapshot(this.root, this.options.filter);
1009
- }
1010
- return buildSnapshot(fs, this.root, this.options.filter);
1011
- }
1012
- ```
1013
-
1014
- **Optimization examples**:
1015
- - **Gitee/GitHub**: Use Git tree API to fetch all file metadata in a single request instead of walking files one by one
1016
- - **IndexedDB**: Use `getAll()` for batch querying instead of individual `stat()` calls
1017
- - **InMemory**: Directly iterate the internal Map (no async I/O overhead)
1018
-
1019
- Backends that do not implement `createSnapshot()` are fully supported — the generic fallback produces identical results.
1020
-
1021
- #### 11.10.2 Separate Source and Target Snapshots
1022
-
1023
- **Previous design** (merged snapshots):
1024
- - `buildInitialSnapshots()` merged source and target into a single map: `sourceSnapshots = new Map([...srcSnap, ...tgtSnap])`
1025
- - `syncBidirectional()` compared `currentMerged` with the cached merged snapshot
1026
- - **Problem**: The merged map lost which filesystem a file belonged to. A file present on target but not source could be "new on target" or "deleted from source" — the merged snapshot couldn't distinguish.
1027
-
1028
- **New design** (separate snapshots):
1029
- - `buildInitialSnapshots()` caches two independent maps: `prevSrcSnap` and `prevTgtSnap`
1030
- - `syncBidirectional()` compares each side independently:
1031
- ```
1032
- srcChanged = !snapshotsEqual(prevSrcSnap, currentSrcSnap)
1033
- tgtChanged = !snapshotsEqual(prevTgtSnap, currentTgtSnap)
1034
- if (!srcChanged && !tgtChanged && prevSrcSnap && prevTgtSnap) → skip sync
1035
- ```
1036
- - Each side's change is detected independently, preserving file-location information
1037
-
1038
- #### 11.10.3 Bidirectional Deletion Propagation
1039
-
1040
- When a file exists on one side but not the other, the sync engine uses previous snapshots to distinguish "created" from "deleted":
1041
-
1042
- ```
1043
- File on target, not on source:
1044
- ├─ Was it on source in the previous snapshot (prevSrcSnap)?
1045
- │ ├─ Yes → file was deleted from source → delete from target too (propagate deletion)
1046
- │ └─ No → file was created on target → copy to source
1047
-
1048
- File on source, not on target:
1049
- ├─ Was it on target in the previous snapshot (prevTgtSnap)?
1050
- │ ├─ Yes → file was deleted from target → delete from source too (propagate deletion)
1051
- │ └─ No → file was created on source → copy to target
1052
- ```
1053
-
1054
- **Without this mechanism**, deleting a file on one side would cause the sync engine to see "the other side still has it → copy it back", effectively undoing the deletion.
1055
-
1056
- **Relationship with tombstones**: The tombstone mechanism (§11.4) and bidirectional deletion propagation operate at different layers and complement each other:
1057
-
1058
- | Mechanism | Layer | Trigger | How It Works |
1059
- |-----------|-------|---------|--------------|
1060
- | Tombstone | `zen-fs-config` (application) | Application calls `deleteFile()` | Writes a `.meta/.deleted/` marker, physically deletes file on all replicas **before** sync runs |
1061
- | Deletion propagation | `zen-fs-sync` (engine) | Sync detects one-side-only file | Compares with previous snapshot to determine if file was created or deleted |
1062
-
1063
- Tombstones handle application-initiated deletions (the common case). Deletion propagation handles deletions that bypass the tombstone flow — e.g., external modifications on the remote backend, or files removed by other sync mechanisms.
1064
-
1065
- #### 11.10.4 `shouldSync()` vs Snapshot Comparison
1066
-
1067
- These two mechanisms are complementary, not interchangeable:
1068
-
1069
- | Mechanism | Purpose | Cost | When Used |
1070
- |-----------|---------|------|-----------|
1071
- | `shouldSync()` | Fast "has anything changed?" boolean | O(1) for remote (ETag/commit check) | `onRemotePoll()` — decide whether to trigger sync at all |
1072
- | Snapshot comparison | "What exactly changed?" detail | O(n) filesystem traversal | `syncBidirectional()` — decide what to copy/delete |
1073
-
1074
- `shouldSync()` is **not** part of the snapshot comparison because:
1075
- 1. `shouldSync()` updates its internal baseline after each call — calling it again during sync would return stale results
1076
- 2. `shouldSync()` returning false doesn't guarantee the snapshot is unchanged — it means the FS's own change detection says nothing changed, which could miss edge cases
1077
- 3. `shouldSync()` returning true doesn't tell us **which** files changed — snapshots are still needed for that
1078
-
1079
- The two-level optimization works as follows:
1080
- 1. **Level 1**: `shouldSync()` in remote poll → if false, skip sync trigger entirely (saves the O(n) snapshot build)
1081
- 2. **Level 2**: Separate snapshot comparison in sync → if both sides unchanged, skip file operations (saves I/O)
1082
-
1083
- ## 12. Data Flow
1084
-
1085
- ### Read Path
1086
- ```
1087
- Application
1088
- → repo.fs.readFileSync('/db/host.json')
1089
- → ZenFS Context (chroot to /app-a/)
1090
- → CachedFileSystem.readFile('/app-a/db/host.json')
1091
- → Cache hit (TTL)? → return cached bytes (0 network)
1092
- → Cache miss/expired? → 304 revalidate with primary backend
1093
- → Deserialize (JSON.parse for .json files)
1094
- → Return typed object
1095
- ```
1096
-
1097
- ### Write Path (auto-synced)
1098
- ```
1099
- Application
1100
- → repo.setConfig('/db/host', { hostname: 'localhost' })
1101
- → Serialize (JSON.stringify)
1102
- → Write config file: /app-a/db/host.json
1103
- → Write version file: /app-a/.db.host.json.version (version++, new hash)
1104
- → CachedFileSystem.writeFile() →穿透 to primary backend → invalidate cache
1105
- → zen-fs-sync watch detects change (poll + debounce)
1106
- → Sync to replicas per sync-rules
1107
- ```
1108
-
1109
- ### Write Path (node-local, synced)
1110
- ```
1111
- Application
1112
- → repo.setNodeConfig('server-1', '/local.json', { ip: '10.0.0.1' })
1113
- → Serialize + write to /nodes/server-1/local.json (local primary)
1114
- → Main bidirectional sync pair replicates /nodes/ to replicas on next poll/flush
1115
- → File is also present on other replicas (sync is not node-scoped)
1116
- ```
1117
-
1118
- ### Publish (one-time sync)
1119
- ```
1120
- Application
1121
- → repo.publishNodeConfig('server-1')
1122
- → Read /nodes/server-1/**/*
1123
- → Create temporary SyncPair with filter: includePrefixes: ['/nodes/server-1/']
1124
- → Execute one sync() call
1125
- → Files pushed to replicas
1126
- → Dispose temporary SyncPair
1127
- ```
1128
-
1129
- ### Mtime Preservation During Sync
1130
-
1131
- **Problem**: When the sync engine copies a file from source to target, it calls `writeFile(path, data)`. The target backend sets its own mtime (typically `Date.now()`), losing the source file's original mtime. This causes the next sync cycle to detect a "modified" file (source mtime ≠ target mtime), triggering unnecessary copies on every sync.
1132
-
1133
- **Solution**: An optional `writeFileWithMtime` method on the `SyncableFS` interface, with automatic fallback to `writeFile` when not implemented:
1134
-
1135
- ```typescript
1136
- interface SyncableFS {
1137
- // ... existing methods ...
1138
-
1139
- /**
1140
- * Optional: write file with precise mtime.
1141
- * If implemented, the sync engine uses this instead of writeFile,
1142
- * passing the source file's mtime so the target can preserve it.
1143
- * Backends that don't support precise mtime should not implement this —
1144
- * the sync engine falls back to plain writeFile.
1145
- */
1146
- writeFileWithMtime?(path: string, data: string | Uint8Array, mtime: number): Promise<void>;
1147
- }
1148
- ```
1149
-
1150
- **Sync engine (`zen-fs-sync`)**: A central helper function handles the fallback:
1151
-
1152
- ```javascript
1153
- async function writeFileWithMtimeFallback(fs, path, data, mtimeMs) {
1154
- if (mtimeMs !== undefined && typeof fs.writeFileWithMtime === "function") {
1155
- await fs.writeFileWithMtime(path, data, mtimeMs);
1156
- } else {
1157
- await fs.writeFile(path, data);
1158
- }
1159
- }
1160
- ```
1161
-
1162
- This helper is used in `copyFile()`, `syncOneWay()`, and `writeFileBoth()`. The source file's mtime is obtained via `stat()` before writing, then passed through to the target.
1163
-
1164
- **Adapters (`zen-fs-config`)**: All three `SyncableFS` adapters implement `writeFileWithMtime`:
1165
-
1166
- | Adapter | Implementation |
1167
- |---|---|
1168
- | `backendToSyncableFS` | Passes `{ mtime }` as options to `backend.writeFile()` — the backend's `writeFile` calls `touch()` with the provided mtime |
1169
- | `zenfsPromisesToSyncableFS` | Calls `promises.writeFile()` then `promises.utimes()` as a fallback (some VFS backends don't support mtime in writeFile) |
1170
- | `cachedFSToSyncableFS` | Passes `{ mtime }` as options to `cached.writeFile()` — mtime flows through to the underlying backend |
1171
-
1172
- **RemoteStorage backend (`zen-fs-remotestoragejs`)**: `writeFileWithMtime` delegates to `writeFile(path, data, { mtime })`, which writes the `.mtime` sidecar file to preserve millisecond-precision mtime (see RemoteStorage DESIGN.md §2 for details).
1173
-
1174
- **Data flow**:
1175
- ```
1176
- Source file: /app-a/db.json (mtime=1700000000123)
1177
- │
1178
- ├─ sync engine: stat("/app-a/db.json") → mtimeMs=1700000000123
1179
- ├─ sync engine: readFile("/app-a/db.json") → data
1180
- ├─ sync engine: writeFileWithMtimeFallback(target, "/app-a/db.json", data, 1700000000123)
1181
- │ ├─ target has writeFileWithMtime? → YES → target.writeFileWithMtime(path, data, 1700000000123)
1182
- │ │ → backend.writeFile(path, data, { mtime: 1700000000123 })
1183
- │ │ → touch(path, { mtimeMs: 1700000000123 })
1184
- │ └─ target has writeFileWithMtime? → NO → target.writeFile(path, data) [fallback]
1185
- │
1186
- └─ Target file: /app-a/db.json (mtime=1700000000123) ← preserved!
1187
- → Next sync: source.mtimeMs === target.mtimeMs → skip (no spurious copy)
1188
- ```
1189
-
1190
- ## 13. Peer Dependencies
1191
-
1192
- | Package | Role | Version | Required |
1193
- |---|---|---|---|
1194
- | `@zenfs/core` | Virtual file system, backends, VFS, Context | >=2.3.0 | Yes |
1195
- | `@zenfs/dom` | IndexedDB backend (browser) | >=1.0.0 | Yes (browser) |
1196
- | `zen-fs-sync` | Cross-backend sync engine | >=0.1.0 | Yes |
1197
- | `zen-fs-cache` | Read caching with ETag/304 revalidation | >=1.0.0 | No (optional) |
1198
-
1199
- ## 14. Extension Points
1200
-
1201
- ### Custom Serializer
1202
- ```typescript
1203
- import { createConfigRepo, type ConfigSerializer } from 'zen-fs-config';
1204
-
1205
- const yamlSerializer: ConfigSerializer = {
1206
- serialize(data: unknown): Uint8Array { ... },
1207
- deserialize(raw: Uint8Array, path: string): unknown { ... },
1208
- canHandle(path: string): boolean { return path.endsWith('.yaml'); }
1209
- };
1210
- ```
1211
-
1212
- ### Custom Conflict Resolver
1213
- ```typescript
1214
- const repo = await createConfigRepo('app-a', {
1215
- ...
1216
- onConflict: async (conflict) => {
1217
- // Custom conflict resolution logic
1218
- // Return merged content, or null to use default strategy
1219
- return customMerge(conflict.sourceContent, conflict.targetContent);
1220
- }
1221
- });
1222
- ```
1223
-
1224
- ### Custom Backend Registry
1225
- ```typescript
1226
- import { registerBackend } from 'zen-fs-config';
1227
-
1228
- registerBackend('CustomStore', async (options) => {
1229
- const { CustomStoreFS } = await import('custom-store-fs');
1230
- return new CustomStoreFS(options);
1231
- });
1232
- ```
1233
-
1234
- ## 15. License
1235
-
1
+ # zen-fs-config — Design Document
2
+
3
+ > **Architecture direction (REQUIREMENTS §9 D2)**: the core uses a **GenericSyncGroup base with two type implementations** — `config-sync` (versioning/tombstone/conflict + hosts `.meta/app-data-groups/`) and `data-sync` (plain file sync, always governed as an app-data-group under config-sync via `ConfigRepo.createAppDataGroup`); **backend-type management and data-backend info persistence stay in the core `zen-fs-config` (UI is presentation-only)**. This document describes the **current implementation (two classes)**; the target architecture follows D2.
4
+
5
+ ## 1. Overview
6
+
7
+ zen-fs-config is a distributed configuration management library built on top of:
8
+ - **ZenFS** (`@zenfs/core`) — Virtual file system with pluggable backends
9
+ - **zen-fs-cache** — Caching layer with ETag/304 revalidation
10
+ - **zen-fs-sync** — Sync engine for mirroring configs across backends
11
+
12
+ It allows multiple application instances (programs) running on different nodes to share configuration through a network of ZenFS backends, with per-app isolation, shared config spaces, node-local config, and conflict safety.
13
+
14
+ zen-fs-config supports two types of **sync groups**, each backed by a multi-backend sync network but serving different purposes:
15
+
16
+ - **Config-sync group** — Synchronizes the configuration repository itself: backend topology, app configs, shared configs, node-local configs. This is the "meta layer."
17
+ - **Data-sync group** — Synchronizes application data only. A data-sync group is always referenced and managed by a config-sync group as an app's data storage layer (via `ConfigRepo.createAppDataGroup`).
18
+
19
+ A config-sync group can reference one or more data-sync groups per app, allowing apps to store bulk data on separate backends (e.g., a different repo/branch under the same account) while keeping configuration management unified.
20
+
21
+ ## 2. Architecture
22
+
23
+ ### 2.1 Three-Layer Stack
24
+
25
+ ```
26
+ Application code
27
+ ↓ (reads/writes via standard node:fs API)
28
+ ConfigRepo (this library)
29
+ ├─ zen-fs-cache → CachedFileSystem (ETag/TTL read cache)
30
+ ├─ ZenFS VFS → Context-isolated fs per app (chroot)
31
+ └─ zen-fs-sync → Change detection + conflict resolution
32
+ ├─ Backend X (replica)
33
+ ├─ Backend Y (replica)
34
+ └─ Backend Z (replica)
35
+ ```
36
+
37
+ ### 2.2 IndexedDB as Local Primary (Offline-First)
38
+
39
+ Every program instance uses **IndexedDB** as its local primary backend. All config reads and writes target IndexedDB directly, ensuring offline availability and fast local access.
40
+
41
+ User-provided backends (Gitee, S3, RemoteStorage, etc.) are added as **replicas** — they receive bi-directional sync with the local IndexedDB but are never the direct target of config operations.
42
+
43
+ ```
44
+ Program A → Primary = IndexedDB (local), Replicas = [Gitee, S3]
45
+ Program B → Primary = IndexedDB (local), Replicas = [Gitee, S3]
46
+ ```
47
+
48
+ This means:
49
+ - Config is always available offline (IndexedDB persists in the browser)
50
+ - Remote backends accelerate multi-device sync, not local access
51
+ - Re-opening the app requires zero backend parameters — IndexedDB + `.meta/backends/` contain everything needed
52
+
53
+ ### 2.3 Self-Describing Configuration
54
+
55
+ Backend topology and sync rules are stored **inside** the configuration repository (in `.meta/`), not passed as external parameters. This means any node that can read the config repo can bootstrap the entire sync network.
56
+
57
+ External input at startup is limited to: **which backend to connect to** and optionally **bootstrap data** (if the repo doesn't exist yet).
58
+
59
+ ### 2.4 Sync Group Types
60
+
61
+ A **sync group** is a set of backends that synchronize with each other. There are two types:
62
+
63
+ ```
64
+ ┌─────────────────────────────────────────────────────────────────┐
65
+ │ Config-Sync Group │
66
+ │ ┌──────────────────────────────────────────────────────┐ │
67
+ │ │ .meta/backends/ ← config-sync backends │ │
68
+ │ │ .meta/app-data-groups/ ← references to data-sync │ │
69
+ │ │ /{appId}/ ← app config data │ │
70
+ │ │ /shared/ ← shared config │ │
71
+ │ │ /nodes/ ← node-local config │ │
72
+ │ └──────────────────────────────────────────────────────┘ │
73
+ │ │ references │
74
+ │ ▼ │
75
+ │ ┌──────────────────────────────────────────────────────┐ │
76
+ │ │ Data-Sync Group (per-app) │ │
77
+ │ │ .meta/backends/ ← data-sync backends │ │
78
+ │ │ / ← app data files │ │
79
+ │ └──────────────────────────────────────────────────────┘ │
80
+ └─────────────────────────────────────────────────────────────────┘
81
+ ```
82
+
83
+ **Config-sync group**:
84
+ - Synced content: `.meta/` (topology), `/{appId}/` (app config), `/shared/`, `/nodes/`
85
+ - Backends: full credentials + storage location (e.g., Gitee: token + owner + repo + branch)
86
+ - Can reference data-sync groups via `.meta/app-data-groups/{appId}/`
87
+
88
+ **Data-sync group**:
89
+ - Synced content: application data files only (no config meta layer)
90
+ - Backends: full credentials + storage location, but typically reusing the same account as a config-sync backend with a different storage target (e.g., same token/owner, different repo/branch)
91
+ - Can be referenced by a config-sync group as an app's data storage (via `ConfigRepo.createAppDataGroup`)
92
+
93
+ **Group type detection**: When connecting to a backend, the library reads `.meta/group-type` to determine the group type. If the file is absent, the backend is treated as a new empty group and the caller decides which type to create.
94
+
95
+ | `.meta/group-type` | Behavior |
96
+ |---|---|
97
+ | `config-sync` | Full system: IndexedDB primary + config replicas + optional data-sync groups |
98
+ | `data-sync` | Lightweight: data backends only, direct read/write, no config meta layer |
99
+ | absent | New empty backend; caller decides group type at creation time |
100
+
101
+ ## 3. File System Structure
102
+
103
+ ### 3.1 Config-Sync Group
104
+
105
+ ```
106
+ /
107
+ ├─ .meta/ [synced to replicas]
108
+ │ ├─ group-type Group type marker: "config-sync"
109
+ │ ├─ backends/ Backend topology (one file per backend)
110
+ │ │ ├─ local-idb.json { id, type, options, description }
111
+ │ │ ├─ gitee-prod.json
112
+ │ │ └─ ...
113
+ │ ├─ app-data-groups/ References to data-sync groups (per app)
114
+ │ │ └─ {appId}/
115
+ │ │ └─ {dataGroupId}.json { groupType: "data-sync", backends: [...] }
116
+ │ ├─ .deleted/ Tombstones for deletion propagation
117
+ │ │ └─ {encoded-path}.json
118
+ │ └─ .conflicts/ Conflict archives (safekeeping)
119
+ │ └─ {timestamp}_{path}/
120
+ │ ├─ meta.json
121
+ │ ├─ source
122
+ │ └─ target
123
+ │
124
+ ├─ {appId}/ [synced: owner → replicas]
125
+ │ ├─ db.json
126
+ │ ├─ cache.json
127
+ │ └─ .db.json.version Sidecar version file
128
+ │
129
+ ├─ shared/ [synced: bi-directional]
130
+ │ ├─ feature-flags.json
131
+ │ ├─ api-version.json
132
+ │ └─ .feature-flags.json.version
133
+ │
134
+ └─ nodes/ [synced — bidirectional, not node-scoped]
135
+ ├─ {nodeId}/
136
+ │ ├─ local.json Node-local config
137
+ │ └─ env.json
138
+ └─ .node-id Current node's ID (auto-generated)
139
+ ```
140
+
141
+ ### 3.2 Data-Sync Group
142
+
143
+ ```
144
+ /
145
+ ├─ .meta/ [synced to replicas]
146
+ │ ├─ group-type Group type marker: "data-sync"
147
+ │ └─ backends/ Data backend topology (one file per backend)
148
+ │ ├─ gitee-data-1.json { id, type, options, description }
149
+ │ └─ gitee-data-2.json
150
+ │
151
+ └─ (application data files) [synced: bi-directional]
152
+ ├─ documents/
153
+ │ ├─ note-1.json
154
+ │ └─ note-2.json
155
+ └─ media/
156
+ └─ config.json
157
+ ```
158
+
159
+ A data-sync group has a much simpler structure: no version sidecars, no tombstones, no conflict archives — just raw data files and a minimal `.meta/` for backend topology and group type identification.
160
+
161
+ ### 3.3 Directory Semantics (Config-Sync Group)
162
+
163
+ | Directory | Sync Direction | Conflict Risk | Purpose |
164
+ |---|---|---|---|
165
+ | `/{appId}/` | Bi-directional (primary ↔ replicas) | Low (single device) | Per-app private config |
166
+ | `/shared/` | Bi-directional | Possible (multiple writers) | Cross-app shared config |
167
+ | `/nodes/` | None (by default) | None | Per-node local config |
168
+ | `/.meta/` | Bi-directional | None (topology files) | Backend topology, tombstones, conflict archives |
169
+
170
+ ### 3.4 Config-to-File Mapping
171
+
172
+ Each config key maps to one file. The mapping is straightforward:
173
+
174
+ - `setConfig('/db/host', { hostname: 'localhost' })` → writes file `/app-a/db/host.json` with content `{"hostname":"localhost"}`
175
+ - `getConfig('/db/host')` → reads file `/app-a/db/host.json`, parses based on extension
176
+ - Path is relative to the app's root (`/{appId}/`), with `.json` extension appended automatically
177
+ - If path already has an extension (e.g., `/readme.md`), the extension is preserved
178
+
179
+ ### 3.5 Serialization
180
+
181
+ The serializer is determined by file extension:
182
+
183
+ | Extension | Serialize | Deserialize |
184
+ |---|---|---|
185
+ | `.json` (default) | `JSON.stringify` | `JSON.parse` |
186
+ | `.yaml` | YAML dump | YAML parse |
187
+ | `.toml` | TOML dump | TOML parse |
188
+ | `.txt` / no struct extension | `String(data)` | Return as string |
189
+
190
+ Users can inject a custom `ConfigSerializer` for other formats.
191
+
192
+ ## 4. Backend Topology (`.meta/backends/*.json`)
193
+
194
+ Each backend is stored as an individual JSON file in `.meta/backends/`. This allows atomic add/remove operations without rewriting the entire topology.
195
+
196
+ **`.meta/backends/local-idb.json`** (always present):
197
+ ```json
198
+ {
199
+ "id": "local-idb",
200
+ "type": "IndexedDB",
201
+ "options": { "storeName": "zen-fs-config-my-app" },
202
+ "description": "Local IndexedDB primary backend"
203
+ }
204
+ ```
205
+
206
+ **`.meta/backends/gitee-prod.json`** (user-added replica):
207
+ ```json
208
+ {
209
+ "id": "gitee-prod",
210
+ "type": "Gitee",
211
+ "options": { "token": "...", "owner": "...", "repo": "...", "branch": "main" },
212
+ "description": "Production Gitee config repo"
213
+ }
214
+ ```
215
+
216
+ The local IndexedDB backend (`local-idb`) is always the primary — all config operations target it directly. All other backends are replicas with bi-directional sync.
217
+
218
+ **Migration**: If a legacy `.meta/backends.json` file exists (pre-0.4.0), it is automatically migrated to individual files on startup.
219
+
220
+ ## 4.1 App Data Groups (`.meta/app-data-groups/{appId}/`)
221
+
222
+ A config-sync group can reference data-sync groups on behalf of specific apps. Each reference is stored as a JSON file under `.meta/app-data-groups/{appId}/`.
223
+
224
+ **`.meta/app-data-groups/my-app/data-store-1.json`**:
225
+ ```json
226
+ {
227
+ "id": "data-store-1",
228
+ "groupType": "data-sync",
229
+ "backends": [
230
+ {
231
+ "id": "gitee-data",
232
+ "type": "Gitee",
233
+ "options": { "token": "...", "owner": "...", "repo": "my-app-data", "branch": "main" },
234
+ "accountBackendId": "gitee-prod",
235
+ "description": "App data on Gitee (reuses gitee-prod account)"
236
+ }
237
+ ]
238
+ }
239
+ ```
240
+
241
+ ### Account Reuse
242
+
243
+ The `accountBackendId` field (optional) references a config-sync backend whose account fields (e.g., `token`, `owner`, `baseUrl` for Gitee/GitHub) are reused. When present, the data-sync backend's `options` only need to specify the **storage location** fields (e.g., `repo`, `branch`); account fields are merged in from the referenced config-sync backend at creation time.
244
+
245
+ When `accountBackendId` is absent or null, the data-sync backend must provide a complete set of options (including credentials).
246
+
247
+ The `accountFields` metadata for each backend type (registered via `BackendMetadata.accountFields`) determines which fields are "account" vs "storage location":
248
+
249
+ | Backend Type | Account Fields | Storage Location Fields |
250
+ |---|---|---|
251
+ | Gitee | `token`, `owner`, `baseUrl` | `repo`, `branch` |
252
+ | GitHub | `token`, `owner`, `baseUrl` | `repo`, `branch` |
253
+ | WebDAV | `url`, `username`, `password` | `rootPath` |
254
+ | RemoteStorage | `userAddress`, `token` | (none) |
255
+
256
+ ## 5. Sync Rules (`.meta/sync-rules.json`)
257
+
258
+ ```json
259
+ {
260
+ "version": 1,
261
+ "rules": [
262
+ {
263
+ "prefix": "/app-a/",
264
+ "direction": "one-way",
265
+ "conflictStrategy": "source-wins",
266
+ "replicas": ["local-idb", "remote-s3"]
267
+ },
268
+ {
269
+ "prefix": "/app-b/",
270
+ "direction": "one-way",
271
+ "conflictStrategy": "source-wins",
272
+ "replicas": ["local-idb", "remote-s3"]
273
+ },
274
+ {
275
+ "prefix": "/shared/",
276
+ "direction": "bi-directional",
277
+ "conflictStrategy": "merge",
278
+ "replicas": ["local-idb", "remote-s3"]
279
+ },
280
+ {
281
+ "prefix": "/nodes/",
282
+ "direction": "none"
283
+ },
284
+ {
285
+ "prefix": "/.meta/",
286
+ "direction": "none"
287
+ }
288
+ ]
289
+ }
290
+ ```
291
+
292
+ - Private app directories (`/{appId}/`): one-way push, no conflict possible
293
+ - Shared directory (`/shared/`): bi-directional, conflict possible, merge strategy
294
+ - Nodes directory (`/nodes/{nodeId}/`): bi-directional via the main sync pair (not node-scoped — every node's config syncs to every replica)
295
+ - Meta directory (`/.meta/`): bi-directional (topology, tombstones, conflict archives)
296
+
297
+ ## 6. Versioning & Change Detection
298
+
299
+ ### 6.1 Sidecar Version Files
300
+
301
+ Each config file has a companion version file:
302
+
303
+ ```
304
+ /app-a/db.json → Config content
305
+ /app-a/.db.json.version → Version metadata
306
+ ```
307
+
308
+ Version file content:
309
+ ```json
310
+ {
311
+ "version": 5,
312
+ "hash": "sha256:a1b2c3d4...",
313
+ "author": "app-a",
314
+ "timestamp": 1689686400000
315
+ }
316
+ ```
317
+
318
+ ### 6.2 Comparison Logic (extends zen-fs-sync's FileSnapshot)
319
+
320
+ | Condition | Action |
321
+ |---|---|
322
+ | hash same | Skip (content unchanged) |
323
+ | hash different, version different | Higher version wins |
324
+ | hash different, version same | **Conflict** → conflict safety mechanism |
325
+ | version/hash missing | Fall back to mtime+size comparison (backward compat) |
326
+
327
+ ### 6.3 Version Increment
328
+
329
+ On each write:
330
+ 1. Read current version file (if exists)
331
+ 2. Increment version by 1
332
+ 3. Compute SHA-256 hash of new content
333
+ 4. Set author to current instance's `{appId}/{nodeId}`
334
+ 5. Write config file first, then version file
335
+
336
+ Crash recovery: on startup, if hash in version file doesn't match actual file content, auto-increment version and update hash.
337
+
338
+ ## 7. Conflict Safety Mechanism
339
+
340
+ When a conflict is detected (same version, different hash on `/shared/` files):
341
+
342
+ ### 7.1 Archive Both Versions
343
+
344
+ Both conflicting versions are saved to `.meta/.conflicts/` before any resolution:
345
+
346
+ ```
347
+ .meta/.conflicts/1689686400000_shared-feature-flags.from-app-a.to-app-b.json
348
+ ```
349
+
350
+ Archive file content:
351
+ ```json
352
+ {
353
+ "conflictPath": "/shared/feature-flags.json",
354
+ "timestamp": 1689686400000,
355
+ "sourceAuthor": "app-a/server-1",
356
+ "targetAuthor": "app-b/server-2",
357
+ "sourceContent": { "darkMode": true, "newFeature": true },
358
+ "targetContent": { "darkMode": false, "newFeature": false },
359
+ "sourceVersion": 3,
360
+ "targetVersion": 3
361
+ }
362
+ ```
363
+
364
+ ### 7.2 Resolution Strategies
365
+
366
+ After archiving, resolve according to the configured strategy:
367
+
368
+ | Strategy | Behavior |
369
+ |---|---|
370
+ | `source-wins` | Source content overwrites target. Target content archived. |
371
+ | `target-wins` | Target content preserved. Source content archived. |
372
+ | `merge` | JSON deep merge. Both originals archived. Non-JSON falls back to source-wins. |
373
+
374
+ ### 7.3 Event Notification
375
+
376
+ zen-fs-sync emits a `conflict` event with full conflict details. Application can:
377
+ - Accept the auto-resolved result
378
+ - Read `.meta/.conflicts/` archives to manually merge
379
+ - Call `configRepo.resolveConflict(conflictId, mergedContent)` to submit a custom merge
380
+
381
+ **Guarantee**: Neither side's content is ever lost. Recovery is always possible from `.meta/.conflicts/`.
382
+
383
+ ## 8. Node-Local Configuration
384
+
385
+ Some configs are specific to a single node and live under `/nodes/{nodeId}/`. They are synced to backends by the main bidirectional sync pair (there is no `direction: "none"` exclusion).
386
+
387
+ ### 8.1 Storage
388
+
389
+ Node-local configs live under `/nodes/{nodeId}/`. There is no `direction: "none"` exclusion for `/nodes/` — the main bidirectional sync pair (`fullFS` ↔ replica, root `/`, no filter) replicates it to every replica exactly like `/{appId}/` or `/shared/` (sync is not node-scoped).
390
+
391
+ ```
392
+ /nodes/server-1/
393
+ ├─ local.json → { "ip": "10.0.0.1", "cpuCount": 8 }
394
+ └─ env.json → { "NODE_ENV": "production" }
395
+ ```
396
+
397
+ ### 8.2 Node ID Source
398
+
399
+ Priority order:
400
+ 1. Explicit parameter: `createConfigRepo('app-a', { nodeId: 'server-1', ... })`
401
+ 2. Environment variable: `process.env.NODE_ID`
402
+ 3. Auto-generated: random ID written to `/nodes/.node-id` on first startup
403
+
404
+ ### 8.3 API
405
+
406
+ ```typescript
407
+ // Write node-local config (writes local primary; synced to replicas on next poll/flush)
408
+ repo.setNodeConfig('server-1', '/local.json', { ip: '10.0.0.1' });
409
+
410
+ // Read node-local config
411
+ const config = repo.getNodeConfig<{ ip: string }>('server-1', '/local.json');
412
+
413
+ // Publish node config to sync backends (one-time, for debugging)
414
+ const result = await repo.publishNodeConfig('server-1');
415
+ // or publish specific files only:
416
+ const result = await repo.publishNodeConfig('server-1', { paths: ['/local.json'] });
417
+
418
+ // Peek at other nodes' published configs (read-only)
419
+ const otherConfig = repo.peekNodeConfig<{ ip: string }>('server-2', '/local.json');
420
+ ```
421
+
422
+ | API | Write Target | Persisted | Synced | Purpose |
423
+ |---|---|---|---|---|
424
+ | `getConfig` / `setConfig` | CachedFS → auto-sync to replicas | Yes | Yes | Normal config |
425
+ | `getNodeConfig` / `setNodeConfig` | Local primary → synced to replicas via main pair (eventual) | Yes | Yes (all `/nodes/*`, not node-scoped) | Node config, namespaced by nodeId |
426
+ | `publishNodeConfig` | Explicit one-shot push to replicas | Yes | Yes (one-shot) | Force immediate sync of node files |
427
+ | `peekNodeConfig` | Read local (synced-in) copy | N/A | N/A | Read another node's config (after it has synced in) |
428
+
429
+ ## 9. ConfigRepo Interface
430
+
431
+ ```typescript
432
+ interface ConfigRepo {
433
+ /** Application ID (e.g., "app-a") */
434
+ readonly appId: string;
435
+ /** Node ID (e.g., "server-1") */
436
+ readonly nodeId: string;
437
+ /** ZenFS-compatible fs object (node:fs API), context-isolated to own directories */
438
+ readonly fs: typeof import('node:fs');
439
+
440
+ /** Load/reload config from raw string (for initial setup) */
441
+ load(rawConfig: string): Promise<void>;
442
+
443
+ /** Read config value */
444
+ getConfig<T>(path: string): T;
445
+
446
+ /** Write config value (auto-synced) */
447
+ setConfig(path: string, data: any): void;
448
+
449
+ /** Read node-local config */
450
+ getNodeConfig<T>(nodeId: string, path: string): T;
451
+
452
+ /** Write node-local config (writes local primary; synced to replicas via main pair) */
453
+ setNodeConfig(nodeId: string, path: string, data: any): void;
454
+
455
+ /** Publish node-local config to sync backends (one-time, for debugging) */
456
+ publishNodeConfig(nodeId: string, options?: {
457
+ paths?: string[];
458
+ }): Promise<SyncResult>;
459
+
460
+ /** Peek at another node's published config (read-only) */
461
+ peekNodeConfig<T>(nodeId: string, path: string): T;
462
+
463
+ /** Manually flush all pending sync */
464
+ flush(): Promise<SyncResult[]>;
465
+
466
+ /** Get sync status for all sync pairs */
467
+ getSyncStatuses(): Map<string, SyncPairStatus>;
468
+
469
+ /** Resolve a conflict with custom merged content */
470
+ resolveConflict(conflictId: string, mergedContent: any): Promise<void>;
471
+
472
+ /** List conflict archives */
473
+ listConflicts(): Promise<ConflictArchive[]>;
474
+
475
+ /** Read backend topology (aggregated from .meta/backends/*.json) */
476
+ getBackends(): Promise<BackendsMeta | null>;
477
+
478
+ /** Write backend topology (writes each backend as individual file) */
479
+ updateBackends(meta: BackendsMeta): Promise<void>;
480
+
481
+ /** Dynamically add a replica backend */
482
+ addBackend(id: string, type: string, options: Record<string, unknown>, description?: string): Promise<void>;
483
+
484
+ /** Dynamically remove a replica backend */
485
+ removeBackend(id: string): Promise<void>;
486
+
487
+ /** Delete a file with tombstone (propagates deletion to all backends) */
488
+ deleteFile(path: string): Promise<void>;
489
+
490
+ /** Sync .meta/ files to all replicas */
491
+ syncMetaToReplicas(): Promise<void>;
492
+
493
+ // --- App Data Storage (data-sync groups) ---
494
+
495
+ /**
496
+ * Create a data-sync group for this app, referencing a config-sync backend's account.
497
+ * The data-sync group gets its own set of backends (typically reusing an account
498
+ * from a config-sync backend but with different storage location like repo/branch).
499
+ *
500
+ * @param id Data group ID (e.g., "data-store-1")
501
+ * @param backends Array of backend descriptors for the data-sync group.
502
+ * Each can optionally specify `accountBackendId` to reuse credentials.
503
+ */
504
+ createAppDataGroup(
505
+ id: string,
506
+ backends: AppDataBackendDescriptor[],
507
+ ): Promise<AppDataGroup>;
508
+
509
+ /**
510
+ * Get an existing data-sync group for this app.
511
+ * Returns a handle with its own fs for reading/writing data files.
512
+ */
513
+ getAppDataGroup(id: string): Promise<AppDataGroup>;
514
+
515
+ /** List all data-sync groups registered for this app. */
516
+ listAppDataGroups(): Promise<AppDataGroupDescriptor[]>;
517
+
518
+ /** Remove a data-sync group (stops sync, removes descriptor). */
519
+ removeAppDataGroup(id: string): Promise<void>;
520
+
521
+ /** Dispose: stop sync, release resources */
522
+ dispose(): Promise<void>;
523
+ }
524
+
525
+ /**
526
+ * A data-sync group handle. Provides direct file system access
527
+ * to the app's data storage, independent of the config-sync layer.
528
+ */
529
+ interface AppDataGroup {
530
+ readonly groupId: string;
531
+ readonly appId: string;
532
+ /** Direct fs for reading/writing data files (chroot to this group's root) */
533
+ readonly fs: typeof import('node:fs');
534
+ /** Get sync status for this data group's sync pairs */
535
+ getSyncStatuses(): Map<string, SyncPairStatus>;
536
+ /** Manually flush pending sync */
537
+ flush(): Promise<SyncResult[]>;
538
+ /** Stop sync and release resources */
539
+ dispose(): Promise<void>;
540
+ }
541
+
542
+ /** Descriptor for a backend within a data-sync group. */
543
+ interface AppDataBackendDescriptor {
544
+ id: string;
545
+ type: string;
546
+ options: Record<string, unknown>;
547
+ /** Optional: reuse account fields from a config-sync backend */
548
+ accountBackendId?: string;
549
+ description?: string;
550
+ }
551
+ ```
552
+
553
+ ## 10. Initialization
554
+
555
+ The recommended entry point is `connect`, which always anchors on a config-sync repo (the host for data groups) and auto-detects the backend type. The lower-level `createConfigRepo` is also available; data groups are always created via `ConfigRepo.createAppDataGroup` (decision A / T5).
556
+
557
+ ### `connect` (recommended — auto-detect)
558
+
559
+ ```typescript
560
+ import { connect } from 'zen-fs-config';
561
+
562
+ // User provides a backend — connect auto-detects config-sync vs data-sync
563
+ const result = await connect('my-app', {
564
+ backendInfo: {
565
+ type: 'Gitee',
566
+ options: { token: '...', owner: '...', repo: '...', branch: 'main' },
567
+ },
568
+ });
569
+
570
+ // result.groupType → "config-sync" or "data-sync"
571
+ // result.repo → ConfigRepo (always — it hosts the data groups)
572
+ // result.dataGroup → config-managed data group (default group on first launch / data-sync connect)
573
+ ```
574
+
575
+ ### Zero-parameter (offline-first)
576
+
577
+ ```typescript
578
+ import { createConfigRepo } from 'zen-fs-config';
579
+
580
+ // No parameters needed — IndexedDB is always created as primary
581
+ const repo = await createConfigRepo('my-app');
582
+
583
+ // Config is immediately available from IndexedDB
584
+ repo.setConfig('/db/host', { hostname: 'localhost', port: 3306 });
585
+ ```
586
+
587
+ ### With initial replica backend
588
+
589
+ ```typescript
590
+ const repo = await createConfigRepo('my-app', {
591
+ // Optional: provide a remote backend as initial replica
592
+ primaryBackendId: 'gitee-prod',
593
+ backendInfo: {
594
+ type: 'Gitee',
595
+ options: { token: '...', owner: '...', repo: '...', branch: 'main' },
596
+ },
597
+ // Optional: customize IndexedDB store name
598
+ idbStoreName: 'my-app-config',
599
+ // Optional: node ID (auto-detected if not provided)
600
+ nodeId: 'server-1',
601
+ });
602
+
603
+ // Later, add more backends dynamically
604
+ await repo.addBackend('s3-backup', 'S3Bucket', {
605
+ bucket: 'app-config',
606
+ region: 'us-east-1',
607
+ }, 'S3 backup');
608
+
609
+ // Remove a backend
610
+ await repo.removeBackend('gitee-prod');
611
+
612
+ // Cleanup
613
+ await repo.dispose();
614
+ ```
615
+
616
+ ### Re-opening (zero parameters)
617
+
618
+ ```typescript
619
+ // On subsequent opens, just pass appId
620
+ // IndexedDB + .meta/backends/ contain all state
621
+ const repo = await createConfigRepo('my-app');
622
+
623
+ // All previously added backends are automatically reconnected
624
+ const backends = await repo.getBackends();
625
+ // backends.backends = [{ id: 'local-idb', ... }, { id: 's3-backup', ... }]
626
+ ```
627
+
628
+ ### Data-Sync Group via Config-Sync (recommended)
629
+
630
+ > Data groups are always managed by a config-sync repo via `createAppDataGroup` (decision A / T5). The snippet below shows the supported path.
631
+
632
+ ```typescript
633
+ import { connect } from 'zen-fs-config';
634
+
635
+ // connect always returns a config-sync repo; data groups live under it
636
+ const result = await connect('my-app', { nodeId: 'node-1' });
637
+ const repo = result.repo!;
638
+
639
+ // Create a data group and attach a data backend — config-sync stores the
640
+ // data group's backend topology in .meta/app-data-groups/{appId}/{id}.json
641
+ const dataGroup = await repo.createAppDataGroup('notes', [
642
+ { id: 'gitee-data', type: 'Gitee', options: { token: '...', owner: '...', repo: 'my-app-data', branch: 'main' } },
643
+ ]);
644
+
645
+ // Read/write data files directly
646
+ await dataGroup.fs.promises.writeFile('/notes/todo.json', JSON.stringify({ task: 'buy milk' }));
647
+ const data = JSON.parse(await dataGroup.fs.promises.readFile('/notes/todo.json', 'utf-8'));
648
+
649
+ // Add more data backends later (multi-backend sync)
650
+ await dataGroup.addBackend('gitee-backup', 'Gitee', {
651
+ token: '...', owner: '...', repo: 'my-app-data-backup', branch: 'main',
652
+ });
653
+
654
+ // Cleanup
655
+ await dataGroup.dispose();
656
+ await repo.dispose();
657
+ ```
658
+
659
+ ### Config-Sync with App Data Group (account reuse)
660
+
661
+ ```typescript
662
+ const repo = await createConfigRepo('my-app', {
663
+ backendInfo: {
664
+ type: 'Gitee',
665
+ options: { token: '...', owner: '...', repo: 'configs', branch: 'main' },
666
+ },
667
+ });
668
+
669
+ // Create a data-sync group that reuses the config backend's account
670
+ // but stores data in a different repo
671
+ await repo.createAppDataGroup('data-store-1', [
672
+ {
673
+ id: 'gitee-data',
674
+ type: 'Gitee',
675
+ accountBackendId: 'gitee-prod', // reuse token + owner from this config backend
676
+ options: { repo: 'my-app-data', branch: 'main' }, // only storage location
677
+ },
678
+ ]);
679
+
680
+ // Get the data group handle for direct file access
681
+ const dataGroup = await repo.getAppDataGroup('data-store-1');
682
+ await dataGroup.fs.promises.writeFile('/cache.json', '{"key":"value"}');
683
+ ```
684
+
685
+ ## 11. Initialization Flow
686
+
687
+ ### 11.1 Config-Sync Group (`createConfigRepo`)
688
+
689
+ ```
690
+ createConfigRepo('my-app', options?)
691
+ │
692
+ ├─ 1. Create IndexedDB backend (always, ID = 'local-idb')
693
+ │ storeName = options.idbStoreName || `zen-fs-config-${appId}`
694
+ │
695
+ ├─ 2. Ensure /.meta/ directory exists
696
+ │
697
+ ├─ 3. Write /.meta/group-type = "config-sync" (if not exists)
698
+ │
699
+ ├─ 4. Migrate legacy .meta/backends.json → .meta/backends/*.json (if exists)
700
+ │
701
+ ├─ 5. If options.backendInfo provided:
702
+ │ ├─ Generate replica ID (options.primaryBackendId or auto)
703
+ │ ├─ Dedup check: same type + options (stable key) already registered?
704
+ │ └─ Write descriptor to .meta/backends/{replicaId}.json (if not duplicate)
705
+ │
706
+ ├─ 6. Read all backend descriptors from .meta/backends/
707
+ │ └─ Dedup: remove duplicates (same type + options, different ID)
708
+ │ ├─ Delete duplicate files on ALL replicas directly
709
+ │ └─ Create tombstone + delete local file
710
+ │
711
+ ├─ 7. Determine nodeId (explicit parameter > auto-generated)
712
+ │
713
+ ├─ 8. Create final ConfigRepo instance (primary = 'local-idb')
714
+ │
715
+ ├─ 9. setupSync: for each replica backend:
716
+ │ ├─ Create backend instance (e.g., Gitee, RemoteStorage)
717
+ │ ├─ Create SyncPair(IndexedDB ↔ replica, bi-directional)
718
+ │ ├─ Register conflict handler
719
+ │ └─ NOTE: Does NOT call watch() yet (see §11.4 for why)
720
+ │
721
+ ├─ 10. Load config cache from IndexedDB (fast, local-only)
722
+ │
723
+ ├─ 11. initialSyncAndDedup() — only if replicas exist:
724
+ │ ├─ unwatchAll() — safety: clear any stale snapshots
725
+ │ ├─ syncAll() — full bidirectional sync (no cached snapshot
726
+ │ │ → every file is compared, remote-only files
727
+ │ │ are pulled to local)
728
+ │ ├─ readAllBackendDescriptors() — dedup duplicates pulled from remote
729
+ │ │ ├─ Delete dup files on ALL replicas directly
730
+ │ │ └─ Create tombstones for deduped descriptors
731
+ │ ├─ processTombstones() — delete deduped files on all replicas
732
+ │ └─ watchAll() — start monitoring for future changes
733
+ │ (snapshots now reflect the fully synced state)
734
+ │
735
+ ├─ 12. syncMetaToReplicas() — background push of .meta/ changes
736
+ │ (watchers already running, this just speeds up initial propagation)
737
+ │
738
+ └─ 13. Return ConfigRepo instance
739
+ ```
740
+
741
+ ### 11.2 Why "Sync Before Watch" (Critical Design Decision)
742
+
743
+ The sync engine (`zen-fs-sync`) uses **snapshot-based change detection**. When `watch()` is called on a SyncPair, it triggers `buildInitialSnapshots()` which:
744
+
745
+ 1. Builds a snapshot of the source (IndexedDB) — walks all files, records `path`, `size`, `mtimeMs`
746
+ 2. Builds a snapshot of the target (remote backend) — same process
747
+ 3. Caches **separate** snapshots: `prevSrcSnap` (source) and `prevTgtSnap` (target)
748
+
749
+ On the next `syncAll()`, `syncBidirectional()` compares each side's current snapshot with its own previous snapshot independently. If both sides are unchanged → **"unchanged" → skip sync entirely**.
750
+
751
+ **The problem**: `buildInitialSnapshots()` only *reads* file metadata — it does NOT copy any files. So if the remote has files that the local doesn't (e.g., duplicate backend descriptors written by another node), the cached snapshots reflect the un-synced state. The subsequent sync sees "snapshots already match this state" and skips — the file is never actually copied to local, and local-only dedup logic never runs.
752
+
753
+ **The fix**: Always perform a full `syncAll()` **before** `watch()`. With no cached snapshot, `syncBidirectional()` does a complete comparison and copies all missing files. After sync completes, `watch()` builds snapshots from the now-consistent state.
754
+
755
+ This pattern is applied in three places:
756
+ - `createConfigRepo()` → `initialSyncAndDedup()` (sync → dedup → watch)
757
+ - `addBackend()` → `syncMetaToReplicas()` then `watch()` (sync → watch)
758
+ - `AppDataGroupImpl.connect()` → `syncAll()` then `watchAll()` (sync → watch)
759
+
760
+ ### 11.3 `flush()` — Manual Sync Trigger
761
+
762
+ ```
763
+ flush()
764
+ │
765
+ ├─ 1. processTombstones()
766
+ │ For each tombstone in /.meta/.deleted/:
767
+ │ ├─ Delete the actual file on primary (in case re-created)
768
+ │ ├─ Delete the actual file on ALL replicas
769
+ │ └─ Delete version sidecars on all replicas
770
+ │
771
+ ├─ 2. syncAll()
772
+ │ For each SyncPair (IndexedDB ↔ replica):
773
+ │ ├─ Build current snapshots of both sides
774
+ │ ├─ Compare with cached snapshot (if any)
775
+ │ ├─ Detect changes: Created / Modified / Deleted
776
+ │ ├─ Resolve conflicts (source-wins strategy)
777
+ │ └─ Copy files in both directions as needed
778
+ │
779
+ ├─ 3. readAllBackendDescriptors() — post-sync dedup
780
+ │ Sync may have pulled duplicate backend descriptors from remote.
781
+ │ Re-run dedup to catch and remove them.
782
+ │ ├─ Delete dup files on ALL replicas directly
783
+ │ └─ Create tombstones for deduped descriptors
784
+ │
785
+ ├─ 4. processTombstones() — process any new tombstones from step 3
786
+ │
787
+ ├─ 5. updateTombstoneConfirmations()
788
+ │ Mark each tombstone as confirmed by all replica backends
789
+ │
790
+ ├─ 6. gcTombstones()
791
+ │ Remove tombstones confirmed by ALL backends in the topology
792
+ │
793
+ └─ Return SyncResult[] (one per sync pair)
794
+ ```
795
+
796
+ ### 11.4 Tombstone-Based Deletion Propagation
797
+
798
+ When a file is deleted via `deleteFile(path)`:
799
+
800
+ ```
801
+ deleteFile('/.meta/backends/old-backend.json')
802
+ │
803
+ ├─ 1. Write tombstone: /.meta/.deleted/++meta__backends__old-backend++json.json
804
+ │ { path, deletedAt, deletedBy, confirmedBy: [primaryBackendId] }
805
+ │
806
+ ├─ 2. Delete the actual file on primary (IndexedDB)
807
+ │
808
+ └─ 3. Delete version sidecar (.old-backend.json.version) on primary
809
+ ```
810
+
811
+ On the next `processTombstones()` (called by `flush()` or `initialSyncAndDedup()`):
812
+
813
+ ```
814
+ For each tombstone:
815
+ ├─ Delete file on primary (in case sync re-created it)
816
+ ├─ Delete file on ALL replicas
817
+ ├─ Delete version sidecar on ALL replicas
818
+ └─ Tombstone file itself is synced to replicas via syncAll()
819
+ → Late-joining replicas see the tombstone and delete the file
820
+ ```
821
+
822
+ **Why tombstones?** Without them, bi-directional sync treats a deleted local file as "missing → needs to be copied from remote". The tombstone explicitly signals "this file was intentionally deleted" so all replicas honor the deletion. Tombstones are garbage-collected after all backends confirm receipt.
823
+
824
+ ### 11.5 Backend Deduplication
825
+
826
+ When `readAllBackendDescriptors()` detects two backends with the same `type` + `options` (using stable key ordering) but different IDs:
827
+
828
+ ```
829
+ Detected: rs-1 and rs-2 have identical type + options
830
+ │
831
+ ├─ 1. Keep the one with the earliest mtime (created first)
832
+ │
833
+ ├─ 2. For each duplicate:
834
+ │ ├─ Delete descriptor file on ALL replicas directly
835
+ │ │ (prevents sync from pulling it back)
836
+ │ ├─ Delete version sidecar on ALL replicas
837
+ │ └─ Create tombstone + delete local file
838
+ │
839
+ └─ 3. Return deduplicated list (duplicates removed)
840
+ ```
841
+
842
+ The stable key function (`backendDedupKey`) sorts object keys recursively, so `{ token: 'a', owner: 'b' }` and `{ owner: 'b', token: 'a' }` produce the same key and are correctly detected as duplicates.
843
+
844
+ ### 11.6 Dynamic Backend Management
845
+
846
+ **`addBackend(id, type, options)`**:
847
+
848
+ ```
849
+ ├─ 1. Dedup check: reject if same type+options already registered
850
+ ├─ 2. Create backend instance
851
+ ├─ 3. Write descriptor to .meta/backends/{id}.json
852
+ ├─ 4. Create SyncPair (IndexedDB ↔ new replica, bi-directional)
853
+ ├─ 5. syncMetaToReplicas() — full sync FIRST (pull + push)
854
+ └─ 6. watch(pairId) — start monitoring AFTER sync completes
855
+ ```
856
+
857
+ **`removeBackend(id)`**:
858
+
859
+ ```
860
+ ├─ 1. Delete descriptor file on the remote backend DIRECTLY
861
+ │ (must happen before removing sync pair — otherwise can't reach remote)
862
+ ├─ 2. Delete version sidecar on remote
863
+ ├─ 3. Create tombstone + delete local descriptor file
864
+ ├─ 4. Remove sync pair (stops watching + disposes)
865
+ ├─ 5. Remove from replicaBackends map
866
+ ├─ 6. Dispose backend instance
867
+ ├─ 7. processTombstones() — propagate deletion to remaining replicas
868
+ └─ 8. flush() — sync + GC tombstones
869
+ ```
870
+
871
+ ### 11.7 Watch Mode (Auto-Sync)
872
+
873
+ After initialization, each SyncPair runs in **watch mode** with hybrid change detection:
874
+
875
+ ```
876
+ watch() triggers:
877
+ │
878
+ ├─ 1. Register onChange callbacks (if backend supports it)
879
+ │ Local backends (IndexedDB) push change notifications
880
+ │ → triggers debounced sync (default 300ms)
881
+ │
882
+ ├─ 2. buildInitialSnapshots()
883
+ │ ├─ BiDirectional: cache separate source and target snapshots
884
+ │ └─ OneWay: cache source snapshot only
885
+ │
886
+ └─ 3. Start poll timers (if backend supports shouldSync)
887
+ ├─ Remote backends poll shouldSync() every pollIntervalMs (default 30min)
888
+ └─ Fallback: if no onChange and no shouldSync, poll every interval
889
+ ```
890
+
891
+ **State guard**: If `unwatch()` is called during `buildInitialSnapshots()` (which is async), the snapshots are discarded — they won't be cached. This prevents stale snapshots from causing sync skips.
892
+
893
+ **Snapshot comparison** in `syncBidirectional()`:
894
+ 1. Build current snapshots of both sides (via `getSnapshot()`)
895
+ 2. Compare each side independently against its own previous snapshot:
896
+ - `srcChanged = !snapshotsEqual(prevSrcSnap, currentSrcSnap)`
897
+ - `tgtChanged = !snapshotsEqual(prevTgtSnap, currentTgtSnap)`
898
+ - If neither changed and both previous snapshots exist → skip sync entirely
899
+ 3. Cache current snapshots as `prevSrcSnap` and `prevTgtSnap` for next comparison
900
+ 4. If either side changed, proceed with full diff and file operations
901
+
902
+ **Key difference from previous design**: The old approach merged source and target snapshots into a single map (`source ∪ target`), which lost information about which filesystem a file belonged to. The new approach keeps them separate, enabling precise per-side change detection and bidirectional deletion propagation (see §11.10).
903
+
904
+ ### 11.8 Standalone Data-Sync Entry — REMOVED
905
+
906
+ The standalone data-sync entry has been **removed** from the public API (decision A / T5). Data groups are created exclusively via `ConfigRepo.createAppDataGroup` and are always owned by a config-sync repo, whose `.meta/app-data-groups/{appId}/{id}.json` is the authoritative source of the data group's backend topology.
907
+
908
+ ### 11.9 Unified Entry Point (`connect`)
909
+
910
+ `createConfigRepo` is the lower-level factory for a config-sync repo. The recommended entry point is `connect`, which always anchors on a config-sync repo and dispatches the backend as follows:
911
+
912
+ ```
913
+ connect('my-app', options?)
914
+ │
915
+ ├─ 1. Determine config-sync repo (create or reuse) — this always hosts the data groups
916
+ │
917
+ ├─ 2. If options.backendInfo provided, read /.meta/group-type
918
+ │
919
+ ├─ "config-sync" → connect the repo to that backend; then load every app
920
+ │ data group from .meta/app-data-groups/{appId}/*.json (UC2)
921
+ │ return { groupType: "config-sync", repo, appDataGroups }
922
+ │
923
+ ├─ "data-sync" → keep the config repo local-only; ensure the default data
924
+ │ group exists and attach the backend to it (UC3). The backend
925
+ │ info is written back into app-data-groups by addBackend.
926
+ │ return { groupType: "data-sync", repo, dataGroup }
927
+ │
928
+ └─ absent (no backendInfo) → local-only config repo + a default data group
929
+ └─ return { groupType: "config-sync" (or options.groupType), repo, dataGroup }
930
+ ```
931
+
932
+ **Usage**:
933
+
934
+ ```typescript
935
+ import { connect } from 'zen-fs-config';
936
+
937
+ // Auto-detect: connects to backend, reads group-type, dispatches accordingly
938
+ const result = await connect('my-app', {
939
+ backendInfo: {
940
+ type: 'Gitee',
941
+ options: { token: '...', owner: '...', repo: '...', branch: 'main' },
942
+ },
943
+ });
944
+
945
+ if (result.groupType === 'config-sync') {
946
+ // result.repo is a ConfigRepo — full config system
947
+ const repo = result.repo;
948
+ repo.setConfig('/db/host', { hostname: 'localhost' });
949
+ } else {
950
+ // result.dataGroup is a config-managed data group (result.repo is its host)
951
+ const dataGroup = result.dataGroup;
952
+ await dataGroup.fs.promises.writeFile('/data.json', '{"key":"value"}');
953
+ }
954
+
955
+ // Explicit override (skip detection, force a specific group type)
956
+ const result = await connect('my-app', {
957
+ backendInfo: { type: 'Gitee', options: {...} },
958
+ groupType: 'data-sync', // force data-sync even if backend has no group-type yet
959
+ });
960
+ ```
961
+
962
+ **Return type**:
963
+
964
+ ```typescript
965
+ interface ConnectResult {
966
+ /** Detected or forced group type */
967
+ groupType: 'config-sync' | 'data-sync';
968
+ /** The config-sync repo — always present; it hosts the data groups */
969
+ repo?: ConfigRepo;
970
+ /** The default app data group (config-managed); present on first launch / data-sync connect */
971
+ dataGroup?: AppDataGroup;
972
+ /** All app data groups loaded for this app (may be empty) */
973
+ appDataGroups?: AppDataGroup[];
974
+ }
975
+ ```
976
+
977
+ **Offline / zero-parameter mode**: When no `backendInfo` is provided, `connect` defaults to `config-sync`, creates an IndexedDB-only config repo, and also creates a default app data group under it via `createAppDataGroup` (same as `createConfigRepo` with no options, plus the default data group).
978
+
979
+ ### 11.10 Snapshot Optimization Design
980
+
981
+ This section describes three interrelated optimizations to the sync engine's snapshot mechanism.
982
+
983
+ #### 11.10.1 FS-Provided `createSnapshot()`
984
+
985
+ The `SyncableFS` interface now includes an optional `createSnapshot()` method:
986
+
987
+ ```typescript
988
+ interface SyncableFS {
989
+ // ... existing methods ...
990
+
991
+ /**
992
+ * Optional: Build a filesystem snapshot.
993
+ * Returns a map of relative path → {size, mtimeMs} for all files under root.
994
+ * Returns null if the filesystem is unreachable.
995
+ *
996
+ * Backends that can provide a more efficient snapshot than the generic
997
+ * walkFiles+stat approach should implement this method.
998
+ */
999
+ createSnapshot?(root: string, filter?: SyncFilter): Promise<Map<string, FileSnapshot> | null>;
1000
+ }
1001
+ ```
1002
+
1003
+ The sync engine's `getSnapshot()` helper dispatches to the FS-provided method when available, falling back to the generic `buildSnapshot()` (walkFiles + stat) otherwise:
1004
+
1005
+ ```typescript
1006
+ private async getSnapshot(fs: SyncableFS): Promise<Map<string, FileSnapshot> | null> {
1007
+ if (fs.createSnapshot) {
1008
+ return fs.createSnapshot(this.root, this.options.filter);
1009
+ }
1010
+ return buildSnapshot(fs, this.root, this.options.filter);
1011
+ }
1012
+ ```
1013
+
1014
+ **Optimization examples**:
1015
+ - **Gitee/GitHub**: Use Git tree API to fetch all file metadata in a single request instead of walking files one by one
1016
+ - **IndexedDB**: Use `getAll()` for batch querying instead of individual `stat()` calls
1017
+ - **InMemory**: Directly iterate the internal Map (no async I/O overhead)
1018
+
1019
+ Backends that do not implement `createSnapshot()` are fully supported — the generic fallback produces identical results.
1020
+
1021
+ #### 11.10.2 Separate Source and Target Snapshots
1022
+
1023
+ **Previous design** (merged snapshots):
1024
+ - `buildInitialSnapshots()` merged source and target into a single map: `sourceSnapshots = new Map([...srcSnap, ...tgtSnap])`
1025
+ - `syncBidirectional()` compared `currentMerged` with the cached merged snapshot
1026
+ - **Problem**: The merged map lost which filesystem a file belonged to. A file present on target but not source could be "new on target" or "deleted from source" — the merged snapshot couldn't distinguish.
1027
+
1028
+ **New design** (separate snapshots):
1029
+ - `buildInitialSnapshots()` caches two independent maps: `prevSrcSnap` and `prevTgtSnap`
1030
+ - `syncBidirectional()` compares each side independently:
1031
+ ```
1032
+ srcChanged = !snapshotsEqual(prevSrcSnap, currentSrcSnap)
1033
+ tgtChanged = !snapshotsEqual(prevTgtSnap, currentTgtSnap)
1034
+ if (!srcChanged && !tgtChanged && prevSrcSnap && prevTgtSnap) → skip sync
1035
+ ```
1036
+ - Each side's change is detected independently, preserving file-location information
1037
+
1038
+ #### 11.10.3 Bidirectional Deletion Propagation
1039
+
1040
+ When a file exists on one side but not the other, the sync engine uses previous snapshots to distinguish "created" from "deleted":
1041
+
1042
+ ```
1043
+ File on target, not on source:
1044
+ ├─ Was it on source in the previous snapshot (prevSrcSnap)?
1045
+ │ ├─ Yes → file was deleted from source → delete from target too (propagate deletion)
1046
+ │ └─ No → file was created on target → copy to source
1047
+
1048
+ File on source, not on target:
1049
+ ├─ Was it on target in the previous snapshot (prevTgtSnap)?
1050
+ │ ├─ Yes → file was deleted from target → delete from source too (propagate deletion)
1051
+ │ └─ No → file was created on source → copy to target
1052
+ ```
1053
+
1054
+ **Without this mechanism**, deleting a file on one side would cause the sync engine to see "the other side still has it → copy it back", effectively undoing the deletion.
1055
+
1056
+ **Relationship with tombstones**: The tombstone mechanism (§11.4) and bidirectional deletion propagation operate at different layers and complement each other:
1057
+
1058
+ | Mechanism | Layer | Trigger | How It Works |
1059
+ |-----------|-------|---------|--------------|
1060
+ | Tombstone | `zen-fs-config` (application) | Application calls `deleteFile()` | Writes a `.meta/.deleted/` marker, physically deletes file on all replicas **before** sync runs |
1061
+ | Deletion propagation | `zen-fs-sync` (engine) | Sync detects one-side-only file | Compares with previous snapshot to determine if file was created or deleted |
1062
+
1063
+ Tombstones handle application-initiated deletions (the common case). Deletion propagation handles deletions that bypass the tombstone flow — e.g., external modifications on the remote backend, or files removed by other sync mechanisms.
1064
+
1065
+ #### 11.10.4 `shouldSync()` vs Snapshot Comparison
1066
+
1067
+ These two mechanisms are complementary, not interchangeable:
1068
+
1069
+ | Mechanism | Purpose | Cost | When Used |
1070
+ |-----------|---------|------|-----------|
1071
+ | `shouldSync()` | Fast "has anything changed?" boolean | O(1) for remote (ETag/commit check) | `onRemotePoll()` — decide whether to trigger sync at all |
1072
+ | Snapshot comparison | "What exactly changed?" detail | O(n) filesystem traversal | `syncBidirectional()` — decide what to copy/delete |
1073
+
1074
+ `shouldSync()` is **not** part of the snapshot comparison because:
1075
+ 1. `shouldSync()` updates its internal baseline after each call — calling it again during sync would return stale results
1076
+ 2. `shouldSync()` returning false doesn't guarantee the snapshot is unchanged — it means the FS's own change detection says nothing changed, which could miss edge cases
1077
+ 3. `shouldSync()` returning true doesn't tell us **which** files changed — snapshots are still needed for that
1078
+
1079
+ The two-level optimization works as follows:
1080
+ 1. **Level 1**: `shouldSync()` in remote poll → if false, skip sync trigger entirely (saves the O(n) snapshot build)
1081
+ 2. **Level 2**: Separate snapshot comparison in sync → if both sides unchanged, skip file operations (saves I/O)
1082
+
1083
+ ## 12. Data Flow
1084
+
1085
+ ### Read Path
1086
+ ```
1087
+ Application
1088
+ → repo.fs.readFileSync('/db/host.json')
1089
+ → ZenFS Context (chroot to /app-a/)
1090
+ → CachedFileSystem.readFile('/app-a/db/host.json')
1091
+ → Cache hit (TTL)? → return cached bytes (0 network)
1092
+ → Cache miss/expired? → 304 revalidate with primary backend
1093
+ → Deserialize (JSON.parse for .json files)
1094
+ → Return typed object
1095
+ ```
1096
+
1097
+ ### Write Path (auto-synced)
1098
+ ```
1099
+ Application
1100
+ → repo.setConfig('/db/host', { hostname: 'localhost' })
1101
+ → Serialize (JSON.stringify)
1102
+ → Write config file: /app-a/db/host.json
1103
+ → Write version file: /app-a/.db.host.json.version (version++, new hash)
1104
+ → CachedFileSystem.writeFile() →穿透 to primary backend → invalidate cache
1105
+ → zen-fs-sync watch detects change (poll + debounce)
1106
+ → Sync to replicas per sync-rules
1107
+ ```
1108
+
1109
+ ### Write Path (node-local, synced)
1110
+ ```
1111
+ Application
1112
+ → repo.setNodeConfig('server-1', '/local.json', { ip: '10.0.0.1' })
1113
+ → Serialize + write to /nodes/server-1/local.json (local primary)
1114
+ → Main bidirectional sync pair replicates /nodes/ to replicas on next poll/flush
1115
+ → File is also present on other replicas (sync is not node-scoped)
1116
+ ```
1117
+
1118
+ ### Publish (one-time sync)
1119
+ ```
1120
+ Application
1121
+ → repo.publishNodeConfig('server-1')
1122
+ → Read /nodes/server-1/**/*
1123
+ → Create temporary SyncPair with filter: includePrefixes: ['/nodes/server-1/']
1124
+ → Execute one sync() call
1125
+ → Files pushed to replicas
1126
+ → Dispose temporary SyncPair
1127
+ ```
1128
+
1129
+ ### Mtime Preservation During Sync
1130
+
1131
+ **Problem**: When the sync engine copies a file from source to target, it calls `writeFile(path, data)`. The target backend sets its own mtime (typically `Date.now()`), losing the source file's original mtime. This causes the next sync cycle to detect a "modified" file (source mtime ≠ target mtime), triggering unnecessary copies on every sync.
1132
+
1133
+ **Solution**: An optional `writeFileWithMtime` method on the `SyncableFS` interface, with automatic fallback to `writeFile` when not implemented:
1134
+
1135
+ ```typescript
1136
+ interface SyncableFS {
1137
+ // ... existing methods ...
1138
+
1139
+ /**
1140
+ * Optional: write file with precise mtime.
1141
+ * If implemented, the sync engine uses this instead of writeFile,
1142
+ * passing the source file's mtime so the target can preserve it.
1143
+ * Backends that don't support precise mtime should not implement this —
1144
+ * the sync engine falls back to plain writeFile.
1145
+ */
1146
+ writeFileWithMtime?(path: string, data: string | Uint8Array, mtime: number): Promise<void>;
1147
+ }
1148
+ ```
1149
+
1150
+ **Sync engine (`zen-fs-sync`)**: A central helper function handles the fallback:
1151
+
1152
+ ```javascript
1153
+ async function writeFileWithMtimeFallback(fs, path, data, mtimeMs) {
1154
+ if (mtimeMs !== undefined && typeof fs.writeFileWithMtime === "function") {
1155
+ await fs.writeFileWithMtime(path, data, mtimeMs);
1156
+ } else {
1157
+ await fs.writeFile(path, data);
1158
+ }
1159
+ }
1160
+ ```
1161
+
1162
+ This helper is used in `copyFile()`, `syncOneWay()`, and `writeFileBoth()`. The source file's mtime is obtained via `stat()` before writing, then passed through to the target.
1163
+
1164
+ **Adapters (`zen-fs-config`)**: All three `SyncableFS` adapters implement `writeFileWithMtime`:
1165
+
1166
+ | Adapter | Implementation |
1167
+ |---|---|
1168
+ | `backendToSyncableFS` | Passes `{ mtime }` as options to `backend.writeFile()` — the backend's `writeFile` calls `touch()` with the provided mtime |
1169
+ | `zenfsPromisesToSyncableFS` | Calls `promises.writeFile()` then `promises.utimes()` as a fallback (some VFS backends don't support mtime in writeFile) |
1170
+ | `cachedFSToSyncableFS` | Passes `{ mtime }` as options to `cached.writeFile()` — mtime flows through to the underlying backend |
1171
+
1172
+ **RemoteStorage backend (`zen-fs-remotestoragejs`)**: `writeFileWithMtime` delegates to `writeFile(path, data, { mtime })`, which writes the `.mtime` sidecar file to preserve millisecond-precision mtime (see RemoteStorage DESIGN.md §2 for details).
1173
+
1174
+ **Data flow**:
1175
+ ```
1176
+ Source file: /app-a/db.json (mtime=1700000000123)
1177
+ │
1178
+ ├─ sync engine: stat("/app-a/db.json") → mtimeMs=1700000000123
1179
+ ├─ sync engine: readFile("/app-a/db.json") → data
1180
+ ├─ sync engine: writeFileWithMtimeFallback(target, "/app-a/db.json", data, 1700000000123)
1181
+ │ ├─ target has writeFileWithMtime? → YES → target.writeFileWithMtime(path, data, 1700000000123)
1182
+ │ │ → backend.writeFile(path, data, { mtime: 1700000000123 })
1183
+ │ │ → touch(path, { mtimeMs: 1700000000123 })
1184
+ │ └─ target has writeFileWithMtime? → NO → target.writeFile(path, data) [fallback]
1185
+ │
1186
+ └─ Target file: /app-a/db.json (mtime=1700000000123) ← preserved!
1187
+ → Next sync: source.mtimeMs === target.mtimeMs → skip (no spurious copy)
1188
+ ```
1189
+
1190
+ ## 13. Peer Dependencies
1191
+
1192
+ | Package | Role | Version | Required |
1193
+ |---|---|---|---|
1194
+ | `@zenfs/core` | Virtual file system, backends, VFS, Context | >=2.3.0 | Yes |
1195
+ | `@zenfs/dom` | IndexedDB backend (browser) | >=1.0.0 | Yes (browser) |
1196
+ | `zen-fs-sync` | Cross-backend sync engine | >=0.1.0 | Yes |
1197
+ | `zen-fs-cache` | Read caching with ETag/304 revalidation | >=1.0.0 | No (optional) |
1198
+
1199
+ ## 14. Extension Points
1200
+
1201
+ ### Custom Serializer
1202
+ ```typescript
1203
+ import { createConfigRepo, type ConfigSerializer } from 'zen-fs-config';
1204
+
1205
+ const yamlSerializer: ConfigSerializer = {
1206
+ serialize(data: unknown): Uint8Array { ... },
1207
+ deserialize(raw: Uint8Array, path: string): unknown { ... },
1208
+ canHandle(path: string): boolean { return path.endsWith('.yaml'); }
1209
+ };
1210
+ ```
1211
+
1212
+ ### Custom Conflict Resolver
1213
+ ```typescript
1214
+ const repo = await createConfigRepo('app-a', {
1215
+ ...
1216
+ onConflict: async (conflict) => {
1217
+ // Custom conflict resolution logic
1218
+ // Return merged content, or null to use default strategy
1219
+ return customMerge(conflict.sourceContent, conflict.targetContent);
1220
+ }
1221
+ });
1222
+ ```
1223
+
1224
+ ### Custom Backend Registry
1225
+ ```typescript
1226
+ import { registerBackend } from 'zen-fs-config';
1227
+
1228
+ registerBackend('CustomStore', async (options) => {
1229
+ const { CustomStoreFS } = await import('custom-store-fs');
1230
+ return new CustomStoreFS(options);
1231
+ });
1232
+ ```
1233
+
1234
+ ## 15. License
1235
+
1236
1236
  MIT