api-dock 0.7.0__tar.gz → 0.8.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. {api_dock-0.7.0 → api_dock-0.8.0}/PKG-INFO +339 -1
  2. {api_dock-0.7.0 → api_dock-0.8.0}/README.md +338 -0
  3. api_dock-0.8.0/api_dock/database_config.py +972 -0
  4. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/example_api_dock_config/config.yaml +8 -0
  5. api_dock-0.8.0/api_dock/example_api_dock_config/databases/config.yaml +82 -0
  6. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/example_api_dock_config/databases/example_db.yaml +14 -0
  7. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/fast_api.py +33 -0
  8. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/flask_api.py +31 -0
  9. api_dock-0.8.0/api_dock/listings.py +346 -0
  10. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/route_mapper.py +60 -13
  11. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/sql_builder.py +498 -14
  12. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/storage_auth.py +133 -36
  13. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/types.py +59 -2
  14. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock.egg-info/PKG-INFO +339 -1
  15. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock.egg-info/SOURCES.txt +5 -0
  16. {api_dock-0.7.0 → api_dock-0.8.0}/pyproject.toml +1 -1
  17. api_dock-0.8.0/tests/test_listings.py +221 -0
  18. api_dock-0.8.0/tests/test_shared_database_config.py +835 -0
  19. api_dock-0.8.0/tests/test_sql_selector.py +277 -0
  20. api_dock-0.7.0/api_dock/database_config.py +0 -437
  21. {api_dock-0.7.0 → api_dock-0.8.0}/LICENSE.md +0 -0
  22. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/__init__.py +0 -0
  23. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/auth.py +0 -0
  24. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/cli.py +0 -0
  25. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/config.py +0 -0
  26. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/config_discovery.py +0 -0
  27. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/encryption.py +0 -0
  28. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock/example_api_dock_config/remotes/example_remote.yaml +0 -0
  29. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock.egg-info/dependency_links.txt +0 -0
  30. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock.egg-info/entry_points.txt +0 -0
  31. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock.egg-info/requires.txt +0 -0
  32. {api_dock-0.7.0 → api_dock-0.8.0}/api_dock.egg-info/top_level.txt +0 -0
  33. {api_dock-0.7.0 → api_dock-0.8.0}/setup.cfg +0 -0
  34. {api_dock-0.7.0 → api_dock-0.8.0}/tests/test_inject_cookies.py +0 -0
  35. {api_dock-0.7.0 → api_dock-0.8.0}/tests/test_proxy_pipeline.py +0 -0
  36. {api_dock-0.7.0 → api_dock-0.8.0}/tests/test_sql_builder.py +0 -0
  37. {api_dock-0.7.0 → api_dock-0.8.0}/tests/test_types.py +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: api_dock
3
- Version: 0.7.0
3
+ Version: 0.8.0
4
4
  Summary: A flexible API gateway that allows you to proxy requests to multiple remote APIs and Databases
5
5
  Author-email: Brookie Guzder-Williams <bguzder-williams@berkeley.edu>
6
6
  License-Expression: BSD-3-Clause
@@ -95,6 +95,7 @@ Here is an example:
95
95
  api_dock_config
96
96
  ├── config.yaml # The default main-config file
97
97
  ├── databases
98
+ │ ├── config.yaml # (optional) shared tables/schemas for all databases
98
99
  │ ├── unversioned_db.yaml # Database config without versioning
99
100
  │ └── versioned_db # Folder containing database configs for different versions
100
101
  │ ├── 0.1.yaml
@@ -231,6 +232,58 @@ The optional `settings` section controls HTTP behavior:
231
232
 
232
233
  - **`timeout`** (default: `10`): Upstream request timeout in seconds, applied to both the streaming and buffered proxy paths. Raise it for slow upstreams (e.g. large aggregation queries) that would otherwise return a 502 on timeout. Set to `null` or `false` to disable the timeout entirely (not recommended — a stalled upstream can hold the connection open indefinitely).
233
234
 
235
+ ### Catalog Endpoints (`expose`)
236
+
237
+ The optional `expose` section adds read-only endpoints that list the models and versions of your configured databases, remotes, or both ("sources"). Listings are **opt-in** — with no `expose` key nothing is added.
238
+
239
+ ```yaml
240
+ # Enable all three defaults: /databases, /remotes, /sources
241
+ expose: true
242
+ ```
243
+
244
+ ```yaml
245
+ GET /databases
246
+ # [{"model": "birdnet", "version": "2.4"},
247
+ # {"model": "birdnet", "version": "3.0"},
248
+ # {"model": "owl", "version": "0.5"}]
249
+
250
+ GET /sources # databases + remotes, combined
251
+ # [{"model": "birdnet", "version": "2.4"}, ..., {"model": "core", "version": "0.5.0"}]
252
+ ```
253
+
254
+ Versions are the config filename stems, so semver like `0.5.0` is preserved; unversioned sources report `version: null`. Enable only what you want, and control the output shape with `dict`:
255
+
256
+ ```yaml
257
+ expose:
258
+ dict: false # return "model/version" strings instead of {model, version} dicts
259
+ databases: true # add /databases
260
+ remotes: true # add /remotes
261
+ # sources omitted → not added
262
+
263
+ # GET /databases -> ["birdnet/2.4", "birdnet/3.0", "owl/0.5"]
264
+ ```
265
+
266
+ Each of `databases` / `remotes` / `sources` accepts several forms:
267
+
268
+ ```yaml
269
+ expose:
270
+ databases: false # do not add the endpoint
271
+
272
+ remotes: "list/remotes" # custom route path (same as true, but at /list/remotes)
273
+
274
+ databases: # explicit list of models (optionally version-filtered)
275
+ - birdnet
276
+ - owl:
277
+ versions: [4.0, 5.0] # ints/floats/strings all match the "4.0"/"5.0" stems
278
+
279
+ sources: # most explicit form
280
+ route: "list/sources"
281
+ include: [birdnet, core] # true (default) | false | list of models
282
+ dict: true # per-endpoint override of the top-level `dict`
283
+ ```
284
+
285
+ Each listing route is served both with and without a trailing slash (e.g. `/sources` and `/sources/` both work), so it doesn't matter which convention your clients use. Custom multi-segment routes (e.g. `list/databases`) take precedence over the `/{remote}/{path}` proxy. If a listing route would shadow a configured remote/database, or an `include` names something that doesn't exist, API Dock emits a startup warning. The exposed routes are also reflected in the root (`/`) metadata's `endpoints`.
286
+
234
287
  ---
235
288
 
236
289
  ## Remote Configurations
@@ -356,6 +409,7 @@ Database configurations are stored in `api_dock_config/databases/` directory. Ea
356
409
  - **tables**: Mapping of table names to file paths (supports S3, GCS, HTTPS, local paths)
357
410
  - **queries**: Named SQL queries for reuse
358
411
  - **routes**: REST endpoints mapped to SQL queries
412
+ - **schema** (optional): the shared schema (from `databases/config.yaml`) this config's `[[table]]` references fall back to. See [Shared Tables and Schemas](#shared-tables-and-schemas-databasesconfigyaml)
359
413
 
360
414
  ### Syntax
361
415
 
@@ -433,6 +487,171 @@ routes:
433
487
  sql: "[[get_permissions]]"
434
488
  ```
435
489
 
490
+ ### Shared Tables and Schemas (`databases/config.yaml`)
491
+
492
+ When several databases or versions read the same tables, define them once in the optional `api_dock_config/databases/config.yaml`. Everything lives under a `database` key: `meta` holds default table metadata, `schema` holds named groups of tables, and every other key is a global table.
493
+
494
+ ```yaml
495
+ # api_dock_config/databases/config.yaml
496
+ database:
497
+ # global tables, available as [[table1]] in any database config
498
+ table1:
499
+ uri: s3://your-bucket/table1.parquet
500
+ table3:
501
+ uri: s3://your-other-bucket/table3.parquet
502
+ region: us-west-1 # a table's own keys override `meta`
503
+ public: false
504
+
505
+ # defaults applied to every table (shared tables and the tables in each version config)
506
+ meta:
507
+ region: us-west-2
508
+ public: true
509
+
510
+ # schemas, available as [[birdnet_2p4.detections]] in any route
511
+ schema:
512
+ birdnet_2p4:
513
+ detections:
514
+ uri: s3://your-bucket/birdnet/2.4/detections.parquet
515
+ birdnet_3p0:
516
+ detections:
517
+ uri: s3://your-bucket/birdnet/3.0/detections.parquet
518
+ public: false
519
+ ```
520
+
521
+ A version config can name the schema it uses with `schema:`. An unqualified `[[name]]` is then looked up in order, first match wins:
522
+
523
+ 1. the version config's own `tables`
524
+ 2. its `schema:` in the shared config
525
+ 3. the shared config's global tables
526
+
527
+ ```yaml
528
+ # api_dock_config/databases/birdnet/2.4.yaml
529
+ name: birdnet
530
+ schema: birdnet_2p4
531
+ tables:
532
+ revisions: s3://your-bucket/birdnet/2.4/revisions.parquet # local to this version
533
+
534
+ routes:
535
+ # [[detections]] isn't in `tables`, so it comes from the birdnet_2p4 schema
536
+ - route: recordings/{{recording_id}}/detections
537
+ sql: SELECT [[detections]].* FROM [[detections]] WHERE [[detections]].recording_id = {{recording_id}}
538
+ ```
539
+
540
+ Any route can reference any schema directly as `[[schema.table]]`, so a different database (e.g. `owl/5.0`) can query `[[birdnet_2p4.detections]]`. To keep a table private to one database/version, define it in that version's `tables` instead. Qualified references are exposed to DuckDB as real views, so you can also use the full name or your own alias in plain SQL:
541
+
542
+ ```yaml
543
+ - route: detections/
544
+ sql: >
545
+ SELECT detections.common_name, COUNT(revisions.id) AS revcount
546
+ FROM [[birdnet_2p4.detections]]
547
+ LEFT JOIN [[revisions]] ON revisions.observation_id = birdnet_2p4.detections.id
548
+ GROUP BY birdnet_2p4.detections.common_name
549
+ ```
550
+
551
+ expands to
552
+
553
+ ```sql
554
+ SELECT detections.common_name, COUNT(revisions.id) AS revcount
555
+ FROM birdnet_2p4.detections
556
+ LEFT JOIN 's3://your-bucket/birdnet/2.4/revisions.parquet' AS revisions ON revisions.observation_id = birdnet_2p4.detections.id
557
+ GROUP BY birdnet_2p4.detections.common_name
558
+ ```
559
+
560
+ Notes:
561
+ - After `FROM`/`JOIN`, `[[schema.table]]` becomes `schema.table` with no alias, so `FROM [[birdnet_2p4.detections]] o` works. Elsewhere it becomes the bare table name (`detections`), because DuckDB doesn't accept `schema.table.*`.
562
+ - Schema and table names used as `[[schema.table]]` must be plain identifiers (letters, digits, underscores).
563
+ - Storage credentials are set per table. Tables whose `region`/`public` differ from the rest get their own S3 secret scoped to their path, so one query can mix regions and public/private buckets.
564
+ - Views are created only for the `[[schema.table]]` tables a query actually references.
565
+
566
+ #### Inline database configs (`slugs`)
567
+
568
+ Simple database/version configs (often just a description and a `schema`) can live in the shared file instead of in their own files. Config files keep working, and the two can be mixed, even for the same database:
569
+
570
+ ```yaml
571
+ # api_dock_config/databases/config.yaml
572
+ slugs:
573
+ - name: birdnet-bullfrog # the database slug in the URL
574
+ version: "2.5" # one version...
575
+ description: American Bullfrog Classifier from Birdnet 2.4
576
+ schema: birdnet_bullfrog_2p5v0p5
577
+ - name: birdnet-apple
578
+ authors: [API Team] # ...or several; keys here are defaults for each version
579
+ versions:
580
+ - version: "1.0"
581
+ description: Apple Classifier 1.0
582
+ schema: birdnet_apple_1p0
583
+ - version: "12.0"
584
+ description: Apple Classifier 12.0
585
+ schema: birdnet_apple_12p0
586
+ - name: notes # no version/versions = an unversioned database
587
+ tables:
588
+ notes: s3://your-bucket/notes.parquet
589
+ ```
590
+
591
+ - Each entry (or each `versions` item) takes the same keys as a database config file: `description`, `authors`, `schema`, `tables`, `routes`, `query_params`, and so on. Shared `routes`/`query_params` (below), including `include`/`exclude`, apply to them like any other database/version.
592
+ - Like file-based databases, a slug is only served if it's listed under `databases:` in the main `config.yaml`.
593
+ - A database's versions are the union of its version files and its `slugs` versions, so `latest`, the `/{database}` versions listing, and the `expose` catalog endpoints all see both. If a file and a slug define the same database/version, the file wins.
594
+ - Quote versions (`version: "2.10"`). Unquoted YAML numbers are floats, so `2.10` would become `"2.1"`.
595
+ - A malformed `slugs` section (missing `name`, both `version` and `versions`, a duplicate version, or a mix of versioned and unversioned entries for one name) returns a 500 "Shared database configuration error".
596
+
597
+ #### Shared routes and query params
598
+
599
+ The shared file can also define top-level `routes` and `query_params`. These are added to **every** database/version, which is handy when each model/version serves the same endpoints over its own `schema`:
600
+
601
+ ```yaml
602
+ # api_dock_config/databases/config.yaml
603
+ database:
604
+ ...
605
+
606
+ routes:
607
+ - route: recordings/{{recording_id}}/detections/
608
+ sql: SELECT [[detections]].* FROM [[detections]] WHERE [[detections]].recording_id = {{recording_id}}
609
+ - route: detections/{{id}}
610
+ sql: SELECT [[detections]].* FROM [[detections]] WHERE [[detections]].id = {{id}}
611
+ - route: not_for_everyone/{{id}}
612
+ sql: SELECT [[other]].* FROM [[other]] WHERE [[other]].id = {{id}}
613
+ exclude: # don't add this route to these slug/versions
614
+ - 'slug1/3.0'
615
+ - slug: slug2
616
+ version: 2.3
617
+ - slug: slug3
618
+ version: '*' # '*' = every version
619
+
620
+ # the same route defined twice: one for everything except birdnet/2.4, one only for it
621
+ - route: detections/
622
+ exclude: ['birdnet/2.4']
623
+ sql: SELECT [[detections]].* FROM [[detections]]
624
+ - route: detections/
625
+ include: ['birdnet/2.4'] # ONLY add this route to these slug/versions
626
+ sql: SELECT [[detections]].*, [[revisions]].id AS revision_id FROM [[detections]] LEFT JOIN [[revisions]] ON [[revisions]].observation_id = [[detections]].id
627
+
628
+ query_params:
629
+ - confidence:
630
+ sql: "[[detections]].confidence >= {{confidence}}"
631
+ - start_time:
632
+ sql: "[[detections]].start_time >= {{start_time}}"
633
+ exclude: ['slug1/3.9']
634
+ - limit:
635
+ sql_append: LIMIT {{limit}}
636
+
637
+ # limit ALL shared routes / query params to these slug/versions
638
+ route_inclusions: ['birdnet', 'owl/5.0']
639
+ query_inclusions: [] # empty or missing = no restriction
640
+
641
+ # opt slug/versions out of ALL shared routes / query params
642
+ route_exclusions: ['legacy_db']
643
+ query_exclusions:
644
+ - slug: slug4
645
+ version: 1.0
646
+ ```
647
+
648
+ Rules:
649
+ - **The version config wins.** Its own routes come first and replace any shared route with the same shape. Shape means the same path segments; `{{param}}` names and leading/trailing slashes are ignored, so `detections/{{id}}` and `/detections/{{detection_id}}/` are the same route. Routes the version config adds on top are kept.
650
+ - Shared `query_params` behave like a version config's top-level `query_params`. They apply to every route, and a param with the same name in the version config (or on a route) overrides the shared one.
651
+ - `include` and `route_inclusions`/`query_inclusions` are the opposite of `exclude` and `route_exclusions`/`query_exclusions`. When given (non-empty), the route or query param is added **only** to the listed slug/versions. A shared item is added only if it passes both the top-level lists and its own `include`/`exclude`.
652
+ - The same route (by shape) or query param (by name) can appear more than once in the shared file. Each database/version gets the first one whose `include`/`exclude` select it, so complementary `include`/`exclude` lists give different databases different versions of an endpoint.
653
+ - `include`, `exclude` and the four top-level lists take a list of `'<slug>/<version>'` strings or `{slug: <slug>, version: <version>}` mappings. `'<slug>'`, `'<slug>/*'` or a missing/`'*'` version match every version, including unversioned databases. Versions compare numerically when possible (`2.3`, `"2.3"`), and `latest` is resolved before matching.
654
+
436
655
  **For more details**, see the [SQL Database Support Wiki](https://github.com/SchmidtDSE/api_dock/wiki/SQL-Database-Support).
437
656
 
438
657
  ---
@@ -672,6 +891,125 @@ Parameters are processed in this order (first match wins for early returns):
672
891
 
673
892
  ---
674
893
 
894
+ ## Conditional SQL Selection
895
+
896
+ Some routes need a *different* base query depending on the request — for example, `?count=true` should return a species histogram (`SELECT … COUNT(*) … GROUP BY …`) rather than rows. `sql_append` can't help (it only adds trailing clauses), and a second `route:` can't either (the path is identical). For this, a route's `sql` may be a **rule list** instead of a string: a first-match-wins decision tree that picks the base query from the presence and value of path, query, and cookie params.
897
+
898
+ Everything downstream is unchanged: the selected base composes with `query_params` WHERE-fragments and `sql_append` exactly as a plain `sql:` string does.
899
+
900
+ ### The histogram example
901
+
902
+ ```yaml
903
+ routes:
904
+ - route: detections
905
+ sql:
906
+ # ?count=<truthy> → species histogram
907
+ - when: count
908
+ then:
909
+ sql: >
910
+ SELECT [[detections]].common_name, [[detections]].scientific_name,
911
+ COUNT(*) AS count
912
+ FROM [[detections]]
913
+ sql_append: GROUP BY [[detections]].common_name, [[detections]].scientific_name
914
+ # otherwise → detection rows
915
+ - else: SELECT [[detections]].* FROM [[detections]]
916
+ query_params:
917
+ - recording:
918
+ sql: "[[detections]].recording_id = {{recording}}"
919
+ multivalue_sql: "[[detections]].recording_id IN {{recording}}"
920
+ - limit:
921
+ sql_append: LIMIT {{limit}} # applies in BOTH modes
922
+ ```
923
+
924
+ ```bash
925
+ GET /db/detections?recording=1&count=true
926
+ # SELECT detections.common_name, detections.scientific_name, COUNT(*) AS count
927
+ # FROM detections WHERE detections.recording_id = '1'
928
+ # GROUP BY detections.common_name, detections.scientific_name
929
+
930
+ GET /db/detections?recording=1
931
+ # SELECT detections.* FROM detections WHERE detections.recording_id = '1'
932
+ ```
933
+
934
+ Note the pipeline order: the selected branch's `sql_append` (the `GROUP BY`) is applied **before** route-level `sql_append` (the shared `LIMIT`), so SQL clause order stays valid.
935
+
936
+ ### Rule forms
937
+
938
+ A `sql` list contains rules evaluated top to bottom; the **first match wins**. Each rule's payload (an *sql node*) is a SQL string, a leaf object `{sql, sql_append}`, or a nested rule list.
939
+
940
+ ```yaml
941
+ sql:
942
+ - when: count # fires when `count` is truthy (shorthand for equals: _truthy)
943
+ then: <sql node>
944
+
945
+ - when: mode
946
+ equals: 'true' # fires only when mode == "true" (case-insensitive)
947
+ then: <sql node>
948
+
949
+ - when: format # value map: different SQL per value
950
+ match:
951
+ species: <sql node> # ?format=species
952
+ recording: <sql node> # ?format=recording
953
+ _truthy: <sql node> # any other truthy value
954
+ _default: <sql node> # any present value not matched above
955
+
956
+ - when: [count, recording] # list: fires when ALL are truthy (AND)
957
+ then: <sql node>
958
+
959
+ - when: [count, something_else] # list + positional case list
960
+ match:
961
+ - values: [_any, x] # something_else == x, count anything
962
+ then: <sql node>
963
+ - values: [_truthy, _absent] # count truthy AND something_else not passed
964
+ then: <sql node>
965
+ - default: <sql node>
966
+
967
+ - else: <sql node> # default (a trailing bare string works too)
968
+ ```
969
+
970
+ Nesting works because a payload can itself be a rule list:
971
+
972
+ ```yaml
973
+ sql:
974
+ - when: some_value
975
+ match:
976
+ '4':
977
+ - when: count
978
+ then: <sql for value 4 with count>
979
+ - else: <sql for value 4>
980
+ _default: <sql for other values>
981
+ - else: <base sql>
982
+ ```
983
+
984
+ ### Value specs
985
+
986
+ | spec | matches when the param… |
987
+ |---|---|
988
+ | `'literal'` (`'4'`, `'true'`) | is present and equals it (case-insensitive) |
989
+ | `_truthy` / `_falsy` | present, and value is / isn't in `{"", "0", "false", "no", "off", "null", "none"}` |
990
+ | `_present` / `_absent` | exists / does not exist |
991
+ | `_any` | wildcard — present or absent (used for a position in a case list) |
992
+ | `_default` | catch-all for any *present* value (value maps only) |
993
+
994
+ In a single-param `match:` map, precedence is order-independent: exact literal > `_falsy`/`_truthy` > `_present`/`_absent` > `_default`. In a positional case list, cases match strictly top-to-bottom.
995
+
996
+ ### No match → URL error
997
+
998
+ If no rule matches and there is no default (`else`, a trailing bare string, or a `_default`/`default` catch-all), the request returns a **400** with `{"error": "No matching query configuration for the given parameters", "http_status": 400}`. Customize it with a terminal `no_match` rule:
999
+
1000
+ ```yaml
1001
+ sql:
1002
+ - when: recording
1003
+ then: SELECT [[detections]].* FROM [[detections]] WHERE recording_id = {{recording}}
1004
+ - no_match:
1005
+ error: "recording is required, or pass count=true for a histogram"
1006
+ http_status: 400
1007
+ ```
1008
+
1009
+ Cookies participate via the `cookies.<name>` key (e.g. `when: cookies.role`, `equals: admin`).
1010
+
1011
+ ---
1012
+
675
1013
  # CLI
676
1014
 
677
1015
  ## Commands