quiverdb 0.10.6 → 0.10.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Binary file
Binary file
Binary file
Binary file
Binary file
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "quiverdb",
3
- "version": "0.10.6",
3
+ "version": "0.10.7",
4
4
  "license": "MIT",
5
5
  "repository": {
6
6
  "type": "git",
package/src/lua-api.ts CHANGED
@@ -93,9 +93,13 @@ midnight.
93
93
  - **Standard library.** Loaded standard libraries: base, string, table, math, coroutine, utf8.
94
94
  That is the pure-computation set — there is no \`os\`, \`io\`, \`debug\`, or \`package\`/\`require\`,
95
95
  and \`dofile\`/\`loadfile\` are removed (string-form \`load\` stays available). Integer division is
96
- the Lua 5.4 \`//\` operator — a language operator, unrelated to \`math\`.
96
+ the Lua 5.4 \`//\` operator — a language operator, unrelated to \`math\`. No \`io\` does **not** mean
97
+ a data file on disk is out of reach: read it with \`db:read_csv\` / \`db:read_csv_stream\` (see
98
+ the CSV file reading section below). Never copy, paste, or re-type a data file's contents into
99
+ the script as literals — read the file.
97
100
  - **Filesystem sandbox.** Every file-touching operation (\`db:export_csv\`, \`db:import_csv\`,
98
- \`db:open_file\`, \`db:bin_to_csv\`, \`db:csv_to_bin\`, \`db:validate_migrations\`, \`expr:save\`) resolves
101
+ \`db:open_file\`, \`db:bin_to_csv\`, \`db:csv_to_bin\`, \`db:validate_migrations\`, \`db:read_csv\`,
102
+ \`db:read_csv_stream\`, \`db:write_csv\`, \`expr:save\`) resolves
99
103
  relative paths against the directory containing the database file and rejects anything outside it
100
104
  (subdirectories are fine; \`..\` escapes and outside absolute paths throw \`Cannot <op>: path '...' escapes the
101
105
  database directory ...\`). On an in-memory database these operations throw
@@ -612,6 +616,153 @@ active\`. Call it outside any \`db:transaction\` / \`db:begin_transaction\` bloc
612
616
 
613
617
  ---
614
618
 
619
+ ## CSV file reading
620
+
621
+ Read a CSV file from disk directly into Lua — the only way to get file data into a script, since
622
+ \`io\` is deliberately absent from the sandbox. \`path\` is sandboxed the same way as every other
623
+ file-touching operation (see Critical rules).
624
+
625
+ \`\`\`lua
626
+ local csv = db:read_csv(path, { separator = ",", header_row = 1 }) -- { header = {...}, rows = {{...}, ...} }
627
+ \`\`\`
628
+
629
+ Every cell arrives as a **string**, with no numeric or date inference — \`"0012"\` stays \`"0012"\`
630
+ and a date stays text. \`csv.header\` is a 1-based array of the file's column names in file order;
631
+ \`csv.rows\` is a 1-based array of rows, each a 1-based array of strings, addressed positionally
632
+ (\`csv.rows[1][1]\`). Rows are never padded to header width — a short row stays short and a field
633
+ past its end is \`nil\`. A 0-byte file throws \`Cannot read_csv: file '<path>' is empty\`; a
634
+ header-only file returns a populated \`header\` and an empty \`rows\`.
635
+
636
+ The options table is optional; its two keys are \`separator\` and \`header_row\`. \`header_row\` is
637
+ 1-based (like every other index here) and defaults to \`1\`; \`header_row = 0\` declares the file has
638
+ no header at all, so \`csv.header\` is absent (\`nil\`, not an empty table) and \`csv.rows[1]\` is the
639
+ file's first line — useful for a file with a junk title row and/or a units row around the real
640
+ header (skip them by naming the header row and slicing \`csv.rows\` in the script). A \`header_row\`
641
+ past the end of the file throws. Passing the separator positionally (\`db:read_csv(path, ";")\`)
642
+ throws \`Cannot read_csv: options must be a table\` instead of silently parsing with a comma; an
643
+ unknown key, a separator that isn't a single character (or is a quote, CR, LF or NUL — none of
644
+ those can be a delimiter), a non-string option key, or a \`header_row\` that isn't a
645
+ non-negative integer also throws.
646
+
647
+ **Reading a file \`db:export_csv\` wrote:** \`export_csv\` emits an Excel-style \`sep=,\` preamble as
648
+ line 1, so its real header is line 2 — read it with \`{ header_row = 2 }\`. With the default
649
+ \`header_row = 1\` the preamble itself becomes a two-column header and the column names come back
650
+ as \`rows[1]\`.
651
+
652
+ \`db:read_csv_stream\` reads the same file through the same parser, row by row, so the process holds
653
+ a bounded window instead of the whole file:
654
+
655
+ \`\`\`lua
656
+ local n = db:read_csv_stream(path, function(row, index, header)
657
+ -- row: same shape as db:read_csv's rows; index: the 1-based ordinal of this DATA row (the
658
+ -- header row is not counted, so skipping index 1 to "skip a header" drops a real record);
659
+ -- header: the same array db:read_csv returns, reachable here so a column can be found by
660
+ -- name before processing row 1.
661
+ return row[1] ~= "" -- returning false stops the read early; a bare comparison as the
662
+ end, { separator = "," }) -- last statement can silently truncate the stream this way
663
+ -- n counts rows FED to the callback, not rows it kept -- a filtering callback logging n as
664
+ -- "imported" would be wrong.
665
+ \`\`\`
666
+
667
+ A Lua error raised inside the callback propagates to the host verbatim, and the file is closed.
668
+
669
+ **Worked example**, over two real files that used to be hand-transcribed into scripts instead of
670
+ read from disk:
671
+
672
+ \`\`\`lua
673
+ -- File 1: BOM + CRLF, a junk title row above the header, a units row below it, apostrophe
674
+ -- thousands separators, and a DD/MM/YYYY date.
675
+ local csv = db:read_csv("ma_energia_residencial.csv", { header_row = 2 })
676
+ -- header_row = 2 skips the block-title junk row (line 1). rows[1] is the units row that sits
677
+ -- BELOW the header (line 3) -- not a reader concern -- so real data starts at rows[2].
678
+ for i = 2, #csv.rows do
679
+ local row = csv.rows[i]
680
+ local dd, mm, yyyy = row[5]:match("(%d%d)/(%d%d)/(%d%d%d%d)")
681
+ local date_key = yyyy .. "-" .. mm
682
+ -- gsub returns TWO values (string, replacement count). tonumber(row[6]:gsub("['%s]", ""))
683
+ -- would hand the count to tonumber as its BASE argument and silently return nil, no error.
684
+ -- The parentheses below truncate the call to one value -- this is the correct form.
685
+ local value = tonumber((row[6]:gsub("['%s]", "")))
686
+ end
687
+
688
+ -- File 2: header is line 1 (the default, no header_row needed), a quoted field containing a
689
+ -- comma, and English month names -- os is unloaded, so there is no date library to lean on.
690
+ local MONTHS = {
691
+ January = 1, February = 2, March = 3, April = 4, May = 5, June = 6,
692
+ July = 7, August = 8, September = 9, October = 10, November = 11, December = 12,
693
+ }
694
+ local gd = db:read_csv("ma_gd_data.csv")
695
+ for i = 1, #gd.rows do
696
+ local row = gd.rows[i]
697
+ -- The date cell is a quoted field containing a comma ("May 1, 2014"); the row still has
698
+ -- exactly 2 fields -- mishandled quoting would have split the date and shifted this value.
699
+ local month_name, _, year = row[1]:match("(%a+) (%d+), (%d+)")
700
+ local date_key = string.format("%d-%02d", tonumber(year), MONTHS[month_name])
701
+ local value = tonumber(row[2]) -- already a plain decimal string, no cleanup needed
702
+ end
703
+ \`\`\`
704
+
705
+ ---
706
+
707
+ ## CSV file writing
708
+
709
+ Write a CSV file to disk — the only way to get data out of a script onto disk, since \`io\` is
710
+ deliberately absent from the sandbox. Streaming only, with no whole-file counterpart: \`db:write_csv\`
711
+ returns a handle, \`w:write_row({...})\` appends one row, \`w:close()\` finishes it. \`path\` is
712
+ sandboxed the same way as every other file-touching operation (see Critical rules).
713
+
714
+ \`\`\`lua
715
+ local w = db:write_csv(path, { separator = ",", header = { "name", "note", "active", "score" } })
716
+ local rows = {
717
+ { "Alpha", "first", true, 42 },
718
+ { "Beta", nil, false, 3.5 }, -- nil is INTERIOR, not the row's last cell: a TRAILING nil
719
+ -- would shorten the row instead of writing an empty cell.
720
+ }
721
+ for _, row in ipairs(rows) do
722
+ w:write_row(row)
723
+ end
724
+ w:close()
725
+ \`\`\`
726
+
727
+ The options table is optional; its only two keys are \`separator\` (a single character, default
728
+ \`,\`) and \`header\` (column names written as the first record, default none — no header row). A
729
+ quote, CR, LF or NUL is rejected as a separator: none of them can be a delimiter, and a file
730
+ written with one could not be read back.
731
+
732
+ Opening \`db:write_csv\` **truncates** an existing file at the target path — there is no overwrite
733
+ guard, so a script can destroy an existing file in the case folder (including the database file
734
+ itself) by writing to its path. This is documented behaviour, not a bug: reopening the same path
735
+ always starts a fresh file. Two writers open on the *same* path at once is refused, though
736
+ (\`Cannot write_csv: file is already open for writing: ...\`) — the second would truncate what the
737
+ first is still buffering. Close the first writer before reopening its path.
738
+
739
+ \`write_row\` after \`close\` throws; \`close\` is idempotent (a second call is a no-op, not an error).
740
+
741
+ A number is written in its shortest round-trip form (\`std::to_chars\`, no synthetic decimal point),
742
+ so a whole float like \`2014.0\` and the integer \`2014\` write identical text — a script that needs
743
+ a decimal point writes the cell as a string. A boolean writes \`1\`/\`0\`, matching the project-wide
744
+ boolean-is-INTEGER write policy.
745
+
746
+ \`nil\` and an empty string are not always the same thing here. An INTERIOR \`nil\` cell (not a row's
747
+ last cell, like \`"Beta"\`'s note above) writes an empty cell, indistinguishable from \`""\` after the
748
+ round trip — CSV has no null. A TRAILING \`nil\`, however, is not a cell at all: Lua stores no key
749
+ for it, so the row's maximum integer key is lower. **With no \`header\`** that makes the row come
750
+ back **one column narrower**, and a script that needs a trailing empty column must write an empty
751
+ string there, not \`nil\`. **With a \`header\`** the padding rule below fills the gap, so the row is
752
+ header-width either way.
753
+
754
+ With a \`header\`, its length is the row width: a \`write_row\` shorter than the header pads with
755
+ empty cells, and a longer one throws, naming the row's ordinal and both counts. Omitting \`header\`
756
+ disables the check entirely — rows of any length are written as-is.
757
+
758
+ A writer never explicitly closed is still flushed and closed when the script's \`run()\` call
759
+ returns — whether or not the script still holds it (a \`local\` that went out of scope and a global
760
+ alike) — so the file is complete and re-readable even without a \`w:close()\` call, and no warning
761
+ is emitted. The writer does not survive that \`run()\`: using the same handle from a later
762
+ \`run()\` throws \`Cannot write_row: writer for '...' is already closed\`.
763
+
764
+ ---
765
+
615
766
  ## Complete example
616
767
 
617
768
  \`\`\`lua