toga-ai 1.0.138 → 1.0.140

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -9,4 +9,4 @@
9
9
  | [Monitoring Framework (Orchestrator + Child Monitors)](features/monitoring-framework.md) | A unified, DB-driven monitoring framework for business-critical data flows (Compass POs, Prudential asset imports, AIG closed claims, …). | worker2/Worker/Monitor.php, worker2/Worker/Monitors/, worker2/Worker/Notification/Email.php, dbchanges2/Core/2026-05-21 - Monitors.sql |
10
10
  | [NetSuite → TOGA Opportunity Sync (API Message Queue + worker2 webhook)](features/netsuite-opportunity-sync.md) | Outbound sync from NetSuite to TOGA for the record types the Forecast2 importer pulls (opportunities first; sales/items/etc. | worker2/Worker/Netsuite.php, worker2/Worker/Netsuite/Opportunity.php, worker2/Controller/Index.php, _underscore/Worker.php, test/@dave/NetSuite/api-message-queue/lib_amq_queue.js, test/@dave/NetSuite/api-message-queue/ue_api_msg_queue_enqueue.js, test/@dave/NetSuite/api-message-queue/ue_amq_drain.js, test/@dave/NetSuite/api-message-queue/ss_amq_drain.js, test/@dave/NetSuite/api-message-queue/DEPLOY_RUNBOOK.md, test/@dave/clickup/backfill_opportunity_numbers.php, test/@dave/clickup/probe_opportunity_fields.php, test/@dave/probe_clickup_desc_match.php, worker/crons/toga2/forecast2/common_import_sales_from_netsuite.php |
11
11
  | [NetSuite → Forecast Open-Orders Sync (salesOrder webhook → OpenOrderItems)](features/netsuite-salesorder-open-orders-sync.md) | Webhook-driven, single-record port of the legacy open-orders importer (TRUE-79142). | worker2/Worker/Netsuite/SalesOrder.php, worker2/Worker/Netsuite.php, test/@dave/probe_salesorder_rest_shape.php, test/@dave/probe_open_order_lines.php, test/@dave/check_so_status.php, test/@dave/check_so_history.php, test/@dave/probe_so_rest_lines.php, worker/crons/toga2/forecast2/import_open_orders.php, worker/crons/toga2/forecast2/common_import_sales_from_netsuite.php |
12
- | [Teams Meeting Transcript Export](features/teams-transcript-export.md) | `_Worker_Team_Transcripts` (action `Team/Transcripts/Export`) polls Microsoft Graph for Teams meeting transcripts produced by a set of organizers, classifies ea | worker2/Worker/Team/Transcripts.php, worker2/Config/production.ini |
12
+ | [Teams Meeting Transcript Export](features/teams-transcript-export.md) | `_Worker_Team_Transcripts` (action `Team/Transcripts/Export`) polls Microsoft Graph for Teams meeting transcripts produced by a set of organizers, classifies ea | worker2/Worker/Team/Transcripts.php, worker2/Config/production.ini, worker2/Database/TeamsTranscriptExports.sql, dbchanges2/Core/2026-06-18a - Teams Transcript Export schedule.sql |
@@ -6,11 +6,13 @@ project: Worker
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-06-12
9
+ updated: 2026-06-18
10
10
  owners: ["ajean"]
11
11
  files:
12
12
  - worker2/Worker/Team/Transcripts.php
13
13
  - worker2/Config/production.ini
14
+ - worker2/Database/TeamsTranscriptExports.sql
15
+ - dbchanges2/Core/2026-06-18a - Teams Transcript Export schedule.sql
14
16
  related:
15
17
  - ../architecture.md
16
18
  - ./creating-worker-actions.md
@@ -20,8 +22,9 @@ related:
20
22
 
21
23
  `_Worker_Team_Transcripts` (action `Team/Transcripts/Export`) polls Microsoft Graph for
22
24
  Teams meeting transcripts produced by a set of organizers, classifies each meeting by
23
- client (from the meeting title), and archives the raw WebVTT to S3. Scheduled via three
24
- `Core.CronJobs` rows (weekdays 10:00 / 13:00 / 17:30 Central).
25
+ client (from the meeting title), and archives the raw WebVTT to S3. Scheduled via a single
26
+ `Core.CronJobs` row firing weekdays at 9:30 / 11:30 / 13:30 / 14:30 / 15:30 / 16:30 / 17:30
27
+ Central (`30 9,11,13,14,15,16,17 * * 1-5`).
25
28
 
26
29
  ## Key files / entry points
27
30
 
@@ -29,6 +32,9 @@ client (from the meeting title), and archives the raw WebVTT to S3. Scheduled vi
29
32
  - `Config/production.ini` `[teams]` section — Entra app creds, organizer source, S3 target.
30
33
  - `[teamsClientAliases]` config section — alias → canonical client name for classification.
31
34
  - Ledger table `Team.TranscriptExports` (DB alias `_underscore::DB_TEAM`, core cluster).
35
+ - `Backfill(bool $dryRun = true, int $limit = 0)` — re-files already-misdated S3 objects (see *Gotchas*).
36
+ - DDL + cron seed: `worker2/Database/TeamsTranscriptExports.sql`; prod schedule change:
37
+ `dbchanges2/Core/2026-06-18a - Teams Transcript Export schedule.sql`.
32
38
 
33
39
  ## How it works
34
40
 
@@ -51,9 +57,17 @@ Per run, after acquiring a Graph client-credentials token:
51
57
  3. `getAllTranscripts(meetingOrganizerUserId='{guid}', startDateTime, endDateTime)`
52
58
  over an incremental window (`MAX(dtCreated)` watermark from the ledger, else `now − lookbackDays`),
53
59
  following `@odata.nextLink`.
54
- 4. Per new transcript (dedup by `transcriptIdentifier`): fetch meeting `subject`+`startDateTime`
55
- (cached per meeting), classify, download VTT (`?$format=text/vtt`), `putObject` to S3, insert
56
- ledger row.
60
+ 4. Per new transcript (dedup by `transcriptIdentifier`): fetch meeting `subject` (cached per
61
+ meeting) for classification, classify, download VTT (`?$format=text/vtt`), `putObject` to S3,
62
+ insert ledger row.
63
+
64
+ **Archive date source (important):** the S3 path date and `dtMeeting` come from the
65
+ transcript's own **`createdDateTime`**, NOT the meeting's `startDateTime`. For a recurring
66
+ meeting, `getMeeting()` (`onlineMeetings/{id}`) returns the **series-anchor** start (the first
67
+ occurrence), so `startDateTime` misfiles every occurrence under one frozen — sometimes future —
68
+ date. `createdDateTime` is stamped minutes after the actual occurrence, so it dates the archive
69
+ correctly. `getMeeting()` is therefore used only for the subject. `dtExported` uses
70
+ `UTC_TIMESTAMP()` (not `NOW()`) so all three timestamps are UTC.
57
71
 
58
72
  **Classification:** case-insensitive substring match of the meeting title against active
59
73
  `Core.Clients` names + `[teamsClientAliases]`, longest needle wins, min length 3. No match → `general`.
@@ -88,6 +102,20 @@ process is identical for all.
88
102
 
89
103
  ## Gotchas / known issues
90
104
 
105
+ - **Recurring meetings return the SERIES-ANCHOR `startDateTime`.** Graph's
106
+ `onlineMeetings/{id}` gives the first-occurrence (or even a future) start for a recurring
107
+ series, not the occurrence that was recorded. Driving the archive date off it misfiled every
108
+ occurrence under one frozen date and produced future-dated folders (e.g. a 6-17 occurrence
109
+ filed under `…/2026-06-03/…`; a "Rebranding Weekly Sync" under `2026-06-19/`). Fixed by
110
+ dating off the transcript's `createdDateTime`. The transcript object does **not** carry the
111
+ occurrence start, so `createdDateTime` (≈ minutes after the meeting) is the best reliable
112
+ signal. The `Backfill()` action re-files objects already written under the wrong date:
113
+ recomputes the correct key purely from ledger columns (no Graph calls), `copyObject` →
114
+ `deleteObject` (idempotent — `NoSuchKey` + `doesObjectExist` treats an already-moved row as a
115
+ no-op), then updates `s3Key` + `dtMeeting`. Run `dryRun:true` first, inspect `sample[]`, then
116
+ `dryRun:false` (optional `limit` for a first batch). `$limit` caps candidates touched, not
117
+ rows scanned, so a run of failures stops at `$limit` instead of sweeping the table.
118
+
91
119
  - **`getAllTranscripts` requires an AAD object-id (GUID), NOT a UPN.** Passing a UPN
92
120
  (e.g. `ajean@togatech.com`) returns the opaque `HTTP 400 BadRequest / "UnknownError"`.
93
121
  This is the single most likely cause of an export failure.
@@ -124,6 +152,10 @@ Policy gap; a `400` only on a UPN means the id was never resolved to a GUID. (A
124
152
 
125
153
  ## Change history
126
154
 
155
+ - 2026-06-18 — Fixed recurring-meeting date misfiling: archive date + `dtMeeting` now come from
156
+ the transcript `createdDateTime`, not the series-anchor `meeting->startDateTime` (worker2 PR
157
+ #82). Added `Backfill()` to re-file existing misdated S3 objects; `dtExported` → `UTC_TIMESTAMP()`.
158
+ Rescheduled to weekdays 9:30–17:30 CT, one consolidated cron row (dbchanges2 PR #398). (ajean)
127
159
  - 2026-06-12 — **Merged (PR #78) and verified in production**: a real `Team/Transcripts/Export`
128
160
  run returned 6/6 organizers, 33 found, 33 exported, 0 errors (classified 30 general / 3 Elite).
129
161
  Added fail-fast on unresolvable organizers (`resolveUserId` → null, loop throws clear error
@@ -29,7 +29,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
29
29
 
30
30
  ## standalone framework
31
31
 
32
- - **togatech** (TOGA Technology Website) — 1 doc(s) → [standalone/apps/togatech/INDEX.md](standalone/apps/togatech/INDEX.md)
32
+ - **togatech** (TOGA Technology Website) — 2 doc(s) → [standalone/apps/togatech/INDEX.md](standalone/apps/togatech/INDEX.md)
33
33
  - **forward** (Forwarder) — 3 doc(s) → [standalone/apps/forward/INDEX.md](standalone/apps/forward/INDEX.md)
34
34
 
35
35
  ## Clients
@@ -3,3 +3,4 @@
3
3
  | Doc | Summary | Files |
4
4
  |-----|---------|-------|
5
5
  | [TOGA Technology Website Architecture](architecture.md) | The public-facing TOGA Technology corporate/marketing website. | togatech/src/main.tsx, togatech/src/App.tsx, togatech/src/routes.tsx, togatech/src/lib/api.ts, togatech/src/lib/contentful.ts, togatech/src/themeConfig/ThemeContext.tsx, togatech/vite.config.ts, togatech/package.json |
6
+ | [SEO / AEO / GEO — Prerendering, Single-Source Meta & Structured Data](features/seo-aeo-geo-prerender.md) | Makes togatech.com visible and citable to search engines **and** AI answer engines (ChatGPT/Perplexity/Claude search, Google AI Overviews). | togatech/vite.config.ts, togatech/scripts/prerender.mjs, togatech/src/routes.config.json, togatech/src/main.tsx, togatech/src/App.tsx, togatech/src/components/SEO/Seo.tsx, togatech/src/components/SEO/JsonLd.tsx, togatech/src/components/SEO/schema.ts, togatech/public/robots.prod.txt, togatech/public/sitemap.xml, togatech/public/llms.txt |
@@ -0,0 +1,62 @@
1
+ ---
2
+ title: SEO / AEO / GEO — Prerendering, Single-Source Meta & Structured Data
3
+ framework: "standalone"
4
+ repo: togatech
5
+ project: TOGA Technology Website
6
+ client: shared
7
+ type: feature
8
+ status: active
9
+ updated: 2026-06-18
10
+ owners: ["ajean"]
11
+ files:
12
+ - togatech/vite.config.ts
13
+ - togatech/scripts/prerender.mjs
14
+ - togatech/src/routes.config.json
15
+ - togatech/src/main.tsx
16
+ - togatech/src/App.tsx
17
+ - togatech/src/components/SEO/Seo.tsx
18
+ - togatech/src/components/SEO/JsonLd.tsx
19
+ - togatech/src/components/SEO/schema.ts
20
+ - togatech/public/robots.prod.txt
21
+ - togatech/public/sitemap.xml
22
+ - togatech/public/llms.txt
23
+ related: []
24
+ ---
25
+
26
+ ## Summary
27
+ Makes togatech.com visible and citable to search engines **and** AI answer engines (ChatGPT/Perplexity/Claude search, Google AI Overviews). The site is a CSR-only Vite/React SPA, so non-JS crawlers saw an empty `<div id="root">`. This feature fixes a production de-indexing bug and adds **build-time prerendering** so every static route ships real HTML, plus single-source per-page meta and JSON-LD structured data. Shipped on branches `fix-seo-deindex-emergency` (PR #29) and `feature-seo-aeo-prerender` (PR #30) — **not yet merged/deployed** as of 2026-06-18.
28
+
29
+ ## Key files / entry points
30
+ - `vite.config.ts` — `robotsPlugin()` now gates on **Vite `mode`** (was `NODE_ENV && VITE_ENV`).
31
+ - `scripts/prerender.mjs` — post-build headless-Chrome (Puppeteer) snapshot; chained into the `build` script.
32
+ - `src/routes.config.json` — single source of static route paths (consumed by the prerender; intended for sitemap too).
33
+ - `src/main.tsx` — `hydrateRoot` when a route was prerendered, else `createRoot`.
34
+ - `src/components/SEO/{Seo.tsx,JsonLd.tsx,schema.ts}` — shared head + JSON-LD; per-page `meta` lives in each page's `viewModel/FIELDS/*.json`.
35
+
36
+ ## How it works
37
+ 1. **robots:** `robotsPlugin` copies `robots.prod.txt` only when `mode === "production"`; beta/gamma/dev get `robots.dev.txt` (`Disallow: /`). `robots.prod.txt` carries an explicit AI-bot allow-list (GPTBot, ClaudeBot, PerplexityBot, …) + `https://` sitemap.
38
+ 2. **prerender:** `npm run build` = `tsc && vite build && node scripts/prerender.mjs`. The script serves `dist/` (reading the original shell once, serving it for all nav routes), snapshots each route in headless Chrome, and writes `dist/<route>/index.html`. Runs in a real browser, so **no SSR/window guards needed**.
39
+ 3. **meta:** each page renders `<Seo {...FIELDS.meta} />` — title/description/keywords/canonical/OG/Twitter from one place per page; canonical/OG URLs derived from one `SITE_URL`. `<JsonLd>` (Helmet `<script type="application/ld+json">`) is baked into the snapshot.
40
+ 4. **structured data:** Organization + WebSite site-wide (AppLayout), per-page BreadcrumbList (Seo), platform SoftwareApplication (OurPlatform), LocalBusiness×5 (from Contact locations), Person (from About team) — all **derived from existing FIELDS** (single source).
41
+
42
+ ## Data model
43
+ None (static marketing site). All copy/meta/NAP/leaders come from `*FIELDS*.json`; structured data is derived from `FOOTERFIELDS.json`, `CONTACTPAGEFIELDS.json` (`map.locations`), and `ABOUTPAGEFIELDS.json` (`teamMembers`).
44
+
45
+ ## Client variations
46
+ None — uniform.
47
+
48
+ ## Gotchas / known issues
49
+ - **De-indexing bug (root cause):** `robotsPlugin` required `NODE_ENV && VITE_ENV==="production"`, but build scripts set neither → every build shipped `robots.dev.txt` (`Disallow: /`). **Do NOT gate robots on `NODE_ENV`** — `vite build` sets it to `production` for *every* mode, which would expose beta/gamma. Gate on Vite `mode`.
50
+ - **`.npmrc` has `ignore-scripts=true`** → the `postbuild` lifecycle hook never fires. Prerender is therefore chained directly in the `build` script, not a `postbuild` hook.
51
+ - **Prerender tooling:** `vite-react-ssg` supports React Router **6 only** (project is RR 7.9); `react-snap` is stale (2022) and won't launch on Node 24. Hence a custom Puppeteer snapshot. Puppeteer's Chrome must be installed at build time: `npx puppeteer browsers install chrome`.
52
+ - **Prerender fail-fast:** a route that never mounts (`#root > *` times out) aborts the build with `process.exit(1)` — otherwise it would snapshot the empty shell (no SEO/JSON-LD) and still exit 0.
53
+ - **🔴 Hosting requirement:** CloudFront/S3 must serve each prerendered `dist/<route>/index.html` for its path; a blanket SPA rewrite to root `index.html` serves the empty shell and defeats the prerender.
54
+ - **og:image dedupe:** static OG/Twitter tags were removed from `index.html` because `<Seo>` emits per-page ones (Helmet appends rather than replacing pre-existing static tags → duplicates otherwise).
55
+ - **Verified data corrections:** real positioning is *integrated IT services + the ERA TOGa Platform*; real socials are LinkedIn `/toga-tech` + YouTube + Instagram (no Twitter/GitHub/Crunchbase); contact is `website@togatech.com` / `+1 (212) 736-0111`; **5** offices. Prior planning docs had several wrong specifics — verify schema/llms.txt against FIELDS before shipping.
56
+ - **Build size:** JS bundle is ~14.8 MB (4.9 MB gzip) — a real CWV/LCP risk; code-splitting is a separate follow-up.
57
+
58
+ ## Change history
59
+ - 2026-06-18 — Initial SEO/AEO/GEO implementation: robots mode-gate + AI allow-list + sitemap/llms; Puppeteer prerender (fail-fast); shared `<Seo>` single-source meta + canonical; JSON-LD (Organization/WebSite/Breadcrumb/SoftwareApplication/LocalBusiness×5/Person×12). PRs #29 + #30, pending merge. (ajean)
60
+
61
+ ## Related docs
62
+ - standalone/apps/togatech/architecture.md (update its Build/deploy section once PR #30 merges — robotsPlugin gating + prerender step).
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "toga-ai",
3
- "version": "1.0.138",
3
+ "version": "1.0.140",
4
4
  "description": "TOGA Technology Team Claude Knowledge System — shared AI coding harness with skills, knowledge base CLI, and project installer for Claude Code.",
5
5
  "keywords": [
6
6
  "claude",