toga-ai 1.0.138 → 1.0.140
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/knowledge/2.0/apps/worker2/INDEX.md +1 -1
- package/knowledge/2.0/apps/worker2/features/teams-transcript-export.md +38 -6
- package/knowledge/INDEX.md +1 -1
- package/knowledge/standalone/apps/togatech/INDEX.md +1 -0
- package/knowledge/standalone/apps/togatech/features/seo-aeo-geo-prerender.md +62 -0
- package/package.json +1 -1
|
@@ -9,4 +9,4 @@
|
|
|
9
9
|
| [Monitoring Framework (Orchestrator + Child Monitors)](features/monitoring-framework.md) | A unified, DB-driven monitoring framework for business-critical data flows (Compass POs, Prudential asset imports, AIG closed claims, …). | worker2/Worker/Monitor.php, worker2/Worker/Monitors/, worker2/Worker/Notification/Email.php, dbchanges2/Core/2026-05-21 - Monitors.sql |
|
|
10
10
|
| [NetSuite → TOGA Opportunity Sync (API Message Queue + worker2 webhook)](features/netsuite-opportunity-sync.md) | Outbound sync from NetSuite to TOGA for the record types the Forecast2 importer pulls (opportunities first; sales/items/etc. | worker2/Worker/Netsuite.php, worker2/Worker/Netsuite/Opportunity.php, worker2/Controller/Index.php, _underscore/Worker.php, test/@dave/NetSuite/api-message-queue/lib_amq_queue.js, test/@dave/NetSuite/api-message-queue/ue_api_msg_queue_enqueue.js, test/@dave/NetSuite/api-message-queue/ue_amq_drain.js, test/@dave/NetSuite/api-message-queue/ss_amq_drain.js, test/@dave/NetSuite/api-message-queue/DEPLOY_RUNBOOK.md, test/@dave/clickup/backfill_opportunity_numbers.php, test/@dave/clickup/probe_opportunity_fields.php, test/@dave/probe_clickup_desc_match.php, worker/crons/toga2/forecast2/common_import_sales_from_netsuite.php |
|
|
11
11
|
| [NetSuite → Forecast Open-Orders Sync (salesOrder webhook → OpenOrderItems)](features/netsuite-salesorder-open-orders-sync.md) | Webhook-driven, single-record port of the legacy open-orders importer (TRUE-79142). | worker2/Worker/Netsuite/SalesOrder.php, worker2/Worker/Netsuite.php, test/@dave/probe_salesorder_rest_shape.php, test/@dave/probe_open_order_lines.php, test/@dave/check_so_status.php, test/@dave/check_so_history.php, test/@dave/probe_so_rest_lines.php, worker/crons/toga2/forecast2/import_open_orders.php, worker/crons/toga2/forecast2/common_import_sales_from_netsuite.php |
|
|
12
|
-
| [Teams Meeting Transcript Export](features/teams-transcript-export.md) | `_Worker_Team_Transcripts` (action `Team/Transcripts/Export`) polls Microsoft Graph for Teams meeting transcripts produced by a set of organizers, classifies ea | worker2/Worker/Team/Transcripts.php, worker2/Config/production.ini |
|
|
12
|
+
| [Teams Meeting Transcript Export](features/teams-transcript-export.md) | `_Worker_Team_Transcripts` (action `Team/Transcripts/Export`) polls Microsoft Graph for Teams meeting transcripts produced by a set of organizers, classifies ea | worker2/Worker/Team/Transcripts.php, worker2/Config/production.ini, worker2/Database/TeamsTranscriptExports.sql, dbchanges2/Core/2026-06-18a - Teams Transcript Export schedule.sql |
|
|
@@ -6,11 +6,13 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-06-
|
|
9
|
+
updated: 2026-06-18
|
|
10
10
|
owners: ["ajean"]
|
|
11
11
|
files:
|
|
12
12
|
- worker2/Worker/Team/Transcripts.php
|
|
13
13
|
- worker2/Config/production.ini
|
|
14
|
+
- worker2/Database/TeamsTranscriptExports.sql
|
|
15
|
+
- dbchanges2/Core/2026-06-18a - Teams Transcript Export schedule.sql
|
|
14
16
|
related:
|
|
15
17
|
- ../architecture.md
|
|
16
18
|
- ./creating-worker-actions.md
|
|
@@ -20,8 +22,9 @@ related:
|
|
|
20
22
|
|
|
21
23
|
`_Worker_Team_Transcripts` (action `Team/Transcripts/Export`) polls Microsoft Graph for
|
|
22
24
|
Teams meeting transcripts produced by a set of organizers, classifies each meeting by
|
|
23
|
-
client (from the meeting title), and archives the raw WebVTT to S3. Scheduled via
|
|
24
|
-
`Core.CronJobs`
|
|
25
|
+
client (from the meeting title), and archives the raw WebVTT to S3. Scheduled via a single
|
|
26
|
+
`Core.CronJobs` row firing weekdays at 9:30 / 11:30 / 13:30 / 14:30 / 15:30 / 16:30 / 17:30
|
|
27
|
+
Central (`30 9,11,13,14,15,16,17 * * 1-5`).
|
|
25
28
|
|
|
26
29
|
## Key files / entry points
|
|
27
30
|
|
|
@@ -29,6 +32,9 @@ client (from the meeting title), and archives the raw WebVTT to S3. Scheduled vi
|
|
|
29
32
|
- `Config/production.ini` `[teams]` section — Entra app creds, organizer source, S3 target.
|
|
30
33
|
- `[teamsClientAliases]` config section — alias → canonical client name for classification.
|
|
31
34
|
- Ledger table `Team.TranscriptExports` (DB alias `_underscore::DB_TEAM`, core cluster).
|
|
35
|
+
- `Backfill(bool $dryRun = true, int $limit = 0)` — re-files already-misdated S3 objects (see *Gotchas*).
|
|
36
|
+
- DDL + cron seed: `worker2/Database/TeamsTranscriptExports.sql`; prod schedule change:
|
|
37
|
+
`dbchanges2/Core/2026-06-18a - Teams Transcript Export schedule.sql`.
|
|
32
38
|
|
|
33
39
|
## How it works
|
|
34
40
|
|
|
@@ -51,9 +57,17 @@ Per run, after acquiring a Graph client-credentials token:
|
|
|
51
57
|
3. `getAllTranscripts(meetingOrganizerUserId='{guid}', startDateTime, endDateTime)`
|
|
52
58
|
over an incremental window (`MAX(dtCreated)` watermark from the ledger, else `now − lookbackDays`),
|
|
53
59
|
following `@odata.nextLink`.
|
|
54
|
-
4. Per new transcript (dedup by `transcriptIdentifier`): fetch meeting `subject
|
|
55
|
-
|
|
56
|
-
ledger row.
|
|
60
|
+
4. Per new transcript (dedup by `transcriptIdentifier`): fetch meeting `subject` (cached per
|
|
61
|
+
meeting) for classification, classify, download VTT (`?$format=text/vtt`), `putObject` to S3,
|
|
62
|
+
insert ledger row.
|
|
63
|
+
|
|
64
|
+
**Archive date source (important):** the S3 path date and `dtMeeting` come from the
|
|
65
|
+
transcript's own **`createdDateTime`**, NOT the meeting's `startDateTime`. For a recurring
|
|
66
|
+
meeting, `getMeeting()` (`onlineMeetings/{id}`) returns the **series-anchor** start (the first
|
|
67
|
+
occurrence), so `startDateTime` misfiles every occurrence under one frozen — sometimes future —
|
|
68
|
+
date. `createdDateTime` is stamped minutes after the actual occurrence, so it dates the archive
|
|
69
|
+
correctly. `getMeeting()` is therefore used only for the subject. `dtExported` uses
|
|
70
|
+
`UTC_TIMESTAMP()` (not `NOW()`) so all three timestamps are UTC.
|
|
57
71
|
|
|
58
72
|
**Classification:** case-insensitive substring match of the meeting title against active
|
|
59
73
|
`Core.Clients` names + `[teamsClientAliases]`, longest needle wins, min length 3. No match → `general`.
|
|
@@ -88,6 +102,20 @@ process is identical for all.
|
|
|
88
102
|
|
|
89
103
|
## Gotchas / known issues
|
|
90
104
|
|
|
105
|
+
- **Recurring meetings return the SERIES-ANCHOR `startDateTime`.** Graph's
|
|
106
|
+
`onlineMeetings/{id}` gives the first-occurrence (or even a future) start for a recurring
|
|
107
|
+
series, not the occurrence that was recorded. Driving the archive date off it misfiled every
|
|
108
|
+
occurrence under one frozen date and produced future-dated folders (e.g. a 6-17 occurrence
|
|
109
|
+
filed under `…/2026-06-03/…`; a "Rebranding Weekly Sync" under `2026-06-19/`). Fixed by
|
|
110
|
+
dating off the transcript's `createdDateTime`. The transcript object does **not** carry the
|
|
111
|
+
occurrence start, so `createdDateTime` (≈ minutes after the meeting) is the best reliable
|
|
112
|
+
signal. The `Backfill()` action re-files objects already written under the wrong date:
|
|
113
|
+
recomputes the correct key purely from ledger columns (no Graph calls), `copyObject` →
|
|
114
|
+
`deleteObject` (idempotent — `NoSuchKey` + `doesObjectExist` treats an already-moved row as a
|
|
115
|
+
no-op), then updates `s3Key` + `dtMeeting`. Run `dryRun:true` first, inspect `sample[]`, then
|
|
116
|
+
`dryRun:false` (optional `limit` for a first batch). `$limit` caps candidates touched, not
|
|
117
|
+
rows scanned, so a run of failures stops at `$limit` instead of sweeping the table.
|
|
118
|
+
|
|
91
119
|
- **`getAllTranscripts` requires an AAD object-id (GUID), NOT a UPN.** Passing a UPN
|
|
92
120
|
(e.g. `ajean@togatech.com`) returns the opaque `HTTP 400 BadRequest / "UnknownError"`.
|
|
93
121
|
This is the single most likely cause of an export failure.
|
|
@@ -124,6 +152,10 @@ Policy gap; a `400` only on a UPN means the id was never resolved to a GUID. (A
|
|
|
124
152
|
|
|
125
153
|
## Change history
|
|
126
154
|
|
|
155
|
+
- 2026-06-18 — Fixed recurring-meeting date misfiling: archive date + `dtMeeting` now come from
|
|
156
|
+
the transcript `createdDateTime`, not the series-anchor `meeting->startDateTime` (worker2 PR
|
|
157
|
+
#82). Added `Backfill()` to re-file existing misdated S3 objects; `dtExported` → `UTC_TIMESTAMP()`.
|
|
158
|
+
Rescheduled to weekdays 9:30–17:30 CT, one consolidated cron row (dbchanges2 PR #398). (ajean)
|
|
127
159
|
- 2026-06-12 — **Merged (PR #78) and verified in production**: a real `Team/Transcripts/Export`
|
|
128
160
|
run returned 6/6 organizers, 33 found, 33 exported, 0 errors (classified 30 general / 3 Elite).
|
|
129
161
|
Added fail-fast on unresolvable organizers (`resolveUserId` → null, loop throws clear error
|
package/knowledge/INDEX.md
CHANGED
|
@@ -29,7 +29,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
|
|
|
29
29
|
|
|
30
30
|
## standalone framework
|
|
31
31
|
|
|
32
|
-
- **togatech** (TOGA Technology Website) —
|
|
32
|
+
- **togatech** (TOGA Technology Website) — 2 doc(s) → [standalone/apps/togatech/INDEX.md](standalone/apps/togatech/INDEX.md)
|
|
33
33
|
- **forward** (Forwarder) — 3 doc(s) → [standalone/apps/forward/INDEX.md](standalone/apps/forward/INDEX.md)
|
|
34
34
|
|
|
35
35
|
## Clients
|
|
@@ -3,3 +3,4 @@
|
|
|
3
3
|
| Doc | Summary | Files |
|
|
4
4
|
|-----|---------|-------|
|
|
5
5
|
| [TOGA Technology Website Architecture](architecture.md) | The public-facing TOGA Technology corporate/marketing website. | togatech/src/main.tsx, togatech/src/App.tsx, togatech/src/routes.tsx, togatech/src/lib/api.ts, togatech/src/lib/contentful.ts, togatech/src/themeConfig/ThemeContext.tsx, togatech/vite.config.ts, togatech/package.json |
|
|
6
|
+
| [SEO / AEO / GEO — Prerendering, Single-Source Meta & Structured Data](features/seo-aeo-geo-prerender.md) | Makes togatech.com visible and citable to search engines **and** AI answer engines (ChatGPT/Perplexity/Claude search, Google AI Overviews). | togatech/vite.config.ts, togatech/scripts/prerender.mjs, togatech/src/routes.config.json, togatech/src/main.tsx, togatech/src/App.tsx, togatech/src/components/SEO/Seo.tsx, togatech/src/components/SEO/JsonLd.tsx, togatech/src/components/SEO/schema.ts, togatech/public/robots.prod.txt, togatech/public/sitemap.xml, togatech/public/llms.txt |
|
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: SEO / AEO / GEO — Prerendering, Single-Source Meta & Structured Data
|
|
3
|
+
framework: "standalone"
|
|
4
|
+
repo: togatech
|
|
5
|
+
project: TOGA Technology Website
|
|
6
|
+
client: shared
|
|
7
|
+
type: feature
|
|
8
|
+
status: active
|
|
9
|
+
updated: 2026-06-18
|
|
10
|
+
owners: ["ajean"]
|
|
11
|
+
files:
|
|
12
|
+
- togatech/vite.config.ts
|
|
13
|
+
- togatech/scripts/prerender.mjs
|
|
14
|
+
- togatech/src/routes.config.json
|
|
15
|
+
- togatech/src/main.tsx
|
|
16
|
+
- togatech/src/App.tsx
|
|
17
|
+
- togatech/src/components/SEO/Seo.tsx
|
|
18
|
+
- togatech/src/components/SEO/JsonLd.tsx
|
|
19
|
+
- togatech/src/components/SEO/schema.ts
|
|
20
|
+
- togatech/public/robots.prod.txt
|
|
21
|
+
- togatech/public/sitemap.xml
|
|
22
|
+
- togatech/public/llms.txt
|
|
23
|
+
related: []
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## Summary
|
|
27
|
+
Makes togatech.com visible and citable to search engines **and** AI answer engines (ChatGPT/Perplexity/Claude search, Google AI Overviews). The site is a CSR-only Vite/React SPA, so non-JS crawlers saw an empty `<div id="root">`. This feature fixes a production de-indexing bug and adds **build-time prerendering** so every static route ships real HTML, plus single-source per-page meta and JSON-LD structured data. Shipped on branches `fix-seo-deindex-emergency` (PR #29) and `feature-seo-aeo-prerender` (PR #30) — **not yet merged/deployed** as of 2026-06-18.
|
|
28
|
+
|
|
29
|
+
## Key files / entry points
|
|
30
|
+
- `vite.config.ts` — `robotsPlugin()` now gates on **Vite `mode`** (was `NODE_ENV && VITE_ENV`).
|
|
31
|
+
- `scripts/prerender.mjs` — post-build headless-Chrome (Puppeteer) snapshot; chained into the `build` script.
|
|
32
|
+
- `src/routes.config.json` — single source of static route paths (consumed by the prerender; intended for sitemap too).
|
|
33
|
+
- `src/main.tsx` — `hydrateRoot` when a route was prerendered, else `createRoot`.
|
|
34
|
+
- `src/components/SEO/{Seo.tsx,JsonLd.tsx,schema.ts}` — shared head + JSON-LD; per-page `meta` lives in each page's `viewModel/FIELDS/*.json`.
|
|
35
|
+
|
|
36
|
+
## How it works
|
|
37
|
+
1. **robots:** `robotsPlugin` copies `robots.prod.txt` only when `mode === "production"`; beta/gamma/dev get `robots.dev.txt` (`Disallow: /`). `robots.prod.txt` carries an explicit AI-bot allow-list (GPTBot, ClaudeBot, PerplexityBot, …) + `https://` sitemap.
|
|
38
|
+
2. **prerender:** `npm run build` = `tsc && vite build && node scripts/prerender.mjs`. The script serves `dist/` (reading the original shell once, serving it for all nav routes), snapshots each route in headless Chrome, and writes `dist/<route>/index.html`. Runs in a real browser, so **no SSR/window guards needed**.
|
|
39
|
+
3. **meta:** each page renders `<Seo {...FIELDS.meta} />` — title/description/keywords/canonical/OG/Twitter from one place per page; canonical/OG URLs derived from one `SITE_URL`. `<JsonLd>` (Helmet `<script type="application/ld+json">`) is baked into the snapshot.
|
|
40
|
+
4. **structured data:** Organization + WebSite site-wide (AppLayout), per-page BreadcrumbList (Seo), platform SoftwareApplication (OurPlatform), LocalBusiness×5 (from Contact locations), Person (from About team) — all **derived from existing FIELDS** (single source).
|
|
41
|
+
|
|
42
|
+
## Data model
|
|
43
|
+
None (static marketing site). All copy/meta/NAP/leaders come from `*FIELDS*.json`; structured data is derived from `FOOTERFIELDS.json`, `CONTACTPAGEFIELDS.json` (`map.locations`), and `ABOUTPAGEFIELDS.json` (`teamMembers`).
|
|
44
|
+
|
|
45
|
+
## Client variations
|
|
46
|
+
None — uniform.
|
|
47
|
+
|
|
48
|
+
## Gotchas / known issues
|
|
49
|
+
- **De-indexing bug (root cause):** `robotsPlugin` required `NODE_ENV && VITE_ENV==="production"`, but build scripts set neither → every build shipped `robots.dev.txt` (`Disallow: /`). **Do NOT gate robots on `NODE_ENV`** — `vite build` sets it to `production` for *every* mode, which would expose beta/gamma. Gate on Vite `mode`.
|
|
50
|
+
- **`.npmrc` has `ignore-scripts=true`** → the `postbuild` lifecycle hook never fires. Prerender is therefore chained directly in the `build` script, not a `postbuild` hook.
|
|
51
|
+
- **Prerender tooling:** `vite-react-ssg` supports React Router **6 only** (project is RR 7.9); `react-snap` is stale (2022) and won't launch on Node 24. Hence a custom Puppeteer snapshot. Puppeteer's Chrome must be installed at build time: `npx puppeteer browsers install chrome`.
|
|
52
|
+
- **Prerender fail-fast:** a route that never mounts (`#root > *` times out) aborts the build with `process.exit(1)` — otherwise it would snapshot the empty shell (no SEO/JSON-LD) and still exit 0.
|
|
53
|
+
- **🔴 Hosting requirement:** CloudFront/S3 must serve each prerendered `dist/<route>/index.html` for its path; a blanket SPA rewrite to root `index.html` serves the empty shell and defeats the prerender.
|
|
54
|
+
- **og:image dedupe:** static OG/Twitter tags were removed from `index.html` because `<Seo>` emits per-page ones (Helmet appends rather than replacing pre-existing static tags → duplicates otherwise).
|
|
55
|
+
- **Verified data corrections:** real positioning is *integrated IT services + the ERA TOGa Platform*; real socials are LinkedIn `/toga-tech` + YouTube + Instagram (no Twitter/GitHub/Crunchbase); contact is `website@togatech.com` / `+1 (212) 736-0111`; **5** offices. Prior planning docs had several wrong specifics — verify schema/llms.txt against FIELDS before shipping.
|
|
56
|
+
- **Build size:** JS bundle is ~14.8 MB (4.9 MB gzip) — a real CWV/LCP risk; code-splitting is a separate follow-up.
|
|
57
|
+
|
|
58
|
+
## Change history
|
|
59
|
+
- 2026-06-18 — Initial SEO/AEO/GEO implementation: robots mode-gate + AI allow-list + sitemap/llms; Puppeteer prerender (fail-fast); shared `<Seo>` single-source meta + canonical; JSON-LD (Organization/WebSite/Breadcrumb/SoftwareApplication/LocalBusiness×5/Person×12). PRs #29 + #30, pending merge. (ajean)
|
|
60
|
+
|
|
61
|
+
## Related docs
|
|
62
|
+
- standalone/apps/togatech/architecture.md (update its Build/deploy section once PR #30 merges — robotsPlugin gating + prerender step).
|
package/package.json
CHANGED