teems 0.3.1 → 0.3.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 5932afeb278892c20e7e3fbb92980da31ff45cd9a528a3e2463de38ee7ea3f70
4
- data.tar.gz: 331a38b48c8bb1d87ff770a0579e1846aa643b5d1d1daf5061c0bbf8f0589264
3
+ metadata.gz: bdc6ea71a9d23f39629fd24d240992925cca1482b53b69df377a84870d66367f
4
+ data.tar.gz: 1d563de02f05c8a120130e3e792b508bbf4d1b2467937951a63f2a3efbb19410
5
5
  SHA512:
6
- metadata.gz: 5d750e3b09de3a6ce07612398e0e00a7eb5ac647aa9cd38d45023cd0605777625a9ec212aa81c111b69118d068ffd3b4bd946280d4194cc1ecd34f40708adf7f
7
- data.tar.gz: 355288361bbfa4e21166b18b248793403a023411e3588be731e205c9fe4f6994684426ae7037a09cf009bdb699e9a17caf932572aea9ab1588449c3298224624
6
+ metadata.gz: 8c903fdae8e9946614e093ab4370e74abc03f7fab0d0a76509da02d07814e0083382f4a8f46e0e9131dd1eb3b16f0c36c5eeb96bb381778679bb92028574ad13
7
+ data.tar.gz: b3e6ff9fa9ec3adb9708a3ab97b30929d53a820e9d6239daaf4d359cee179c32448b306a7ac3ab26f336aa3c723402868a1206711637af791a974aa0e3f0e832
data/CHANGELOG.md CHANGED
@@ -1,5 +1,15 @@
1
1
  # Changelog
2
2
 
3
+ ## [0.3.2] - 2026-09-29
4
+
5
+ ### Added
6
+ - `teems transcripts sync` discovers calendar Teams meetings and downloads available WebVTT transcripts into a private XDG data directory. The first run scans 30 days; subsequent runs scan seven, with an idempotent manifest and catch-up after time away. `--date`, `--since`, and `--dry-run` are supported. Recordings and audio are not downloaded.
7
+ - Transcript sync writes speaker-turn Markdown copies to `~/.local/share/teems/transcripts-md/` for local search indexes such as qmd.
8
+ - Transcript sync can run a `transcripts.post_sync_command` from `config.json` (for example, a qmd index refresh) after a sync that changed Markdown. The command gets the changed paths in `TEEMS_TRANSCRIPTS_*` environment variables, is bounded by `post_sync_timeout` (default 300 seconds), and warns rather than failing the sync. `--no-post-sync` skips it for one run.
9
+ - Transcript sync keeps one transcript per recording, so meetings that were restarted mid-session no longer lose the later recording's transcript. Existing per-meeting downloads are reused when identical rather than duplicated.
10
+ - `teems meeting --json` prints call events, recording links, and transcript markers as JSON (the option was documented but ignored).
11
+ - `teems meeting --transcript --recording-url URL` downloads the transcript for a specific recording in the meeting instead of the first one.
12
+
3
13
  ## [0.3.1] - 2026-09-02
4
14
 
5
15
  ### Added
data/README.md CHANGED
@@ -87,11 +87,48 @@ teems meeting <thread-id> --transcript -o ~/Downloads # Download transcript (VT
87
87
  teems meeting <thread-id> --recording -o ~/Downloads # Download recording (MP4)
88
88
  teems meeting <thread-id> --recording --transcript -o ~/Downloads # Both, with embedded subtitles
89
89
  teems meeting <event-id> # By calendar event ID (AAMk...)
90
+ teems meeting <event-id> --date 2026-09-29 --json # Call events + recording links as JSON
91
+ teems meeting <event-id> --date 2026-09-29 --transcript --recording-url <url> # A specific recording
90
92
  teems meeting "https://teams.microsoft.com/..." # By Teams URL or recap link
91
93
  ```
92
94
 
93
95
  Recording download requires `ffmpeg` (`brew install ffmpeg`) and downloads via DASH streaming with 5 parallel threads. No browser required.
94
96
 
97
+ #### Transcript sync
98
+
99
+ ```bash
100
+ teems transcripts sync --dry-run # Preview meetings and recording counts
101
+ teems transcripts sync # First run: 30 days; later: 7 days
102
+ teems transcripts sync --date 2026-09-28 # Retry a specific day
103
+ teems transcripts sync --since 30 # Explicit historical backfill
104
+ teems transcripts sync --no-post-sync # Skip the configured post-sync command once
105
+ ```
106
+
107
+ This saves **only WebVTT transcripts** in `~/.local/share/teems/transcripts/` on the machine running the command, separate from live-caption `meeting-capture` files. Each sync also writes speaker-turn Markdown copies to `~/.local/share/teems/transcripts-md/` (regenerated when missing or older than the VTT) for local search tools such as [qmd](https://github.com/tobi/qmd). Each recording keeps its own transcript, so a meeting restarted mid-session yields one VTT per recording. A private manifest in `~/.local/state/teems/transcript-sync.json` makes repeats idempotent and catches up after time away (up to 30 days). Data and state directories are mode 0700; transcripts and the manifest are mode 0600. The calendar scan retries unavailable transcripts within the lookback window; errors return a nonzero exit code. To run every evening on a GFE, schedule `teems transcripts sync` there via launchd.
108
+
109
+ Coverage is **calendar Teams meetings with a saved and accessible recording transcript**, not every meeting attended: ad-hoc calls, meetings with no saved transcript/recording, and inaccessible organizer-owned artifacts cannot be recovered by this route. No audio/video is downloaded, and nothing is copied to another machine.
110
+
111
+ To search the Markdown locally with qmd, keep it in a dedicated index so refreshes don't re-scan other collections:
112
+
113
+ ```bash
114
+ qmd --index teems-transcripts collection add ~/.local/share/teems/transcripts-md --name teems-transcripts --mask '**/*.md'
115
+ qmd --index teems-transcripts update && qmd --index teems-transcripts embed
116
+ qmd --index teems-transcripts query "what did we decide about monitoring?"
117
+ ```
118
+
119
+ To refresh that index automatically, set a post-sync command in `~/.config/teems/config.json`:
120
+
121
+ ```json
122
+ {
123
+ "transcripts": {
124
+ "post_sync_command": "qmd --index teems-transcripts update && qmd --index teems-transcripts embed",
125
+ "post_sync_timeout": 900
126
+ }
127
+ }
128
+ ```
129
+
130
+ The command runs via `sh -c` only after a sync writes or regenerates Markdown; not on `--dry-run`, when nothing changed, or with `--no-post-sync`. It receives `TEEMS_TRANSCRIPTS_CHANGED` (newline-separated Markdown paths), `TEEMS_TRANSCRIPTS_CHANGED_COUNT`, `TEEMS_TRANSCRIPTS_MARKDOWN_DIR`, and `TEEMS_TRANSCRIPTS_DIR`, and is killed after `post_sync_timeout` seconds (default 300). A failure or timeout prints a warning but does not change the sync's exit code.
131
+
95
132
  ### People
96
133
 
97
134
  ```bash
data/lib/teems/cli.rb CHANGED
@@ -48,6 +48,7 @@ module Teems
48
48
  'channels' => Commands::Channels,
49
49
  'chats' => Commands::Chats,
50
50
  'meeting' => Commands::Meeting,
51
+ 'transcripts' => Commands::Transcripts,
51
52
  'messages' => Commands::Messages,
52
53
  'sync' => Commands::Sync,
53
54
  'who' => Commands::Who,
@@ -9,6 +9,7 @@ module Teems
9
9
  ['channels', 'List joined teams and channels'],
10
10
  ['chats', 'List recent chats'],
11
11
  ['meeting', 'View meeting details, chat, transcripts, and recordings'],
12
+ ['transcripts', 'Sync saved meeting transcripts to this machine'],
12
13
  ['messages', 'Read messages from a channel or chat'],
13
14
  ['sync', 'Sync chat history locally'],
14
15
  ['who', "Look up a user's profile"],
@@ -83,6 +84,7 @@ module Teems
83
84
  teems channels List all channels
84
85
  teems chats List recent chats
85
86
  teems messages <channel-id> Read messages from a channel
87
+ teems transcripts sync Download recent meeting transcripts
86
88
  teems who Show your profile
87
89
  teems who john Search for a user
88
90
  teems org Show your org chart
@@ -24,10 +24,12 @@ module Teems
24
24
  --chat Show meeting chat messages
25
25
  --date YYYY-MM-DD Pick a single occurrence of a recurring meeting by date
26
26
  (filters call events, recordings, and transcripts)
27
+ --recording-url URL With --transcript, use this recording's transcript
28
+ (a meeting restarted mid-session has several)
27
29
  -o, --output-dir Directory for downloads (default: current directory)
28
30
  -v, --verbose Show debug output
29
31
  -q, --quiet Suppress output
30
- --json Output as JSON
32
+ --json Output meeting details (call events, recordings) as JSON
31
33
  -h, --help Show this help
32
34
 
33
35
  EXAMPLES:
@@ -39,6 +41,7 @@ module Teems
39
41
  teems meeting 19:meeting_abc123@thread.v2 --recording -o ~/Downloads
40
42
  teems meeting 19:meeting_abc123@thread.v2 --recording --transcript -o ~/Downloads
41
43
  teems meeting "<recurring-url>" --date 2026-05-04 --audio --no-video --transcript
44
+ teems meeting AAMkAGVmMDEz... --date 2026-05-04 --json # List recordings
42
45
  teems meeting AAMkAGVmMDEz... # By calendar event ID
43
46
  HELP
44
47
 
@@ -51,6 +54,7 @@ module Teems
51
54
  '--no-video' => ->(opts, _args) { opts[:no_video] = true },
52
55
  '--chat' => ->(opts, _args) { opts[:chat] = true },
53
56
  '--date' => ->(opts, args) { opts[:date] = args.shift },
57
+ '--recording-url' => ->(opts, args) { opts[:recording_url] = args.shift },
54
58
  '-o' => ->(opts, args) { opts[:output_dir] = args.shift },
55
59
  '--output-dir' => ->(opts, args) { opts[:output_dir] = args.shift }
56
60
  }.freeze
@@ -604,11 +608,29 @@ module Teems
604
608
  end
605
609
  end
606
610
 
611
+ # Machine-readable meeting details; `teems transcripts sync` uses these to enumerate recordings.
612
+ module MeetingJsonSummary
613
+ private
614
+
615
+ def output_meeting_json(target, classified)
616
+ output_json(thread_id: target[:thread_id],
617
+ call_events: classified[:call_events].map { |event| json_call_event(event) },
618
+ recordings: classified[:recordings].map { |rec| rec.slice(:time, :url, :call_id) },
619
+ transcripts: classified[:transcripts].map { |item| item.slice(:time) })
620
+ 0
621
+ end
622
+
623
+ def json_call_event(event)
624
+ event.slice(:time, :call_id, :duration).merge(participants: event[:participants].map { |part| part[:name] })
625
+ end
626
+ end
627
+
607
628
  # View meeting details, chat, transcripts, and recordings
608
629
  class Meeting < Base
609
630
  include MeetingTargetResolver
610
631
  include MeetingMessageParser
611
632
  include MeetingDisplay
633
+ include MeetingJsonSummary
612
634
  include MeetingChatDisplay
613
635
  include MeetingCallFilter
614
636
  include MeetingDateFilter
@@ -691,6 +713,7 @@ module Teems
691
713
  media = media_output_spec
692
714
  return download_media_with_transcript(target, classified, media) if media
693
715
  return download_transcript(target, classified) || 0 if @options[:transcript]
716
+ return output_meeting_json(target, classified) if @options[:json]
694
717
 
695
718
  display_meeting_summary(target, classified)
696
719
  0
@@ -138,24 +138,39 @@ module Teems
138
138
  end
139
139
  end
140
140
 
141
+ # Chooses which recording's transcript to fetch (--recording-url, else the first link).
142
+ module MeetingRecordingSelection
143
+ private
144
+
145
+ def first_recording_url(classified)
146
+ @options[:recording_url] || recording_urls(classified).first
147
+ end
148
+
149
+ def unknown_recording_url?(classified)
150
+ requested = @options[:recording_url]
151
+ requested && !recording_urls(classified).include?(requested)
152
+ end
153
+
154
+ def recording_urls(classified) = classified[:recordings].filter_map { |rec| rec[:url] }
155
+ end
156
+
141
157
  # Downloads meeting transcripts via SharePoint API (no Safari required)
142
158
  module MeetingTranscript
159
+ include MeetingRecordingSelection
143
160
  include EmbedPageParser
144
161
  include MeetingFilename
145
162
 
146
163
  private
147
164
 
148
165
  def download_transcript(target, classified)
166
+ return error('--recording-url is not a recording in this meeting') if unknown_recording_url?(classified)
167
+
149
168
  sharing_url = target[:fileUrl] || first_recording_url(classified)
150
169
  return error('No recording sharing link found for transcript download') unless sharing_url
151
170
 
152
171
  execute_transcript_pipeline(sharing_url)
153
172
  end
154
173
 
155
- def first_recording_url(classified)
156
- classified[:recordings].filter_map { |rec| rec[:url] }.first
157
- end
158
-
159
174
  def execute_transcript_pipeline(sharing_url)
160
175
  info('Fetching transcript via SharePoint...')
161
176
  embed_url = fetch_embed_url(sharing_url)
@@ -0,0 +1,466 @@
1
+ # frozen_string_literal: true
2
+
3
+ require 'English'
4
+ require 'date'
5
+ require 'open3'
6
+ require 'tmpdir'
7
+
8
+ module Teems
9
+ module Commands
10
+ # Fetches saved Teams transcripts for calendar meetings. Transcript retrieval still
11
+ # goes through `teems meeting`, keeping its SharePoint/auth implementation in one place.
12
+ class Transcripts < Base
13
+ HELP = <<~HELP
14
+ teems transcripts - Sync saved meeting transcripts onto this machine
15
+
16
+ USAGE:
17
+ teems transcripts sync [--since DAYS | --date YYYY-MM-DD] [--dry-run] [--no-post-sync]
18
+
19
+ OPTIONS:
20
+ --since DAYS Calendar lookback (default 7; first run 30)
21
+ --date YYYY-MM-DD Only scan one calendar date
22
+ --dry-run List meetings and recording counts without downloading
23
+ --no-post-sync Skip the configured post-sync command for this run
24
+ -q, --quiet Suppress progress (not errors)
25
+
26
+ WebVTTs: ~/.local/share/teems/transcripts/
27
+ Markdown (for local search, e.g. qmd): ~/.local/share/teems/transcripts-md/
28
+ Sync state: ~/.local/state/teems/transcript-sync.json
29
+ Only meetings on your Teams calendar with a saved, accessible recording
30
+ transcript can be retrieved; each recording's transcript is kept (a meeting
31
+ restarted mid-session has several). Does not alter meeting-capture files.
32
+
33
+ POST-SYNC COMMAND:
34
+ Set "transcripts": {"post_sync_command": "..."} in ~/.config/teems/config.json to
35
+ run a shell command (e.g. a qmd index refresh) after a sync that changed Markdown.
36
+ It receives TEEMS_TRANSCRIPTS_CHANGED (newline-separated Markdown paths),
37
+ TEEMS_TRANSCRIPTS_CHANGED_COUNT, TEEMS_TRANSCRIPTS_MARKDOWN_DIR, and
38
+ TEEMS_TRANSCRIPTS_DIR. "post_sync_timeout" (seconds, default 300) bounds it.
39
+ A failing or timed-out command is reported but does not fail the sync.
40
+ HELP
41
+
42
+ def execute
43
+ validation = validate_options
44
+ return validation if validation
45
+ unless positional_args == ['sync']
46
+ return error('Usage: teems transcripts sync [--since DAYS | --date YYYY-MM-DD]')
47
+ end
48
+ return error('--since must be between 1 and 366') unless (1..366).cover?(@options.fetch(:since, 7))
49
+ return error('Invalid --date (expected YYYY-MM-DD)') if @options[:invalid_date]
50
+
51
+ TranscriptSyncEngine.new(@options, output, hook: transcript_settings).run
52
+ end
53
+
54
+ protected
55
+
56
+ def handle_option(arg, pending)
57
+ case arg
58
+ when '--since' then @options[:since] = Integer(pending.shift, exception: false)
59
+ when '--date' then parse_date_option(pending.shift)
60
+ when '--dry-run' then @options[:dry_run] = true
61
+ when '--no-post-sync' then @options[:no_post_sync] = true
62
+ else super
63
+ end
64
+ end
65
+
66
+ def transcript_settings
67
+ settings = config['transcripts']
68
+ settings.is_a?(Hash) ? settings : {}
69
+ end
70
+
71
+ def parse_date_option(value)
72
+ @options[:date] = Date.iso8601(value)
73
+ rescue ArgumentError, TypeError
74
+ @options[:invalid_date] = true
75
+ end
76
+
77
+ def help_text = HELP
78
+ end
79
+
80
+ # Private on-disk storage, separate from meeting-capture.
81
+ module TranscriptSyncFiles
82
+ private
83
+
84
+ def prepare_private_directories
85
+ [@output_dir, @state_dir].each do |dir|
86
+ FileUtils.mkdir_p(dir, mode: 0o700)
87
+ File.chmod(0o700, dir)
88
+ end
89
+ end
90
+
91
+ def log_locked
92
+ log 'Another transcript sync is running; skipping.'
93
+ 0
94
+ end
95
+
96
+ def load_manifest
97
+ File.file?(@manifest_path) ? JSON.parse(File.read(@manifest_path)) : { 'events' => {} }
98
+ end
99
+
100
+ def downloaded?(key, manifest)
101
+ prior = manifest.fetch('events', {})[key]
102
+ prior && valid_vtt?(File.join(@output_dir, File.basename(prior.fetch('file'))))
103
+ end
104
+
105
+ def persist_download(date, event, key, file, manifest)
106
+ adopted = adopt_legacy_download(date, event, file, manifest)
107
+ target_name = adopted || store_download(file, key)
108
+ manifest['events'][key] = { 'file' => target_name, 'downloaded_at' => Time.now.iso8601,
109
+ 'date' => date.iso8601, 'subject' => event_name(event) }
110
+ save_manifest(manifest)
111
+ log "#{date}: #{adopted ? 'already had' : 'saved'} #{target_name}"
112
+ end
113
+
114
+ def store_download(file, key)
115
+ target_name = "#{File.basename(file, '.vtt')}--#{key[0, 10]}.vtt"
116
+ target = File.join(@output_dir, target_name)
117
+ File.rename(file, target) unless valid_vtt?(target)
118
+ File.chmod(0o600, target)
119
+ target_name
120
+ end
121
+
122
+ # Earlier syncs kept one transcript per event. Reuse that file when it is
123
+ # this recording's transcript instead of saving a duplicate.
124
+ def adopt_legacy_download(date, event, file, manifest)
125
+ key = legacy_key(date, event)
126
+ legacy = manifest['events'][key]
127
+ return unless legacy
128
+
129
+ existing = File.join(@output_dir, File.basename(legacy.fetch('file')))
130
+ return unless valid_vtt?(existing) && FileUtils.identical?(existing, file)
131
+
132
+ manifest['events'].delete(key)
133
+ File.basename(existing)
134
+ end
135
+
136
+ def valid_vtt?(file)
137
+ File.file?(file) && File.size(file) > 10 && File.open(file, 'rb') { |io| io.read(6) == 'WEBVTT' }
138
+ end
139
+
140
+ def valid_download?(status, files)
141
+ status.success? && files.length == 1 && valid_vtt?(files.first)
142
+ end
143
+
144
+ def save_manifest(manifest)
145
+ write_private(@manifest_path, "#{JSON.pretty_generate(manifest)}\n")
146
+ end
147
+
148
+ def write_private(path, content)
149
+ temp = "#{path}.tmp.#{$PROCESS_ID}"
150
+ File.open(temp, File::WRONLY | File::CREAT | File::TRUNC, 0o600) { |file| file.write(content) }
151
+ File.rename(temp, path)
152
+ File.chmod(0o600, path)
153
+ ensure
154
+ File.delete(temp) if temp && File.exist?(temp)
155
+ end
156
+ end
157
+
158
+ # Speaker-turn Markdown copies of each WebVTT, suitable for a local search index.
159
+ module TranscriptSyncMarkdown
160
+ private
161
+
162
+ def sync_markdown(manifest)
163
+ FileUtils.mkdir_p(@markdown_dir, mode: 0o700)
164
+ File.chmod(0o700, @markdown_dir)
165
+ results = manifest.fetch('events', {}).values.map { |entry| export_markdown(entry) }
166
+ written = results.count(:written)
167
+ log "Markdown: updated #{written} transcript(s)" if written.positive?
168
+ results
169
+ end
170
+
171
+ def export_markdown(entry)
172
+ vtt = File.join(@output_dir, File.basename(entry.fetch('file')))
173
+ return :missing unless valid_vtt?(vtt)
174
+
175
+ target = File.join(@markdown_dir, "#{File.basename(vtt, '.vtt')}.md")
176
+ return :current if markdown_current?(target, vtt)
177
+
178
+ write_private(target, markdown_for(vtt, entry))
179
+ @changed_markdown << target
180
+ :written
181
+ rescue StandardError => e
182
+ markdown_failure(entry, e)
183
+ end
184
+
185
+ def markdown_current?(target, vtt) = File.file?(target) && File.mtime(target) >= File.mtime(vtt)
186
+
187
+ def markdown_failure(entry, error)
188
+ failure "Markdown export failed for #{File.basename(entry['file'].to_s)} (#{error.class}: #{error.message})"
189
+ :failed
190
+ end
191
+
192
+ def markdown_for(vtt, entry)
193
+ stem = File.basename(vtt, '.vtt').sub(/--\h{10}\z/, '')
194
+ Formatters::TranscriptMarkdown.new(
195
+ File.read(vtt, encoding: 'bom|utf-8'),
196
+ title: entry['subject'] || stem.sub(/\A\d{4}-\d{2}-\d{2} - /, '').sub(/-\d{8}_\d{6}UTC\z/, ''),
197
+ date: entry['date'] || transcript_date(stem),
198
+ source: File.basename(vtt)
199
+ ).render
200
+ end
201
+
202
+ def transcript_date(stem)
203
+ return stem[0, 10] if stem.match?(/\A\d{4}-\d{2}-\d{2} - /)
204
+
205
+ stem.match(/-(\d{4})(\d{2})(\d{2})_\d{6}UTC\z/)&.captures&.join('-')
206
+ end
207
+ end
208
+
209
+ # Calendar discovery is intentionally limited to scheduled Teams meetings.
210
+ module TranscriptSyncCalendar
211
+ private
212
+
213
+ def calendar_events(date)
214
+ stdout, stderr, status = Open3.capture3(teems_executable, 'cal', '--date', date.iso8601, '--json')
215
+ raise "teems cal failed (exit #{status.exitstatus}): #{redacted_error(stderr)}" unless status.success?
216
+
217
+ return [] if stdout.strip == 'No events found'
218
+
219
+ events = JSON.parse(stdout)
220
+ raise 'teems cal returned a non-array response' unless events.is_a?(Array)
221
+
222
+ events.select { |event| teams_event?(event) }
223
+ end
224
+
225
+ def teams_event?(event)
226
+ return false unless event.is_a?(Hash)
227
+ return false if event['is_cancelled'] || event['is_all_day'] || event['response_status'] == 'declined'
228
+
229
+ event['id'].to_s.start_with?('AAMk') &&
230
+ event['online_meeting_url'].to_s.start_with?('https://teams.microsoft.com/')
231
+ end
232
+
233
+ def teems_executable = ENV.fetch('TEEMS_EXECUTABLE', 'teems')
234
+
235
+ def redacted_error(text)
236
+ lines = text.to_s.scrub.lines.map(&:strip).reject { |line| line.empty? || line.include?('warning:') }
237
+ message = lines.find { |line| line.start_with?('Error:') } || lines.first || 'unknown error'
238
+ message.gsub(%r{https?://\S+}, '[URL]').gsub(/Bearer\s+\S+/i, 'Bearer [REDACTED]')[0, 220]
239
+ end
240
+ end
241
+
242
+ # One transcript per recording: a meeting restarted mid-session has several.
243
+ module TranscriptSyncRecordings
244
+ NO_TRANSCRIPT = Regexp.union(/No recording sharing link found/i,
245
+ /No transcripts found for this recording/i,
246
+ /No meeting activity found for/i)
247
+
248
+ private
249
+
250
+ def sync_event(date, event, manifest)
251
+ urls = recording_urls(date, event)
252
+ return unavailable?(date, event) if urls.empty?
253
+ return preview_event?(date, event, urls, manifest) if @options[:dry_run]
254
+
255
+ urls.map { |url| sync_recording?(date, event, url, manifest) }.all?
256
+ rescue StandardError => e
257
+ failure "#{date}: failed: #{event_name(event)} (#{e.class}: #{e.message})"
258
+ false
259
+ end
260
+
261
+ def recording_urls(date, event)
262
+ stdout, stderr, status = Open3.capture3(teems_executable, 'meeting', event.fetch('id'),
263
+ '--date', date.iso8601, '--json')
264
+ return recordings_from(stdout) if status.success?
265
+ return [] if NO_TRANSCRIPT.match?("#{stderr} #{stdout}")
266
+
267
+ raise "teems meeting failed (exit #{status.exitstatus}): #{redacted_error("#{stderr} #{stdout}")}"
268
+ end
269
+
270
+ def recordings_from(json)
271
+ recordings = JSON.parse(json).fetch('recordings', [])
272
+ recordings.select { |rec| rec['url'] }.sort_by { |rec| rec['time'].to_s }.map { |rec| rec['url'] }.uniq
273
+ end
274
+
275
+ def recording_key(date, event, url) = Digest::SHA256.hexdigest("#{date.iso8601}:#{event.fetch('id')}:#{url}")
276
+
277
+ def legacy_key(date, event) = Digest::SHA256.hexdigest("#{date.iso8601}:#{event.fetch('id')}")
278
+
279
+ def sync_recording?(date, event, url, manifest)
280
+ key = recording_key(date, event, url)
281
+ downloaded?(key, manifest) || download_recording?(date, event, url, key, manifest)
282
+ end
283
+
284
+ def preview_event?(date, event, urls, manifest)
285
+ pending = urls.count { |url| !downloaded?(recording_key(date, event, url), manifest) }
286
+ pending -= 1 if pending.positive? && downloaded?(legacy_key(date, event), manifest)
287
+ log "#{date}: candidate: #{event_name(event)} (#{urls.length} recording(s), #{pending} not yet saved)"
288
+ true
289
+ end
290
+
291
+ def download_recording?(date, event, url, key, manifest)
292
+ Dir.mktmpdir('teems-transcript-', @output_dir) do |temp_dir|
293
+ stdout, stderr, status = Open3.capture3(teems_executable, 'meeting', event.fetch('id'),
294
+ '--date', date.iso8601, '--transcript',
295
+ '--recording-url', url, '-o', temp_dir)
296
+ files = Dir.glob(File.join(temp_dir, '*.vtt'))
297
+ return download_failure?(date, event, status, "#{stderr} #{stdout}") unless valid_download?(status, files)
298
+
299
+ persist_download(date, event, key, files.first, manifest)
300
+ end
301
+ true
302
+ end
303
+
304
+ def download_failure?(date, event, status, message)
305
+ return unavailable?(date, event) if NO_TRANSCRIPT.match?(message)
306
+
307
+ failure "#{date}: failed: #{event_name(event)} (exit #{status.exitstatus}; #{redacted_error(message)})"
308
+ false
309
+ end
310
+
311
+ def unavailable?(date, event)
312
+ log "#{date}: unavailable via teems: #{event_name(event)}"
313
+ true
314
+ end
315
+ end
316
+
317
+ # Optional user command run after a sync changes Markdown, e.g. to refresh a qmd index.
318
+ # A failing or slow command is reported but never fails the sync itself.
319
+ module TranscriptSyncHook
320
+ DEFAULT_HOOK_TIMEOUT = 300
321
+
322
+ private
323
+
324
+ def run_post_sync_hook
325
+ command = @hook['post_sync_command'].to_s.strip
326
+ return if command.empty? || @options[:no_post_sync] || @changed_markdown.empty?
327
+
328
+ log "Post-sync: running command for #{@changed_markdown.length} changed transcript(s)"
329
+ report_hook(*execute_hook(command))
330
+ rescue StandardError => e
331
+ @output.warn("Post-sync command could not run (#{e.class}: #{e.message})")
332
+ end
333
+
334
+ # Runs in its own process group so a timeout can stop the whole pipeline.
335
+ def execute_hook(command)
336
+ Open3.popen2e(hook_env, 'sh', '-c', command, pgroup: true) do |stdin, stdout, wait|
337
+ stdin.close
338
+ reader = Thread.new { stdout.read }
339
+ finished = wait.join(hook_timeout)
340
+ stop_hook(wait.pid) unless finished
341
+ [finished ? wait.value : nil, reader.value]
342
+ end
343
+ end
344
+
345
+ def stop_hook(pid)
346
+ Process.kill('KILL', -pid)
347
+ rescue Errno::ESRCH
348
+ nil
349
+ end
350
+
351
+ def hook_env
352
+ { 'TEEMS_TRANSCRIPTS_CHANGED' => @changed_markdown.join("\n"),
353
+ 'TEEMS_TRANSCRIPTS_CHANGED_COUNT' => @changed_markdown.length.to_s,
354
+ 'TEEMS_TRANSCRIPTS_MARKDOWN_DIR' => @markdown_dir,
355
+ 'TEEMS_TRANSCRIPTS_DIR' => @output_dir }
356
+ end
357
+
358
+ def hook_timeout = [@hook['post_sync_timeout']].grep(Numeric).find(&:positive?) || DEFAULT_HOOK_TIMEOUT
359
+
360
+ def report_hook(status, hook_output)
361
+ return log('Post-sync: done') if status&.success?
362
+
363
+ reason = status ? "exit #{status.exitstatus}" : "timed out after #{hook_timeout}s"
364
+ @output.warn("Post-sync command failed (#{reason}): #{hook_tail(hook_output)}")
365
+ end
366
+
367
+ def hook_tail(text)
368
+ lines = text.to_s.scrub.lines.map(&:strip).reject(&:empty?).last(3)
369
+ lines.empty? ? 'no output' : lines.join(' | ')[0, 300]
370
+ end
371
+ end
372
+
373
+ # Local manifest and replay engine. Never transfers transcripts to another machine.
374
+ class TranscriptSyncEngine
375
+ include TranscriptSyncFiles
376
+ include TranscriptSyncCalendar
377
+ include TranscriptSyncMarkdown
378
+ include TranscriptSyncRecordings
379
+ include TranscriptSyncHook
380
+
381
+ DEFAULT_LOOKBACK = 7
382
+ INITIAL_LOOKBACK = 30
383
+
384
+ def initialize(options, output, hook: {})
385
+ @options = options
386
+ @output = output
387
+ @hook = hook
388
+ @changed_markdown = []
389
+ data_home = ENV.fetch('XDG_DATA_HOME', File.join(Dir.home, '.local', 'share'))
390
+ state_home = ENV.fetch('XDG_STATE_HOME', File.join(Dir.home, '.local', 'state'))
391
+ @output_dir = File.join(data_home, 'teems', 'transcripts')
392
+ @markdown_dir = File.join(data_home, 'teems', 'transcripts-md')
393
+ @state_dir = File.join(state_home, 'teems')
394
+ @manifest_path = File.join(@state_dir, 'transcript-sync.json')
395
+ end
396
+
397
+ def run
398
+ File.umask(0o077)
399
+ return scan(load_manifest) if @options[:dry_run]
400
+
401
+ prepare_private_directories
402
+ File.open(File.join(@state_dir, 'transcript-sync.lock'), File::RDWR | File::CREAT, 0o600) do |lock|
403
+ return log_locked unless lock.flock(File::LOCK_EX | File::LOCK_NB)
404
+
405
+ scan(load_manifest)
406
+ end
407
+ end
408
+
409
+ private
410
+
411
+ def log(message)
412
+ @output.puts(message) unless @options[:quiet]
413
+ end
414
+
415
+ def failure(message)
416
+ @output.error(message)
417
+ end
418
+
419
+ def dates_to_scan(manifest)
420
+ return [@options[:date]] if @options[:date]
421
+
422
+ days = lookback_days(manifest)
423
+ log "Scanning #{days} day(s) through #{Date.today} on this machine"
424
+ ((Date.today - days + 1)..Date.today).to_a
425
+ end
426
+
427
+ def lookback_days(manifest)
428
+ days = @options.fetch(:since, DEFAULT_LOOKBACK)
429
+ return days unless days == DEFAULT_LOOKBACK
430
+
431
+ previous = manifest['last_scan_date']
432
+ return INITIAL_LOOKBACK unless previous
433
+
434
+ # Catch up after time away, capped at the requested initial window.
435
+ (Date.today - Date.iso8601(previous) + 1).to_i.clamp(days, INITIAL_LOOKBACK)
436
+ end
437
+
438
+ def scan(manifest)
439
+ results = dates_to_scan(manifest).map { |date| scan_date(date, manifest) }
440
+ unless @options[:dry_run]
441
+ persist_scan(manifest, results.include?(nil))
442
+ results << !sync_markdown(manifest).include?(:failed)
443
+ run_post_sync_hook
444
+ end
445
+ results.all? ? 0 : 1
446
+ end
447
+
448
+ def scan_date(date, manifest)
449
+ results = calendar_events(date).map { |event| sync_event(date, event, manifest) }
450
+ results.all?
451
+ rescue StandardError => e
452
+ failure "#{date}: calendar error (#{e.class}: #{e.message})"
453
+ nil
454
+ end
455
+
456
+ def persist_scan(manifest, calendar_failed)
457
+ # Failed artifacts stay in the log and get seven-day retries. A calendar
458
+ # failure leaves the wider scan pending so historical days aren't lost.
459
+ manifest['last_scan_date'] = Date.today.iso8601 unless @options[:date] || calendar_failed
460
+ save_manifest(manifest)
461
+ end
462
+
463
+ def event_name(event) = event['subject'].to_s.gsub(/[\r\n]/, ' ')[0, 100]
464
+ end
465
+ end
466
+ end
@@ -0,0 +1,93 @@
1
+ # frozen_string_literal: true
2
+
3
+ require 'json'
4
+
5
+ module Teems
6
+ module Formatters
7
+ # Converts Teams WebVTT transcripts into speaker-turn Markdown for local search indexes.
8
+ # Consecutive cues from the same speaker are merged so each paragraph carries context.
9
+ class TranscriptMarkdown
10
+ Cue = Data.define(:start, :speaker, :text)
11
+
12
+ TIMING = /\A(?<start>(?:\d+:)?\d{2}:\d{2}\.\d{3})\s+-->/
13
+ VOICE = /<v(?:\.[^\s>]+)?\s+([^>]+)>/
14
+ ENTITIES = {
15
+ '&amp;' => '&', '&lt;' => '<', '&gt;' => '>', '&quot;' => '"', '&#39;' => "'",
16
+ '&apos;' => "'", '&nbsp;' => ' ', '&lrm;' => '', '&rlm;' => ''
17
+ }.freeze
18
+ ENTITY_PATTERN = Regexp.union(ENTITIES.keys)
19
+ MAX_TURN_CHARS = 1500
20
+
21
+ def initialize(vtt, title:, date: nil, source: nil)
22
+ @vtt = vtt.to_s
23
+ @title = single_line(title).then { |value| value.empty? ? 'Teams transcript' : value }
24
+ @date = date
25
+ @source = source
26
+ end
27
+
28
+ def render
29
+ lines = front_matter + ["# #{@title}", '', ['Teams transcript', @date].compact.join(' · '), '']
30
+ turns.each { |turn| lines.push(format_turn(turn), '') }
31
+ "#{lines.join("\n").rstrip}\n"
32
+ end
33
+
34
+ def turns
35
+ cues.each_with_object([]) { |cue, turns| append_cue(turns, cue) }
36
+ end
37
+
38
+ private
39
+
40
+ def append_cue(turns, cue)
41
+ last = turns.last
42
+ if last && last[:speaker] == cue.speaker && last[:text].length < MAX_TURN_CHARS
43
+ last[:text] << ' ' << cue.text
44
+ else
45
+ turns << { start: cue.start, speaker: cue.speaker, text: +cue.text }
46
+ end
47
+ end
48
+
49
+ def cues
50
+ normalized = @vtt.scrub.delete_prefix("\uFEFF").gsub(/\r\n?/, "\n")
51
+ normalized.split(/\n{2,}/).filter_map { |block| parse_cue(block) }
52
+ end
53
+
54
+ def parse_cue(block)
55
+ lines = block.lines.map(&:chomp)
56
+ timing_index = lines.index { |line| TIMING.match?(line) }
57
+ return unless timing_index
58
+
59
+ build_cue(lines[timing_index][TIMING, :start], lines[(timing_index + 1)..].join(' '))
60
+ end
61
+
62
+ def build_cue(start, raw)
63
+ text = clean(raw)
64
+ return if text.empty?
65
+
66
+ Cue.new(start: start, speaker: raw[VOICE, 1]&.then { |name| clean(name) }, text: text)
67
+ end
68
+
69
+ def clean(text) = single_line(decode(text.gsub(/<[^>]*>/, '')))
70
+
71
+ def decode(text) = text.gsub(ENTITY_PATTERN, ENTITIES)
72
+
73
+ def single_line(text) = text.to_s.gsub(/\s+/, ' ').strip
74
+
75
+ def front_matter
76
+ [
77
+ '---',
78
+ "title: #{@title.to_json}",
79
+ ("date: #{@date}" if @date),
80
+ 'source: teams-transcript',
81
+ ("vtt: #{@source.to_json}" if @source),
82
+ '---',
83
+ ''
84
+ ].compact
85
+ end
86
+
87
+ def format_turn(turn)
88
+ speaker = turn[:speaker] || 'Unknown speaker'
89
+ "**#{speaker}** (#{turn[:start].sub(/\.\d+\z/, '')}): #{turn[:text]}"
90
+ end
91
+ end
92
+ end
93
+ end
data/lib/teems/version.rb CHANGED
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module Teems
4
- VERSION = '0.3.1'
4
+ VERSION = '0.3.2'
5
5
  end
data/lib/teems.rb CHANGED
@@ -78,6 +78,7 @@ module Teems
78
78
  autoload :MessageFormatter, 'teems/formatters/message_formatter'
79
79
  autoload :MarkdownFormatter, 'teems/formatters/markdown_formatter'
80
80
  autoload :CalendarFormatter, 'teems/formatters/calendar_formatter'
81
+ autoload :TranscriptMarkdown, 'teems/formatters/transcript_markdown'
81
82
  end
82
83
 
83
84
  # CLI commands implementing user-facing functionality
@@ -95,6 +96,7 @@ module Teems
95
96
  autoload :Ooo, 'teems/commands/ooo'
96
97
  autoload :Org, 'teems/commands/org'
97
98
  autoload :Meeting, 'teems/commands/meeting'
99
+ autoload :Transcripts, 'teems/commands/transcripts'
98
100
  autoload :Status, 'teems/commands/status'
99
101
  end
100
102
 
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: teems
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.3.1
4
+ version: 0.3.2
5
5
  platform: ruby
6
6
  authors:
7
7
  - Eric Boehs
@@ -49,12 +49,14 @@ files:
49
49
  - lib/teems/commands/org.rb
50
50
  - lib/teems/commands/status.rb
51
51
  - lib/teems/commands/sync.rb
52
+ - lib/teems/commands/transcripts.rb
52
53
  - lib/teems/commands/who.rb
53
54
  - lib/teems/formatters/calendar_formatter.rb
54
55
  - lib/teems/formatters/format_utils.rb
55
56
  - lib/teems/formatters/markdown_formatter.rb
56
57
  - lib/teems/formatters/message_formatter.rb
57
58
  - lib/teems/formatters/output.rb
59
+ - lib/teems/formatters/transcript_markdown.rb
58
60
  - lib/teems/models/account.rb
59
61
  - lib/teems/models/channel.rb
60
62
  - lib/teems/models/chat.rb