@capgo/capacitor-speech-recognition 8.0.9 → 8.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -35,8 +35,8 @@ The most complete doc is available here: https://capgo.app/docs/plugins/speech-r
35
35
  ## Install
36
36
 
37
37
  ```bash
38
- npm install @capgo/capacitor-speech-recognition
39
- npx cap sync
38
+ bun add @capgo/capacitor-speech-recognition
39
+ bunx cap sync
40
40
  ```
41
41
 
42
42
  ## Usage
@@ -66,6 +66,97 @@ await SpeechRecognition.stop();
66
66
  await partialListener.remove();
67
67
  ```
68
68
 
69
+ ## On-device recognition mode
70
+
71
+ This plugin now supports an opt-in on-device recognition path behind the explicit
72
+ `useOnDeviceRecognition` flag.
73
+
74
+ ### What it is
75
+
76
+ The default path keeps the long-standing recognizer flow for backward compatibility.
77
+ `useOnDeviceRecognition` switches to a newer local speech pipeline when the platform supports it:
78
+
79
+ - On iOS 26+, it uses Apple's `SpeechAnalyzer` / `SpeechTranscriber` stack.
80
+ - On recent Android versions, it uses the on-device `SpeechRecognizer` path.
81
+
82
+ ### Why you might want it
83
+
84
+ - Better alignment with the latest native speech APIs.
85
+ - Improved on-device model handling on supported platforms.
86
+ - A cleaner rollout path if you want to adopt newer speech stacks without changing every user immediately.
87
+
88
+ ### Why it is opt-in
89
+
90
+ Even when a new stack is technically available, changing recognition behavior silently can affect:
91
+
92
+ - transcript wording
93
+ - punctuation behavior
94
+ - partial-result timing
95
+ - product metrics and user expectations
96
+
97
+ That is why the plugin keeps the legacy recognizer by default and requires an explicit flag for the new path.
98
+
99
+ ### Recommended rollout
100
+
101
+ 1. Check generic speech support with `available()`.
102
+ 2. Check the on-device path with `isOnDeviceRecognitionAvailable()`.
103
+ 3. Enable `useOnDeviceRecognition` only when that second check returns `true`.
104
+ 4. Roll it out gradually if your app depends on stable transcripts or analytics.
105
+
106
+ ### Example
107
+
108
+ ```ts
109
+ import { SpeechRecognition } from '@capgo/capacitor-speech-recognition';
110
+
111
+ await SpeechRecognition.requestPermissions();
112
+
113
+ const { available } = await SpeechRecognition.available();
114
+ if (!available) {
115
+ throw new Error('Speech recognition is not available on this device.');
116
+ }
117
+
118
+ const { available: onDeviceRecognitionAvailable } =
119
+ await SpeechRecognition.isOnDeviceRecognitionAvailable({
120
+ language: 'en-US',
121
+ });
122
+
123
+ await SpeechRecognition.start({
124
+ language: 'en-US',
125
+ partialResults: true,
126
+ useOnDeviceRecognition: onDeviceRecognitionAvailable,
127
+ });
128
+ ```
129
+
130
+ ### When not to use it yet
131
+
132
+ Stay on the default path if:
133
+
134
+ - you need unchanged behavior for existing users
135
+ - you have not validated transcripts for your target locale
136
+ - you want identical production behavior across older and newer OS versions
137
+
138
+ ### Platform notes
139
+
140
+ - iOS uses the newer on-device path only on iOS 26+ and only for locales Apple exposes through the newer speech stack.
141
+ - Android uses the on-device recognizer only in inline mode. `popup: true` keeps using the system dialog and is not compatible with `useOnDeviceRecognition`.
142
+ - On Android, a supported on-device language may require a model download before recognition can begin.
143
+
144
+ ## Push-to-talk and session events
145
+
146
+ This plugin also supports a push-to-talk oriented flow built around three APIs:
147
+
148
+ - `setPTTState({ held })` lets your UI tell the plugin when the button is pressed or released.
149
+ - `forceStop()` stops the active session immediately and emits the last cached partial result with `forced: true` when available.
150
+ - `getLastPartialResult()` lets you read back the latest cached transcript at any point.
151
+
152
+ `continuousPTT` is the experimental cross-platform mode that keeps a held push-to-talk session alive by restarting recognition as speech segments finalize. Android and iOS both support this restart flow for inline/native recognition.
153
+
154
+ The plugin also emits deterministic session lifecycle events so UIs can react cleanly:
155
+
156
+ - `listeningState` now carries `state`, `sessionId`, `reason`, and optional `errorCode` in addition to the legacy `status`.
157
+ - `error` is emitted for every native recognizer error instead of relying only on promise rejections.
158
+ - `readyForNextSession` signals when native resources are torn down and the plugin is ready for another start.
159
+
69
160
  ### iOS usage descriptions
70
161
 
71
162
  Add the following keys to your app `Info.plist`:
@@ -78,8 +169,12 @@ Add the following keys to your app `Info.plist`:
78
169
  <docgen-index>
79
170
 
80
171
  * [`available()`](#available)
172
+ * [`isOnDeviceRecognitionAvailable(...)`](#isondevicerecognitionavailable)
81
173
  * [`start(...)`](#start)
82
174
  * [`stop()`](#stop)
175
+ * [`forceStop(...)`](#forcestop)
176
+ * [`getLastPartialResult()`](#getlastpartialresult)
177
+ * [`setPTTState(...)`](#setpttstate)
83
178
  * [`getSupportedLanguages()`](#getsupportedlanguages)
84
179
  * [`isListening()`](#islistening)
85
180
  * [`checkPermissions()`](#checkpermissions)
@@ -89,6 +184,8 @@ Add the following keys to your app `Info.plist`:
89
184
  * [`addListener('segmentResults', ...)`](#addlistenersegmentresults-)
90
185
  * [`addListener('partialResults', ...)`](#addlistenerpartialresults-)
91
186
  * [`addListener('listeningState', ...)`](#addlistenerlisteningstate-)
187
+ * [`addListener('error', ...)`](#addlistenererror-)
188
+ * [`addListener('readyForNextSession', ...)`](#addlistenerreadyfornextsession-)
92
189
  * [`removeAllListeners()`](#removealllisteners)
93
190
  * [Interfaces](#interfaces)
94
191
  * [Type Aliases](#type-aliases)
@@ -111,6 +208,33 @@ Checks whether the native speech recognition service is usable on the current de
111
208
  --------------------
112
209
 
113
210
 
211
+ ### isOnDeviceRecognitionAvailable(...)
212
+
213
+ ```typescript
214
+ isOnDeviceRecognitionAvailable(options?: Pick<SpeechRecognitionStartOptions, "language"> | undefined) => Promise<SpeechRecognitionAvailability>
215
+ ```
216
+
217
+ Checks whether the platform's newer on-device recognition path is available for the selected locale.
218
+
219
+ This is the capability check you should use before enabling `useOnDeviceRecognition`.
220
+ A `true` result means the current device, OS version, and locale can use the newer
221
+ on-device path for that platform.
222
+
223
+ Returns `false` when the device only supports the legacy recognizer path.
224
+
225
+ Platform SDK docs:
226
+ iOS: [Speech](https://developer.apple.com/documentation/speech)
227
+ Android: [SpeechRecognizer](https://developer.android.com/reference/android/speech/SpeechRecognizer)
228
+
229
+ | Param | Type |
230
+ | ------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
231
+ | **`options`** | <code><a href="#pick">Pick</a>&lt;<a href="#speechrecognitionstartoptions">SpeechRecognitionStartOptions</a>, 'language'&gt;</code> |
232
+
233
+ **Returns:** <code>Promise&lt;<a href="#speechrecognitionavailability">SpeechRecognitionAvailability</a>&gt;</code>
234
+
235
+ --------------------
236
+
237
+
114
238
  ### start(...)
115
239
 
116
240
  ```typescript
@@ -120,7 +244,11 @@ start(options?: SpeechRecognitionStartOptions | undefined) => Promise<SpeechReco
120
244
  Begins capturing audio and transcribing speech.
121
245
 
122
246
  When `partialResults` is `true`, the returned promise resolves immediately and updates are
123
- streamed through the `partialResults` listener until {@link stop} is called.
247
+ streamed through the `partialResults` listener until the session ends.
248
+
249
+ The default path keeps the legacy recognizer behavior for backward compatibility.
250
+ Pass `useOnDeviceRecognition: true` only after checking
251
+ {@link SpeechRecognitionPlugin.isOnDeviceRecognitionAvailable}.
124
252
 
125
253
  | Param | Type |
126
254
  | ------------- | --------------------------------------------------------------------------------------- |
@@ -142,6 +270,56 @@ Stops listening and tears down native resources.
142
270
  --------------------
143
271
 
144
272
 
273
+ ### forceStop(...)
274
+
275
+ ```typescript
276
+ forceStop(options?: ForceStopOptions | undefined) => Promise<void>
277
+ ```
278
+
279
+ Force stops the current session.
280
+
281
+ On Android, this first tries a normal stop and then falls back to destroy/recreate after `timeout`.
282
+ On iOS, the current session is stopped immediately.
283
+
284
+ If a partial transcript is cached, it is emitted through the `partialResults` listener with `forced: true`.
285
+
286
+ | Param | Type |
287
+ | ------------- | ------------------------------------------------------------- |
288
+ | **`options`** | <code><a href="#forcestopoptions">ForceStopOptions</a></code> |
289
+
290
+ --------------------
291
+
292
+
293
+ ### getLastPartialResult()
294
+
295
+ ```typescript
296
+ getLastPartialResult() => Promise<LastPartialResult>
297
+ ```
298
+
299
+ Gets the last cached partial transcription result.
300
+
301
+ **Returns:** <code>Promise&lt;<a href="#lastpartialresult">LastPartialResult</a>&gt;</code>
302
+
303
+ --------------------
304
+
305
+
306
+ ### setPTTState(...)
307
+
308
+ ```typescript
309
+ setPTTState(options: PTTStateOptions) => Promise<void>
310
+ ```
311
+
312
+ Updates the current push-to-talk button state.
313
+
314
+ Use this together with `continuousPTT` or with a custom hold-to-talk flow.
315
+
316
+ | Param | Type |
317
+ | ------------- | ----------------------------------------------------------- |
318
+ | **`options`** | <code><a href="#pttstateoptions">PTTStateOptions</a></code> |
319
+
320
+ --------------------
321
+
322
+
145
323
  ### getSupportedLanguages()
146
324
 
147
325
  ```typescript
@@ -283,6 +461,42 @@ Listen for changes to the native listening state.
283
461
  --------------------
284
462
 
285
463
 
464
+ ### addListener('error', ...)
465
+
466
+ ```typescript
467
+ addListener(eventName: 'error', listenerFunc: (event: SpeechRecognitionErrorEvent) => void) => Promise<PluginListenerHandle>
468
+ ```
469
+
470
+ Listen for recognition errors.
471
+
472
+ | Param | Type |
473
+ | ------------------ | ------------------------------------------------------------------------------------------------------- |
474
+ | **`eventName`** | <code>'error'</code> |
475
+ | **`listenerFunc`** | <code>(event: <a href="#speechrecognitionerrorevent">SpeechRecognitionErrorEvent</a>) =&gt; void</code> |
476
+
477
+ **Returns:** <code>Promise&lt;<a href="#pluginlistenerhandle">PluginListenerHandle</a>&gt;</code>
478
+
479
+ --------------------
480
+
481
+
482
+ ### addListener('readyForNextSession', ...)
483
+
484
+ ```typescript
485
+ addListener(eventName: 'readyForNextSession', listenerFunc: (event: SpeechRecognitionReadyEvent) => void) => Promise<PluginListenerHandle>
486
+ ```
487
+
488
+ Listen for the recognizer becoming ready for another session.
489
+
490
+ | Param | Type |
491
+ | ------------------ | ------------------------------------------------------------------------------------------------------- |
492
+ | **`eventName`** | <code>'readyForNextSession'</code> |
493
+ | **`listenerFunc`** | <code>(event: <a href="#speechrecognitionreadyevent">SpeechRecognitionReadyEvent</a>) =&gt; void</code> |
494
+
495
+ **Returns:** <code>Promise&lt;<a href="#pluginlistenerhandle">PluginListenerHandle</a>&gt;</code>
496
+
497
+ --------------------
498
+
499
+
286
500
  ### removeAllListeners()
287
501
 
288
502
  ```typescript
@@ -304,6 +518,23 @@ Removes every registered listener.
304
518
  | **`available`** | <code>boolean</code> |
305
519
 
306
520
 
521
+ #### SpeechRecognitionStartOptions
522
+
523
+ Configure how the recognizer behaves when calling {@link SpeechRecognitionPlugin.start}.
524
+
525
+ | Prop | Type | Description |
526
+ | ---------------------------- | -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
527
+ | **`language`** | <code>string</code> | Locale identifier such as `en-US`. When omitted the device language is used. |
528
+ | **`maxResults`** | <code>number</code> | Maximum number of final matches returned by native APIs. Defaults to `5`. |
529
+ | **`prompt`** | <code>string</code> | Prompt message shown inside the Android system dialog (ignored on iOS). |
530
+ | **`popup`** | <code>boolean</code> | When `true`, Android shows the OS speech dialog instead of running inline recognition. Defaults to `false`. |
531
+ | **`partialResults`** | <code>boolean</code> | Emits partial transcription updates through the `partialResults` listener while audio is captured. |
532
+ | **`addPunctuation`** | <code>boolean</code> | Enables native punctuation handling where supported (iOS 16+). |
533
+ | **`useOnDeviceRecognition`** | <code>boolean</code> | Opt in to the platform's newer on-device recognition path when available. On iOS 26+, this uses Apple's `SpeechAnalyzer` / `SpeechTranscriber` pipeline. On recent Android versions, this uses the on-device `SpeechRecognizer` path. It is intentionally opt-in so existing apps keep the legacy flow unless they choose to roll out the new behavior. Use {@link SpeechRecognitionPlugin.isOnDeviceRecognitionAvailable} before enabling it in production. Platform SDK docs: iOS: [Speech](https://developer.apple.com/documentation/speech), [SpeechAnalyzer](https://developer.apple.com/documentation/speech/speechanalyzer), [SpeechTranscriber](https://developer.apple.com/documentation/speech/speechtranscriber) Android: [SpeechRecognizer](https://developer.android.com/reference/android/speech/SpeechRecognizer) Defaults to `false`. |
534
+ | **`allowForSilence`** | <code>number</code> | Allow a number of milliseconds of silence before splitting the recognition session into segments. Required to be greater than zero and currently supported on Android only. |
535
+ | **`continuousPTT`** | <code>boolean</code> | EXPERIMENTAL: Keep a PTT session alive across silence by restarting recognition while the button stays held. This restart behavior is implemented for Android inline recognition and iOS native recognition. |
536
+
537
+
307
538
  #### SpeechRecognitionMatches
308
539
 
309
540
  | Prop | Type |
@@ -311,19 +542,33 @@ Removes every registered listener.
311
542
  | **`matches`** | <code>string[]</code> |
312
543
 
313
544
 
314
- #### SpeechRecognitionStartOptions
545
+ #### ForceStopOptions
315
546
 
316
- Configure how the recognizer behaves when calling {@link SpeechRecognitionPlugin.start}.
547
+ Options for {@link SpeechRecognitionPlugin.forceStop}.
548
+
549
+ | Prop | Type | Description |
550
+ | ------------- | ------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
551
+ | **`timeout`** | <code>number</code> | Android only: timeout in milliseconds before forcing stop via destroy/recreate. On iOS, the current session is stopped immediately and this value is ignored. Defaults to `1500`. |
552
+
553
+
554
+ #### LastPartialResult
555
+
556
+ Result from {@link SpeechRecognitionPlugin.getLastPartialResult}.
557
+
558
+ | Prop | Type | Description |
559
+ | --------------- | --------------------- | --------------------------------------------------------------- |
560
+ | **`available`** | <code>boolean</code> | Whether a partial result is currently cached. |
561
+ | **`text`** | <code>string</code> | The most recent transcript text known to the native recognizer. |
562
+ | **`matches`** | <code>string[]</code> | All current match alternatives when available. |
317
563
 
318
- | Prop | Type | Description |
319
- | --------------------- | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
320
- | **`language`** | <code>string</code> | Locale identifier such as `en-US`. When omitted the device language is used. |
321
- | **`maxResults`** | <code>number</code> | Maximum number of final matches returned by native APIs. Defaults to `5`. |
322
- | **`prompt`** | <code>string</code> | Prompt message shown inside the Android system dialog (ignored on iOS). |
323
- | **`popup`** | <code>boolean</code> | When `true`, Android shows the OS speech dialog instead of running inline recognition. Defaults to `false`. |
324
- | **`partialResults`** | <code>boolean</code> | Emits partial transcription updates through the `partialResults` listener while audio is captured. |
325
- | **`addPunctuation`** | <code>boolean</code> | Enables native punctuation handling where supported (iOS 16+). |
326
- | **`allowForSilence`** | <code>number</code> | Allow a number of milliseconds of silence before splitting the recognition session into segments. Required to be greater than zero and currently supported on Android only. |
564
+
565
+ #### PTTStateOptions
566
+
567
+ Options for {@link SpeechRecognitionPlugin.setPTTState}.
568
+
569
+ | Prop | Type | Description |
570
+ | ---------- | -------------------- | ----------------------------------------- |
571
+ | **`held`** | <code>boolean</code> | Whether the PTT button is currently held. |
327
572
 
328
573
 
329
574
  #### SpeechRecognitionLanguages
@@ -372,25 +617,77 @@ Raised whenever a segmented result is produced (Android only).
372
617
 
373
618
  Raised whenever a partial transcription is produced.
374
619
 
375
- | Prop | Type |
376
- | ------------- | --------------------- |
377
- | **`matches`** | <code>string[]</code> |
620
+ | Prop | Type | Description |
621
+ | --------------------- | --------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
622
+ | **`matches`** | <code>string[]</code> | Current recognition matches when the native recognizer reports them. This can be omitted for forced or accumulated-only payloads. |
623
+ | **`accumulated`** | <code>string</code> | Accumulated transcription from earlier continuous PTT cycles. |
624
+ | **`accumulatedText`** | <code>string</code> | Final accumulated text including the current result. |
625
+ | **`isRestarting`** | <code>boolean</code> | `true` when the plugin is restarting recognition inside a continuous PTT session. |
626
+ | **`forced`** | <code>boolean</code> | `true` when the payload was emitted by `forceStop()`. |
378
627
 
379
628
 
380
629
  #### SpeechRecognitionListeningEvent
381
630
 
382
631
  Raised when the listening state changes.
383
632
 
384
- | Prop | Type |
385
- | ------------ | ----------------------------------- |
386
- | **`status`** | <code>'started' \| 'stopped'</code> |
633
+ The original `status` field is preserved for backward compatibility and is present
634
+ on the binary `started` / `stopped` states.
635
+
636
+ | Prop | Type | Description |
637
+ | --------------- | --------------------------------------------------------------------- | ---------------------------------------------------------- |
638
+ | **`state`** | <code><a href="#listeningfinitestate">ListeningFiniteState</a></code> | Finite state of the recognition session. |
639
+ | **`sessionId`** | <code>number</code> | Unique identifier for the current listening session. |
640
+ | **`reason`** | <code><a href="#listeningreason">ListeningReason</a></code> | Why this state transition occurred. |
641
+ | **`errorCode`** | <code>string</code> | Error code when the transition is caused by an error. |
642
+ | **`status`** | <code>'started' \| 'stopped'</code> | Backward-compatible binary state used by earlier releases. |
643
+
644
+
645
+ #### SpeechRecognitionErrorEvent
646
+
647
+ Raised whenever native recognition reports an error.
648
+
649
+ | Prop | Type |
650
+ | --------------- | ------------------- |
651
+ | **`code`** | <code>string</code> |
652
+ | **`message`** | <code>string</code> |
653
+ | **`sessionId`** | <code>number</code> |
654
+
655
+
656
+ #### SpeechRecognitionReadyEvent
657
+
658
+ Emitted after native resources have been torn down and the plugin is ready for another session.
659
+
660
+ | Prop | Type |
661
+ | --------------- | ------------------- |
662
+ | **`sessionId`** | <code>number</code> |
387
663
 
388
664
 
389
665
  ### Type Aliases
390
666
 
391
667
 
668
+ #### Pick
669
+
670
+ From T, pick a set of properties whose keys are in the union K
671
+
672
+ <code>{
392
673
  [P in K]: T[P];
393
674
  }</code>
675
+
676
+
394
677
  #### PermissionState
395
678
 
396
679
  <code>'prompt' | 'prompt-with-rationale' | 'granted' | 'denied'</code>
397
680
 
681
+
682
+ #### ListeningFiniteState
683
+
684
+ Finite state values for the recognition session lifecycle.
685
+
686
+ <code>'startingListening' | 'started' | 'stoppingListening' | 'stopped'</code>
687
+
688
+
689
+ #### ListeningReason
690
+
691
+ Why a listening state transition happened.
692
+
693
+ <code>'userStart' | 'userStop' | 'forceStop' | 'results' | 'silence' | 'error' | 'unknown'</code>
694
+
398
695
  </docgen-api>
@@ -12,6 +12,8 @@ public interface Constants {
12
12
  String END_OF_SEGMENT_EVENT = "endOfSegmentedSession";
13
13
  String LISTENING_EVENT = "listeningState";
14
14
  String PARTIAL_RESULTS_EVENT = "partialResults";
15
+ String ERROR_EVENT = "error";
16
+ String READY_FOR_NEXT_SESSION_EVENT = "readyForNextSession";
15
17
  String RECORD_AUDIO_PERMISSION = Manifest.permission.RECORD_AUDIO;
16
18
  String LANGUAGE_ERROR = "Could not get list of languages";
17
19
  }