Measured, not promised
Speed matters. So do the words.
Earlier internal pilot: 119.22 seconds of mixed English, Hindi, and German speech; 228 reference words; M4 Mac running macOS 26.6.2. These are end-to-end transcription timings from a small pilot, not controlled comparisons or iPhone results.
Ultra + Tiny + Whisper Small
- Processing
- 31.47 seconds
- Word error rate
- 31.58%
After bounded context recovery; recovered an empty short German segment.
Ultra + Tiny + Apple → Small
- Processing
- 20.88 seconds
- Word error rate
- 26.75%
Earlier run with different cache/order conditions. Some uncertain speech still used Whisper.
Word error rate (WER) is lower when better. These results are not a speed or accuracy promise. An earlier repeat failed on an empty 2.96-second crop; recovery now retries with nearby context. Language boundaries and scripts still need improvement. No new physical-iPhone, one-hour bulk, or speaker-diarization accuracy benchmark is established by these runs.
A fast model can still miss the words.
Dolphin pilot: six public FLEURS clips, three Hindi (37.92 seconds, 93 reference words) and three Mandarin (32.52 seconds, 81 reference characters). Apple M4, macOS 26.6.2; sherpa-onnx 1.13.8 CPU with two threads. Timings below are warm decode only, outside Meetly; they exclude model load and downloads.
Base · Hindi
- Warm decode
- 0.378 seconds
- Word error rate
- 50.54%
Rejected for Hindi: substantial omissions.
Small · Hindi
- Warm decode
- 0.945 seconds
- Word error rate
- 38.71%
Rejected for Hindi: omissions remain.
Base · Mandarin
- Warm decode
- 0.303 seconds
- Character error rate
- 18.52%
Prototype candidate, not qualified.
Small · Mandarin
- Warm decode
- 1.692 seconds
- Character error rate
- 17.28%
Research comparison, not qualified.
Base model load: 0.170 seconds, peak process memory 517.5 MB. Small: 0.483 seconds, 841.9 MB. These are process measurements, not mobile memory guarantees.
Not a Whisper speedup claim: no same-fixture Whisper comparison was completed for this pilot. Concurrent machine work and execution order were uncontrolled; Small’s slower Mandarin repeat illustrates that limitation. No iPhone or bulk benchmark. Scores retain number-format differences and do not convert between simplified and traditional script. Speaker diarization was not measured. Hindi stays on the selected established fallback.
OmniASR: better Hindi, weaker Mandarin in this pilot.
The same six public clips, M4 CPU, two threads, sherpa-onnx 1.13.8. These are standalone-runtime measurements, not end-to-end Meetly timings.
OmniASR · Hindi
- Warm decode
- 3.391 seconds
- Word error rate
- 21.51%
20 errors / 93 reference words; CER 7.26%. Promising, not qualified.
OmniASR · Mandarin
- Warm decode
- 2.884 seconds
- Character error rate
- 37.04%
30 errors / 81 reference characters; worse than Dolphin here.
Model load 0.396 seconds; peak process memory about 1.28 GiB. The Hindi result covers just three read-speech clips. No iPhone, code-switching, or same-fixture Whisper comparison was completed for this six-clip pilot. Arabic follow-up results appear below. Both repetitions produced identical text. Do not treat advertised multilingual coverage as validated routing.
Arabic, Japanese, and Korean: different tradeoffs.
Nine more public FLEURS clips, three per language. Same M4 CPU and two-thread runtime; one model reused across the batch. Warm decode excludes load. Arabic WER and Japanese/Korean CER are shown; strict Unicode normalization retains diacritics and number differences.
Dolphin Base · Arabic
- Warm decode
- 0.525 seconds
- Word error rate
- 40.58%
38.34s audio; CER 16.81%. Research result, not a qualified route.
Dolphin Base · Japanese
- Warm decode
- 0.475 seconds
- Character error rate
- 19.75%
37.50s audio; CER 19.75%. Research result, not a qualified route.
Dolphin Base · Korean
- Warm decode
- 0.321 seconds
- Character error rate
- 16.54%
32.22s audio; CER 16.54%. Research result, not a qualified route.
OmniASR · Arabic
- Warm decode
- 3.460 seconds
- Word error rate
- 34.78%
38.34s audio; CER 8.85%. Research result, not a qualified route.
OmniASR · Japanese
- Warm decode
- 3.354 seconds
- Character error rate
- 27.16%
37.50s audio; CER 27.16%. Research result, not a qualified route.
OmniASR · Korean
- Warm decode
- 2.882 seconds
- Character error rate
- 11.81%
32.22s audio; CER 11.81%. Research result, not a qualified route.
Omni improves Arabic and Korean character scores here, while Dolphin is better on Japanese. Arabic word errors remain substantial; Korean spacing inflates word-level errors. Peak process memory reached 571 MB for Dolphin and 1.69 GB for Omni; loads were 0.212 and 0.481 seconds. No physical iPhone, conversational, or code-switch validation. These experiments do not enable new routes or establish a universal best model.
SenseVoice: strong Mandarin pilot, not an enabled route.
Nine public FLEURS clips, same three Mandarin, Japanese, and Korean clips. M4 CPU two threads, sherpa-onnx 1.13.8, INT8 ONNX export, automatic language and inverse text normalization enabled. This measures the CPU export, not FluidAudio Core ML or iPhone.
SenseVoice · Mandarin
- Warm decode
- 0.568 seconds
- Character error rate
- 0.00%
0/81 character errors across three clips, 32.52s audio. Research only.
SenseVoice · Japanese
- Warm decode
- 0.705 seconds
- Character error rate
- 8.02%
13/162 character errors across three clips, 37.50s audio. Research only.
SenseVoice · Korean
- Warm decode
- 0.549 seconds
- Character error rate
- 16.54%
21/127 character errors across three clips, 32.22s audio. Research only.
Mandarin zero errors covers only 81 reference characters, not general perfect accuracy. Japanese CER 8.02%; Korean CER 16.54% and whitespace WER 75%. Model load 0.538s; peak process memory 1.17 GB. Commercial-use clarification for the custom FunASR weight license is pending; no SenseVoice route is shipped or enabled. A maintainer discussion updated on 22 September is not a final licensing clearance.
60-minute English · Debug scalability check
One 60.24-minute job completed in the Mac Debug harness at e5387f11: the same 33.78-second English clip repeated 107 times. Ultra + Tiny; all transcription routes stayed on Parakeet, with no fallback ASR.
| Measurement | Time |
|---|
| Audio duration | 3614.46s |
|---|
| Total Auto time | 708.31s |
|---|
| Tiny language detection (722 calls) | 439.25s |
|---|
| Parakeet transcription | 249.08s |
|---|
| Speech detection | 16.85s |
|---|
| Checkpoint resume (identical result) | 0.65s |
|---|
Debug build, single job, repeated source audio, uncontrolled host load. This checks long-job completion and resume, not independent accuracy, bulk queues, Release speed, or physical iPhone behavior. No WER/CER is assigned. Language detection remains the largest measured stage; avoiding fallback does not eliminate classifier cost.
Standalone Parakeet versus Auto: speed and segmentation both matter.
Same Mac native services and public fixtures. Standalone runs include service scheduling/loading and then a repeated warm pass, without language detection. Auto v3 runs include language classification and selected Whisper fallback. Cache and execution order were not controlled.
| Case / mode | Time | WER |
|---|
| mono-enparakeet-v3 · Standalone · first pass | 1.47s | 10.96% |
|---|
| mono-enparakeet-v3 · Standalone · warm repeat | 0.90s | 10.96% |
|---|
| switch-en-deparakeet-v3 · Standalone · first pass | 1.69s | 13.92% |
|---|
| switch-en-deparakeet-v3 · Standalone · warm repeat | 1.05s | 13.92% |
|---|
| mono-enparakeet-ultra · Standalone · first pass | 1.91s | 21.92% |
|---|
| mono-enparakeet-ultra · Standalone · warm repeat | 0.87s | 21.92% |
|---|
| switch-en-deparakeet-ultra · Standalone · first pass | 1.90s | 5.06% |
|---|
| switch-en-deparakeet-ultra · Standalone · warm repeat | 1.05s | 5.06% |
|---|
| mono-env3 · Auto / Tiny / Whisper Small | 3.84s | 12.33% |
|---|
| switch-en-hiv3 · Auto / Tiny / Whisper Small | 11.92s | 40.38% |
|---|
| switch-en-dev3 · Auto / Tiny / Whisper Small | 6.53s | 7.59% |
|---|
On this 33.78-second English fixture, standalone v3 had 10.96% WER and Ultra 21.92%; segmented Auto Ultra had 4.11%. The English/German fixture favored standalone Ultra instead. This is not a global v3-versus-Ultra accuracy ranking. Standalone avoids routing overhead, but a faster transcript is not necessarily a more complete transcript.
A larger classifier did not automatically improve routing.
Early native Mac pilot with Ultra primary and Small fallback. Tiny baseline plus Base/Small classifier runs; same public English, Hindi, mixed English–Hindi, and seven-language fixtures. These are full Auto timings, not detector-only speed.
| Case / detector | Auto time | Error / coverage |
|---|
| mono-enTiny | 3.63s | 4.1% WER / 98.4% |
|---|
| mono-hiTiny | 18.19s | 57.0% WER / 90.9% |
|---|
| switch-en-hiTiny | 11.05s | 37.5% WER / 90.0% |
|---|
| switch-seven-languagesTiny | 28.37s | 14.5% CER / 69.1% |
|---|
| mono-enBase | 120.72s | 4.1% WER / 98.4% |
|---|
| mono-hiBase | 14.71s | 81.7% WER / 35.9% |
|---|
| switch-en-hiBase | 21.39s | 41.3% WER / 69.1% |
|---|
| switch-seven-languagesBase | Failed | Not scored |
|---|
| mono-enSmall | 4.15s | 4.1% WER / 98.4% |
|---|
| mono-hiSmall | 17.17s | 72.0% WER / 32.4% |
|---|
| switch-en-hiSmall | 18.95s | 47.1% WER / 86.3% |
|---|
| switch-seven-languagesSmall | Failed | Not scored |
|---|
Tiny remains the default. Hindi source-utterance language coverage was about 90.9% with Tiny, 35.9% with Base, and 32.4% with Small in these runs. Coverage includes internal utterance silence and counts unknown/mixed as not correct; it is not frame-annotated detector accuracy. Base’s first English run (120.72s) included cold compilation/loading: later cases must not be compared as an isolated model-speed ranking. Both larger classifiers failed on the seven-language fixture. Small sample and decoding errors prevent a general ranking.
iOS simulator · 14-case Auto batch
All 14 cases completed on the iOS 27 arm64 simulator, source e5289bcf. Parakeet v3 + Tiny detection + Custom → Whisper Small. Seven monolingual recordings and seven stitched switches use the same hashed public corpus as the Mac batch. This is a Mac-hosted simulator with 24 GiB host memory, not a physical iPhone test. Hindi therefore passed the memory eligibility check.
| Case / engines | Time | Error rate |
|---|
| mono-enparakeet-tdt-v3 | 11.01s | 10.96% WER |
|---|
| mono-hicustom, whisperkit | 17.11s | 24.73% WER |
|---|
| mono-zhcustom, whisperkit | 6.97s | 19.75% CER |
|---|
| mono-arwhisperkit | 41.57s | 42.03% WER |
|---|
| mono-jawhisperkit | 68.55s | 18.52% CER |
|---|
| mono-kowhisperkit | 42.66s | 52.50% WER |
|---|
| mono-deparakeet-tdt-v3, whisperkit | 17.72s | 6.45% WER |
|---|
| switch-en-hicustom, parakeet-tdt-v3, whisperkit | 25.13s | 24.04% WER |
|---|
| switch-en-zhcustom, parakeet-tdt-v3, whisperkit | 16.33s | 7.04% CER |
|---|
| switch-en-arparakeet-tdt-v3, whisperkit | 33.36s | 35.29% WER |
|---|
| switch-en-japarakeet-tdt-v3, whisperkit | 20.85s | 8.08% CER |
|---|
| switch-en-koparakeet-tdt-v3, whisperkit | 24.76s | 27.54% WER |
|---|
| switch-en-deparakeet-tdt-v3 | 12.22s | 8.86% WER |
|---|
| switch-seven-languagescustom, parakeet-tdt-v3, whisperkit | 122.39s | 12.19% CER |
|---|
App-reported elapsed times include routing and transcription. Separate model preparation took 74.27s; caches and host contention were not controlled. Mac results used Ultra, so these are not a direct platform speed comparison. Mixed-language CER includes all reference characters. Completion does not establish accuracy, thermal behavior, background execution, or memory safety on an iPhone.
Native Auto regression · Final 42-case Mac batch
41 of 42 cases completed; the seven-language Custom-policy case failed. M4 Mac, isolated native app compiled from e5387f11. Public FLEURS fixtures: English, Hindi, Mandarin, Arabic, Japanese, Korean, and German recordings plus seven stitched language-switch cases, each run with three fallback policies. Parakeet Ultra + Tiny classifier; selected Whisper Small fallback. Times include detection, local model load, transcription, and checkpoints, excluding app launch. New process and job per case; filesystem/Core ML caches not flushed.
| Case / fallback | Time | Error rate |
|---|
| mono-enWhisper | 3.82s | 4.11% WER |
|---|
| mono-hiWhisper | 19.75s | 55.91% WER |
|---|
| mono-zhWhisper | 7.22s | 17.28% CER |
|---|
| mono-arWhisper | 9.44s | 40.58% WER |
|---|
| mono-jaWhisper | 7.45s | 20.37% CER |
|---|
| mono-koWhisper | 7.11s | 50.00% WER |
|---|
| mono-deWhisper | 7.40s | 6.45% WER |
|---|
| switch-en-hiWhisper | 11.12s | 37.50% WER |
|---|
| switch-en-zhWhisper | 7.07s | 7.51% CER |
|---|
| switch-en-arWhisper | 10.80s | 30.59% WER |
|---|
| switch-en-jaWhisper | 6.50s | 4.62% CER |
|---|
| switch-en-koWhisper | 8.43s | 23.19% WER |
|---|
| switch-en-deWhisper | 5.61s | 3.80% WER |
|---|
| switch-seven-languagesWhisper | 25.77s | 14.26% CER |
|---|
| mono-enApple → Whisper | 3.34s | 4.11% WER |
|---|
| mono-hiApple → Whisper | 6.02s | 31.18% WER |
|---|
| mono-zhApple → Whisper | 6.65s | 17.28% CER |
|---|
| mono-arApple → Whisper | 9.29s | 44.93% WER |
|---|
| mono-jaApple → Whisper | 7.72s | 20.37% CER |
|---|
| mono-koApple → Whisper | 7.12s | 50.00% WER |
|---|
| mono-deApple → Whisper | 8.96s | 4.84% WER |
|---|
| switch-en-hiApple → Whisper | 9.21s | 24.04% WER |
|---|
| switch-en-zhApple → Whisper | 10.27s | 7.51% CER |
|---|
| switch-en-arApple → Whisper | 18.27s | 36.47% WER |
|---|
| switch-en-jaApple → Whisper | 14.07s | 4.62% CER |
|---|
| switch-en-koApple → Whisper | 10.17s | 23.19% WER |
|---|
| switch-en-deApple → Whisper | 6.75s | 3.80% WER |
|---|
| switch-seven-languagesApple → Whisper | 36.09s | 17.15% CER |
|---|
| mono-enCustom → Whisper | 3.82s | 4.11% WER |
|---|
| mono-hiCustom → Whisper | 9.47s | 24.73% WER |
|---|
| mono-zhCustom → Whisper | 5.14s | 18.52% CER |
|---|
| mono-arCustom → Whisper | 10.45s | 36.23% WER |
|---|
| mono-jaCustom → Whisper | 24.12s | 20.37% CER |
|---|
| mono-koCustom → Whisper | 8.06s | 47.50% WER |
|---|
| mono-deCustom → Whisper | 8.72s | 6.45% WER |
|---|
| switch-en-hiCustom → Whisper | 11.99s | 14.42% WER |
|---|
| switch-en-zhCustom → Whisper | 7.45s | 2.82% CER |
|---|
| switch-en-arCustom → Whisper | 15.05s | 29.41% WER |
|---|
| switch-en-jaCustom → Whisper | 7.65s | 4.62% CER |
|---|
| switch-en-koCustom → Whisper | 11.06s | 23.19% WER |
|---|
| switch-en-deCustom → Whisper | 7.66s | 3.80% WER |
|---|
| switch-seven-languagesCustom → Whisper | Failed | Not scored |
|---|
In the English-only fixture, all three policies used Parakeet without a fallback transcription pass (3.34–3.83 seconds for 33.78 seconds of audio). This is a small Mac pilot, not an iPhone promise. Apple policy uses Apple only for eligible installed locales; other segments use Whisper. Custom policy uses only its enabled language/device routes and otherwise Whisper. Mixed-language CER includes English characters: it is not a Mandarin-only score. Sequential, contended execution and cache differences prevent an isolated speed ranking. Failure means no complete transcript and no invented error rate. No physical-iPhone claim.
- Initial batch: all 14 Custom cases failed at model setup because of a path bug; these had no ASR accuracy score.
- After the path fix (3290a8de): 12 of 14 Custom cases completed; standalone Mandarin and the seven-language case still failed.
- The next fix gave short speech fragments wider context without changing ownership boundaries. This final batch preserves all cases, including failures, instead of selecting only successful outputs.
Speaker counts matter as much as the aggregate error.
Two independent, human-annotated AMI ES2004a English meeting crops; 180 seconds each. Production offline diarizer through an isolated Mac test harness, pyannote.metrics 4.0.0, 0.25-second collar, overlap included, optimal speaker mapping.
60–240 second crop
- Diarization error rate
- 6.255%
- Speakers found / reference
- 1 / 2
The minority speaker was not recovered as a separate cluster. A low overall error masks that omission. Native processing: 2.131 seconds.
180–360 second crop
- Diarization error rate
- 24.435%
- Speakers found / reference
- 2 / 4
Chosen from reference annotations before inference. Both shorter speakers had zero recall; speaker-count coverage failed. Native processing: 2.201 seconds.
iOS simulator · 60–240 second crop
- Diarization error rate
- 6.255%
- Speakers found / reference
- 1 / 2
Same human-reference crop and scoring protocol. Minority speakers remain missed. Mac-hosted simulator; not physical iPhone performance. Native processing: 9.988 seconds.
iOS simulator · 180–360 second crop
- Diarization error rate
- 24.435%
- Speakers found / reference
- 2 / 4
Same human-reference crop and scoring protocol. Minority speakers remain missed. Mac-hosted simulator; not physical iPhone performance. Native processing: 8.692 seconds.
These are actual speaker-diarization errors, separate from transcription WER/CER. Timings include native model validation/loading. Two English windows do not qualify multilingual diarization or robust multi-speaker separation. The simulator repeats used the same production diarizer and no enrollment; separate simulator setup took 22.39s. These are repeats of two crops, not four independent samples.