Everything you need to follow a meeting in another language
This is what AI Translator can do today, exactly as it works in the app. Anything that is only in the paid plans is marked clearly.
Features at a glance
Live translation
5 languages, auto-detect
Subtitle bar
Drag, lock, resize the text
Transcript
History, export (Pro)
Glossary
Names and terms (Pro)
Audio source
Any meeting app, no bot
Shortcuts and tray
Control without leaving the meeting
Models and hardware
Standard and Lite packs
Performance
Measured, with conditions
AI translation without putting your meeting in the cloud
Two AI models, one for speech recognition and one for translation, run directly on your computer. No step sends your audio or meeting content to a server or a cloud AI service.
- Your computermeeting audio
- Internet
- Provider serverplus an AI on the cloud
- Internet
- Subtitlesback to you
The meeting audio leaves your computer and is processed somewhere you do not control.
- Audioon your computer
- AI running on your computerWhisper + Hy-MT2
- Subtitleson your screen
No trip to the cloud: the audio, the transcript and the translation stay on your computer.
Two AI models running locally
Whisper recognizes speech and Hy-MT2 translates. Both are downloaded once and then run on your machine using the GPU (Metal on macOS) or the CPU.
Low latency because there is no detour
There is no round trip to a server, so the translation appears right after the speaker finishes a sentence. Median under 1.1 seconds on a Mac M4 Pro; the test conditions are in the Performance section below.
Conversation data stays on your computer
Audio lives only in RAM; the transcript and translation stay on your machine. History is off by default and, if you turn it on, is encrypted on your computer.
Listen, recognize, translate and show subtitles in one loop
AI Translator captures the audio playing on your computer, splits it into sentences, recognizes the speech, translates it and shows the result on the subtitle bar. On a Mac M4 Pro the full translation appears about a second (median) after the speaker finishes a sentence; other computers may be slower, see the measurements.
- Five languages for both the source audio and the translation: English, 中文, 日本語, 한국어, Tiếng Việt. We plan to add more languages in the future (no schedule yet)
- Automatic detection of the spoken language among the ones you tick under Languages spoken in the meeting, or lock a single one in Source language when you know what will be spoken
- A sentence that is already in the language you want to read is shown as it is, not translated again
- A sentence that is not final yet appears dimmer, then is replaced by the complete sentence when the speaker carries on
- Filters out the "phantom" sentences that speech recognition tends to invent over music or silence
- Recognized Chinese text is converted to Simplified characters
One direction: from the meeting to you.
AI Translator translates the audio coming out of your computer into your language. It does not translate your own voice into the meeting yet. If both sides install the app, each person sees subtitles of the other side.
A floating bar you can customize, out of your meeting's way
The subtitle bar is its own borderless window with a translucent background. It stays on top and never takes focus from your meeting app, so you can keep typing in chat or clicking buttons as usual.
- Drag it to move it, drag an edge or corner to resize it; its position and size are remembered for each screen
- Text size 14–48 px, 5 text colors (white, yellow, green, light blue, orange), 5 background colors, background opacity 0–100 %
- Lock: mouse clicks pass through the bar and no buttons get in the way; unlock it with a shortcut or from the tray menu
- Keeps the last 1,000 sentences; scroll back with the mouse wheel or a shortcut, and the Latest button takes you back to the current sentence
- The original text in small type above the translation (on by default, can be turned off)
- Small indicators in the corner: listening with sound detected, loading models, falling behind, less than 5 minutes of translation left

Any meeting app, no bot, no plugin
Because the app captures system audio, AI Translator works with anything that makes sound: Zoom, Microsoft Teams, Google Meet, Zalo PC, webinars, videos and online courses.
- macOS: listen to the whole system (except the app itself) or to a single app that is playing sound, so notification sounds from other apps are not translated by mistake
- Windows (when released): the default playback device, or a device you choose
- The source reopens automatically when you switch playback devices, for example when you plug in headphones or connect Bluetooth
- Pause that ends a sentence is adjustable from 50 to 800 ms: shorter gives you subtitles sooner, longer cuts fewer sentences in half
On macOS, the first time you press Start the system asks for the system audio recording permission. The app does not use the microphone. See how to grant the audio permission.
Keep what you heard, your way
The transcript of the current session works on every plan. Saving history and exporting to a file are Pro features.
- Transcript: every line has a time, the original sentence and the translation; search and copy all on every plan
- Export to TXT, SRT or Markdown; for SRT you choose whether the text is the translation or the original Pro
- Session history: review, open and delete sessions one by one, or all at once Pro
- History is off by default. When you turn it on, it is stored on your computer, encrypted, with the key kept in the Keychain (macOS) or Credential Manager (Windows)
- Delete all data with one button (after a confirmation), on every plan
Read the guide to history and export
Hints for names and terms, passed to the translator
Add source → target pairs for product names, partner names and industry terms. When a sentence contains one of them, the app passes it to the translator as a hint.
- Up to 500 terms, with CSV import and export (two columns, UTF-8)
- Matching ignores case; Chinese, Japanese and Korean terms also match inside words
- Up to 20 relevant entries are used for each sentence
Terms are hints for the translator, so correct use is not guaranteed every time. The app says so too.
Control it without leaving the meeting
Five global shortcuts work even when AI Translator is not the window in front. You can change each one in Settings › Shortcuts.
| Action | macOS | Windows (when released) |
|---|---|---|
| Start or stop translating | ⌃ + ⌥ + T | Ctrl + Alt + T |
| Show or hide subtitles | ⌃ + ⌥ + H | Ctrl + Alt + H |
| Lock or unlock subtitles | ⌃ + ⌥ + L | Ctrl + Alt + L |
| Scroll subtitles up (older sentences) | ⌃ + ⌥ + PageUp | Ctrl + Alt + PageUp |
| Scroll subtitles down (newer sentences) | ⌃ + ⌥ + PageDown | Ctrl + Alt + PageDown |
On macOS, ⌃ is Control and ⌥ is Option. A shortcut needs at least one of Ctrl, Alt or Cmd/Win. The Windows version is not released yet.
The menu bar icon (macOS) or system tray icon (Windows, when released) lets you start or stop translating, show or hide and lock the subtitles, and open the main window. Closing the window only hides the app in the tray; to quit completely, choose Quit.
Two model packs, and the app recommends the one that fits
The recognition and translation models are downloaded once and then run entirely on your computer. The app checks your RAM, free disk space and graphics card to recommend a pack.
| Detail | Standard pack | Lite pack |
|---|---|---|
| Download | about 2.5 GB | about 1.3 GB |
| Recognition | Whisper large-v3-turbo | Whisper small |
| Translation | Hy-MT2-1.8B (Q8_0) | Hy-MT2-1.8B (Q4_K_M) |
| Engine RAM (M4 Pro) | about 2.9 GiB | about 1.8–1.9 GiB |
| The app recommends it for | A Mac with 16 GB or more; Windows with 16 GB or more and a discrete card with 6 GB of VRAM | Computers with 8 GB up to under 16 GB, or Windows without a capable discrete card |
The Lite pack recognizes speech less clearly than Standard in Vietnamese, Japanese, Korean and Chinese. If you listen to those languages a lot, choose the Standard pack.
- macOS
- macOS 14.2 or later, Apple Silicon (M1 or newer)
No version for Intel Macs - RAM
- At least 8 GB, 16 GB recommended
- Disk
- At least 1 GB free on top of the size of the model being downloaded
- Windows
- Windows 10/11 64-bit, CPU with AVX2: not released yet
For the Standard pack, a discrete Vulkan-capable graphics card with 6 GB or more of VRAM is recommended
Real measurements, with the conditions
We publish only what we measured, on the machine we measured it on, and we say which machines we have not tested.
| Measurement (Mac M4 Pro 24 GB, macOS 26, Metal GPU) | Standard pack | Lite pack |
|---|---|---|
| Median (p50) delay: from the speaker finishing a sentence to the complete translation appearing | 0.76–1.03 s | 0.61–0.84 s |
| p90 delay | 0.94–1.34 s | 0.73–1.14 s |
| First translated word appears (p50) | 0.63–0.69 s | 0.52–0.57 s |
| RAM of the two engine processes | 2.9 GiB | 1.8–1.9 GiB |
Long sessions
One continuous session of 5 h 23 min (a lecture video in English, on an ad-hoc-signed release build): 6,019 segments, no errors, median delay 0.48 s. In a 2-hour stability test, no process was restarted.
Translation and recognition quality
In an internal test on 320 sentences in five directions involving Vietnamese, Hy-MT2-1.8B scored 0.837 on COMET: clearly above MADLAD-3B (0.779) and NLLB-600M (0.736), and on par with HY-MT1.5-1.8B (0.833). We only compared open-source models with each other. On clean read speech, Vietnamese recognition in the Standard pack has a word error rate of 8.7 %. That is read speech, not real meeting conversation.
Translation and recognition quality by direction
COMET scores were measured through the app's own translation path (llama-server with the same prompt), 100 sentences per direction (Vietnamese → Chinese, Japanese, Korean: 40 sentences). COMET is a relative score from 0 to 1, higher is better; it is not a percentage accuracy.
| Direction (text) | Standard pack | Lite pack |
|---|---|---|
| English → Vietnamese | 0.842 | 0.841 |
| 中文 → Vietnamese | 0.829 | 0.831 |
| 日本語 → Vietnamese | 0.830 | 0.815 |
| 한국어 → Vietnamese | 0.834 | 0.822 |
| Vietnamese → English | 0.821 | 0.822 |
| Vietnamese → 中文 | 0.836 | 0.821 |
| Vietnamese → 日本語 | 0.847 | 0.845 |
| Vietnamese → 한국어 | 0.851 | 0.842 |
| Speech recognition (error rate, lower is better) | Standard pack | Lite pack |
|---|---|---|
| English (word errors) | 5.4 % | 6.6 % |
| Vietnamese (word errors) | 8.7 % | 22.5 % |
| 中文 (character errors) | 5.6 % | 9.6 % |
| 日本語 (character errors) | 4.5 % | 13.1 % |
| 한국어 (character errors) | 4.1 % | 8.2 % |
How to read this: the translation test set leans toward everyday spoken language, so it is only a rough proxy for meeting speech; recognition was measured on clean read speech (about 15 minutes per language), not real meeting conversations, and narrowband Bluetooth headsets raise the error rate. All eight directions above involve Vietnamese; the other 12 directions among English, 中文, 日本語 and 한국어 run but have no quality score yet. The Lite pack is noticeably less accurate at recognizing Vietnamese, Japanese, Korean and Chinese.
What we have not measured, and do not promise
All the figures above were measured on a single Mac M4 Pro, with a 300 ms pause to end a sentence (the app's current default is 50 ms; the 5-hour session used the default). We have not measured: a base Mac M1, 8 GB computers, discrete Windows graphics cards, battery and power use, accuracy on real meeting conversation or over Bluetooth headsets, and translation quality for directions that do not involve Vietnamese (they work, but we have no quality score to publish). A preliminary test on one Windows laptop with integrated graphics showed that the Standard pack did not meet our latency target.
Built on open-source components, running on your computer
AI Translator is a commercial product; the components below are third-party open-source projects, listed in full on the About screen of the app.
Whisper + whisper.cpp
OpenAI's multilingual speech recognition (MIT license), run through whisper.cpp.
Hy-MT2-1.8B + llama.cpp
Tencent's translation model (Apache 2.0 license) in GGUF format, run through llama.cpp.
Silero VAD
Detects speech so sentences can be cut at the right place. MIT license, runs in pure Rust inside the app.
Tauri 2, Rust, React
A Rust core with a React interface. The two engines run in separate processes, so a GPU driver failure cannot take the app down.
Metal and Vulkan
GPU acceleration: Metal on macOS, Vulkan on Windows. It falls back to the CPU if the GPU fails.
SQLCipher
History, if you turn it on, is encrypted with SQLCipher; the random key is kept in your operating system's key store.
Which feature is in which plan
| Feature | Free (trial) | Monthly | Yearly |
|---|---|---|---|
| Live translated subtitles, 5 languages | Yes | Yes | Yes |
| Customizable subtitle bar, shortcuts, tray | Yes | Yes | Yes |
| View, search and copy the transcript | Yes | Yes | Yes |
| Glossary (up to 500 terms) | No | Yes | Yes |
| Save session history | No | Yes | Yes |
| Export to TXT, SRT, Markdown | No | Yes | Yes |
| Translation time | 30 min per day, for 10 days | 50 hours per 30 days | Unlimited, for 365 days |
| Price (VND) | 0 ₫ | 50,000 ₫ | 500,000 ₫ |
Try it on your real meetings
A free 10-day trial with 30 minutes a day. No card, no account.