Everything you need to follow a meeting in another language

This is what AI Translator can do today, exactly as it works in the app. Anything that is only in the paid plans is marked clearly.

On-device AI

AI translation without putting your meeting in the cloud

Two AI models, one for speech recognition and one for translation, run directly on your computer. No step sends your audio or meeting content to a server or a cloud AI service.

How most cloud translation tools work
  1. Your computermeeting audio
  2. Internet
  3. Provider serverplus an AI on the cloud
  4. Internet
  5. Subtitlesback to you

The meeting audio leaves your computer and is processed somewhere you do not control.

AI Translator
  1. Audioon your computer
  2. AI running on your computerWhisper + Hy-MT2
  3. Subtitleson your screen

No trip to the cloud: the audio, the transcript and the translation stay on your computer.

Two AI models running locally

Whisper recognizes speech and Hy-MT2 translates. Both are downloaded once and then run on your machine using the GPU (Metal on macOS) or the CPU.

Low latency because there is no detour

There is no round trip to a server, so the translation appears right after the speaker finishes a sentence. Median under 1.1 seconds on a Mac M4 Pro; the test conditions are in the Performance section below.

Conversation data stays on your computer

Audio lives only in RAM; the transcript and translation stay on your machine. History is off by default and, if you turn it on, is encrypted on your computer.

Live translation

Listen, recognize, translate and show subtitles in one loop

AI Translator captures the audio playing on your computer, splits it into sentences, recognizes the speech, translates it and shows the result on the subtitle bar. On a Mac M4 Pro the full translation appears about a second (median) after the speaker finishes a sentence; other computers may be slower, see the measurements.

  • Five languages for both the source audio and the translation: English, 中文, 日本語, 한국어, Tiếng Việt. We plan to add more languages in the future (no schedule yet)
  • Automatic detection of the spoken language among the ones you tick under Languages spoken in the meeting, or lock a single one in Source language when you know what will be spoken
  • A sentence that is already in the language you want to read is shown as it is, not translated again
  • A sentence that is not final yet appears dimmer, then is replaced by the complete sentence when the speaker carries on
  • Filters out the "phantom" sentences that speech recognition tends to invent over music or silence
  • Recognized Chinese text is converted to Simplified characters
AI Translator main screen with the Languages card: the language to translate into and the languages spoken in the meeting
The Languages card: choose the language you want to read and the languages that may be spoken in the meeting.

One direction: from the meeting to you.

AI Translator translates the audio coming out of your computer into your language. It does not translate your own voice into the meeting yet. If both sides install the app, each person sees subtitles of the other side.

Subtitle bar

A floating bar you can customize, out of your meeting's way

The subtitle bar is its own borderless window with a translucent background. It stays on top and never takes focus from your meeting app, so you can keep typing in chat or clicking buttons as usual.

  • Drag it to move it, drag an edge or corner to resize it; its position and size are remembered for each screen
  • Text size 14–48 px, 5 text colors (white, yellow, green, light blue, orange), 5 background colors, background opacity 0–100 %
  • Lock: mouse clicks pass through the bar and no buttons get in the way; unlock it with a shortcut or from the tray menu
  • Keeps the last 1,000 sentences; scroll back with the mouse wheel or a shortcut, and the Latest button takes you back to the current sentence
  • The original text in small type above the translation (on by default, can be turned off)
  • Small indicators in the corner: listening with sound detected, loading models, falling behind, less than 5 minutes of translation left
Subtitle bar with yellow text on a navy background at a large font size
The same bar with yellow text, a navy background and a size of 26.
Subtitles settings: font size, text color, background color, background opacity and the option to show the original text
Settings › Subtitles: every change shows on the bar immediately.
Audio settings listing the apps that are playing sound so you can choose one as the source
Settings › Audio: listen to the whole system or to just one app (macOS).
Audio source

Any meeting app, no bot, no plugin

Because the app captures system audio, AI Translator works with anything that makes sound: Zoom, Microsoft Teams, Google Meet, Zalo PC, webinars, videos and online courses.

  • macOS: listen to the whole system (except the app itself) or to a single app that is playing sound, so notification sounds from other apps are not translated by mistake
  • Windows (when released): the default playback device, or a device you choose
  • The source reopens automatically when you switch playback devices, for example when you plug in headphones or connect Bluetooth
  • Pause that ends a sentence is adjustable from 50 to 800 ms: shorter gives you subtitles sooner, longer cuts fewer sentences in half

On macOS, the first time you press Start the system asks for the system audio recording permission. The app does not use the microphone. See how to grant the audio permission.

Transcript, history and export

Keep what you heard, your way

The transcript of the current session works on every plan. Saving history and exporting to a file are Pro features.

Transcript with time, original sentence and translation, a search box, a Copy all button and an export format picker
The transcript: search, copy, export to TXT, SRT or Markdown.
  • Transcript: every line has a time, the original sentence and the translation; search and copy all on every plan
  • Export to TXT, SRT or Markdown; for SRT you choose whether the text is the translation or the original Pro
  • Session history: review, open and delete sessions one by one, or all at once Pro
  • History is off by default. When you turn it on, it is stored on your computer, encrypted, with the key kept in the Keychain (macOS) or Credential Manager (Windows)
  • Delete all data with one button (after a confirmation), on every plan
History screen listing saved sessions with date and time, minutes, number of sentences and a preview
History: sessions saved on your computer.
Glossary with pairs of source terms and translations and buttons to import and export CSV
Glossary: add, edit, delete, import and export CSV.
Glossary Pro

Hints for names and terms, passed to the translator

Add source → target pairs for product names, partner names and industry terms. When a sentence contains one of them, the app passes it to the translator as a hint.

  • Up to 500 terms, with CSV import and export (two columns, UTF-8)
  • Matching ignores case; Chinese, Japanese and Korean terms also match inside words
  • Up to 20 relevant entries are used for each sentence

Terms are hints for the translator, so correct use is not guaranteed every time. The app says so too.

Shortcuts and system tray

Control it without leaving the meeting

Five global shortcuts work even when AI Translator is not the window in front. You can change each one in Settings › Shortcuts.

ActionmacOSWindows (when released)
Start or stop translating⌃ + ⌥ + TCtrl + Alt + T
Show or hide subtitles⌃ + ⌥ + HCtrl + Alt + H
Lock or unlock subtitles⌃ + ⌥ + LCtrl + Alt + L
Scroll subtitles up (older sentences)⌃ + ⌥ + PageUpCtrl + Alt + PageUp
Scroll subtitles down (newer sentences)⌃ + ⌥ + PageDownCtrl + Alt + PageDown

On macOS, ⌃ is Control and ⌥ is Option. A shortcut needs at least one of Ctrl, Alt or Cmd/Win. The Windows version is not released yet.

The menu bar icon (macOS) or system tray icon (Windows, when released) lets you start or stop translating, show or hide and lock the subtitles, and open the main window. Closing the window only hides the app in the tray; to quit completely, choose Quit.

Shortcuts settings showing the five default shortcuts on macOS
Settings › Shortcuts: press Change, then the new key combination.
Models and your computer

Two model packs, and the app recommends the one that fits

The recognition and translation models are downloaded once and then run entirely on your computer. The app checks your RAM, free disk space and graphics card to recommend a pack.

DetailStandard packLite pack
Downloadabout 2.5 GBabout 1.3 GB
RecognitionWhisper large-v3-turboWhisper small
TranslationHy-MT2-1.8B (Q8_0)Hy-MT2-1.8B (Q4_K_M)
Engine RAM (M4 Pro)about 2.9 GiBabout 1.8–1.9 GiB
The app recommends it forA Mac with 16 GB or more; Windows with 16 GB or more and a discrete card with 6 GB of VRAMComputers with 8 GB up to under 16 GB, or Windows without a capable discrete card

The Lite pack recognizes speech less clearly than Standard in Vietnamese, Japanese, Korean and Chinese. If you listen to those languages a lot, choose the Standard pack.

macOS
macOS 14.2 or later, Apple Silicon (M1 or newer)
No version for Intel Macs
RAM
At least 8 GB, 16 GB recommended
Disk
At least 1 GB free on top of the size of the model being downloaded
Windows
Windows 10/11 64-bit, CPU with AVX2: not released yet
For the Standard pack, a discrete Vulkan-capable graphics card with 6 GB or more of VRAM is recommended
Performance

Real measurements, with the conditions

We publish only what we measured, on the machine we measured it on, and we say which machines we have not tested.

Measurement (Mac M4 Pro 24 GB, macOS 26, Metal GPU)Standard packLite pack
Median (p50) delay: from the speaker finishing a sentence to the complete translation appearing0.76–1.03 s0.61–0.84 s
p90 delay0.94–1.34 s0.73–1.14 s
First translated word appears (p50)0.63–0.69 s0.52–0.57 s
RAM of the two engine processes2.9 GiB1.8–1.9 GiB

Long sessions

One continuous session of 5 h 23 min (a lecture video in English, on an ad-hoc-signed release build): 6,019 segments, no errors, median delay 0.48 s. In a 2-hour stability test, no process was restarted.

Translation and recognition quality

In an internal test on 320 sentences in five directions involving Vietnamese, Hy-MT2-1.8B scored 0.837 on COMET: clearly above MADLAD-3B (0.779) and NLLB-600M (0.736), and on par with HY-MT1.5-1.8B (0.833). We only compared open-source models with each other. On clean read speech, Vietnamese recognition in the Standard pack has a word error rate of 8.7 %. That is read speech, not real meeting conversation.

Translation and recognition quality by direction

COMET scores were measured through the app's own translation path (llama-server with the same prompt), 100 sentences per direction (Vietnamese → Chinese, Japanese, Korean: 40 sentences). COMET is a relative score from 0 to 1, higher is better; it is not a percentage accuracy.

Direction (text)Standard packLite pack
English → Vietnamese0.8420.841
中文 → Vietnamese0.8290.831
日本語 → Vietnamese0.8300.815
한국어 → Vietnamese0.8340.822
Vietnamese → English0.8210.822
Vietnamese → 中文0.8360.821
Vietnamese → 日本語0.8470.845
Vietnamese → 한국어0.8510.842
Speech recognition (error rate, lower is better)Standard packLite pack
English (word errors)5.4 %6.6 %
Vietnamese (word errors)8.7 %22.5 %
中文 (character errors)5.6 %9.6 %
日本語 (character errors)4.5 %13.1 %
한국어 (character errors)4.1 %8.2 %

How to read this: the translation test set leans toward everyday spoken language, so it is only a rough proxy for meeting speech; recognition was measured on clean read speech (about 15 minutes per language), not real meeting conversations, and narrowband Bluetooth headsets raise the error rate. All eight directions above involve Vietnamese; the other 12 directions among English, 中文, 日本語 and 한국어 run but have no quality score yet. The Lite pack is noticeably less accurate at recognizing Vietnamese, Japanese, Korean and Chinese.

What we have not measured, and do not promise

All the figures above were measured on a single Mac M4 Pro, with a 300 ms pause to end a sentence (the app's current default is 50 ms; the 5-hour session used the default). We have not measured: a base Mac M1, 8 GB computers, discrete Windows graphics cards, battery and power use, accuracy on real meeting conversation or over Bluetooth headsets, and translation quality for directions that do not involve Vietnamese (they work, but we have no quality score to publish). A preliminary test on one Windows laptop with integrated graphics showed that the Standard pack did not meet our latency target.

Technology

Built on open-source components, running on your computer

AI Translator is a commercial product; the components below are third-party open-source projects, listed in full on the About screen of the app.

Whisper + whisper.cpp

OpenAI's multilingual speech recognition (MIT license), run through whisper.cpp.

Hy-MT2-1.8B + llama.cpp

Tencent's translation model (Apache 2.0 license) in GGUF format, run through llama.cpp.

Silero VAD

Detects speech so sentences can be cut at the right place. MIT license, runs in pure Rust inside the app.

Tauri 2, Rust, React

A Rust core with a React interface. The two engines run in separate processes, so a GPU driver failure cannot take the app down.

Metal and Vulkan

GPU acceleration: Metal on macOS, Vulkan on Windows. It falls back to the CPU if the GPU fails.

SQLCipher

History, if you turn it on, is encrypted with SQLCipher; the random key is kept in your operating system's key store.

By plan

Which feature is in which plan

FeatureFree (trial)MonthlyYearly
Live translated subtitles, 5 languagesYesYesYes
Customizable subtitle bar, shortcuts, trayYesYesYes
View, search and copy the transcriptYesYesYes
Glossary (up to 500 terms)NoYesYes
Save session historyNoYesYes
Export to TXT, SRT, MarkdownNoYesYes
Translation time30 min per day, for 10 days50 hours per 30 daysUnlimited, for 365 days
Price (VND)0 ₫50,000 ₫500,000 ₫

See the full pricing details

Try it on your real meetings

A free 10-day trial with 30 minutes a day. No card, no account.