Why AI Translator exists
The first idea was to take an existing offline translation app for Android (an open-source one) and use it for calls and meetings. We hit a hard limit of the operating system almost at once. Android lets third-party apps capture the audio of media and games, but not the audio of phone calls or VoIP calls, and during a call an app's microphone usually hears only silence.
A computer is different. Windows offers WASAPI loopback and macOS (from version 14.2) offers Core Audio process taps; both let an app capture the sound the computer itself is playing. So the project moved to the desktop: an app that listens to system audio, recognizes speech, translates it and shows subtitles, and works with any meeting app without a bot or a plugin.
While researching (September 2026), we found that translated subtitles in the large meeting apps tend to sit in higher paid tiers and run in the cloud, and most third-party tools do the same. We wanted a different option: everything processed on your machine, any meeting app, a focus on Vietnamese, and, because no server does the translating, no extra infrastructure cost for each minute you use. Here is a comparison of offline and cloud translation.
We chose the translation model by measurement. In an internal test on 29 September 2026 (320 sentences of text from the WMT24++ set, five translation directions, run on a Mac M4 Pro), Hy-MT2-1.8B scored 0.837 COMET, while the three other translation models tested under the same conditions scored between 0.736 and 0.833. That test used text, not real speech, and we did not compare against any cloud service.