My server pays for a voice AI call it never hears
Falar is a speaking tutor for European Portuguese. You talk to Ana, a teacher from Lisbon, and she answers out loud, corrects you, and remembers which mistakes you keep making. It has been in internal testing on Google Play since 1 October.
The architecture decision that shapes everything else: the phone talks to OpenAI's Realtime API directly, over WebRTC. My backend creates the call and holds the API key, but the audio never passes through it. That keeps the round trip near one second, and it is the only way I found that lets the model actually hear how a learner pronounces a word rather than read a transcript of it.
It also means the most expensive thing in the system runs on a device I do not control, on a connection I cannot see, billed to my account.
