Real-time voice AI
Real-time voice practice through my own WebSocket relay
Speaking practice that feels like a real conversation: the student talks, the AI listens and answers out loud, inside the scenario.
Animated recreation
The problem
Google retired two audio models and SpeakUp broke on every platform. The model that stays in character only works with an API key, and an API key can never go to the browser.
What I built
- A WebSocket relay in Node.js, running in Docker on my VPS, that holds the key and bridges each student to Gemini Live.
- One live session per student, with the scenario, role and level set at the start of the call.
- 24 scenarios in 4 categories, from ordering at a café to a job interview.
- Each scenario character speaks with its own voice, set per character in the relay. The final feedback can be spoken in any cloned voice (ElevenLabs); today it uses the teacher's.
- Works on desktop, Android and iPhone, including the installed app.
Architecture
- 01
Browser
Microphone in, audio out
- 02
WebSocket relay
Node.js on my VPS, holds the key
- 03
Gemini Live
Native audio model, stays in role
- 04
Voice out
Each character speaks with its own voice
Results
- conversation scenarios
- 24conversation scenarios
- scenario categories
- 4scenario categories
- live session per student, isolated
- 1live session per student, isolated
- API keys exposed to the browser
- 0API keys exposed to the browser
Stack
- Node.js
- ws (WebSocket)
- Google GenAI (Gemini Live)
- Docker
- EasyPanel
My role
I diagnosed the break, designed the relay and shipped it to production. Frontend, relay and deployment are mine.
Why there is no public access
SpeakUp runs inside the paid platform and handles students' voices, which are personal data under GDPR and LGPD. The animation above is a recreation with a fictional exchange.
I can show it running live on a call.
I open the real system, with personal data protected, and answer anything you want to know about how it is built.
mirandamauro93@gmail.com