You have read the job description three times. You have a list of STAR-method stories. You know the company’s tech stack. But when the interviewer asks the first question, your mind goes blank and the words come out wrong.
Interview preparation is not a reading problem — it is a speaking problem. And until now, most AI interview tools asked you to type your answers, which is nothing like the real thing.
GUÍA, Xeito’s AI interview coach, now supports voice input and spoken questions. You hear the question read aloud, speak your answer out loud, and get scored feedback — the same way a real interview works, minus the stakes.
How GUÍA works: three phases, one session
GUÍA is not a chatbot that invents generic questions. It runs a structured three-phase flow designed to simulate what actually happens when you interview for a specific role at a specific company.
Phase 1 — Prep (10 seconds of research)
Tell GUÍA the role you are targeting, the company name, and optionally paste the job posting URL. GUÍA researches the company — pulling real information from the posting and the company’s careers page — and generates a prep report:
- Predicted questions tailored to the role and company, categorised by type (behavioural, technical, situational)
- STAR-method examples showing how to structure your answers
- Technical topics you should review before the interview
- Salary negotiation data and red flags to watch for
This is not a template. The questions are grounded in what GUÍA actually found about the company. If the job posting mentions “event-driven architecture,” expect a question about event-driven architecture.
Phase 2 — Interview (10 minutes of practice)
Press Start Interview and the predicted questions from your prep become the actual interview. Each question appears with its type badge (behavioural, technical, situational), a difficulty indicator, and coaching tips.
With voice enabled, you hear the question read aloud by GUÍA’s voice — choose between two voice options in your preferences — and speak your answer into the microphone. Your words are transcribed in real time and appear in the answer field, where you can review or edit before submitting. Prefer typing? The text input is always there.
Every answer gets immediate AI feedback: a score from 0 to 100, what you did well, what to improve, and a suggestion for a stronger response. After a few seconds, the next question appears automatically.
This is where the voice input matters. Typing “I led a cross-functional team of eight engineers to deliver the migration on time” is easy. Saying it clearly while maintaining composure is the actual skill you need in the interview room. GUÍA lets you practise that skill.
Phase 3 — Debrief (30 seconds of coaching)
After the final question, GUÍA synthesises everything into a coaching debrief:
- Overall score with a readiness label (e.g. “Strong candidate — ready for the real thing”)
- Strongest and weakest answers with explanations of why they scored high or low
- Prioritised improvements — the two or three things that would most improve your performance
- Summary narrative — a paragraph-length coaching note you can re-read before the real interview
The debrief is where the value compounds. Run three sessions for the same role and you will see your weakest areas shrink as you internalise the feedback.
Why voice changes everything
Typing interview answers trains you to write well under pressure. That is a useful skill — for take-home assignments. It is not the skill you need when a human is looking at you over a video call and waiting for you to speak.
Voice input closes that gap:
- Pacing: you learn how long your answers actually take when spoken, not when typed
- Filler words: “um,” “like,” and trailing silence are invisible in text but obvious in speech — practising aloud makes you aware of them
- Confidence: hearing yourself answer a tough behavioural question builds the muscle memory that keeps you calm in the real interview
- Realism: the combination of hearing the question and speaking the answer is closer to a video interview than any text-based tool
Voice output adds the other half. Hearing the question read aloud — instead of reading it on screen — forces you to listen and process before you respond, exactly as you would with a real interviewer.
Privacy: how voice data is handled
Voice recognition runs through your browser’s built-in Web Speech API. The audio stream is processed locally or by your browser vendor’s speech service — Xeito’s servers never receive the audio. Only the final transcribed text is sent to Xeito’s API for scoring, and that API runs on EU-hosted infrastructure (Neon Postgres in Frankfurt, Cloudflare Workers in EU data centres).
Voice output (hearing questions read aloud) is synthesised server-side. Xeito’s API sends only the question text to an external speech-synthesis provider (based outside the EU) — never your answers, never your personal data. The resulting audio is streamed back through Xeito’s EU-hosted API.
Getting started
- Open your Xeito dashboard and go to Mock Interview
- Enter the role, company, and optionally the job posting URL
- Choose the interview type (general, technical, behavioural, or system design) and length
- Click Get Interview Prep — GUÍA researches the company and generates your prep report
- Press Start Interview — allow microphone access when prompted
- Speak your answers, review the real-time transcription, and submit
- After the final question, read your Debrief and plan your next session
GUÍA is part of Xeito Premium. Every free account includes credits to try it — no payment details required.
The interview is a skill, not a knowledge test
You already know your experience. You already know the technologies. What you need is practice delivering that knowledge under pressure, out loud, to another entity that is evaluating you in real time.
That is what GUÍA with voice gives you. Research in 10 seconds, practice in 10 minutes, coaching in 30 seconds. Repeat until the real interview feels routine.
Sources
- Web Speech API — MDN Web Docs — browser speech recognition used for voice input
- SpeechRecognition browser compatibility — MDN — supported browsers for voice input
- STAR Interview Method — The Muse — the STAR framework referenced in GUÍA’s prep reports