Audio Survey: How to Collect Spoken Answers
Search for an audio survey and you mostly find research papers about sound. If what you mean is a questionnaire where people answer out loud, there are four ways to run one. This is how they differ, and how to decide.
What people mean by “audio survey”
Type the phrase into Google and the first page is mostly academic surveys of audio — enhancement systems, listening habits, acoustic models. Nobody looking for those is trying to run a questionnaire. The meaning this page is about is narrower: a survey in which at least some answers are spoken instead of typed. You ask a question in text, the respondent taps a button and talks, and you get the answer back as audio, as a transcript, or both.
The people asking that question are usually in one of two situations. They have an open-ended question that gets one-line typed answers and want richer ones. Or their respondents type badly — on a phone, in a second language, with limited literacy — and speaking is simply the easier channel. Either way the practical question is the same: how do I get spoken answers into the survey I already have?
Four ways to collect audio responses in a survey
There are only four architectures. They differ in where the recording happens and what the respondent has to leave behind to do it.
| Approach | What the respondent does | What it costs you |
|---|---|---|
| 1. A native audio question type | Taps record inside the survey | Whatever limits and plan tier that platform attaches to the feature |
| 2. A recording widget inside your existing survey | Taps record inside the survey | Pasting a script into a platform that lets you run your own JavaScript |
| 3. Audio file upload | Records on their phone, then uploads the file | Friction for the respondent, and transcribing and matching files yourself |
| 4. An AI-moderated interview or phone line | Leaves your questionnaire and talks to a separate system | A different instrument, hosted elsewhere, with its own export |
1. A native audio question
If your platform has one, use it. Typeform, for example, ships a Video and Audio question type with automatic transcripts, but each answer is capped at two minutes and the feature sits behind its Growth or Talent plans. Whether your platform has anything comparable is a documentation question, and we keep a checked-against-the-docs table in Which survey platforms support voice responses. Read it before you build anything; for some platforms the answer ends there.
2. A recording widget inside the survey you already have
Most professional survey platforms do not record voice, but several let you add your own HTML and JavaScript to the page the respondent sees. That is enough for a recorder to live inside the question: the respondent never leaves, the questionnaire logic and quotas stay untouched, and the result is written back into the same field, so it comes out in your normal export. Voice Capture is this kind of tool — a snippet you paste into Alchemer, QuestionPro, FormAssembly or Sawtooth Lighthouse Studio, or into any web form that accepts custom HTML (Qualtrics support is coming soon). Setup guides are at /integrations.
The widget route has one boundary worth stating early: platforms that strip custom JavaScript — SurveyMonkey is the common one — cannot host it at all.
3. Audio file upload
Some forms let respondents attach a file. For audio this means recording with a separate app, finding the file, and uploading it: three steps where the widget has one. Completion rates drop at each step, and you inherit a folder of files with no transcript, which you then have to transcribe and join back to the right respondent by ID. It works for a small, motivated panel. It does not scale to a general-population sample.
4. An AI-moderated interview or a phone line
A different product category entirely: the respondent leaves your questionnaire for a conversation run by a separate system. That is the right choice when you want probing and follow-up questions, and the wrong one when you want a spoken open end inside a structured survey with quotas and skip logic. The trade-off is laid out in AI-moderated interviews vs voice in your survey.
Do you need the recording, or only what was said?
This is the decision most guides skip, and it changes which approach is open to you.
Most market research open ends are coded, quoted and counted as text. For that, the transcript is the deliverable and the audio is a by-product. If that is your case, keeping audio files is a liability as much as an asset: recordings of a person’s voice are personal data, they need a retention policy, and they have to be deleted on request.
The audio itself matters when the sound is the data — tone and emotion analysis, a highlight reel of real customer voices for a stakeholder presentation, or a regulatory requirement to retain the original. Then you need a tool that stores recordings.
Voice Capture is built for the first case and is honest about the second: audio is transcribed and discarded, and only the text is stored, in EU-based infrastructure. If you need to keep the recordings, it is not the right tool, and one of the file-based approaches above is.
Designing an audio survey question that works
A spoken question fails differently from a typed one, and most of the failures are in the prompt, not the technology.
- Ask for one thing. “Tell us what you liked, what you didn’t, and what you would change” produces a rambling answer that is hard to code. One question, one topic.
- Warn about the microphone prompt. The browser asks for permission once, and that prompt is where most respondents drop out. A line of survey text before the question — you can answer this one by speaking; your browser will ask for microphone access — removes most of the friction.
- Keep typing available. Not everyone can or wants to speak, and a public place is a bad place to be forced to. Offer voice as an option and let a text answer and a voice answer to the same question land in the same column.
- Set the length expectation. Voice Capture’s maximum answer length depends on the plan — 90 seconds on the free trial, 180 on Essential, 300 on Professional — so phrase the prompt to fit. A question that needs a five-minute monologue is an interview, not a survey item.
- Soft-launch. Send a tenth of the sample, read the transcripts yourself, and check the prompt produces the kind of answer you expected before the rest goes out.
What the data looks like afterwards
Spoken answers are longer than typed ones — on our own production data about twice as long, 23.1 words on average against 10.1 typed across 693 voice and 6,476 typed responses, not the order-of-magnitude gap some vendors advertise. The sample and the caveats are on /methodology. They are also transcribed, not typed, so check a sample for systematically mis-heard brand names and category jargon; a brand that comes back wrong the same way every time is a find-and-replace, not a problem. The longer write-up of the whole workflow, from fieldwork to coding, is in the complete guide to voice surveys.
Frequently asked questions
Can I add audio responses to a survey I already built?
Usually yes, if your platform lets you run custom JavaScript on the respondent page. You add a script once and point it at the question that should accept a spoken answer. If your platform strips custom JavaScript, you need either its native audio question (if it has one) or a different platform.
Do respondents need to install anything?
No. Voice Capture uses the browser’s standard microphone permission — no app, no plugin — and works on iOS Safari and Android Chrome. If a browser cannot record, the widget hides itself and the text box keeps working, so nobody hits a dead end.
What languages can an audio survey handle?
Voice Capture transcribes 36 languages. On the free trial the language locks to one after your first project; paid packs include all of them.
Are the recordings stored?
Not by Voice Capture: audio is transcribed and discarded, and only the text is kept. If you need the original recordings, use a tool that stores files and plan the retention and deletion policy that comes with them.
How much does it cost?
One credit is one transcribed response, and credits never expire. The free trial includes 250 credits; the Essential pack is $99 for 1,000 credits. Current packs are on the pricing page.
Where to start
Decide first whether you need the recording or only the words, then check whether your platform already has a native audio question. If it does not but it accepts custom JavaScript, the setup guides take a survey you already have and add a spoken option to one question. The free trial is enough to soft-launch one.
Ready to add voice to your surveys?
Start free — no credit card required. Setup takes 2 minutes.
Try Voice Capture Free