About this tool
Convert spoken language into written text.
Voice to Text turns speech into written text using the browser's built-in Web Speech Recognition API, showing the words live as you talk and then translating the finished transcript into one of nine target languages — Hindi, English, Spanish, French, Bengali, Arabic, Russian, Japanese or German. It is for anyone who would rather dictate than type: students capturing a thought, support staff drafting a reply, or someone speaking one language and needing the message in another. Recognition starts and stops on a single mic button, and the translated output has a one-click copy.
Open Voice to Text on AltFTool — it loads instantly in your browser.
Pick a target from the Translate to dropdown — Hindi, English, Spanish, French, Bengali, Arabic, Russian, Japanese or German.
Tap the mic and start speaking; the mic turns into a square you click to stop.
Watch the words build up under Original Text, then read Translated Output and press the copy button on that panel.
Translation fires automatically when you stop speaking, so you never have to copy the transcript into a second tool.
Interim results are enabled, so the text updates as you speak instead of appearing only after you finish the sentence.
Change the target from the dropdown and the existing transcript is re-translated, so you can produce the same message in a second language immediately.
Chrome and other Chromium-based browsers such as Edge, plus Safari, which expose the Web Speech Recognition API (as webkitSpeechRecognition). Firefox does not enable it by default, so the microphone button will not produce a transcript there.
Nine: Hindi, English, Spanish, French, Bengali, Arabic, Russian, Japanese and German. Hindi is the default target, and the source language is auto-detected rather than chosen from a list.
Not entirely. Speech recognition is handled by your browser, and most desktop browsers do that by sending audio to the vendor's cloud service; the transcript is then sent to the MyMemory translation API to be translated. Do not dictate passwords, medical details or other confidential information into it.
Recognition is set to non-continuous mode, so it ends after a natural pause in speech rather than recording indefinitely. Tap the mic again for the next passage, and note that starting a new session clears the previous transcript and translation.