A transcript is not a subtitle
Most "audio to subtitle" tools give you a transcript and call it a caption file. But a transcript is only the words.
A good subtitle also needs timing, readable line breaks, and punctuation that matches the speaker's rhythm.
Without that, viewers get a wall of text or captions that appear at the wrong time. Mitsuko is built to turn audio into timed subtitle text, not just a raw transcript.
What the gap looks like: two real examples
Example 1 - casual creator speech
| Type | Line |
|---|---|
| Spoken creator sample | Okay, so we're gonna export this and fix the timing later. |
| Raw transcript style | Okay so we are going to export this and fix the timing later |
| Subtitle-ready text (Mitsuko) | Okay, we'll export this and fix the timing later. |
The raw transcript drops the comma, keeps the filler "so," expands "we're gonna" to "we are going to," and has no closing punctuation.
That is how speech often comes out of a mouth, but it is not suitable for a subtitle.
The subtitle-ready version keeps the casual voice, tightens the phrasing, restores punctuation, and creates a line a viewer can read while it is on screen.
Example 2 - a spoken lesson
| Type | Line |
|---|---|
| Spoken lesson sample | This formula matters because it tells us when the reaction slows down. |
| Unpolished transcript | This formula matters because it tells us when reaction slows down |
| Timed subtitle draft (Mitsuko) | This formula matters because it tells us when the reaction slows down. |
Even on a clean line, the transcript loses the article "the" and the period. Small things matter when text has to be read quickly under a video.
The subtitle draft restores readable punctuation and phrasing, so the line is ready for timed captions and translation without a manual cleanup pass.
The audio-first subtitle flow
- Upload single or multiple audio. One recording or a batch - interviews, lessons, episodes, voiceovers.
- Review timed subtitle text result. Word-level timestamps with natural pacing produce cues that are in sync and readable, not a flat transcript dumped into a file.
- Translate or export the subtitle file. Review the captions, then either export the subtitle file as-is or translate it into another language inside Mitsuko.
Why timestamps matter more than word count
The tempting metric for transcription is "how many words did it get right."
For subtitles, the better metric is timing quality. Are the cues aligned to speech? Are the lines short enough to read while they are shown? Does the text appear when the speaker says it?
Mitsuko produces word-level timestamps and uses them to build cues with natural pacing.
That means the subtitle is readable and in sync, not just verbatim. A transcript with perfect words and broken timing is still a bad subtitle.
Who this is for
- Creators starting from raw video who have no subtitle file to begin with.
- Course teams that need captions first before they can even think about translation.
- Agencies receiving media without subtitle files - you transcribe once, then translate.
From audio to translated subtitles in one place
The point of doing transcription and translation in the same tool is that the subtitle file you review is the same file you translate.
You do not need to export a transcript, clean it up in a second tool, re-import it, and then translate. You review the timed captions, then translate them directly into 100+ languages.
Related workflows
- YouTube subtitle translator - localize the captions you just generated.
- Batch subtitle translation - transcribe and translate multiple recordings at once.
- Subtitle localization for agencies - for teams receiving media without subtitle files.
- Pricing - how transcription and translation are billed.
Upload your audio, review the timed subtitle text, then translate or export - all without leaving the workflow.