How to Extract Subtitles & Text from a TikTok Video (4 Ways)
Text in a TikTok (or Douyin) video comes from two places: the captions burned into the frame, and the words the creator actually speaks — and the two don’t always match. Extraction methods split along the same line: read text from the frames (OCR), or transcribe text from the audio (speech recognition). This guide ranks the four common methods by time cost and accuracy, so you can choose by how often you actually do this.
Method 1: Type it yourself (free, slowest)
The primitive method: slow the video to 0.5x and type along. It’s completely free and forces a close listen; it’s also brutally slow — one minute of voiceover takes 5–8 minutes to type, and long videos are simply impractical. Reserve it for the occasional video whose content deserves sentence-by-sentence chewing.
Two tricks help: play on your phone and type on your computer instead of juggling one screen, and type straight through without punctuation first, then go back to add breaks — noticeably faster overall.
Method 2: CapCut caption recognition, then export
Save the video locally first (possible when the creator allows downloads), import it into CapCut, run Text → Auto captions to generate a caption track, then copy the lines out or export a transcript. Chinese recognition accuracy is decent — this is today’s most common free route.
The weakness is the length of the chain: download, import, recognize, export — several app switches, and videos with downloads disabled need screen recording first. Numbers, English words, and proper nouns misrecognize often enough that anything bound for real notes needs a human pass.
Method 3: Screenshots + OCR for on-frame captions
If all you need is the captions burned into the frame (quote-card videos, for instance), screenshot frame by frame and run OCR — WeChat’s Extract Text, your phone gallery’s built-in recognition, or any OCR tool. Single frames process fast, and printed-style captions read accurately.
But OCR only captures what’s on the frame — spoken lines without captions are lost; long videos mean dozens of screenshots processed one by one, with lines easily dropped during stitching. Good for pulling a few quotable lines, wrong for full transcripts.
Method 4: An auto-transcription tool (paste a link, get text)
The fourth category: give the tool a video link and the speech comes back as text. Attar, for example, is a Chrome extension — copy the TikTok or Douyin share link and paste it in, or click “Import current video (free)” on the video page, and you get a free preview of the recognized voiceover. English videos are supported, Chinese recognition is a particular strength, and no download is needed.
Attar’s point isn’t just transcription: once the preview checks out, click “Generate script (uses 1 credit)” for actionable takeaways, a viral breakdown, and script-modeling references, saved as a local Markdown file. You get 3 free generations per month, with importing and previewing never counted; the annual membership (about $13.99/yr) includes 360 credits, and a $1.49 top-up pack adds 30.
Choosing: one comparison list
Choose by frequency. One or two videos a month: CapCut is plenty. Only need the on-frame quotes: screenshots plus OCR. Breaking down videos every week and turning them into notes as you go: an auto-transcription tool saves the most time.
- Typing yourself: free; about 5–8 minutes per minute of content; accuracy is on you
- CapCut captions: free; long chain; decent Chinese accuracy; proofread numbers and names
- Screenshot OCR: fast per frame; captions only; heavy lifting on long videos
- Auto transcription (e.g. Attar): paste a link; free preview; a credit only when generating the structured note
FAQ
The video blocks downloads — can I still get the text?
Yes. The audio-transcription route doesn’t depend on downloading: screen-record and run CapCut recognition, or more directly with Attar — paste the share link into the extension and the voiceover is recognized without saving any video file locally.
On-frame captions and the voiceover disagree — which wins?
Depends on purpose. For studying the writing, trust the voiceover — captions are often compressed and rewritten. For quote cards, use the on-frame text. The rigorous approach extracts both and compares; the differences themselves reveal what the creator chose to emphasize.
Does Attar charge for text extraction?
Importing a video and previewing the recognition are free, and you get 3 full free generations each month. Only clicking “Generate script” to produce the structured note uses 1 credit. Membership runs about $13.99 a year for 360 credits; a $1.49 top-up pack adds 30.
Can I batch-extract multiple videos at once?
Attar doesn’t batch-generate: paste several links at once and only the first is processed; live-stream and collection links aren’t supported either. For multiple videos, paste and process them one at a time.