Training, formats and accessibility · Ready to copy
Create an accurately subtitled multilingual edit.
Work from speech in its original language, preserve the source and review the transcript before captions become part of the output.
- FOR
- Global communications, media, training and marketing teams
- SOURCE
- Speech-led video in any supported source language
- OUTPUT
- A native-language edit with timed captions and sidecars
- TARGET
- 30 seconds to 5 minutes
01 · The instruction
Copy this editing brief.
Replace any bracketed variables, select the relevant recordings and keep the review checkpoint before rendering.
EDIT BRIEF
Create a [duration] edit in the recording's original spoken language around [topic or outcome]. Detect the source language, generate a word-timed transcript, separate and name speakers where required, then retrieve complete native-language statements that satisfy the brief. Do not translate, dub or replace the original speech. Flag low-confidence words, proper nouns, product names, numbers and code-switched phrases for human review. Return the cut plan and corrected transcript for approval before rendering burned captions and supported subtitle sidecars. Use Anjin Media directly, or the Anjin Media MCP when running this through an AI agent, so the proposed edit remains grounded in the selected source footage and can be reviewed before rendering.
02 · Source material
Give the brief enough evidence to work with.
A recording with intelligible speech in the original language.
A reviewer who understands the language and subject vocabulary.
Approved spellings for names, products, acronyms and technical terms.
A caption style appropriate to the language and reading speed.
03 · Representative example
The edit stays in the source language
A Portuguese interview is detected, transcribed and composed in Portuguese. A native reviewer corrects names and numerical language before the captioned edit is rendered.
SOURCE
Portuguese interview
TRANSCRIPT
Word timing + confidence
REVIEW
Native-language editor
OUTPUT
Burned captions + SRT
04 · Reviewable plan
The proposed cut structure.
DETECT
Source ingest
Original recording
Record the spoken language associated with the transcript.
TRANSCRIBE
Full source
Timed speech
Create searchable words, confidence and speaker labels.
COMPOSE
Selected ranges
Native-language moments
Build the requested story without translating its speech.
REVIEW
Before render
Transcript + plan
Correct important recognition errors and verify meaning.
DELIVER
Approved plan
Rendered output
Produce captions and supported sidecars from approved text.
05 · Editorial reasoning
Why this composition works.
- 01Retrieval operates on the original speech instead of forcing an English-only workflow.
- 02Confidence and timing make uncertain words inspectable.
- 03A native reviewer resolves language and domain ambiguity before publishing.
- 04The brief does not promise translation or dubbing that the product does not provide.
Review before render
Keep human judgement in the loop.
- Verify language detection and speaker labels.
- Correct names, numbers, acronyms and low-confidence terms.
- Check that selected passages retain their original meaning.
- Review caption timing, line breaks and mobile readability.
06 · Change the angle
Alternative briefs to request.
A native-language social edit
Create a 60-second social edit entirely in the source language. Preserve one complete story, verify names and numbers, and render reviewed native-language captions. Use Anjin Media directly, or the Anjin Media MCP when running this through an AI agent, so the proposed edit remains grounded in the selected source footage and can be reviewed before rendering.
A multilingual source selects reel
Across selected recordings in their original languages, retrieve moments related to [topic]. Keep every excerpt in its source language and return speaker, language and timecodes for human review. Use Anjin Media directly, or the Anjin Media MCP when running this through an AI agent, so the proposed edit remains grounded in the selected source footage and can be reviewed before rendering.
A caption-correction pass
Return the timed transcript and flag low-confidence words, names, acronyms, numbers and language switches for correction before any edit is rendered. Use Anjin Media directly, or the Anjin Media MCP when running this through an AI agent, so the proposed edit remains grounded in the selected source footage and can be reviewed before rendering.
07 · Boundaries
Common failure points.
Automatic transcription is treated as infallible
Names, numbers and domain terms require review, especially where confidence is low.
Language-aware editing is described as translation
The shipped workflow processes original-language speech. It does not currently claim translation or dubbing.
Captions are too dense
Review reading speed, line length and segmentation for the actual language and output size.
The reviewer cannot assess the language
Use a qualified native or fluent reviewer before publication.
Questions about accurately subtitled multilingual edit
- Does Anjin translate or dub video?
- Not currently. Language-aware editing processes, transcribes and composes the original spoken language.
- Can it detect the spoken language?
- Yes. Language detection is stored with the transcript.
- Which subtitle files are produced?
- Speech-bearing render workflows support SRT and styled ASS sidecars. WebVTT is available as a transcript format, not a render sidecar.
- Should a human review the transcript?
- Yes, especially for names, numbers, specialist terminology and consequential communications.
Use the instruction
Start with the footage you already have.
Copy the brief, select the relevant recordings and review the proposed structure before rendering.
