AI video composition
Composition,
not clipping.
A clip is detected. A composition is made. Describe the edit you need - subject, audience, length, format - and Anjin Media retrieves the right moments from across your footage, cuts the repetition and returns one finished, source-grounded edit.
This is prompt-driven video editing taken to its conclusion: the brief is the interface, and the deliverable is an edit - not a shortlist of clips.
One video or an entire archive.
AI video composition product demonstration
- 1
- An edit that executes a written brief: moments retrieved from across the source footage - rarely adjacent, never repeated - ordered to carry one argument, rendered at a set duration and aspect ratio.
- cf.
- clip, n. - a single moment, detected and excised. Context not included.
Five layers apart.
A clipper and a composer look identical for the first minute: footage in, video out. The difference is everything that happens in between - and it runs five layers deep.
| Layer | Conventional AI tools | AI video composition |
|---|---|---|
| Clipping | Detection. The tool scores the footage for spikes and hands back whichever moment lit up. Whether it says anything is your problem. | Execution. A brief goes in; several moments come back - pulled from wherever they sit in the footage, ordered to make one point. |
| Search | Filenames, tags, and whatever metadata somebody remembered to type. | The footage itself. Every word is transcribed with word-level timecodes, so a brief can reach any sentence said on camera. |
| Editing | A timeline, a scrub bar, and your afternoon. | A cut plan you read and revise in plain language - then a rendered edit at exactly the duration and aspect you set. |
| Automation | Batch presets inside one editor, still driven by hand. | The whole workflow over REST and MCP. Code and agents create edits the way a person does: brief in, edit out. |
| Enterprise | The archive is a storage bill with a search box on top. | Footage is indexed once and stays addressable. Every brief after that starts from what you have already loaded. |
The fifth layer is a page of its own - what an archive plan includes →
One brief. Three voices. Four years.
A communications team needs sixty seconds for LinkedIn on how the company's diversity position has evolved. Nobody scrubs footage; somebody writes a paragraph.
“A 60-second film for LinkedIn on how our position on diversity has evolved: open on the commitment we made in 2022, let leadership’s own words carry the change, and end on where we stand now. Senior voices only. No repetition.”
None of those four moments sat next to each other. Two are the same voice years apart; all four were found by what was said, not by filename. Chronology is set to strict, which holds each session’s moments in the order they were recorded; the sequence across the four sessions is the plan’s own argument about the change the brief describes - and since the plan is the checkpoint, striking a segment or asking for a different opening costs nothing before the render.
What composition holds itself to.
Say the outcome.
The brief states what the edit must do - subject, audience, length, format. Retrieval, selection and assembly are the system's work, and a brief can run to 2,000 characters: enough to direct, not just to request.
Only your footage.
Nothing is generated: no synthetic voice, no stock, no invented sentences. Every segment comes from your sources and keeps its source timecodes to prove it.
Review before spend.
A composition arrives first as a cut plan - readable, revisable in plain language, free. Rendering is the one moment that touches your allowance.
One capability, every surface.
The app, the REST API and the MCP server call the same engine - a person, a pipeline and an agent all get the identical edit.
Spoken in the request.
These are not settings buried in menus - they are fields on one request, named here exactly as the API spells them.
target_duration_sLength
How long the finished edit should run - anywhere from ten seconds to ten minutes. Every plan is assembled against it.
10–600 s · required
duration_tolerance_pctGive
How far from target the edit may land. Held tight, the length is near-exact; loosened, a strong moment gets to finish its sentence.
5–30% · default 15
paceRhythm
How briskly it cuts. Relaxed leaves air around what is said; tight closes the pauses down and keeps the edit moving.
relaxed · standard · tight
chronologyOrder
Flexible lets the argument decide the running order. Strict holds a session's moments in the order they were recorded; where a brief draws on several sessions, the sequence between them is still the plan's.
flexible · strict
aspectsFormat
One edit, up to three deliveries: widescreen at 1920×1080, square at 1080×1080, vertical at 1080×1920 - each rendered from the same plan.
16:9 · 1:1 · 9:16
crop_modeVertical framing
How widescreen footage becomes vertical. Bars keeps the whole frame, letterboxed; stack fills the frame with two crops - one per person - where the shot supports it.
bars · stack - 9:16 only
The request holds more - a minimum segment length, pause handling, a planner tier - every field documented with the API.
Every edit shows its working.
The cut plan is the paper trail. Before anything renders, every segment names its source video, its transcript text and its source timecodes - the edit is checkable line by line before it exists.
The render ships with its working attached: SRT and WebVTT captions, an EDL in JSON naming every segment’s source and timecodes, and a QC report on the render itself. For any second of the output, you can answer where it came from and what was said around it.
Captions
Subtitles of the finished edit, timed to it - SubRip and WebVTT.
Edit decision list
Every segment's source and timecodes, machine-readable - ready for a conform or an audit.
Quality report
The checks the render passed, alongside the file itself.
Questions
- What is prompt-driven video editing?
- Editing where the instruction is a written brief rather than timeline work. You describe what the edit must say, who it is for, how long it runs and what format it ships in; the system retrieves the right moments from your footage, composes them and renders the result. On Anjin Media the brief can run to 2,000 characters - enough to direct, not just to request.
- How is AI video composition different from AI clipping?
- A clipper detects promising moments and hands them back one at a time; making them mean something is still manual work. Composition executes the brief end to end: it selects several moments - rarely adjacent ones - removes repetition, orders them into a single argument and returns a finished edit with every segment's source attached.
- Can I control the length and format of the edit?
- Yes - length is a required part of every request: a target between 10 seconds and 10 minutes, with a tolerance you set between 5 and 30 percent. Formats are 16:9, 1:1 and 9:16, up to three per edit, and vertical output either keeps the full frame letterboxed or fills the frame with stacked crops.
- Does it generate footage or voiceover?
- No. A composition contains nothing that was not in your sources - no synthetic voice, no stock, no invented sentences. Every segment carries its source timecodes, so this is checkable rather than promised.
- Do I see the edit before it costs anything?
- Yes. Every composition arrives first as a cut plan: the segment list with sources, transcript text and source timecodes. Reading it, editing it and revising it in plain language are free - your allowance is only touched when you render.
