Anjin Media

AI video composition

Composition,
not clipping.

A clip is detected. A composition is made. Describe the edit you need - subject, audience, length, format - and Anjin Media retrieves the right moments from across your footage, cuts the repetition and returns one finished, source-grounded edit.

This is prompt-driven video editing taken to its conclusion: the brief is the interface, and the deliverable is an edit - not a shortlist of clips.

One video or an entire archive.

AI video composition product demonstration

com·po·si·tion/ˌkɒm.pəˈzɪʃ.ən/noun
1
An edit that executes a written brief: moments retrieved from across the source footage - rarely adjacent, never repeated - ordered to carry one argument, rendered at a set duration and aspect ratio.
cf.
clip, n. - a single moment, detected and excised. Context not included.
The difference

Five layers apart.

A clipper and a composer look identical for the first minute: footage in, video out. The difference is everything that happens in between - and it runs five layers deep.

Conventional AI video tools compared with AI video composition across five layers: clipping, search, editing, automation and enterprise.
LayerConventional AI toolsAI video composition
ClippingDetection. The tool scores the footage for spikes and hands back whichever moment lit up. Whether it says anything is your problem.Execution. A brief goes in; several moments come back - pulled from wherever they sit in the footage, ordered to make one point.
SearchFilenames, tags, and whatever metadata somebody remembered to type.The footage itself. Every word is transcribed with word-level timecodes, so a brief can reach any sentence said on camera.
EditingA timeline, a scrub bar, and your afternoon.A cut plan you read and revise in plain language - then a rendered edit at exactly the duration and aspect you set.
AutomationBatch presets inside one editor, still driven by hand.The whole workflow over REST and MCP. Code and agents create edits the way a person does: brief in, edit out.
EnterpriseThe archive is a storage bill with a search box on top.Footage is indexed once and stays addressable. Every brief after that starts from what you have already loaded.

The fifth layer is a page of its own - what an archive plan includes →

A worked example

One brief. Three voices. Four years.

A communications team needs sixty seconds for LinkedIn on how the company's diversity position has evolved. Nobody scrubs footage; somebody writes a paragraph.

THE BRIEFPOST /v1/cuts

“A 60-second film for LinkedIn on how our position on diversity has evolved: open on the commitment we made in 2022, let leadership’s own words carry the change, and end on where we stand now. Senior voices only. No repetition.”

target_duration_s 60aspects ["1:1"]chronology "strict"

None of those four moments sat next to each other. Two are the same voice years apart; all four were found by what was said, not by filename. Chronology is set to strict, which holds each session’s moments in the order they were recorded; the sequence across the four sessions is the plan’s own argument about the change the brief describes - and since the plan is the checkpoint, striking a segment or asking for a different opening costs nothing before the render.

Principles

What composition holds itself to.


Say the outcome.

The brief states what the edit must do - subject, audience, length, format. Retrieval, selection and assembly are the system's work, and a brief can run to 2,000 characters: enough to direct, not just to request.


Only your footage.

Nothing is generated: no synthetic voice, no stock, no invented sentences. Every segment comes from your sources and keeps its source timecodes to prove it.


Review before spend.

A composition arrives first as a cut plan - readable, revisable in plain language, free. Rendering is the one moment that touches your allowance.


One capability, every surface.

The app, the REST API and the MCP server call the same engine - a person, a pipeline and an agent all get the identical edit.

Controls

Spoken in the request.

These are not settings buried in menus - they are fields on one request, named here exactly as the API spells them.

target_duration_s

Length

How long the finished edit should run - anywhere from ten seconds to ten minutes. Every plan is assembled against it.

10–600 s · required

duration_tolerance_pct

Give

How far from target the edit may land. Held tight, the length is near-exact; loosened, a strong moment gets to finish its sentence.

5–30% · default 15

pace

Rhythm

How briskly it cuts. Relaxed leaves air around what is said; tight closes the pauses down and keeps the edit moving.

relaxed · standard · tight

chronology

Order

Flexible lets the argument decide the running order. Strict holds a session's moments in the order they were recorded; where a brief draws on several sessions, the sequence between them is still the plan's.

flexible · strict

aspects

Format

One edit, up to three deliveries: widescreen at 1920×1080, square at 1080×1080, vertical at 1080×1920 - each rendered from the same plan.

16:9 · 1:1 · 9:16

crop_mode

Vertical framing

How widescreen footage becomes vertical. Bars keeps the whole frame, letterboxed; stack fills the frame with two crops - one per person - where the shot supports it.

bars · stack - 9:16 only

The request holds more - a minimum segment length, pause handling, a planner tier - every field documented with the API.

Provenance

Every edit shows its working.

The cut plan is the paper trail. Before anything renders, every segment names its source video, its transcript text and its source timecodes - the edit is checkable line by line before it exists.

The render ships with its working attached: SRT and WebVTT captions, an EDL in JSON naming every segment’s source and timecodes, and a QC report on the render itself. For any second of the output, you can answer where it came from and what was said around it.

SRT · VTT

Captions

Subtitles of the finished edit, timed to it - SubRip and WebVTT.

EDL · JSON

Edit decision list

Every segment's source and timecodes, machine-readable - ready for a conform or an audit.

QC

Quality report

The checks the render passed, alongside the file itself.

Questions

What is prompt-driven video editing?
Editing where the instruction is a written brief rather than timeline work. You describe what the edit must say, who it is for, how long it runs and what format it ships in; the system retrieves the right moments from your footage, composes them and renders the result. On Anjin Media the brief can run to 2,000 characters - enough to direct, not just to request.
How is AI video composition different from AI clipping?
A clipper detects promising moments and hands them back one at a time; making them mean something is still manual work. Composition executes the brief end to end: it selects several moments - rarely adjacent ones - removes repetition, orders them into a single argument and returns a finished edit with every segment's source attached.
Can I control the length and format of the edit?
Yes - length is a required part of every request: a target between 10 seconds and 10 minutes, with a tolerance you set between 5 and 30 percent. Formats are 16:9, 1:1 and 9:16, up to three per edit, and vertical output either keeps the full frame letterboxed or fills the frame with stacked crops.
Does it generate footage or voiceover?
No. A composition contains nothing that was not in your sources - no synthetic voice, no stock, no invented sentences. Every segment carries its source timecodes, so this is checkable rather than promised.
Do I see the edit before it costs anything?
Yes. Every composition arrives first as a cut plan: the segment list with sources, transcript text and source timecodes. Reading it, editing it and revising it in plain language are free - your allowance is only touched when you render.

Ask for more than a clip.

One video or an entire archive.