Anjin Media

Video intelligence

Separate the voices.
Keep the conversation clear.

Anjin assigns speech to distinct speakers, lets your team name them and carries that identity through transcript search and editorial planning.

One video or an entire archive.

Speaker diarisation product demonstration

Speakers

SpeakersDIARISATION COMPLETE
J

Jocko

SPEAKER 1

46%
M

Matt

SPEAKER 2

31%
J

Jeff

SPEAKER 3

23%
27:44Jockoand it just locked in

What is speaker diarisation for video?

The separation of recorded speech into consistent speaker-labelled segments.

Speaker labels are stored at transcript level and group speakers can be named through the API. Naming is snapshotted into planning work to keep results stable.

Controls and evidence

See what the system used and produced.

Useful automation leaves something a person can inspect. Anjin keeps the edit connected to its brief, source media and output settings.

SPEAKER DIARISATIONVERIFIABLE OUTPUT
diar_label
The stable detected label for a voice in a source group.
speaker_name
A human-readable name supplied by your team.
transcript
Timed words associated with the detected speaker.
snapshot
Names captured for consistent asynchronous planning.

Why it matters

More control, less timeline work.

01

Search by voice

Understand who made a statement as well as what was said.

02

Clarify cut plans

Review multi-speaker edits without anonymous transcript blocks.

03

Improve captions

Carry clearer speaker context into downstream editorial outputs.

Fit and boundaries

Use the capability where it genuinely helps.

Anjin is strongest on speech-led recorded material where the editorial job can be described clearly and every selected moment needs to remain verifiable. It supports human judgement with searchable evidence, a reviewable plan and structured outputs.

A strong fit

Recorded knowledge with a story inside it

Interviews, podcasts, webinars, discussions and archive programmes benefit when the useful material is distributed across a long recording or several selected sessions.

Keep elsewhere

Final craft and image-led montage

Detailed colour, sound design, motion graphics and wordless visual storytelling remain finishing tasks for a professional editor and their preferred creative tools.

Evaluation checklist

What should you verify before adopting speaker diarisation?

CHECK 01

diar_label

Confirm that the workflow exposes this clearly: the stable detected label for a voice in a source group.

CHECK 02

speaker_name

Confirm that the workflow exposes this clearly: a human-readable name supplied by your team.

CHECK 03

transcript

Confirm that the workflow exposes this clearly: timed words associated with the detected speaker.

CHECK 04

snapshot

Confirm that the workflow exposes this clearly: names captured for consistent asynchronous planning.

Questions about speaker diarisation

What does speaker diarisation do?
It answers who spoke when by assigning speech segments to consistent speaker labels within the recording.
Is diarisation the same as voice identification?
No. Diarisation separates voices. A person or system still needs to attach a real name to a detected speaker label.
Can speakers be renamed?
Yes. Anjin provides speaker naming controls so plans can use recognisable names instead of generic labels.
Does it work with several camera angles?
Yes. Speaker information belongs to the synchronised source group and can inform planning across its available angles.

Use speaker diarisation in your next edit.

Start with footage you already have and a clear description of the result you need.

One video or an entire archive.