Speaker Diarization with n8n: Label Who Said What on a Call
AgentLeverage Team
10/9/2026

You have a recording of a vendor call with three people on it. The transcript your meeting app gave you says Speaker 1 on every line, or no speaker at all. Somebody agreed to send a quote by the 15th, and you cannot tell from the text who it was.
This n8n workflow lets you label speakers in a call recording from any file you can upload. You open an n8n form, attach the file, type how many people were on the call and their names, and submit. When the run finishes, the browser downloads a .txt transcript with a timestamp and a name on every turn. AgentLeverage Speaker Separation does the speaker diarization. A short Code node puts your names on the labels. The template is label-speakers-call-recording.json, and it runs on the n8n-nodes-agentleverage community node, version 0.1.0 on npm.
What the flow returns
The workflow takes one call recording and returns a plain-text file named after it, such as harbor-supply-discovery-call-labeled-transcript.txt. Each line has a [mm:ss] timestamp, a speaker name, and what that person said. Back-to-back segments from the same voice merge into one turn, so the file reads like a script. A speaker you did not name stays "Speaker A", "Speaker B", and so on. If you want the background on how diarization works in general, read our speaker diarization guide. This post is the n8n how-to.
A real run in self-hosted n8n 2.42.5 with n8n-nodes-agentleverage 0.1.0. Every node succeeded on a 48-second, three-speaker test call.
This is the file that came back from that call, unedited:
[00:00] Samantha: Thanks for making time, both of you. I want to confirm what Harbor Supply needs before I send a proposal.
[00:06] Daniel: Sure. The main problem is our weekly vendor calls. Nobody can tell who agreed to what afterwards.
[00:13] Karen: And from the finance side, I need the price changes the vendors mention with the name of who said them. Got it. How many calls is that each week?
[00:22] Daniel: About 12. Most are 30 minutes, and they are recorded in Google Meet.
[00:28] Karen: Budget is approved for this quarter, but I need a written quote by the 15th.
[00:32] Samantha: I can do that. I will send the quote and a sample transcript by Friday.
[00:36] Daniel: One more thing. 2 of our vendors talk over each other a lot, so check how that comes out.
[00:43] Samantha: Noted. I will include one of those calls in the sample.Every name is right except one line. At 00:13, "Got it. How many calls is that each week?" is Samantha, but it landed at the end of Karen's turn. The review section below covers how to catch that.
Install the community node
n8n-nodes-agentleverage is on npm as 0.1.0, published with provenance. The source is Life-With-Data/n8n-agentleverage. Self-hosted n8n can install it from the UI or the CLI.
n8n UI
- Open Settings > Community Nodes.
- Select Install.
- Enter
n8n-nodes-agentleverage. - Agree to the risk warning and install.
Self-hosted CLI
cd ~/.n8n/nodes
npm install n8n-nodes-agentleverageRestart n8n after a CLI install. Then add a credential named Agent Leverage API. The URL is https://www.agentleverage.co and the API key is a token from Settings → API tokens in AgentLeverage. n8n Cloud only lists verified community nodes, so use self-hosted n8n for now.
Import the workflow from URL
In n8n, choose File → Import from URL and paste the template address. After import, open Upload Audio, Speaker Separation, and Wait for Transcript and pick your Agent Leverage credential on each. The other four nodes are stock n8n and need no credential.
https://www.agentleverage.co/n8n/label-speakers-call-recording.jsonWalk the nodes
The template has seven nodes in one line: Form, Upload Audio, Speaker Separation, Wait for Transcript, Label Transcript, Transcript File, and Download Transcript. Three are AgentLeverage nodes on one credential. The rest are stock n8n, so you can read and change every step.
1. Form
The Form trigger has three fields. Recording is required and accepts MP3, WAV, M4A, OGG, WEBM, or FLAC. Number of speakers and Speaker names are optional. Names are comma-separated, in the order people first speak.
The form n8n serves for the template, filled in for the three-speaker test call.
2. Upload Audio
Upload Audio stores the file in your AgentLeverage workspace and returns audioFileUrl. Its file name and type expressions read the Recording field whether n8n hands it over as an object or an array, so it keeps the real file name on n8n 2.x.
3. Speaker Separation
Speaker Separation starts a job from audioFileUrl. Under Additional Fields, Speakers Expected reads Number of speakers from the form and is skipped when the field is blank. Fill it in when you know the count. Left blank, a three-person call can come back as two speakers, with two voices merged under one label.
4. Wait for Transcript
Wait for Transcript polls the job every 5 seconds for up to 10 minutes. The finished job lists each speaker with talk time and segment count, plus every segment with its speaker, start and end time, and text.
5. Label Transcript
Label Transcript is a stock Code node. It applies your names to the speaker labels in the order each voice first speaks, merges back-to-back segments from one speaker, and formats the lines. Here is the full source:
// Turn Speaker Separation segments into a labeled transcript.
// Names from the form are applied in the order each voice first speaks.
const job = $json.job;
const segments = job.output?.segments ?? [];
const raw = $('Form').first().json['Speaker names'] || '';
const names = raw.split(',').map((n) => n.trim()).filter(Boolean);
const order = [];
for (const s of segments) if (!order.includes(s.speaker)) order.push(s.speaker);
const labelFor = (id) => names[order.indexOf(id)] || `Speaker ${id}`;
const stamp = (t) => {
const m = Math.floor(t / 60);
const s = Math.floor(t % 60);
return `${String(m).padStart(2, '0')}:${String(s).padStart(2, '0')}`;
};
// Merge back-to-back segments from the same speaker into one turn.
const turns = [];
for (const s of segments) {
const last = turns[turns.length - 1];
if (last && last.speaker === s.speaker) last.text += ' ' + s.text.trim();
else turns.push({ speaker: s.speaker, start: s.startTime, text: s.text.trim() });
}
const lines = turns.map((t) => `[${stamp(t.start)}] ${labelFor(t.speaker)}: ${t.text}`);
return [{
json: {
speakers: order.map((id) => ({ id, label: labelFor(id) })),
turns: turns.length,
transcript: lines.join('\n'),
},
}];The node output shows which label got which name. On the test call, A became Samantha, B became Daniel, and C became Karen, across 8 turns.
Label Transcript output from the same run.
6. Transcript File and Download Transcript
Transcript File is a Convert to File node that writes transcript to <recording>-labeled-transcript.txt. Download Transcript is a Form Ending node set to return that file, so the browser tab that submitted the form downloads it when the run ends.
What you still review
A labeled transcript from this workflow is a working draft. Read it once before you rely on who said what.
- Short replies can stick to the previous speaker. When someone answers right after another person with no pause, their line can land in the previous turn. That is what happened at 00:13 above. Skim quick back-and-forth spots.
- Names follow order of first speech. If the second person to talk is not the second name you typed, every label after that is wrong. Check the first two or three turns.
- Labels are not identity recognition. Speaker Separation tells voices apart. It does not know who anyone is. The names come only from the form.
- If you would rather rename speakers on the AgentLeverage job itself, the node also has a Speaker Separation > Update Labels operation. This template does not use it.
Cost and limits
Speaker Separation costs 2 credits per minute of audio, rounded up, so a 48-second call costs 2 credits. New accounts start with 30 credits and no card. See pricing. The Wait for Transcript node times out after 10 minutes. For a long recording, raise timeoutMs on that node.
Next steps with the labeled job
If you want a recap, decisions, and action items from the same call: turn the Speaker Separation job into minutes in n8n
If you have one file and do not need n8n: split the audio by speaker in the browser
Good to know
- Set Number of speakers when you know it. Without it, similar voices can merge into one label.
- The transcript is not a certified legal or medical record. For a court, regulator, or patient chart, use the process that record already requires.
- n8n Cloud will not list the node until n8n verifies it. Self-hosted n8n installs it from npm today.
Frequently asked questions
What is speaker diarization with n8n?
Speaker diarization splits a recording into turns and tags each turn with the voice that spoke it. In n8n, this template sends the file to AgentLeverage Speaker Separation through the n8n-nodes-agentleverage node, then uses a Code node to put your names on the labels and return a .txt file.
Can I install n8n-nodes-agentleverage from npm?
Yes. The package is n8n-nodes-agentleverage 0.1.0. Install it from Settings > Community Nodes > Install in self-hosted n8n, or run npm install n8n-nodes-agentleverage in ~/.n8n/nodes and restart n8n.
Does it know people's names?
No. Speaker Separation labels voices A, B, C. The workflow maps the names you type in the form to those labels in the order each voice first speaks. Leave Speaker names blank and the file keeps "Speaker A" and so on.
Try it yourself
Speaker Separation
Upload a recording. This speaker separation labels multi-speaker audio. Rename Speaker 1 and download isolated stems plus a timestamped transcript.
Start with 30 free credits — no card required.
Related Articles
Call Recording to Minutes in n8n
Upload a call recording in n8n, run Speaker Separation, then Minutes from that job. Install n8n-nodes-agentleverage. You still rename Speaker A and review.
The same sequence with a simple prompt in Claude or ChatGPT
Paste a prompt in Claude or ChatGPT that runs Speaker Separation, then Minutes from that job. You still rename Speaker A and review.
Move past Speaker 1 chat paste with Speaker Separation
Chat has the recap. The file still says Speaker 1. Label, rename, export txt/srt. Guest Speaker Separation is 2/day. Signed-in orgs pay 2 credits/min.
Get weekly AI tips
Practical AI productivity tips every week. No fluff.