Speaker Identification: Who Said What
Learn how speaker diarization works by plan in VoxScriber: automatic AI naming plus how to rename speakers and separate transcriptions by who said what.
With speaker detection enabled, the transcription comes out organized by who is speaking, with timestamps per segment. Then, the AI also tries to figure out the name and role of each participant on its own — and you can rename anyone manually.
How to Enable
When uploading the file, enable the speaker detection option. If you know how many people speak in the audio, enter the expected number — this helps the engine separate voices more accurately.
What Each Plan Offers
| Plan | Speaker detection |
|---|---|
| Free | Enabled, but the result shows up to 2 speakers |
| Lite and above | Full, with no speaker limit |
On the free plan, detection runs normally, but the interface shows at most 2 distinct speakers. From Lite onward, all detected speakers appear.
For meetings and interviews with many voices where precise attribution is essential, the Ultra engine (Pro plans or higher) is the specialist in speaker separation. See Choosing the engine.
Automatic AI Naming
Immediately after transcription, the AI analyzes the content and tries to identify the name and role of each speaker (for example, "Dr. Ana — interviewer"). This feature:
- Runs automatically for all plans with diarization;
- Does not consume any AI quota;
- Is skipped when there are fewer than 2 speakers;
- Uses your interface language for the roles.
Names suggested by the AI never overwrite names you defined yourself — what you typed always wins.
Renaming Speakers Manually
Open the transcription
In your library, open the desired transcription.
Click on the speaker's name
Click on the speaker's label (for example, "Speaker 1") and type the real name.
Done — saved instantly
Renaming speakers saves immediately, without needing the save button. The new names appear in exports (PDF, DOCX, TXT) and in the meeting minutes.
Renaming speakers saves instantly, but text edits do not autosave — use the Save button. See Editing the transcription.
Multichannel Audio
Recordings with multiple channels (for example, telephony with one channel per participant) can be combined with speaker detection: labels come out as "1A", "2A" (channel + speaker), preserving which channel each utterance came from.
Frequently Asked Questions
Why do only 2 speakers appear on the free plan?
Detection runs fully, but the free plan shows up to 2 distinct speakers. Upgrading to Lite or above unlocks all detected speakers — including on transcriptions already created.
The AI gets someone's name wrong — what do I do?
Click the name and fix it. Your edit overrides the AI's suggestion and applies to all exports.
Does automatic naming consume my AI quotas?
No. It is free and automatic, separate from the summary, chat, and minutes quotas.
Do I need to tell how many people speak in the audio?
It's not required, but it helps: entering the expected number of speakers improves voice separation.
Related Articles
- Choosing the transcription engine
- Accuracy and how to improve it
- Editing the transcription
- AI resources
Still not solved? Open a ticket and our team will help you.