en
3 min read Transcription

Speaker Identification: Who Said What

Learn how speaker diarization works by plan in VoxScriber: automatic AI naming plus how to rename speakers and separate transcriptions by who said what.

With speaker detection enabled, the transcription comes out organized by who is speaking, with timestamps per segment. Then, the AI also tries to figure out the name and role of each participant on its own — and you can rename anyone manually.

How to Enable

When uploading the file, enable the speaker detection option. If you know how many people speak in the audio, enter the expected number — this helps the engine separate voices more accurately.

What Each Plan Offers

PlanSpeaker detection
FreeEnabled, but the result shows up to 2 speakers
Lite and aboveFull, with no speaker limit

On the free plan, detection runs normally, but the interface shows at most 2 distinct speakers. From Lite onward, all detected speakers appear.

For meetings and interviews with many voices where precise attribution is essential, the Ultra engine (Pro plans or higher) is the specialist in speaker separation. See Choosing the engine.

Automatic AI Naming

Immediately after transcription, the AI analyzes the content and tries to identify the name and role of each speaker (for example, "Dr. Ana — interviewer"). This feature:

  • Runs automatically for all plans with diarization;
  • Does not consume any AI quota;
  • Is skipped when there are fewer than 2 speakers;
  • Uses your interface language for the roles.

Names suggested by the AI never overwrite names you defined yourself — what you typed always wins.

Renaming Speakers Manually

1

Open the transcription

In your library, open the desired transcription.

2

Click on the speaker's name

Click on the speaker's label (for example, "Speaker 1") and type the real name.

3

Done — saved instantly

Renaming speakers saves immediately, without needing the save button. The new names appear in exports (PDF, DOCX, TXT) and in the meeting minutes.

Renaming speakers saves instantly, but text edits do not autosave — use the Save button. See Editing the transcription.

Multichannel Audio

Recordings with multiple channels (for example, telephony with one channel per participant) can be combined with speaker detection: labels come out as "1A", "2A" (channel + speaker), preserving which channel each utterance came from.

Frequently Asked Questions

Why do only 2 speakers appear on the free plan?

Detection runs fully, but the free plan shows up to 2 distinct speakers. Upgrading to Lite or above unlocks all detected speakers — including on transcriptions already created.

The AI gets someone's name wrong — what do I do?

Click the name and fix it. Your edit overrides the AI's suggestion and applies to all exports.

Does automatic naming consume my AI quotas?

No. It is free and automatic, separate from the summary, chat, and minutes quotas.

Do I need to tell how many people speak in the audio?

It's not required, but it helps: entering the expected number of speakers improves voice separation.

Still not solved? Open a ticket and our team will help you.