Answer

How do AI notetakers know who's speaking?

Short answer

Most of them don't work it out on their own — they separate the voices and get the names from somewhere else. Deciding who spoke is really two problems: separation, which is acoustic (grouping the audio into distinct voices), and identification, which is not (attaching a human name to each voice). Sound alone can establish that two different people spoke; nothing in the sound reveals that one of them is Priya. So every tool needs an outside source of names, and which one it has is decided by where it sits. Tools inside the meeting — a notetaker bot in the participant list, or the platform's own built-in AI — read the participant roster, and often get one clean audio stream per person, so they get both halves almost free. Tools that hear one mixed stream, including anything capturing your computer's system audio, a recorder on a table, or a phone, do the acoustic work themselves and have no roster to read. That is exactly what "Speaker 1" means: the separation worked and the naming had nothing to work from. The fix that works on every tool is the low-tech one — say people's names out loud during the call.

Last updated September 1, 2026

Most of them don’t work it out on their own. They separate the voices, and then they get the names from somewhere else. That distinction explains almost everything about why speaker labels behave the way they do — and the single most useful thing to know is that a name on a transcript line comes from the meeting, not from the audio.

Two problems, not one

“Who’s speaking” is two jobs that get done by completely different means:

Sound can establish that two different people spoke. It can never, on its own, reveal who they are. So every tool in this category needs an external source of names, and there are only four: the meeting platform’s participant roster, a voice profile someone enrolled in advance, a name said out loud in the conversation itself, or a human typing it in afterwards.

Speaker 1 is what you see when separation worked and identification had nothing to work from. It’s not a missing setting or a bug — it’s a tool being honest that it heard audio and nothing else.

Where the tool sits decides what it knows

This falls straight out of the four ways notetakers capture a meeting. The methods that live inside the meeting get identity handed to them; the methods that live on your machine have to earn it.

Capture methodHow it separates voicesWhere names come fromTypically shows
Notetaker bot in the callOften per-participant streams from the platform, so barely any acoustics neededParticipant roster — the bot is a participant and can see the listReal names
Platform-native AI (Teams, Zoom, Meet, Webex)Inside the pipeline; one stream per participantThe account signed into each connectionReal names, most reliably of any method
Caption-scraping extensionDoesn’t separate anything — reads text the platform already labelledThe platform did it before rendering the captionReal names
System-audio capture (no bot)Acoustic clustering on one mixed streamNo roster available — names must be enrolled, spoken, or typedNumbered labels, unless told otherwise
Recorder in a room, or a phoneAcoustic clustering on one mixed stream, in bad acousticsSame — nothing to readNumbered labels

The pattern is worth stating plainly, because it cuts against the direction this site usually argues: on attribution specifically, being in the meeting is a real advantage. A bot gets names for free because it’s sitting in the participant list, and platform AI gets them more reliably still because it’s inside the software that authenticated everyone. A tool that never joins the call has traded that away. That’s a genuine cost of bot-free capture, and it’s covered honestly in how accurate are AI meeting notes as well.

Two bugs that look identical and aren’t

When speaker labels are wrong, they’re wrong in one of two ways, and knowing which one you have tells you whether it’s fixable.

Separation failed. The count is wrong. One person appears as Speaker 2 for the first half and Speaker 5 for the second, because they switched from a headset to laptop speakers, moved away from the mic, or simply got quieter. Or two similar voices on similar microphones collapse into one label. Crosstalk drives both, since the boundary between speakers is precisely what’s destroyed when people talk over each other. This is an audio problem and responds to audio fixes.

Identification failed. The voices are separated perfectly and the names are wrong anyway — and this one surprises people, because it happens on tools that do have a roster. The reason is a fact worth keeping: a participant list is a list of connections, not a list of people. One connection labelled “Conf Room 4” is three humans sharing a table microphone. A dial-in appears as a phone number. Two colleagues watching from one laptop are, as far as the platform is concerned, one person. Somebody joins from a partner company’s shared account and every line they say is attributed to a name that isn’t theirs. The roster is accurate about who connected; the transcript quietly treats it as a claim about who spoke.

That second failure is the one that reaches other people, because it produces a transcript that looks authoritative and reads wrong.

What actually fixes it

Ranked by how much they help, and the first one is not close:

  1. Say names out loud. “Priya, does that work on your side?” puts the identity into the content, where every tool can read it regardless of how it captured the audio. It’s the one form of attribution that survives every capture method, it works retroactively when you’re reading the transcript yourself, and it helps the humans dialled in who can’t see who’s talking either.
  2. One person, one device, one connection. The shared conference-room laptop is the most common cause of tangled labels anywhere. If a meeting’s record matters, have people in the room join individually on mute, or accept that the room is one speaker.
  3. Rename once, early. Most tools propagate a rename across every line from that speaker. Ten seconds, and it should happen before you forward anything.
  4. Wear headphones. Slightly counterintuitively, this improves attribution for tools capturing your own machine: with speakers, your microphone picks up your voice and the call bleeding through the room, smearing the two together. With headphones they stay cleanly separate — see can an AI notetaker hear my meeting if I wear headphones.
  5. Enrol a voice profile, if the tool offers it and you’re comfortable with it. It’s the only way a tool with no roster can produce a real name unprompted.

Nothing on this list is exotic, and the top two are behavioural rather than technical — which is a fair summary of the whole problem.

Why the labels carry more weight than they look like they do

Attribution is what turns a transcript into a record. It’s also load-bearing for the layer above: action-item detection is only ever as good as the speaker labels underneath it, so a misattributed commitment doesn’t stay a formatting nuisance — it becomes a task assigned to the wrong person in a summary that gets forwarded to people who weren’t in the room and can’t correct it. Of all the ways an AI summary goes wrong, misattribution is among the most consequential, because a wrong name attached to a decision looks exactly as confident as a right one.

The in-person and hybrid case is the extreme version: one distant microphone, several voices, real reverberation, and no per-participant streams to fall back on. A room is where both halves of this problem fail at once.

Where Canary sits on this

Canary is a real-time, bot-free meeting summarizer. It captures your computer’s system audio (no bot in the call, no plugin, no virtual audio device) and shows a live, multi-resolution rolling summary — from what’s being said right now to the whole call — so you can catch up the instant your name is called. It runs on macOS, Windows, and Linux.

Honestly, that architecture is on the harder side of this particular problem. Reading the audio your computer is already playing means one mixed stream and no participant roster to consult, the same position a room recorder is in — better audio, because remote voices arrive as cleanly as they reached your speakers, but the same absence of a name list. A notetaker bot or the platform’s own AI has a structural advantage here and it would be silly to pretend otherwise. If a verbatim, reliably-attributed transcript is the artifact you need, Teams live transcription and the transcription-first tools like Otter are pointed at that job.

There is one speaker it can always be certain about, though, and it’s the one you care about most often: you. Your own voice arrives via your microphone, on a separate path from everyone else’s, so “you” and “not you” is a distinction that comes free — which is what makes your commitments and your questions the part of the record that holds up best.

The larger point is about timing rather than technique. The reason anyone wants speaker labels is almost always to answer a question they’d have known the answer to if they’d been following: who just asked me that? who objected? whose deadline is that? Labels are a reconstruction aid, and how badly you need them scales with how long after the fact you’re reading. Twenty seconds after a sentence, you still have the meeting around you — the tiles, the voice you recognise, the person still sitting there to be asked. Two days later, all you have is the transcript, and the label is the only thing standing between you and a guess. Closing that gap is the whole idea behind reading a summary during the call rather than after it: the best fix for attribution is the same as the best fix for accuracy, which is to look while the meeting is still there to correct you.

One thing worth thinking about

Better speaker labels make the record more useful and also more consequential, and that cuts both ways. An unlabelled transcript is a blur. A labelled one is quotable material attached to named people who were speaking casually, thinking out loud, disagreeing in a first draft — and it’s far easier to forward, search, and take out of context. When you tell people you’re using an AI notetaker, what they typically picture is notes being taken; “and each sentence will be attributed to you by name” is a different sentence, and on a sensitive call it’s the one actually worth saying. The same applies to voice enrolment, which creates an identity record rather than a meeting record — see who can see my AI meeting notes and do AI notetakers train on my meeting data for where those end up. None of this argues for worse labels. It argues for being as careful with an accurate record of other people’s words as you’d want them to be with yours.

Bottom line

Separating voices is an audio problem; naming them isn’t one. Tools sitting inside the meeting read the roster and get names nearly free, tools listening to one mixed stream get numbers until you tell them otherwise, and both can still be wrong — because a participant list describes connections rather than people. If the record matters, say names out loud, give people their own connections, and rename once before you forward anything. And if what you actually need is to know who just said the thing you missed, the cheapest answer has never been a better label. It’s a shorter gap between the sentence and the moment you read it — which is the argument for bot-free notes that arrive during the call rather than after it.

Frequently asked questions

Why does my transcript say Speaker 1 and Speaker 2 instead of names?

Because the tool could tell the voices apart but had no list of names to match them to. Numbered labels are the honest output of a tool that heard audio and nothing else — they are not a bug or a missing setting, and no amount of better audio will turn them into names. Names have to come from outside the sound: the meeting platform's participant roster (which only a tool sitting inside the meeting can read), a voice profile you enrolled in advance, someone's name being spoken out loud in the conversation, or you typing it in once afterwards. Most tools that produce numbered labels let you rename a speaker once and apply it to every line, which takes about ten seconds and is worth doing before you forward the transcript to anyone.

Why does one person show up as two different speakers?

That's a separation failure rather than a naming failure, and it usually means the person's voice reached the tool in two different-sounding ways. Switching from a headset to laptop speakers mid-call, moving away from the microphone, a connection quality change, or simply speaking much more quietly in the second half can all produce audio that clusters as a new voice. The reverse happens too: two people with similar voices, on similar microphones, can be merged into one label. Crosstalk makes both worse, because the boundary between speakers is exactly what gets destroyed when people talk over each other. If a transcript matters, the cheapest insurance is one person per device and one device per connection — a shared laptop in a conference room is the single most common cause of tangled speaker labels.

Can an AI notetaker recognise my voice specifically?

Some tools offer voice enrolment, where you record a short sample once and the tool matches that voiceprint against future meetings so your lines get your name even when there's no roster. It works reasonably well and it is the only way a tool hearing one mixed stream can put a real name on a voice without being told. It's also worth thinking about before you turn it on, because a stored voiceprint is a durable identity record rather than a transcript, and enrolling other people's voices is a decision you'd be making on their behalf. Availability and storage terms vary by tool, so check the specific product rather than assuming — and see who can see my AI meeting notes for where these records actually live.