Do AI notetakers train on my meeting data?
Usually not in the way people picture it — your meeting is very unlikely to end up in the weights of a general-purpose AI model — but that's the narrowest version of the question, and it's not the one worth asking. 'Training' bundles three different things: training a model on your words, using your content to 'improve the service' (a much broader phrase that can cover evaluation, tuning, and a human reviewing samples), and simply retaining your transcript to operate and debug the product. Most people are protected from the first and would object most to the second and third, which is exactly where policy language is vaguest. The bigger structural point: no notetaker runs its own speech and language models end to end, so the answer depends as much on the providers it calls as on the vendor itself — which is why a named subprocessor list is the single most informative thing on a privacy page. And the same vendor often answers differently depending on which door you came in, since consumer free tiers and business or API tiers are commonly governed by different terms. Check three things: the subprocessor list, the retention period, and whether an 'improve our AI' setting is on by default.
Last updated August 31, 2026
Probably not, in the sense you mean it. Your Tuesday standup is very unlikely to have been absorbed into the weights of a general-purpose AI model. Reputable notetakers say so explicitly, and the model providers they build on generally commit to the same thing on their commercial tiers.
But that’s the narrowest reading of the question, and answering only that reading is how vendors say “we don’t train on your data” while leaving the part you’d actually mind untouched.
“Training” is three different things wearing one word:
- Training a model on your words — your meeting’s content becoming part of what some model learned. This is what people picture, and it’s the one they’re most likely to be protected from.
- “Improving the service” — a much broader phrase that can cover building evaluation sets, tuning models, debugging quality complaints, and, in some products, a person reading a sample of real content to check how the system performed.
- Just holding onto it — retention for operating the product: logs, caches, abuse monitoring, a support engineer who can reach your transcript to fix a bug.
Only the first is training. But the second and third are where the exposure usually lives and where the policy language is vaguest — and if what you’re really asking is “could a stranger end up reading what we said in that call,” those are your answers, not the first one. Worth separating them before you go looking.
The chain your meeting actually travels
Here’s the part that reframes the question. No notetaker runs the whole pipeline itself. Turning speech into a live summary takes a speech-to-text model and a language model, and essentially every tool in this category calls someone else’s. So “does this vendor train on my data” is really two questions — what the vendor does, and what it has contracted for with the companies it hands your content to.
| Layer | What it sees | The question worth asking | Where the answer lives |
|---|---|---|---|
| The notetaker vendor | Everything: audio (if kept), transcript, summaries, your account | Does it train on customer content, and for how long does it keep anything? | Privacy policy + retention section |
| Its speech-to-text provider | The audio, or a stream of it | Is it on a tier that doesn’t train on submitted content? Is audio retained after transcription? | The vendor’s subprocessor list |
| Its language-model provider | The transcript text | Same question, one layer up — this is the layer that sees your words as words | The vendor’s subprocessor list |
| The meeting platform (if you used its built-in AI) | The whole meeting, from inside | What did your organization’s agreement with that platform actually cover? | Your company’s contract, not a public page |
| Wherever the notes go next | Whatever you paste, forward, or upload | What are that product’s terms? | Not the notetaker’s policy at all |
Two rows deserve more attention than they get. The subprocessor list is the most informative thing on a privacy page and almost nobody reads it: a vendor that names the companies processing your meetings is telling you who else has them, and a vendor that won’t name them is telling you something too. And the last row is the one this page exists to point at, because it’s the layer where the careful policy you just verified stops applying.
The split that decides most answers
If you take one practical thing from this page: the same product frequently has two different answers depending on which door you came in.
Across the software industry, consumer free tiers and business, enterprise, or API tiers of the same service are commonly governed by different terms — broader rights over submitted content on the free consumer side, stricter commitments on the business side. It’s not a trick; it’s the ordinary economics of a product with no subscription revenue, which the site’s what “free” actually costs walks through in more detail. But it means “does X train on my data” is unanswerable without also naming the plan, and a policy summary you read somewhere may have been describing the other tier.
This is the specific reason to check rather than infer, and to re-check for the plan you’re actually on. Terms in this area are also being revised often enough that a confident answer from a year ago isn’t one now.
The training question is downstream of the retention question
Here’s the part that doesn’t appear on anyone’s pricing page. What can be trained on is bounded by what exists.
A tool that keeps a full audio or video recording of every meeting you’ve ever had is holding the largest possible corpus of your working life. A tool that keeps the transcript is holding less. A tool that keeps only summaries is holding least. The commitments each of them makes may be identical — but the commitments are promises, and the retention is a fact.
That distinction matters more than it sounds. A promise not to train is a promise about the future handling of data you’ve already handed over, made by a company that can be acquired, change its terms with notice, or be compelled to produce what it holds. Retention limits are the only part of the answer that survives all three. Data that no longer exists can’t be repurposed by anyone, including by whoever owns the company next year. If you want a durable answer rather than a current one, ask what’s kept and for how long — can I get meeting notes without recording the call covers what each capture method leaves behind.
Which also explains a structural difference between products that’s easy to miss. For a tool whose value proposition is the searchable archive — every call you’ve had, forever, in a dashboard — the corpus is the product, and keeping everything isn’t a policy choice so much as the business. For a tool whose value is delivered during the meeting, the useful work is finished when the call ends, and there’s much less reason to accumulate. Neither is dishonest. They just have different amounts of your data sitting around, which turns out to be the thing the training question is downstream of.
The exposure that isn’t the notetaker
Once your notes exist, they’re text — and text moves.
The most likely path from your meeting into somebody’s training pipeline probably isn’t the notetaker at all. It’s the transcript pasted into a free consumer chatbot to draft the follow-up email, the summary forwarded into a doc in a tool nobody vetted, the recording uploaded to a transcription site someone found. Each of those is governed by that product’s terms, not by the notetaker’s — and the person doing it has usually just spent no time at all on the question you’re currently spending time on.
The job interviews page flags a sharp version of this: pasting a candidate’s transcript into a general chatbot moves a named person’s unguarded words somewhere nobody told them about. The everyday version is milder and far more common. It isn’t an argument against using AI on a transcript — that’s genuinely useful, and ChatGPT is a good place to work with one after the fact. It’s an argument for doing it on a tier whose terms you’ve actually read, and for noticing that a meeting transcript is a denser concentration of other people’s names, numbers, and half-formed opinions than almost anything else you handle.
You’re answering this for everyone who was in the room
Every other privacy question about meeting notes is at least partly about you. This one isn’t, and that’s the thing worth sitting with.
A transcript is not your data. It’s four other people’s words with their names attached, captured in a setting where they were thinking out loud. When you leave an “improve our AI” setting switched on, or paste the transcript somewhere convenient, you are making a data-processing decision on behalf of people who were never shown the setting. They agreed — if they were asked at all — to notes being taken. That’s not the same as agreeing to a specific vendor, two model providers, a thirty-day log retention, and whatever you do with the file afterwards.
Norms have mostly caught up to “I’m using an AI notetaker.” They haven’t caught up to the sentence after it. You don’t need to read a subprocessor list aloud at the top of a standup, but for anything sensitive — customer calls, personnel conversations, anything under an NDA — the honest disclosure is slightly longer than the one people say, and naming the tool is the cheap way to give it: “I’m running [tool] for notes” lets anyone who cares go look, which “I’m taking AI notes” does not. See how to tell participants you’re using an AI notetaker for the wording, and note that consent rules are about the capture rather than the storage, so none of this replaces asking first.
The same logic points at the practical rule for the settings you do control: when you’re deciding on behalf of other people, default to the more conservative option, because they can’t.
Where Canary sits
Canary is a real-time, bot-free meeting summarizer. It captures your computer’s system audio (no bot in the call, no plugin) and shows a live, multi-resolution rolling summary — from what’s being said right now to the whole call — so you can catch up the instant your name is called.
On this specific question, the honest answers are:
- Your content isn’t training material. Speech-to-text and summarization run on third-party API tiers that don’t train on your content by default, and those providers are named in the privacy policy — so you can check the layer below rather than take the top layer’s word for it.
- Raw audio isn’t kept. Audio is streamed in short chunks for transcription and then discarded. What’s stored is the text transcript and the rolling summaries, encrypted at rest.
- Capture is on-demand. Nothing runs unless you start it (or use a calendar auto-start you opted into), so there’s no ambient corpus accumulating in the background.
- Deletion is real. Free-tier notes are purged automatically after a week; account deletion is permanent.
And the boundary, stated plainly because a page arguing for scrutiny shouldn’t exempt itself: Canary is not an offline product. Audio and transcript text pass through cloud transcription and language models on the way to becoming a summary, exactly as they do for essentially every hosted notetaker, bot-free ones included. Not sending a bot into the call removes a participant holding a copy of the conversation; it doesn’t remove the pipeline. If a fully on-device tool is a hard requirement, the local-model path in the free notetaker answer is the only architecture that meets it, and it’s real work.
Bottom line
Your meetings are probably not training a public AI model — and that was the easy question. The useful ones are what your vendor means by “improve the service,” which companies are named on its subprocessor list, how long anything is kept, whether the plan you’re on is governed by the terms you read, and where the transcript goes after it exists. The retention answer is the one that ages best, because it’s the only part that doesn’t depend on a promise. And whatever you decide, remember you’re deciding for everyone who was on the call.
Frequently asked questions
Does Otter, Fireflies, or Fathom train on my meetings?
That's a question to answer from the vendor's current policy rather than from any third party, including this page — these terms differ by plan and get revised, and a claim about them is out of date the moment it's published. The 60-second check works on any of them: open the privacy policy and search the page for 'train,' 'improve,' 'retention,' and 'subprocessor.' You want an explicit statement that customer content isn't used to train models, a named list of the speech-to-text and language-model providers that actually process the content, a stated retention period, and a clear note on whether the defaults differ between free and paid plans. The pattern worth knowing across the whole software industry is that consumer free tiers and business or enterprise tiers of the same product are frequently governed by different terms, so confirm the answer for the plan you're actually on. If a vendor won't answer these plainly, that's your answer.
Is it safe to paste a meeting transcript into ChatGPT or another chatbot?
It's the most common way meeting content leaves a carefully-worded policy, and it's worth a moment's thought before you do it. Once the transcript is text on your clipboard, it's governed by whatever product you paste it into — not by your notetaker's terms — and consumer chatbot tiers commonly reserve broader rights over submitted content than business tiers of the same product do. The content itself also deserves a second look: a meeting transcript is other people's unguarded words, usually with their names attached, and often includes customer names, salary figures, or unannounced plans that nobody expected to travel. None of that makes it forbidden — working through a transcript with an AI assistant afterwards is genuinely useful. It means the decision is yours to make deliberately: use a business or team tier where the terms are stricter, check whether your organization has an approved tool, and strip names and identifiers from anything sensitive first.
Can I opt out of my meeting data being used to train AI?
Often yes, and it's usually a setting rather than a support request — but the default matters more than the availability, because most people never open the page. Look in account or workspace privacy settings for anything phrased as 'help improve our AI,' 'share data for product improvement,' or 'allow human review,' and note that on a company-provisioned tool this is frequently an administrator-level control rather than a personal one, so the answer may not be yours to change. Two caveats keep expectations right: opting out generally applies going forward, not retroactively to what you've already sent; and an opt-out is a promise about future use, so a short retention period is the sturdier protection — data that no longer exists can't be repurposed by anyone.