What Is an AI Scribe and How Does It Work in a Consultation?

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for What Is an AI Scribe and How Does It Work in a Consultation?
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for What Is an AI Scribe and How Does It Work in a Consultation?

An AI scribe is software that listens to a consultation through a phone or computer microphone, turns the conversation into a transcript, and uses a language model to draft a structured clinical note for the clinician to check, edit and sign. It documents; it does not diagnose. The clinician remains responsible for every word that goes into the record.

The reason to understand the mechanics is that the draft looks finished. Its errors are small and plausible: a dropped "no" that turns "no bladder changes" into "bladder changes", a left knee that becomes a right knee, an examination finding the clinician never actually performed. Each stage of the pipeline produces its own kind of mistake, so knowing the stages tells you where to look when you review.

Follow me on Instagram@sagnikteaches

The consultation, stage by stage

Most AI scribes, whether they are general products such as Heidi or Nabla, enterprise systems such as Microsoft Dragon Copilot, or a feature built into practice software, follow the same sequence.

Connect on LinkedInSagnik Bhattacharya

Before the patient sits down

  1. Template choice. The clinician picks, or the system defaults to, a note template: a general consultation, a SOAP note (subjective, objective, assessment, plan), a specialty letter. The template decides the headings the draft will use.
  2. Context. Tools linked to the clinical record may pull in the patient's name, age and problem list. Standalone tools start with nothing and know only what is said aloud.
  3. Consent. The patient is told the consultation will be transcribed by an AI tool and asked if that is all right. The scribe does not do this for you.

During the consultation

  1. Capture. The clinician taps start. The microphone records the conversation, either "ambiently" (the whole natural conversation) or as dictation (the clinician speaking the note after the patient leaves). Audio is usually sent to the vendor's servers, either streamed as it is recorded or uploaded when the clinician stops.
  2. Speech recognition. A speech-to-text model converts the audio into words. Medical vocabulary, drug names and numbers are the hard parts.
  3. Speaker separation. Often called diarisation, this labels which words came from the clinician and which from the patient (and any relative or interpreter). It matters because "I've had chest pain" means something different from "have you had chest pain?"

After the patient leaves

  1. Drafting. A large language model reads the labelled transcript and writes the note under the template's headings, removing small talk and repetition. This usually takes seconds.
  2. Review and edit. The clinician reads the draft against their memory of the consultation, corrects it, and signs it. This is the only step that makes the note trustworthy.
  3. Export. The signed note is copied or pushed into the clinical record. Deeper integrations write straight into the right fields; simpler tools rely on copy and paste.
  4. Deletion or retention. The audio is deleted or kept according to the vendor's settings, and the transcript is kept for a set period or deleted. More on this below, because it varies more than people expect.

What goes in and what comes out: a short example

An illustrative extract from a transcript, lightly cleaned, for a patient with back pain:

Subscribe on YouTube@codingliquids

Clinician: So how long has this been going on?
Patient: About five days. I was moving boxes at the weekend.
Clinician: Any pain going down your legs?
Patient: No, just here in the low back, more on the right.
Clinician: Any numbness, or any changes with your bladder or bowels?
Patient: No, nothing like that.
Clinician: What have you been taking?
Patient: Ibuprofen, it takes the edge off.
Clinician: OK, I'm just going to press here... that's tender on the right side. Can you lift your leg for me, straight... both sides fine, no pain down the leg. Reflexes are normal and sensation is normal in both legs.

And the drafted note (illustrative), using a SOAP template:

S: Low back pain for 5 days after lifting boxes. Right-sided, no
   radiation to legs. No numbness. No bladder or bowel changes.
   Ibuprofen gives partial relief.
O: Tenderness right lumbar paraspinal region. Straight leg raise
   negative bilaterally. Reflexes and sensation normal in lower limbs.
A: Mechanical low back pain. No red flag features.
P: [from the rest of the consultation] Continue simple analgesia,
   keep active, review in 2 weeks if not improving. Safety-netting
   advice given on leg numbness and bladder or bowel changes.

Notice two things. First, the assessment "no red flag features" is the model summarising; the clinician must agree with it before signing, because it is a clinical conclusion. Second, the objective findings only appear because the clinician said them aloud. Had the leg raise been done in silence, the scribe would either have left it out or, worse, filled the gap with a plausible "neurological examination normal".

Where the errors come from at each stage

This is the part worth pinning next to the screen. Each example is illustrative but typical of what published reviews of AI scribes describe: omissions, errors and hallucinated content.

StageTypical errorExampleHow to spot it
Speech recognitionMisheard drug name or number"Amlodipine 5 mg" becomes "10 mg"Check every dose, date and measurement
Speech recognitionDropped negation"No bladder changes" becomes "bladder changes"Read red-flag lines word by word
Speaker separationClinician's question recorded as patient's symptom"Any chest pain?" becomes "reports chest pain"Ask: did the patient say this, or did I ask it?
Speaker separationRelative's words credited to patientDaughter's worry recorded as patient's complaintCheck who reported what in the history
DraftingInvented finding"Chest clear on auscultation" when the chest wasn't examinedCheck every objective finding was done
DraftingWrong sideRight knee becomes left kneeCheck laterality in every line
DraftingOmissionSafety-net advice given but not recordedCheck the plan matches what you said
DraftingOver-confident summary"Likely viral" when you said "could be viral"Check the assessment says what you meant

The invented finding is the most dangerous because it reads like good documentation. A realistic way it shows up: a clinician signs twenty notes at the end of a clinic, and weeks later a colleague reads "abdomen soft, non-tender" in a note for a patient whose abdomen was never examined. Nothing flagged it, because it looked exactly like what a thorough clinician would have written.

A 60-second review before signing

Reviewing a note properly does not mean re-reading it from the top as if it were an essay. It means checking the lines most likely to be wrong, in the same order every time, so the habit survives a busy clinic. A routine that fits in about a minute:

  1. Red flags and negatives. Read every "no" and "denies" line word by word. These are where a dropped word changes the meaning.
  2. Numbers. Doses, frequencies, durations, measurements, dates. Compare each with what you said.
  3. Sides and sites. Left or right, upper or lower, which tooth, which joint.
  4. Examination. For each finding listed, ask: did I actually do that? Delete anything you did not.
  5. Who said it. Check that symptoms came from the patient, not from your questions or a relative's worries.
  6. Assessment and plan. Make sure the wording matches your certainty ("possible", "likely", "confirmed") and that safety-net advice you gave is recorded.

An illustrative correction from that routine: the draft said "Plan: increase sertraline to 100 mg". The clinician had actually said "we might think about increasing it at the next review". Step six caught it. Signed as drafted, the note would have recorded a medication change that never happened, and the next clinician to read it might have acted on it.

When the room makes transcription harder

Scribes are tested in quiet rooms with two clear voices. Real consultations are often messier, and accuracy drops in predictable situations:

  • Several voices. A parent answering for a child, a partner chipping in, an interpreter repeating everything. Speaker separation struggles with three or more voices, so check the history for who said what with extra care.
  • Background noise. A crying baby, a busy corridor, a fan. Put the phone or microphone between you and the patient, not on the far side of the desk.
  • Quiet or muffled speech. Masks, a patient who speaks softly, a clinician turned towards the screen. Repeat key facts back ("so that's three weeks, yes?"), which helps the transcript and the patient.
  • Sensitive moments. A patient who wants to say something off the record. Pause the recording visibly and say that you have, then restart it.

In an illustrative interpreted consultation, the scribe recorded the interpreter's first-person translations as the clinician's words, so the history read "I have had headaches for a month" under the clinician's name. The fix was simple once spotted: the clinician now says "the interpreter is translating for the patient" at the start, and reads the history section first when reviewing any interpreted note.

What happens to the audio, transcript and note

A scribe creates three things, and each has its own lifespan. The audio recording, the transcript and the drafted note. Vendors handle them differently, and the defaults change:

  • Heidi says it does not keep the audio.
  • Nabla discards the audio and keeps transcripts for 14 days by default.
  • SimplePractice, which sells a note taker for therapists, switched its default in June 2026: new users now start with de-identified transcript retention turned on. It is a useful reminder that a setting you checked at sign-up can change later.

Before any patient is recorded, get written answers to these questions from the vendor: where the audio is processed and stored; when the audio is deleted; how long the transcript is kept and whether you can shorten that; whether any of it is used to train or improve models; which vendor staff can reach it; and whether the vendor will sign the data-processing agreement your data-protection rules require. Health information gets extra protection under data-protection law, so ask your data-protection adviser if any answer is unclear. The broader list is in the patient data and AI confidentiality checklist.

Ambient scribe, dictation and a human scribe compared

Ambient AI scribeAI dictationHuman scribe
What it capturesThe whole conversationOnly what the clinician says afterwardsWhat the scribe hears and understands
Clinician effort during consultAlmost none, beyond saying findings aloudNone during, a minute or two afterNone
Main riskMisattribution and invented findingsOmitting what the clinician forgets to sayCost and availability
Patient experienceClinician looks at the patient, not the screenSame during; clinician talks to the device afterA third person in the room
Best forHistory-heavy consultationsProcedures, short reviews, noisy roomsHigh-volume specialist clinics

Many clinicians end up using both AI modes: ambient for new problems, dictation for quick follow-ups. Many AI scribe products offer both.

What an AI scribe does not do

  • It does not diagnose. Some products add coding suggestions, letters or prompts about next steps. Anything that suggests a diagnosis or treatment is a different and more tightly regulated category of software in many places; ask the vendor what classification its features have where you practise.
  • It does not see. A rash, a limp or a wince exists in the note only if someone says it aloud.
  • It does not know the history unless it is linked to the record. A standalone scribe will not know the patient had the same problem last year.
  • It does not keep an evidential recording in most setups, because the audio is deleted. If you need a recording of what was said, that is a separate decision.
  • It does not replace consent. You still tell the patient, and you still stop recording if they ask.

How patients are told, and what they ask

A short, plain explanation at the start works better than a form. An illustrative script:

"Before we start, I use a computer tool that listens to our conversation and types up my notes, so I can look at you rather than the keyboard. I check and correct everything it writes. The recording is deleted once the note is done. Is that all right with you? If you'd rather I didn't use it, that's completely fine."

The most common patient questions are "who hears it?", "is it kept?" and "can you turn it off for this bit?". Have answers ready, and make pausing the recording easy mid-consultation. Detailed wording for posters, forms and scripts is in consent wording for AI note-taking, and the wider rollout is covered in rolling out an AI scribe without losing patient trust.

Speaking so the scribe gets it right

Small changes in how the clinician talks make a large difference to the draft. Compare two illustrative examination segments.

Before: "OK, let me just have a look... mm... and the other one... fine." The scribe has nothing to work with and may invent "examination unremarkable".

After: "Right knee: mild swelling, no redness, flexion to about 110 degrees, ligaments stable. Left knee normal for comparison." Every finding, side and number is in the transcript, and the patient hears what you found, which many like.

Three habits help most. Say the side and the number every time. Say findings as you find them, not in your head. And close with a spoken summary ("So, to sum up, the plan is...") because the plan section of the note will lean heavily on it.

Trying one fairly in your own clinic

A two-week trial tells you more than any demo. Pick one or two scribes, use them on a mix of consultation types, and score every note before editing it. An illustrative scoring sheet from one clinician's first week:

MeasureScribe AScribe B
Notes drafted4238
Notes needing no correction116
Notes with a factual error (dose, side, finding)49
Invented findings found13
Average review time per note1 min 50 s2 min 40 s
Patients who declined21

Scribe B looked faster in its demo, but the error count decides it. Scoring before editing is the important discipline, because once a note is corrected you forget how wrong it was. Whether the time saved justifies the cost for a small clinic is its own calculation, set out in whether an AI scribe is worth it for a small private clinic, and prices are compared in what small clinics pay for AI scribes. Heidi, for example, has a free tier alongside its paid plans, which makes a no-cost first trial possible; check its current plan limits before you start.

Whichever tool you choose, keep counting errors after the trial ends. A scribe that was accurate in week one can drift when the vendor updates its models, and a monthly sample of ten notes read against the transcript is the cheapest way to notice.

More questions about AI scribes

Does an AI scribe work for telephone and video consultations?

Most do. For video, the scribe usually runs on the same computer and captures both sides through the system audio or the microphone and speakers. For telephone, some tools offer a dial-in or speakerphone mode. Check how the tool captures the far side of the call, because a scribe that only hears the clinician produces a one-sided note.

Can the scribe handle a consultation in another language?

Many scribes transcribe several languages and some draft the note in the clinician's language from a conversation in another. Accuracy varies by language and accent, and interpreted consultations with three voices are harder. Test with your real patient mix before relying on it, and check medication names and numbers with extra care.

Is the audio recording kept as part of the medical record?

Usually not. Many scribes delete the audio soon after the note is drafted, and some keep the transcript for a short period. The signed note is the record. If you want recordings kept, that is a separate decision with its own consent and storage rules, so check the vendor's retention settings and your own record-keeping policy.

Further reads

Sources: Microsoft Dragon Copilot product pages; Heidi pricing and product pages; Nabla and Heidi data retention statements; published reviews of ambient AI scribes in clinical documentation research.

Thinking about an AI scribe for your clinic?

On a 1:1 call we can look at how notes are written in your clinic today, set up a fair two-week trial of one or two scribes, and work out the consent and data questions before any patient is recorded.

Book a 1:1 call with me