yapyapNotes
Guide/

Why your transcript puts the wrong name on the wrong line

You open a transcript of a four-person meeting and it lists two speakers. Or you open a one-hour call and it lists eighty-seven. Neither number is close, and both are common.

This is worth understanding before you send anyone a write-up, because the words are usually right and the names on them are the part that fails.

The words and the names are two different jobs

Turning speech into text is a solved problem. Audio goes in, text comes out, and there is one correct answer to score against. Any tool you download does this well.

Working out who spoke is a different job with no correct answer to check against. The software has never heard these people before. Nobody tells it how many of them are in the room. It listens for voices that sound alike, groups them, and the group count falls out of a threshold someone had to pick. That threshold is one number, and it has to cover your quiet colleague, a noisy café, and a parliament debate.

The same fifteen seconds on two tracks. The sentence boundaries and the speaker boundaries do not land on the same second.
The two jobs cut the same audio in different places.

Five minutes that fix most of it

Do these before the call, not after.

  1. Put the microphone closer to the table than to one person. A voice picked up across a room clusters badly.
  2. Record the call rather than the room when you can. Your microphone and the far end arrive as separate channels, and no software can confuse the two.
  3. Ask people not to talk over each other for the first minute. Early clean speech is what the labels are built from.
  4. Name the speakers once in yapyap after the first recording. Names carry forward to later recordings of the same people.
  5. Skim the label list before you export. A speaker holding two percent of a meeting is usually a mistake, not a person.

What we got wrong, in case you hit it

For months, almost every recording in yapyap came back with exactly two speakers. We were measuring how many people talked at the same time, which is nearly always two, and calling that the speaker count. It has nothing to do with how many people are in the room.

The fix over-corrected. A 55-minute recording produced 87 speaker labels, most of them under one percent of the audio. Version 0.8.3 tightened the grouping and added a rule that a speaker must hold at least 2% of the talking time. The same recording went from 102 labels to 3.

The honest limit

On broadcast recordings we measured one person's own voice varying more than two different people's voices differ from each other. One speaker's segments sat 0.8 apart. Two parliament speakers sat 0.52 apart. No setting separates those two cases. Ours is a tuned guess that fails gently in both directions.

So treat a speaker label as a strong suggestion. Read it before you quote it, and fix it in the transcript when it is wrong. That correction sticks.

The one place the guessing stops is a recorded call. Your side and the far side arrive on separate channels, which is a physical fact rather than a threshold. yapyap keeps them apart and never merges across them. Most tools mix the two into one track first and throw that certainty away.

yapyap records, transcribes, and summarizes on the hardware you already own. No account, no subscription.

Download yapyap

macOS · Windows · Linux

Free to try. No card, no account.
Then €69, once. Yours forever, every update included.
Your recordings never leave your computer.