Speakers and voice detection
How LimePulse finds who is talking in a podcast episode, seats up to 4 speakers, and what to do when an intro jingle is mistaken for a person.
Upload an episode and the studio works out who is talking and when, from the audio alone. There is nothing to label and no script to paste.
How many speakers can it handle?
Up to 4 voices per episode, which covers a host, a co-host and guests. If more people are in the room than that, the quietest are folded into the ones they most resemble rather than dropped, so the visuals keep moving.
It found more voices than there are people
Usually an intro. A jingle with a voiceover, an advert read by someone who is not on the show, or a clip played mid-episode all contain real voices and the analysis is right to notice them. What matters is that they do not take a seat on the stage, so a voice that speaks for only a few seconds across a whole episode is listed separately rather than seated.
How accurate is it?
Precise to a fraction of a second on clean, well-recorded conversation. It is weaker on heavy crosstalk, on a very noisy room, and on two people whose voices genuinely sound alike. It is good enough that a whole episode usually needs no correction at all.
Setting up each speaker
Click a speaker, add a photo, type a name and a role, then choose how they come alive: a breathing halo, studio meters, a radial equalizer, a waveform, a clean frame and more. The motion is built for speech rather than music, so a profile eases in when someone starts, follows the energy of their voice while they talk, and settles when they stop. Two people talking over each other both light up.
Does it transcribe the episode?
No. It works out who is speaking and when, not what they said. That is a different job and it is not what the visuals need.
Last updated 2026-09-15
LimePulse is not affiliated with, sponsored or endorsed by Spotify AB or Apple Inc. Product names are used only to describe what the software makes files for.