Podcast: the video podcast studio
Upload an episode and get a video podcast. Automatic speaker detection finds up to 4 voices, seats each host on an animated stage, and the visuals move with whoever is talking.
Podcasts are audio, but the platforms that grow them are video. Podcast, the video podcast studio, turns one upload into an episode you can put on YouTube or deliver to Spotify, without opening an editor.
How does speaker detection work?
Upload the episode as it is, mixed, mono or stereo. The studio listens to the whole recording and works out how many people are talking, when each of them starts and stops and where they overlap, from the voices alone. There is no script to paste and nothing to label. It seats up to 4 speakers, which covers a host, a co-host and guests.
More detail, including what happens when an intro jingle or an advert is mistaken for a person, is on speakers and voice detection.
What does the finished episode look like?
- Each speaker has a photo, a name and a role
- Their portrait animates while they talk and settles when they stop
- An animated scene or your own photo sits behind them
- Titles, logos and images sit on top, in layers you control
What does it export?
1920 x 1080 at 50 frames a second, with your original episode audio untouched, for episodes up to 3 hours.
Related reading
- The Podcast product page
- Turn your podcast into video: AI that knows who is talking
- Video podcast specs for Spotify and YouTube
Last updated 2026-09-15
LimePulse is not affiliated with, sponsored or endorsed by Spotify AB or Apple Inc. Product names are used only to describe what the software makes files for.