Your episode, with a face. Upload it and walk away.
Podcasts are audio, and the platforms that grow them want video. Drop in the mixed episode as it is: Podcast, the video podcast studio, works out how many people are talking and when, seats each of them on an animated stage, and the person speaking is the one who moves.
For podcasters who want a YouTube and Spotify video version of every episode without editing one.
A stage, a background, the speakers, and titles on top. The one talking animates; the others settle.
What you get
- Input: One mixed audio file, mono or stereo
- Speakers: Up to 4, detected from the audio alone
- Length: Up to 3 hours
- Output: 1920 x 1080 MP4 at 50 fps, original audio passed through
- Captions: None. It works from the voices, not the words
- Watermark: None on any export
From audio file to video episode
- Upload the episode. The mixed file as it is. Up to 3 hours, mono or stereo.
- Confirm who is who. The analysis comes back with the voices it found. Give each seat a name, a role and a photo.
- Dress it and render. Pick the stage, the background and the titles, then render 1920 x 1080 at 50 fps with your original audio.
What makes it different from a waveform on a photo
It knows who is talking. Not just that somebody is. The detection separates the voices and tracks them through the whole episode, including where people overlap.
Nothing to label. No script, no transcript, no separate tracks, no tagging. One mixed file is the entire input.
A real stage, not a slideshow. Speakers sit in a designed scene with a background, titles and logos, in layers you control.
Your audio, untouched. The episode audio is passed through as it is. Nothing is re-encoded on the way to the video.
Sized for where podcasts grow. 1920 x 1080 at 50 fps, the shape YouTube wants and the shape a video podcast is delivered in.
Long episodes are normal. Up to 3 hours. A two-hour conversation is the ordinary case, not the edge case.
What an episode looks like
One host on an animated scene
Host and guest, the speaker who talks lights up
Three voices, seated by how much they talk
Up to four seats on one stage
Works with
Music Visualizer: For a music show, the visualizer makes the video instead.
Compose: Titles, logos and props on the podcast stage work the way they do in Compose.
Frequently asked questions
Do I need separate tracks for each speaker?
No. One mixed file is the input. Separate tracks are not used and not asked for.
How many speakers can it handle?
Up to 4, seated by how much they actually talk, so the host and co-host get the seats rather than a voice from an advert.
What if it hears a jingle or an advert as a speaker?
Voices that only appear briefly are treated differently from the people carrying the episode, and you can see and correct what it found before rendering.
Does it transcribe the episode?
No. It works from the voices rather than the words, which is why it does not need to understand the language. There are no captions on the video.
Can I use my own photos for the speakers?
Yes. Each seat takes a photo, a name and a role.
Read more
- Turn your podcast into video with speaker detection
- Podcast, documented
- Video podcast specs for Spotify and YouTube