Your episode, with a face. Upload it and walk away.
Podcasts are audio, and the platforms that grow them want video. Drop in the mixed episode as it is and tell Podcast, the video podcast studio, how many people are talking. It finds who speaks when, seats each of them on an animated stage, and the person speaking is the one who moves.
No software needed. Everything runs in your browser.
For podcasters who want a YouTube and Spotify video version of every episode without editing one.
What you get
- Input: One mixed MP3 or WAV, mono or stereo, up to 200 MB
- Episode length: 30 seconds to 3 hours; the video runs to the end of the episode
- Speakers: Up to 4; you give the head count, the studio finds who speaks when
- Seating: In order of talk time; brief voices, such as an intro voice-over, get no seat
- Output: 1920 x 1080, 16:9, 50 fps, H.264 MP4 with AAC audio
- Speaker looks: Halo, Equalizer, Bloom, Meters, Waveform, Ribbon, Frame, Voiceprint, Cadence and Aura
- Scenes: Animated scenes with palette, colour, speed, intensity and detail controls, or your own photo, clear or blurred
- Speaker photos: PNG or JPEG up to 20 MB; without one, the speaker's initials show
- On top: Text in 200 fonts, images and logos, subscribe and follow prompts, QR codes
- Captions: None; it follows voices, not words
- Watermark: None; the picture carries only what you put on it
- Rendering: Unlimited on the Creator plan, on our servers; download it or send it to YouTube
How to make a video podcast from an audio file
- Upload the episode. Upload the mixed episode as an MP3 or WAV, between 30 seconds and 3 hours long and up to 200 MB. Mono and stereo both work. Both limits apply, so for a long episode use an MP3: a WAV of a full episode is usually over the size limit.
- Say how many people are talking. The studio asks how many people are in the episode, up to 4, then listens to the whole recording and finds who speaks when. The speakers land on the stage when it finishes, the one who talks most in the first seat.
- Tell it who is who. A card for each voice shows how long they talk and plays a few seconds of their voice, so you can match voices to people by ear. Give each one a name, a role, a photo and a look, then choose the scene behind them.
- Add titles and render. Add an episode title, your logo or a subscribe prompt, and press play to check it against the real audio. Then render: the 1920 x 1080 MP4 at 50 fps lands in your library, ready to download or publish to YouTube.
How to turn an audio podcast into a video podcast
A video podcast does not have to be filmed. Most shows start as audio, and the simplest way to put one on YouTube or Spotify as video is to give each voice a face and let the picture follow the conversation. Here is what the platforms ask for, the options you have, and how Podcast, the video podcast studio, does it from one mixed file.
Why put an audio podcast on video
YouTube's upload page takes video files, not audio, so an MP3 on its own is refused. The only way in without making a video yourself is to connect your RSS feed and let YouTube build one from your show art. Spotify takes video episodes too, and a video episode delivered there stays audio-only everywhere else your feed goes, so listeners on other apps hear the show exactly as before.
Video does not mean a camera. Both platforms treat an episode with animated speaker artwork as a video episode, so nobody has to be filmed, lit or dressed for it. What matters is that the picture gives a viewer something to follow: who is talking, and when.
The ways to make the video
The quickest is a static image: your show art held on screen for the whole episode, with the audio underneath. That is what YouTube builds from an RSS feed. It gets the show uploaded, but nothing moves, so it gives a viewer nothing to watch and you nothing to design.
A step up is a waveform on a photo, the audiogram look. It moves, but the same way for every voice, so on a two-person show the picture cannot tell the host from the guest. Filming is the other extreme: real faces, and a camera, lighting and an edit for every episode.
In between sits a speaker-aware stage: a designed scene with a portrait for each person, where the one talking is the one who moves. To build one, something has to know who speaks when, which otherwise means separate tracks or timestamps typed in by hand. Podcast works it out from the finished mix.
Start with the mixed episode
Upload the episode as it went out: one MP3 or WAV, mono or stereo, up to 3 hours long and 200 MB. There are no separate tracks to prepare. Before it listens, the studio asks one question, how many people are in the episode, with a key for each head count up to 4. It asks because counting voices is the part an analysis cannot do reliably on its own when two people sound alike and talk over each other.
Then it listens to the whole recording and marks when each person starts and stops, including the moments two of them overlap, so both can move at once. Seats go in order of talk time, so the person who carries the show sits first. A voice heard only briefly, like a voice-over on the intro jingle or an advert read by someone who is not on the show, is counted as a brief voice and left off the stage.
The analysis is saved with the upload, so reopening the project reads it back instead of listening to the whole episode again.
Give every voice a face
When the analysis finishes, a card for each voice shows how long that person talks and plays a few seconds from their longest stretch of speech, so you can tell who is who by ear. Then each speaker gets a name, a role such as Host or Guest, a photo and a look. A speaker without a photo shows their initials.
The looks are built for speech rather than music: Halo, Equalizer, Bloom, Meters, Waveform, Ribbon, Frame, Voiceprint, Cadence and Aura. A portrait eases in when its speaker starts, follows the energy of their voice while they talk, and settles when they stop, so a pause looks like a pause.
Behind them goes a scene. Choose one of the animated scenes and set its palette, colours, motion speed, intensity and detail, or use your own photo, shown clear or blurred, with a darken control so names stay readable. On top, add the episode title in any of 200 fonts, your logo, a subscribe or follow prompt for YouTube or Instagram, or a QR code that links to the show.
Check it against the real episode
Press play and the episode itself plays through the stage. Each portrait moves from the same voice analysis the render reads, so a guest who cuts in late in the preview cuts in at the same moment in the finished file. Scrub to any point, move the speakers around and change their size while it plays. Every change is saved to the project automatically.
For the next episode, open the same project and pick the new audio. The scene and titles stay, and your speakers keep their names, photos and looks. The new voices are seated onto them in order of talk time, and the setup opens again so you can listen to each one and fix any name or photo that landed on the wrong voice.
Export for YouTube and Spotify
Every episode renders at 1920 x 1080, 16:9, at 50 frames a second, as an H.264 MP4 with AAC audio. That is YouTube's full HD size, the 16:9 shape Spotify recommends for video podcasts, and a frame rate on Spotify's list (24, 25, 30, 50 or 60 fps).
Landscape is also the right shape for a whole episode on YouTube, which files vertical and square videos of 3 minutes or less as Shorts. The render runs on our servers, so you can close the tab while it works, and the finished file arrives in your library with no watermark.
From the library, download it for Spotify, or send it to YouTube without a second upload: connect your channel once, fill in the title, description, tags and category, and the file goes up from our servers. It arrives as Private unless you choose otherwise, which leaves you time to watch it through and schedule it in YouTube Studio.
What makes it more than a waveform on a photo
It knows who is talking. Not only that someone is speaking, but which of your speakers it is. The analysis follows each voice through the whole episode and marks where people talk over each other, so the right portrait moves at the right moment, and both move when two people talk at once.
One mixed file, nothing to tag. No separate tracks, no script and no transcript. Tell the studio how many people are in the episode, up to 4, and it works out who speaks when from the voices alone.
Intros and adverts stay off the stage. A voice heard only briefly, like a voice-over on the intro jingle or an advert read by someone who is not on the show, is listed separately instead of taking a seat. The seats go to the people who carry the conversation, in order of how much they talk.
Preview with the real episode. Press play and your episode plays through the stage, with each speaker moving from the same voice analysis the render reads. Scrub to any minute and check a guest's first answer before you render.
A designed stage, in layers. Pick an animated scene and tune its palette, speed and intensity, or put your own photo behind the conversation, clear or blurred and darkened for readability. Titles in 200 fonts, your logo, a subscribe or follow prompt and a QR code sit on top, in a layers list you can reorder and hide.
Made for YouTube and Spotify. Every episode renders at 1920 x 1080, 16:9, at 50 fps: YouTube's full HD size, and a frame rate on Spotify's list for video podcasts. The render runs on our servers, so an episode of up to 3 hours never ties up your computer, and the file carries no watermark.
Works with
Music Visualizer: For a music show, the visualizer makes the video instead.
Compose: Titles, logos and props on the podcast stage work the way they do in Compose.
Frequently asked questions
How do I turn an audio podcast into a video?
Open Podcast and upload the mixed episode as an MP3 or WAV, then say how many people are in it. The studio works out who speaks when and gives each voice a seat. Add a name, a photo and a look for each speaker, pick a scene and a title, and render. The MP4 is waiting in your library when it finishes.
Can I upload an MP3 podcast to YouTube?
Not as it is. YouTube's uploader only takes video files, so a bare MP3 is refused, and the RSS route turns your show art into a still-image video. Podcast makes a real video from the same MP3, with your speakers moving, which you can download or publish to your channel from your library.
Do I need separate tracks for each speaker?
No. One mixed file is the whole input, mono or stereo. The voices are told apart from the mix, so separate tracks are not used and not asked for.
Why does it ask how many people are talking?
Because counting voices is the one thing the analysis cannot do reliably on its own when people sound alike and talk over each other. Told the head count, up to 4, it listens for that many voices instead of guessing how many there are. It asks before every analysis, so each episode gets its own answer.
What if an intro jingle or an advert is heard as a speaker?
A voice heard only briefly across the episode, such as a voice-over on the intro or an advert read by someone who is not on the show, counts as a brief voice and gets no seat. The setup tells you how many brief voices were skipped and how long they spoke in total, so you can see what was left off.
Do I need to be on camera for a video podcast?
No. YouTube and Spotify both count an episode with animated speaker artwork as a video episode. Each speaker here is a photo you choose, or their initials if you leave the photo out.
Does it add captions or a transcript?
No. It works out who is speaking and when, not what they said, so there are no captions on the video and no language to set.
Will the video meet Spotify's video podcast specs?
Spotify recommends 16:9 at full HD or higher, H.264 or H.265 in an MP4, at 24, 25, 30, 50 or 60 fps. Podcast renders 1920 x 1080, 16:9, at 50 fps as an H.264 MP4, which matches each of those, and an episode here runs up to 3 hours, inside Spotify's length guidance. Download the file from your library and upload it as a video episode.
Can I make vertical clips for Shorts, Reels or TikTok?
Not in this studio. Podcast renders the whole episode in 16:9, the shape of YouTube's main player and the one Spotify recommends for video episodes.
Can I reuse my setup for the next episode?
Yes. Open the project and pick the new episode's audio. The scene and titles stay, and your speakers keep their names, photos and looks; the new voices are seated onto them by talk time, and the setup opens again so you can listen and fix any name or photo that landed on the wrong voice.
Is Podcast included in the Creator plan?
Yes. There is one plan, $9.99 a month or $99.99 a year, and it covers Podcast alongside every other studio, with unlimited renders and no watermark.
Read more
- Podcast, documented
- Speakers and voice detection
- Podcast video specs for YouTube and Spotify
- Publishing to YouTube
- Subscribe and follow animations
- Rendering and downloading
- Turn your podcast into video