Keyword research should happen before the recording, not after it. Search engines read the text layer around an episode, so the page, title, description, transcript and show notes, and not the audio. If the keyword work happens only afterwards, the target terms sit in the metadata but not in the transcript, which is the densest part of the content. Building the keyword architecture into the research brief means the host speaks those terms naturally on the day.
Podcast SEO guidance agrees on one thing almost without exception: the text layer is what search engines actually read. The episode page, the title, the description, the transcript, the show notes. Those are the parts that get crawled and indexed, and the audio waveform is not. Google's structured data documentation is explicit that markup helps it understand page content and qualify for richer search features, while guaranteeing nothing about ranking on its own. The audio needs a well-built text wrapper around it to be findable at all.
Given that, the usual workflow reads perfectly sensibly on paper. Record the episode, generate a transcript, then work the show notes, title and description over for the terms you want to rank for. What it does, though, is treat the transcript as raw material you can fit search intent onto afterwards, and the transcript is only ever a record of what the host knew and said when they sat down. Do the keyword research only after the recording and the terms you want end up living in your metadata, and not necessarily in the densest part of what you made, which is what the host actually said.
If you only do post-recording SEO, you often end up writing around a conversation that never targeted a real query in the first place.
What the research says about text versus audio
The current evidence supports transcripts and detailed episode pages being the main discoverability layer for podcasts in web search, because they are what exposes the words, entities and long-tail phrases inside an episode to a search engine. Audio indexed by a platform like Spotify runs through that platform's own internal search, which is a separate thing from open web search. For open web visibility, meaning the kind that puts an episode in front of somebody typing into Google, the text wrapper is what matters. The industry guides say the same thing consistently: one episode page per episode, a clear title, a detailed description, a full transcript, relevant show notes.
One honest caveat belongs here. Direct comparisons between pre-recording SEO and post-recording SEO in a podcast context do not appear to exist in the academic literature, at least not as of the research done for this article. The case for doing it beforehand is an inference from adjacent findings, and a strong one, but it is not a proven head-to-head result. What is well supported is that keyword density is not a direct ranking factor, and Semrush is explicit about that, and that topical completeness and natural language matter more than repeating a term mechanically.
I have not found controlled trials comparing episodes planned for SEO before recording against episodes optimised only after publication. What holds the pre-recording argument up is mechanism evidence from speech cognition research plus SEO fundamentals, and not a podcast-specific ranking comparison. I am stating what the evidence supports and nothing beyond it.
The speech and cognition angle
Cognitive research on language production consistently finds that a concept you have actually internalised is easier to say naturally than a term you are trying to slot in on purpose. Speaking also reinforces recall, in that producing the language strengthens the memory trace of the material. Put that against podcasting and it means a host who genuinely took a set of target terms on board during preparation is more likely to have them come out on their own across the recording. If those terms only exist in show notes bolted on afterwards, the transcript itself is thinner.
That difference is not a small one. The strongest SEO signal a transcript can carry is natural, frequent use of the target language right across the running length, rather than a cluster of terms sitting in a show notes section the host never actually said out loud. A 45-minute episode with a full transcript is a lot of indexable text. Prepare that episode with a clear search intent and the relevant terms, entities and variations turn up in it on their own. Do not, and the terms end up in a thin layer of copy wrapped around a conversation that went somewhere else.
How I do it in practice
Before an episode goes into production I work out the primary query and two or three secondary semantic targets it should serve. Those go into the brief alongside the research findings. What the host gets is a document holding the verified claims, the argument architecture, and inside that architecture the language I want turning up naturally in the recording. Not a list of terms to insert anywhere. Part of how the argument is framed.
After the recording the transcript gets cleaned up as editorial content rather than left as raw machine output, so readable speaker labels, section breaks, a summary someone would actually use. The episode page gets built with the primary target in the title, a descriptive lead in the show notes, timestamps and structured data markup. Both phases are necessary and they are doing different jobs. The work before the recording decides what gets said. The work after it builds the page Google reads.
What this means for the transcript quality
There is a simple test for this. Read the transcript of an episode that was prepared with a clear search target next to one that wasn't. In the first the topic and the terminology around it turn up naturally the whole way through, because the host knew what they were talking about before they opened their mouth. In the second the topic is probably in the title and the show notes, while the transcript wanders and hedges and circles back, because the argument was never set. Post-production can clean the audio up. It cannot tighten reasoning that was never there.
What the current evidence supports best is planning the topic and the search intent before recording, speaking naturally to that intent, then publishing a strong text wrapper afterwards. Pre-recording and post-recording SEO were never competing approaches. They handle different parts of the same problem, and both are worth doing properly. The mistake is expecting the pass you do afterwards to make up for the work you skipped beforehand.