Download the data

Two tiers. The rankings and coverage tables are open to anyone. The transcript archive is third-party copyrighted text, so it sits behind an access key.

Open — metadata

Show rankings, feed URLs, episode counts, and per-show transcript coverage. No transcript text, so nothing here is restricted.

finance-podcasts-ranked.csv280 KiBopen
metadata-csv.tar.gz8.1 MiBopen

Access key required — transcript archives

Two archives, plus the 43 MB episode-level index (too large for this site's 25 MiB per-file limit, so it lives with them). Enter the key to get the file list and links.

corpus-transcripts.tar.gz1.9 GB 153,261 transcripts as clean prose, timestamps stripped
corpus-vtt.tar.gz11.3 GB 207,637 raw WebVTT captions, timestamps intact
finance-podcasts-transcripts-index.csv43 MiB one row per episode: show, title, date, source, word count

Take the first if you want the words. Take the second only if you need to know when each line was spoken — audio alignment, clip extraction. They cover the same material.

The archive contains transcripts from roughly 580 publishers plus YouTube captions. It is shared for research use; the underlying text remains its owners' copyright.

Reassembling an archive

The upload tool caps a single object at 300 MiB, so each archive ships as ~280 MB parts — 8 for the text, 42 for the captions. Downloads support resume, so a dropped part restarts where it stopped.

cat corpus-vtt.tar.gz.part-* > corpus-vtt.tar.gz
shasum -a 256 corpus-vtt.tar.gz   # compare with MANIFEST.txt
tar xzf corpus-vtt.tar.gz

Check each part's size against MANIFEST.txt first — a truncated part still concatenates cleanly and only fails at extraction.