Two tiers. The rankings and coverage tables are open to anyone. The transcript archive is third-party copyrighted text, so it sits behind an access key.
Show rankings, feed URLs, episode counts, and per-show transcript coverage. No transcript text, so nothing here is restricted.
| finance-podcasts-ranked.csv | 280 KiB | open |
| metadata-csv.tar.gz | 8.1 MiB | open |
Two archives, plus the 43 MB episode-level index (too large for this site's 25 MiB per-file limit, so it lives with them). Enter the key to get the file list and links.
| corpus-transcripts.tar.gz | 1.9 GB | 153,261 transcripts as clean prose, timestamps stripped |
| corpus-vtt.tar.gz | 11.3 GB | 207,637 raw WebVTT captions, timestamps intact |
| finance-podcasts-transcripts-index.csv | 43 MiB | one row per episode: show, title, date, source, word count |
Take the first if you want the words. Take the second only if you need to know when each line was spoken — audio alignment, clip extraction. They cover the same material.
The archive contains transcripts from roughly 580 publishers plus YouTube captions. It is shared for research use; the underlying text remains its owners' copyright.
The upload tool caps a single object at 300 MiB, so each archive ships as ~280 MB parts — 8 for the text, 42 for the captions. Downloads support resume, so a dropped part restarts where it stopped.
cat corpus-vtt.tar.gz.part-* > corpus-vtt.tar.gz shasum -a 256 corpus-vtt.tar.gz # compare with MANIFEST.txt tar xzf corpus-vtt.tar.gz
Check each part's size against MANIFEST.txt first — a truncated part still concatenates cleanly and only fails at extraction.