PodcastCleanup

Frequently asked questions

Is my audio uploaded to a server?

No. Transcription runs entirely in your browser using the Whisper model (via WebGPU, with a slower fallback on older browsers). There is no audio upload endpoint — privacy is architectural, not a policy promise. The only network call is optional: generating the publish pack sends transcript text (never audio) to a language model.

How accurate is the transcription?

The "Accurate" model (Whisper small) handles typical podcast audio — music intros, mild accents, two speakers — well. The "Fast" model is fine for clear single-speaker recordings. Heavy crosstalk, noisy rooms, and strong accents will produce more errors; for those, paste a transcript from your recording platform instead and use the cleanup layer.

How long does a 60-minute episode take?

A few minutes on a modern laptop with WebGPU (Chrome, Edge, Safari, Firefox current versions). On older machines the fallback path works but can take 10-20 minutes. The first run also downloads the model (40-250MB, one time — it is cached by your browser afterwards).

What file formats and sizes are supported?

MP3, M4A, WAV, AAC, OGG, FLAC, and MP4/WebM (audio track), up to 300MB. For long episodes, a compressed MP3 export decodes fastest and keeps memory use low.

Is this a free Descript alternative?

For transcription, transcript cleanup, and publish planning — yes, and it is genuinely free because there is no server cost to recover. For audio editing, multitrack, and overdub, no: Descript edits audio, this tool plans text. Many creators use both: transcribe and plan here, cut in your editor.

Can I export SRT captions?

Yes — when you transcribe audio here, timestamps are preserved and you can download an .srt file for captions. Pasted transcripts without timestamps export as clean text and Markdown only.

Will the cleanup change what my guest actually said?

It only removes filler words in the discourse-marker position (a standalone "like," between commas), immediate repeats ("the the"), and stray punctuation. It never rewrites or paraphrases. Toggle "show what was removed" to audit every single cut — deletions are struck through in red.

What is the Descript free plan limit?

As of July 2026, Descript’s free plan includes 60 transcription minutes per month, 100 one-time AI credits, and 720p export with a watermark — less than one typical episode of transcription. Paid tiers run $16-50/month billed annually. See our full Descript pricing breakdown for what each tier limits and when the free plan is actually enough.

Does it work for interviews, lectures, or YouTube videos?

Yes — anything spoken-word in English works: interviews, courses, webinars, YouTube audio. Multi-language support depends on the Whisper model and is not officially supported in this version.