AI & Media5 min read•October 8, 2026

How to Transcribe Audio and Video Files in Your Browser with an On-Device Whisper Model

Learn how to pick a Whisper model size and language, transcribe a recording without uploading it, and export the result as TXT or SRT subtitles.

S
SmartToolPack Media Desk✓ Verified
Reviewed by the GreenCode AI & Media Editorial Desk
Last Updated: October 8, 2026
Disclaimer: Calculations and results provided by this tool are mathematical estimates for informational and educational purposes only. They do not constitute professional financial, tax, or legal advice.
Featured Utility Tool

Audio & Video Transcription

Open Audio & Video Transcription

Introduction

Meeting recordings, lectures and interviews are easier to search and quote once they exist as text. The Audio & Video Transcription tool decodes a media file with the Web Audio API, runs a Whisper speech recognition model inside a web worker and returns a transcript with timestamps. The audio never leaves the device; the only network activity is the one-time download of the model weights from Hugging Face, which the browser then caches.

Step-by-Step Instructions

  1. Drop a file onto Drop Audio or Video File Here or click Choose File; the card then shows the size, duration and whether it is audio or video.
  2. Pick an AI Model radio option: Whisper Tiny (about 77 MB), Base (about 148 MB) or Small (about 488 MB), trading speed for accuracy.
  3. Choose the Spoken Language from the dropdown, or leave it on Auto-Detect if you are unsure.
  4. Click Start transcription and keep the tab open while the model downloads and the Transcribing Audio progress runs.
  5. Review the Plain Text or With Timestamps tab, then use Copy transcript, Download .TXT or Download .SRT.

Key Use Cases

  • Subtitle Files: Generate an SRT file for a short video and load it into an editor or player.
  • Meeting Notes: Turn a recorded call into searchable text without sending the audio to a cloud service.
  • Interview Quotes: Pull exact wording from a voice memo with timestamps to locate the moment.

Frequently Asked Questions

Which file formats work?

The picker accepts any audio or video type, and the page lists MP3, WAV, MP4, WebM, OGG, FLAC, AAC, M4A and MKV. Decoding relies on the browser's own media decoder, so a container or codec the browser cannot play fails with an Audio decoding failed message. Converting the file to MP3 or WAV first usually resolves that.

Is there a length limit?

There is no hard cutoff, but the whole file is decoded into memory as 16 kHz mono audio and transcribed in the tab. Files of 30 minutes or more trigger a warning that the run takes roughly as long as the recording, and files of 90 minutes or more show a stronger warning because they often exhaust memory on phones and laptops. Splitting long recordings is recommended.

How much data does the model download use, and is it repeated?

The first run downloads the selected model from Hugging Face: about 77 MB for Tiny, 148 MB for Base and 488 MB for Small. The browser caches the weights, so later runs with the same model start without a download. Switching to a different model size triggers a download for that size.