Audio and Video Transcriber

Transcribe one local audio or video file in your browser with pinned Whisper Tiny models. Review timestamped text, then download TXT, SRT, or WebVTT without uploading the recording to Access Free Tools.

Use the Audio and Video Transcriber

Private browser transcription

Turn local audio or video into editable captions

The recording stays on this device. Only the pinned speech model downloads after you start.

Up to 60 minutes Up to 250 MB

No media upload. Decoding, speech recognition, editing, and exports happen in this tab. The model files come from Hugging Face.

Step 1

Choose a recording

Drop one audio or video file hereMP3, WAV, M4A, FLAC, OGG, MP4, MOV, WebM, or MKV

Step 2

Choose speech settings

English uses the smaller English-only model. Auto and every other choice use the multilingual model. Auto is a model guess, so choose the language when you know it.

Processing mode

No WebGPU adapter was detected. The tool will use its q8 WebAssembly compatibility path.

Local job

Choose a recording to inspect

idle
0% complete0:00 elapsed0 caption segments
Full-body smoke-kawaii mascot turning a local audio and video timeline into editable caption rows beside a privacy shield.
Audio and Video Transcriber artwork shows a local recording becoming timestamped text, with browser processing and caption downloads kept visible.View in the smoke-kawaii gallery
No media uploadEnglish + multilingualEditable timestampsTXT, SRT, and WebVTT

How to use the Audio and Video Transcriber

  1. Choose one permitted audio or video file up to 250 MB and 60 minutes. The browser checks its container and audio tracks before downloading a model.
  2. Choose the dialogue audio track when more than one is available, then select English, Auto, or a named language.
  3. Use WebAssembly for the broadest compatibility, confirm permission, and press Transcribe recording. Desktop is recommended above 15 minutes.
  4. Review each completed section, click timestamps to listen again, and correct names, numbers, dates, specialist words, or unclear speech.
  5. Copy the checked transcript or download TXT, SRT, or WebVTT before resetting or closing the tab.

What people use it for

Create a draft transcript from your own interview, lesson, meeting, voice note, or video recording.

Build editable SRT or WebVTT caption files for permitted media.

Recover completed transcript sections when a long browser job is cancelled or a later block fails.

Check several audio tracks in a video before choosing the dialogue track.

Transcribe English with the smaller English model or choose a named language for multilingual speech.

Keep private media on the device instead of sending it to a speech API or upload server.

Quick examples

English voice note

Choose a 6-minute MP3, select English, confirm permission, and transcribe

Editable timestamped text plus TXT, SRT, and WebVTT downloads

Caption a local video

Choose a 20-minute MP4, select its dialogue track, and use WebAssembly mode

Caption segments appear as each five-minute block finishes

Spanish recording

Choose Spanish before starting a permitted M4A recording

The multilingual Whisper Tiny model produces a timestamped draft for manual review

Interrupted long recording

Stop after two completed sections of a longer recording

The completed partial transcript remains editable and downloadable

Need the guide or a nearby tool?

Need a slower walkthrough, a related tool, or the full library? These links keep you close to the task you started.

Frequently asked questions

Plain-language answers about supported media, browser decoding, model downloads, partial transcripts, caption exports, privacy, permission, and accuracy limits.

Does the transcriber upload my audio or video?

No. The selected file, decoded audio, transcript, filename, and language choice stay in this browser tab. They are not uploaded to Access Free Tools. The pinned model files download from Hugging Face only after you start transcription.

Which audio and video files can I transcribe?

The tool can inspect MP3, WAV, M4A or AAC, FLAC, OGG or Opus, MP4, MOV, WebM, and MKV containers. The audio codec inside the file must also be decodable by your current browser. An unsupported codec is rejected before the model download.

What do the main Audio and Video Transcriber inputs mean?

Choose the local recording, then select its dialogue audio track when several tracks are present. English uses the smaller English model; Auto or a named non-English language uses the multilingual model. WebAssembly is the compatibility choice, while WebGPU is an optional beta path on supported devices.

What are the file and recording limits?

The beta accepts one file up to 250 MB and 60 minutes. Desktop is recommended above 15 minutes because decoding and speech recognition use device memory and processor time.

Which model does the browser transcriber use?

English uses the pinned onnx-community Whisper Tiny English timestamped model. Auto and named non-English choices use the pinned multilingual timestamped model. Both use quantized browser files and run through Transformers.js.

How should I read the Audio and Video Transcriber result?

Treat the timestamped text as an editable first draft, not a certified transcript. Click a timestamp to compare each important section with the local recording before copying or exporting it.

What should I double-check before trusting the Audio and Video Transcriber transcript?

Check names, numbers, dates, accents, technical terms, overlapping voices, quiet speech, music, and noisy sections. Whisper can mishear speech or produce plausible words that were not spoken.

Can I edit and download a partial transcript?

Yes. Each completed five-minute section appears immediately. If you stop the job or a later section fails, completed caption segments stay available for editing, copying, and TXT, SRT, or WebVTT download.

What is the difference between SRT and WebVTT?

Both store timed captions. SRT is widely accepted by editors and video platforms. WebVTT is designed for web video and starts with a WEBVTT header. TXT contains the words without caption timing syntax.

Why does the tool offer WebAssembly and WebGPU?

The q8 WebAssembly path is the dependable compatibility choice. WebGPU can be faster on a supported device but remains experimental for this workload, so the tool detects an adapter, labels the option as beta, and can fall back to WebAssembly.

Can this tool identify different speakers?

No. Version one does not provide speaker identification or diarization. It also does not translate, record a microphone, import URLs, process batches, or render captions back into a video.

What happens when I reset or close the page?

The browser revokes the local media URL and terminates its workers. The transcript is not stored by Access Free Tools, so download the files you want before resetting, refreshing, or closing the tab.