English voice note
Choose a 6-minute MP3, select English, confirm permission, and transcribeEditable timestamped text plus TXT, SRT, and WebVTT downloads
Transcribe one local audio or video file in your browser with pinned Whisper Tiny models. Review timestamped text, then download TXT, SRT, or WebVTT without uploading the recording to Access Free Tools.
Private browser transcription
The recording stays on this device. Only the pinned speech model downloads after you start.
No media upload. Decoding, speech recognition, editing, and exports happen in this tab. The model files come from Hugging Face.
Step 1
Local job

Create a draft transcript from your own interview, lesson, meeting, voice note, or video recording.
Build editable SRT or WebVTT caption files for permitted media.
Recover completed transcript sections when a long browser job is cancelled or a later block fails.
Check several audio tracks in a video before choosing the dialogue track.
Transcribe English with the smaller English model or choose a named language for multilingual speech.
Keep private media on the device instead of sending it to a speech API or upload server.
Editable timestamped text plus TXT, SRT, and WebVTT downloads
Caption segments appear as each five-minute block finishes
The multilingual Whisper Tiny model produces a timestamped draft for manual review
The completed partial transcript remains editable and downloadable
Need a slower walkthrough, a related tool, or the full library? These links keep you close to the task you started.
Plain-language answers about supported media, browser decoding, model downloads, partial transcripts, caption exports, privacy, permission, and accuracy limits.
No. The selected file, decoded audio, transcript, filename, and language choice stay in this browser tab. They are not uploaded to Access Free Tools. The pinned model files download from Hugging Face only after you start transcription.
The tool can inspect MP3, WAV, M4A or AAC, FLAC, OGG or Opus, MP4, MOV, WebM, and MKV containers. The audio codec inside the file must also be decodable by your current browser. An unsupported codec is rejected before the model download.
Choose the local recording, then select its dialogue audio track when several tracks are present. English uses the smaller English model; Auto or a named non-English language uses the multilingual model. WebAssembly is the compatibility choice, while WebGPU is an optional beta path on supported devices.
The beta accepts one file up to 250 MB and 60 minutes. Desktop is recommended above 15 minutes because decoding and speech recognition use device memory and processor time.
English uses the pinned onnx-community Whisper Tiny English timestamped model. Auto and named non-English choices use the pinned multilingual timestamped model. Both use quantized browser files and run through Transformers.js.
Treat the timestamped text as an editable first draft, not a certified transcript. Click a timestamp to compare each important section with the local recording before copying or exporting it.
Check names, numbers, dates, accents, technical terms, overlapping voices, quiet speech, music, and noisy sections. Whisper can mishear speech or produce plausible words that were not spoken.
Yes. Each completed five-minute section appears immediately. If you stop the job or a later section fails, completed caption segments stay available for editing, copying, and TXT, SRT, or WebVTT download.
Both store timed captions. SRT is widely accepted by editors and video platforms. WebVTT is designed for web video and starts with a WEBVTT header. TXT contains the words without caption timing syntax.
The q8 WebAssembly path is the dependable compatibility choice. WebGPU can be faster on a supported device but remains experimental for this workload, so the tool detects an adapter, labels the option as beta, and can fall back to WebAssembly.
No. Version one does not provide speaker identification or diarization. It also does not translate, record a microphone, import URLs, process batches, or render captions back into a video.
The browser revokes the local media URL and terminates its workers. The transcript is not stored by Access Free Tools, so download the files you want before resetting, refreshing, or closing the tab.