VoxScriber
100% free · No signup · 99 languages

Free Arabic audio transcription in your browser

VoxScriber Nano understands Arabic (العربية) and processes locally on your device. 10-minute limit in free mode.

Runs VoxScriber Nano (open-source) in your browser — local AI, up to 10 min per file, basic accuracy (~85%). For professional use, try Premium.

🔒 Local AI💰 100% free📝 10 min per file

Transcription runs locally in your browser. You can optionally share the result with us (optional, with consent) to help improve the service. Limit: 10 min per file, ~85% accuracy.

See Premium →

Free vs Premium — see the difference

Free (browser) Premium (cloud)
File limit 10 min 10 horas
Accuracy ~85% >95%
Speaker diarization
Word-level timestamps
Video support (MP4/MOV)
Export formats TXT, SRT, VTT DOCX, PDF, JSON…
Speed (1h of audio) ~2 min / 1h ~2 min / 1h
Privacy 100% local ☁️ + 🔒
🔒

Local AI

Transcription runs in your browser. Sharing with our servers is optional (requires consent).

Fast and local

AI processing runs directly in your browser — no waiting in queues.

🌍

99 languages

Automatically detects the language of your audio.

💻

No signup needed

Start transcribing instantly, no account required.

How it works

1

Upload or record audio

Drag an MP3, WAV, M4A, or OGG file, or use your microphone directly.

2

AI runs on your device

Whisper AI downloads once and stays cached. No wait on your next visit.

3

Copy or download the text

Results appear in seconds. Download as .txt or copy with one click.

How well does Whisper handle Arabic?

Whisper transcribes Modern Standard Arabic most reliably, written right-to-left without diacritics (tashkeel) — matching how Arabic is normally written. Dialect support varies: Egyptian and Levantine perform reasonably, while Maghrebi (Darija) is the weakest. Code-switching with French or English inside a sentence works but reduces accuracy.

Where Arabic audio usually comes from

WhatsApp voice notes — the dominant voice medium across the Arab world — plus Quranic recitation and sermons, university lectures, and TV or news clips.

Related languages you can also transcribe for free: Hebrew · Turkish · French · All 20 free transcription languages

How accurate is browser transcription?

Browser transcription runs OpenAI's Whisper model directly on your device using WebAssembly. We offer three model sizes, and accuracy depends on which one you pick:

  • Nano (~40MB) — The default. Around 85% accuracy on clear speech. Best for quick notes, voice messages and drafts. The only model that runs on iOS.
  • Mini (~150MB) — Roughly 90% accuracy. A good middle ground if your device has 4GB+ of RAM and you need cleaner output.
  • Plus (~500MB) — The most accurate local option, approaching 93% on clear audio. Slower to download and run; best on desktop machines with 8GB+ of RAM.

What lowers accuracy for any local model: background noise, multiple people talking over each other, heavy accents, and low-bitrate recordings such as compressed voice notes. If you need professional accuracy above 95%, word-level timestamps or speaker labels, that requires cloud models — see the comparison above.

Browser vs cloud transcription: which one do you need?

Browser transcription is the right tool when privacy matters most or the audio is short: nothing is uploaded, there is nothing to delete afterwards, and it costs nothing. The trade-off is speed and precision — your CPU processes roughly one hour of audio in twenty minutes, and the local model skips speaker labels and word-level timing.

Cloud transcription is the right tool when you are working: meetings, interviews, lectures, legal recordings. Dedicated GPUs turn an hour of audio into text in about two minutes with over 95% accuracy, label up to 30 different speakers, accept files up to 10 hours long, and export to DOCX, PDF and JSON on top of the subtitle formats.

A practical rule of thumb: if you would be comfortable reading the recording aloud in a cafe, the cloud's speed and accuracy win. If the audio is sensitive — a medical consultation, a confidential meeting, a private voice note — the browser tool keeps everything on your machine and still gives you a usable transcript in minutes. Many of our users combine both: quick private notes in the browser, professional work in the cloud.

See Premium plans →

Supported audio formats

Upload MP3, WAV, M4A, OGG, OPUS, FLAC or WEBM — anything your browser can decode. Common sources work out of the box: WhatsApp voice notes (OPUS), iPhone voice memos (M4A), Android recorder files, Zoom recordings (M4A/MP4), Telegram voice messages (OGG) and podcast files (MP3). Video containers like MP4 and MOV are decoded for their audio track when the browser supports the codec. If a file fails to load, the usual cause is an unusual codec inside a common container — converting it to MP3 first solves it in almost every case.

Need a different format first? Use our free converters: free MP3 / WAV / OGG / AAC audio converter

🚀 Premium

Need more? Try Premium

For professional use — speaker diarization, long files, AI analysis and full export formats.

🎭

Speaker diarization

Automatically identifies who is speaking in each segment. Perfect for meetings, interviews and podcasts.

⏱️

Files up to 10 hours

The local model supports up to 10 min. Premium handles files up to 10 hours long.

🧠

Summary, sentiment & topics

AI analyzes the content and generates executive summaries, sentiment analysis and topic extraction.

📄

Full export options

Export to SRT, VTT, DOCX, JSON and PDF — ideal for subtitles, documents and automation.

See Premium plans →

Frequently asked questions

Is my audio uploaded to a server?

Transcription runs 100% locally in your browser. After completing, you can optionally share the audio and transcript with us (consent checkbox) to help improve the service — but it's entirely optional and never happens without your consent.

Which audio formats are supported?

MP3, WAV, M4A, OGG, FLAC, WEBM, MP4, MOV and any format your browser can decode. Video containers are auto-converted.

How long can my audio file be?

Up to 10 minutes per file in free mode. For longer audio, our Premium plan supports up to 10 hours.

Which languages are supported?

The Whisper model supports 99 languages including English, Portuguese, Spanish, French, German, Japanese, Arabic and many more. Detection is automatic.

Do I need to install anything?

No. It works directly in your browser. The AI model (~40MB) downloads once and stays cached for future visits.

Is the transcription saved anywhere?

By default, no — results stay only in your browser. If you check the consent box, the audio and transcript are sent to our servers and deleted after 7 days. You can decline consent at any time.

What is the difference vs. Premium?

The free mode uses the VoxScriber Nano model (4-bit quantized, q4) running locally: 10-min limit per file, ~85% accuracy, no speaker diarization, and timestamps only at segment level (~30s chunks — not per word). Premium uses cloud models (AssemblyAI + Whisper Large float32): >95% accuracy, diarization up to 30 speakers, word-level timestamps, files up to 10h, MP4/MOV/MKV video support, and exports to DOCX, PDF, and JSON. Speed: 1h of audio takes ~20min on your local CPU vs ~2min on Premium's dedicated GPU.

Does it work on mobile?

Yes, but performance depends on your device. On low-RAM smartphones, transcription may be slower.

Is it really free?

Yes. The browser transcriber is genuinely free with no trial period, no watermark and no signup. We make money from the Premium cloud plans, not from the free tool.

Does my audio leave my device?

No — transcription runs locally via WebAssembly. The only exception is if you explicitly tick the optional consent checkbox to share a recording with us.

Is there a file size limit?

The practical limit is duration (10 minutes per file) and your device's memory, not megabytes. A 10-minute MP3 is typically 10-20MB and works fine on most devices.

How long does transcription take?

With the Nano model, expect roughly 1-2x the audio duration on a modern laptop — a 5-minute file takes about 5-10 minutes. The first run adds a one-time model download of ~40MB.

Can I export subtitles (SRT)?

Yes — free exports include .txt, .srt and .vtt with segment timestamps. For word-level timestamp precision and DOCX/PDF/JSON exports, see Premium.

Can I transcribe several files at once?

Yes — you can queue up to 5 files and they are processed one after another in your browser. Premium removes the queue limit and processes files in parallel in the cloud.

Why does the first transcription take longer?

On your first visit the AI model is downloaded and compiled by your browser. It is then cached, so every later transcription starts immediately.

Does it work offline?

Partially — once the model is cached, the transcription itself needs no connection. You still need to be online to load the page itself.

Does it transcribe dialects or only Modern Standard Arabic?

Both, but MSA is most accurate; Egyptian and Levantine work reasonably, Maghrebi less so.

Does it add tashkeel (diacritics)?

No — output is standard undiacritized Arabic, as it is normally written.

Is right-to-left display supported?

Yes — the transcript renders right-to-left and exports preserve it.

Free transcription in 20 languages

Whisper supports 99 languages with automatic detection, and we maintain a dedicated page for each of the 20 most-requested languages, with notes on how the model handles that specific language. Pick yours below — the transcriber pre-selects the right language for better accuracy.