How VoxScriber Transcription Works
Understand the complete journey of your audio in VoxScriber: validation, queue, AI engines, speaker identification, and when cycles are billed.
When you send an audio or video file, it goes through an automatic pipeline: file validation, processing queue, transcription by an AI engine, and post-processing (title, speakers, timestamps). All you need to do is upload the file — and you only pay in cycles when the transcription is successfully completed.
The steps, in order
File validation
Every file is inspected on the server to confirm that there is actual audio inside it and to extract the exact duration. Audio from video formats is extracted automatically, and less common formats are converted to MP3 before moving on — you don't need to convert anything before uploading.
Processing queue
Your file enters a queue with plan-based priority: free accounts have low priority, Lite and Advanced have normal priority, and Pro/Premium/Enterprise have high priority. Even during peak times, normal-priority jobs are automatically promoted after a waiting period, so no one gets stuck in the queue.
Transcription by the AI engine
The default engine for all plans is Premium, highly accurate and optimized for Brazilian Portuguese. Depending on your plan, you can also choose Standard (great for noisy audio and the cheapest in cycles), Ultra (maximum quality with precise speaker separation), or Nano (runs in your browser and uses no cycles). If an engine fails due to instability, the system automatically tries an alternative engine.
Automatic post-processing
As soon as the text is ready, the AI generates a title for the transcription in the language of your interface and, when more than one person is speaking, tries to identify the name and role of each speaker — without using any quota. Timestamps are saved for exporting captions and using the synchronized player.
When cycles are charged
Cycles are only debited after the transcription is successfully completed, in the same operation that saves the result — billed and delivered. If the transcription fails, nothing is debited, and the automatic retry system kicks in on its own.
The cost depends on the engine: Standard uses 1 cycle per minute, Premium 4 cycles per minute, Ultra 10 cycles per minute, and Nano is free (0 cycles). Fractions are rounded up.
Short audios (up to 5 minutes) usually finish in less than half a minute on the Premium engine. Files of 1 hour or more may take a few minutes. You can close the tab: processing continues on the server.
What if something goes wrong?
Temporary failures (service instability, network, request limits) trigger automatic retries — first immediate, then scheduled at increasing intervals, always using the most reliable engine. Your file is kept for 7 days while there are pending retries.
File errors (invalid format, corrupted file) are not retryable: in these cases, check the original file and upload it again. See the full list in Common errors and solutions.
Frequently asked questions
Do I need to keep the page open during processing?
No. Processing happens on the server. You can close the tab and come back later — the transcription will be in your library. The exception is the Nano engine, which runs in your browser and requires the tab to be open.
Am I charged if the transcription fails?
No. Cycles are only debited when the transcription is completed and delivered. Failures do not generate charges.
Which engine is used if I don't choose one?
Premium, the default engine for all plans, including the free one.
How long is the audio file kept?
The audio is available for 30 days after transcription (for example, for the synchronized player). The transcription text does not expire by default.
Related articles
Still stuck? Open a ticket and our team will help you.