Frequently asked questions

If the answer you need is not here, write to us and we will add it.

Getting started

What is Scrivio?

A web application that turns recorded audio and video into written transcripts. It separates the speakers automatically and gives you a full editor to correct, annotate, tag, share and export the result. It runs entirely in Switzerland.

Who is it for?

Anyone with recordings worth reading: researchers, students and lecturers, journalists, podcasters, teams that record their meetings, and people transcribing family or archive material. There is no single intended profession.

How much does it cost to start?

Nothing. Scrivio is in public beta and free, with 60 audio-minutes a month. You need an email address; you do not need a credit card, now or later.

What files can I upload?

Most audio and video formats — mp3, wav, m4a, flac, ogg, mp4, mov, mkv and others. Video files have their audio extracted automatically, so you do not need to convert anything yourself. The free plan accepts files up to 500 MB.

How long does a transcription take?

Usually a fraction of the length of the recording, but it depends on the file and on how busy the queue is. The app shows a live estimate and updates by itself; you can close the tab and we will email you when it is finished.

Quality and editing

How accurate is it?

Very good on clear recordings and noticeably worse on bad ones — background noise, heavy crosstalk, distance from the microphone and strong accents all cost accuracy. Treat the output as a strong first draft that needs a read-through, not as a finished document. That is exactly why the editor exists.

Which languages are supported?

A wide range of spoken languages, selected at upload so the model transcribes natively rather than translating.

Arabic, Basque, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Filipino, Finnish, French, Galician, Georgian, German, Greek, Hebrew, Hindi, Hungarian, Italian, Japanese, Korean, Latvian, Malayalam, Norwegian, Norwegian Nynorsk, Persian, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish, Telugu, Turkish, Ukrainian, Urdu, Vietnamese

How well does speaker detection work?

Well on recordings where people take turns and are reasonably separated. It struggles with heavy crosstalk and with voices that sound alike. You can always correct the assignment yourself, and renaming a speaker applies everywhere at once.

Can I fix mistakes?

Yes, and that is the point. Edit any word in place, split a segment, merge two, rename speakers, highlight passages and leave comments. Changes save automatically as you type — there is no save button to forget.

What can I export?

Plain text and SRT subtitles, with subtitle line lengths already tuned for on-screen reading. Export as often as you like; there is no limit and no paywall on getting your own content out.

Privacy and data

Where is my data stored?

On servers operated by Infomaniak in Switzerland. Compute, database, file storage and outbound email are all Swiss, under the revised Swiss Data Protection Act and the GDPR.

Do you use my recordings to train models?

No. Not your audio, not your transcripts, not your corrections, and we have no arrangement with any third party that would permit it. Your material is processed to produce your transcript and for nothing else.

Who can see my transcripts?

You, and anyone you explicitly add as an owner. Access is enforced on every request rather than only in the interface. Our administrators can see account metadata for support and abuse handling; we do not read transcripts as a matter of course.

Do I need consent from the people in my recording?

Very possibly, and that is your responsibility rather than ours. Recording law varies by country and by context, and the people speaking in your file never agreed to anything with us. By uploading you confirm you have the right to do so — see the Acceptable Use policy.

What exactly happens when I delete something?

Deleting a transcript removes the audio and the generated transcript files from disk; the transcript text stays in your account record for usage accounting, which we state plainly rather than glossing over. Deleting your account removes everything, text included, and is immediate.

Limits and plans

What are the free plan limits?

60 audio-minutes per month, any number of files, up to 500 MB and 60 minutes per file, and one transcription processing at a time. Everything else — speaker detection, the editor, annotations, exports — is included with no restriction.

What if I go over?

New uploads pause until the monthly reset. Nothing already transcribed is affected and nothing is deleted. We email you at 80% so you can plan around it, and your usage panel shows the running total and the reset date.

How long do you keep my files?

90 days from the last time you opened a transcript. Opening it resets the clock, so anything you use stays. We warn you by email 14 days and 7 days before deletion, and you can export at any time.

When is the Pro plan available?

Not yet. It is displayed so you can see where we are going. There is no sign-up list and nothing to do — when it launches it will simply appear in your account, and nothing is charged until you actively choose to subscribe.

Do I get minutes back if I delete a file?

No. Minutes are consumed when the transcription runs, since that is when the computing happens. If deleting refunded them, the monthly budget would be resettable at will and would not mean anything.

Still stuck?

Write to us. We read everything and we answer.