Voice AI FAQ

Live Transcription

How do I get the best accuracy? Choose the spoken language before you start; a session cannot begin without it. Add names and jargon under Terms. Use a headset or a room microphone rather than a laptop’s built-in one where you can.

How do I transcribe an online meeting (BigBlueButton, Zoom, MS Teams)? Run the meeting in a Chrome or Edge tab, choose Browser tab as the audio source, pick that tab and tick Also share tab audio. Only sound is captured; your microphone is added so your own contributions appear too. No software is needed.

In other browsers, or for a desktop meeting app, route the loudspeaker into a virtual microphone:

  • Windows: install VB-CABLE. In the meeting app set the speaker to CABLE Input; in the browser choose CABLE Output as the microphone.
  • macOS: install BlackHole. In Audio MIDI Setup create a Multi-Output Device with your speakers and BlackHole, make it the default output, and choose BlackHole as the microphone.
  • Linux: load the PulseAudio loopback module (pactl load-module module-loopback) and route the meeting app’s monitor to the browser’s input with pavucontrol.

Can I translate live? Yes, German → English and English → German. Choosing a direction sets the spoken language. Other languages can be transcribed but not translated yet.

How do speaker labels work? Switch on Speaker IDs before starting. Up to four speakers are told apart; click a label to give the speaker a name, which also appears on the live link. Speaker IDs add a few seconds of delay per sentence and use extra GPU time, so leave them off for a single speaker or a large group.

The words keep changing while I read. Can I stop that? Turn on Steady text. The text still being recognised is hidden and only finished sentences are shown, a few seconds behind but without movement.

How do I show the captions to others? Sharing is off by default. During a session press Create live link and send the link; viewers see the captions read-only, with any names and corrections you add, and need no account, so guests from outside can follow. Stop sharing or ending the session makes every link stop working immediately; no link lasts longer than four hours in any case. For a projector or second screen use the detached live view.

Can I get minutes? Yes. After stopping, open the document menu, give the meeting a title, choose the language and press Generate & download. The minutes are a draft produced by GWDG Chat AI; check names, figures and decisions against the transcript.

“Server is currently busy”: what now? Each GPU serves 5 live sessions at once. Wait a minute and try again. If your institution plans regular use for many people, write to info@kisski.de so capacity can be planned.

Uploaded Files

Which model should I choose? Whisper large-v2 (default) is the all-rounder with translation, speaker labels and subtitles. Whisper large-v3 is slightly more accurate on some languages. Qwen3-ASR 1.7B is strong on Chinese and other Asian languages but does not translate. Cohere Transcribe needs the language set, gives no timestamps and no speaker labels. The interface hides options a model cannot honour.

Which files can I upload, and how large? mp3, wav, m4a, ogg, opus, flac, aac, mp4, mkv, mov and webm, up to 2 GB. Larger files: convert to FLAC first, for example ffmpeg -i input.wav -vn -acodec flac output.flac, or split the recording.

My job shows failed. Why? The reason is shown next to the job. The most common cause is a damaged or truncated upload; re-export the recording and upload again. If the reason is unclear, send the job’s file name to support@gwdg.de.

What are vocabulary hints? A short list of names and domain terms that the model is told to prefer. Available for Whisper and Qwen3-ASR. They are stored beside your upload and deleted with it.

Can I use Voice AI from a script? Yes, through the SAIA API key; the endpoint and examples are documented under SAIA – Voice to text. Batch processing with a queue is planned; see the roadmap.

Data Privacy

Are my conversations or usage data used for AI training? No. Nothing you upload or say is used to train models.

Are my audio files and conversations stored on your servers? Uploaded files are kept only until the job has run, then deleted. Results (text, SRT, VTT, JSON) stay on the data mover node for 30 days, or until you delete them. Live audio is processed in memory and never written to disk. The live transcript stays in your browser; only if you create a live link is a copy held in server memory, and it is dropped the moment you stop sharing or end the session.

What data does Voice AI keep? Usage records: user name, time stamp, service used and model, for capacity planning and accounting. Vocabulary hints and terms are deleted with the job or kept only in your browser.

Where does the summary or the minutes get generated? By GWDG Chat AI, hosted at GWDG. The transcript is sent there for the request and not retained.

Availability and Cost

My institution wants to roll Voice AI out to many users. Can you handle the load? Contact info@kisski.de so that capacity can be planned; for live transcription the number of simultaneous sessions is what matters.

Is Voice AI free? Yes for Academic Cloud users.

Is there training material? A short video and a PDF guide per scenario are in preparation. An online walkthrough for the affected staff of an institution can be arranged through support@gwdg.de.