Voice AI

Voice AI is KISSKI’s AI-based transcription, translation and captioning service, running on the HPC infrastructure of the GWDG. It offers a choice of open speech-recognition models, speaker labelling, subtitles and AI summaries, and is free of charge for Academic Cloud users.

Access. You need an Academic Cloud account (federated login from most German universities, or a self-registered account). Open the service, sign in, and it is ready; there is nothing to install.

Two Ways to Use It

  1. Audio file transcription and translation: upload a recording (up to 2 GB) and download the result.
  2. Live transcription (beta): real-time captions, subtitles and minutes for a meeting or lecture, from your browser.

How it works. You sign in through Academic Cloud single sign-on; a proxy forwards only authenticated requests to the Voice AI or Voice Live AI web service. Uploads are handed to the KISSKI platform, where the job is scheduled on a GPU compute node. Results are written to the data mover node, from which you download them. All components run at GWDG.

Audio File Transcription and Translation

Upload a recording, pick a model and options, and download the transcript when the job has run on the GPU cluster. Jobs are queued; typical waiting time is minutes, longer at peak hours.

Accepted input. mp3, wav, m4a, ogg, opus, flac, aac and the video containers mp4, mkv, mov and webm, up to 2 GB per file. A damaged or truncated file fails with a clear message rather than producing a shortened transcript, and the uploaded file is deleted automatically.

Models. Choose per job; the interface only offers what the selected model supports.

ModelTranslate to EnglishSpeaker labelsSRT / VTTInput languageVocabulary hints
Whisper large-v2 (default)yesyesyesoptional, 58 languagesyes
Whisper large-v3yesyesyesoptional, 59 languagesyes
Qwen3-ASR 1.7Bnoyesyesoptional, 30 languagesyes
Cohere Transcribe 03-2026nonono (no timestamps)required, 14 languagesno

Features

  • Output as plain text, SRT, VTT (two-line subtitles, 42 characters per line) or JSON with segments, speakers and timings.
  • Speaker labels (diarization); each new speaker starts a new subtitle block.
  • Vocabulary hints: names and domain terms typed into the job so the model spells them correctly.
  • AI summary of the transcript and a chat box to ask questions about it, through GWDG Chat AI.
  • Job list with status in queue, finished or failed with the reason; results downloadable for 30 days; delete at any time.
  • Interface in English and German, light and dark mode, works on small screens.

Steps for an Uploaded File

  1. Choose the model.
  2. Choose the language of the recording (optional for Whisper and Qwen, required for Cohere).
  3. Choose transcription or translation to English, the output format, and whether to label speakers.
  4. Optionally add vocabulary hints.
  5. Upload the file. The job appears in your list.
  6. Download the result when the status is finished; request a summary or open the chat if you like.

API. Transcription is also available programmatically through the SAIA API key; see SAIA – Voice to text for the endpoint and examples. Queued batch processing through the API is on the roadmap.

Live Transcription (Beta)

Live transcription shows what is being said as captions within about a second, optionally translated, directly in the browser. It is designed to be run by one person during a meeting, lecture or conference talk, with no software to install.

Audio Source

  • Microphone: your own voice, or a room or speaker microphone connected to your computer.
  • Browser tab (Chrome or Edge): the sound of a tab running your BigBlueButton, Zoom or Teams meeting. Only audio is captured, never video; your microphone is added so your own voice is included too. This replaces the virtual audio cable for most users; the cable route remains in the FAQ for other browsers.

Options Before You Start

OptionWhat it does
Spoken languageRequired. 15 languages including German, English, French, Spanish, Italian, Dutch, Polish, Turkish, Russian, Arabic, Chinese, Japanese, Korean, Hindi, Portuguese
ModeTranscription, Translate German → English or Translate English → German. Translation is currently offered for German ↔ English only
Speaker IDsLabels who is speaking, for up to four people; you can name the speakers by clicking a label. Costs a few seconds of delay per sentence; switch off for one speaker or large groups
TermsNames, places and jargon the recogniser should spell correctly; remembered in your browser
Steady textHides text still being recognised so nothing on screen moves; a few seconds behind, but calm to read
AppearanceFont size, line spacing, font colour and a dyslexia-friendly font; light and dark mode

While the Session Runs

  • Document view for the host: the full transcript with search, inline corrections, speaker naming and marks for decisions, actions and notes.
  • Detached live view: a separate window showing the last sentences in large type, for a second screen or a projector.
  • Live link, off by default: press Create live link to mirror the transcript to a read-only web page for the audience, including the names and corrections you make. Viewers need no login, so external guests can read along. Stop sharing or ending the session invalidates every link at once, and no link lives longer than four hours. Until you create a link, nothing about the session leaves your browser.

After the Session

  • Give the meeting a title and body; export the transcript as Markdown or plain text, or copy it.
  • Generate minutes drafts structured minutes in the language you choose, through GWDG Chat AI, and downloads them. Check names, figures and decisions against the transcript before you circulate them.
  • A punctuation and casing pass tidies the finished transcript automatically.

Capacity. Each GPU serves 5 sessions at a time. If the service is full you see Server is currently busy; try again a little later.

Steps for a Live Session

  1. Sign in and open the live tool.
  2. Choose the audio source and the spoken language (or a translation direction).
  3. Switch on Speaker IDs and add terms if useful; turn on Steady text if the audience reads along.
  4. Press Start Session. Create a live link if the audience should read along, or open the detached view.
  5. Press Stop Session. Export the transcript or generate minutes.

Privacy in Short

All processing takes place on GWDG hardware in Göttingen. Uploaded files are deleted after processing; results stay for 30 days or until you delete them. Live audio is processed in memory and never written to disk; the live transcript exists only in your browser, unless you create a live link, in which case a copy is held in server memory only while sharing is on. Summaries and minutes are produced by GWDG Chat AI, also hosted at GWDG. Usage (user name, time, service, model) is logged for capacity planning and accounting. Nothing is used to train models. See the FAQ and the data privacy notice.

Author

Narges Lux

Support

Questions and feedback: support@gwdg.de. Institutions planning wider roll-out: info@kisski.de.