Published April 19, 2022 | Version 1.1

Transcribing audio data: overview and transcripts of several automatic transcription tools

  • 1. Utrecht University

Description

Throughout institutions, audio recordings are being made regularly. To be able to further process these recordings, the audio often needs to be transcribed. In order to avoid having to transcribe the audio manually, there is a wealth of tools available for doing so automatically. In this record, we present an overview of several often-used tools to automatically transcribe pre-recorded audio data, including their features, costs, and security.

To check the quality of the tool, we also recorded an audio fragment in Dutch that we ran through all tools in this overview in March of 2022. This original audio fragment (Test_interview_20220203.mp3), the cleaned-up transcription (Test_interview_cleaned_transcript.odt) and each tool’s raw transcript of the audio fragment (Test_interview_[name-tool]_raw_[date-run]) are included in this record as well. The raw transcripts were downloaded as .docx or .txt files and the .docx files saved as .odt. No edits to the transcripts were made before saving them, except an incidental removal of a personal email address or hyperlink.

The overview contains information and transcripts of following transcription tools:

  • Amberscript
  • HappyScribe
  • Kaldi
  • NVIVO transcription
  • Sonix
  • SpokenOnline
  • Transcribe
  • Trint
  • Microsoft Word 365 Online

About

This overview was created through a collaboration between Utrecht University’s Research Data Management (RDM) Support and the DataHub SSH programme situated at the faculty of Humanities.

The details in the overview have last been updated April 19, 2022. Please note that at the time you are downloading these files, the quality of the (Dutch) speech-to-text conversion may have been improved by the respective supplier.

Files

Audio_transcription_tool_overview.pdf

Files (3.7 MB)

Name Size Download all
md5:0fb3df9a110aefd38a3dd084e15c8766
137.4 kB Preview Download
md5:84c71610873cc8917d0214ae0948b20b
3.4 MB Preview Download
md5:e508bf319a558a91125fa8d8fb5365d5
8.2 kB Download
md5:e13e55323d660cfe995a1c9a81569062
9.1 kB Download
md5:ba25109621dbbedc7b364254fd7e5e97
7.2 kB Download
md5:73de82bde867a2a2e6fca78ddf757289
2.1 kB Preview Download
md5:6b83c90907562fa5f1119687d572a5d1
24.0 kB Preview Download
md5:74b81b173ff4fa2489e8ff1a9081ded2
7.5 kB Download
md5:90995fede3928d5e849bacbf8311107f
7.4 kB Download
md5:e26f4b9490c7d684db3b33913b870088
5.7 kB Download
md5:56eaf4615c8cd428d43b44e0f46218aa
2.1 kB Preview Download
md5:3a561c5f46bb39bc3582775067ed22e4
6.7 kB Download
md5:7d77c1ab6b99a522d0ee446b7958367e
8.9 kB Download

Additional details

Related works

Subjects

transcription
http://purl.org/coar/resource_type/6NC7-GK9S