Voice Recognition

deepspeech.pytorch – Implementation of DeepSpeech2 using Baidu Warp-CTC

deepspeech.pytorch is an implementation of DeepSpeech2 using Baidu Warp-CTC. The software creates a network based on the DeepSpeech2 architecture, trained with the CTC activation function.

Key Features

  • Train DeepSpeech, configurable RNN types and architectures with multi-GPU support.
  • Language model support using kenlm (WIP currently).
  • Multiple dataset downloaders, support for AN4, TEDLIUM, Voxforge and Librispeech. Datasets can be merged, and support for custom datasets is included.
  • Noise injection (dynamic) for online training to improve noise robustness.
  • Audio augmentation to improve noise robustness. This applies small changes to the tempo and gain when loading audio to increase robustness.
  • Easy start/stop capabilities in the event of crash or hard stop during training.
  • Visdom/Tensorboard support for visualizing training graphs.

This software has the following dependencies: python-levenshtein, torch, visdom, wget, librosa, and tqdm.

The project also offers a set of pre-trained networks for evaluation usage.

Website: github.com/SeanNaren/deepspeech.pytorch
Support:
Developer: Sean Naren
License: MIT License

deepspeech.pytorch is written in Python. Learn Python with our recommended free books and free tutorials.


Related Software

Speech Recognition Tools
WhisperAutomatic speech recognition (system trained on 680,000 hours of data
whisper.cppRun Whisper locally with fast, efficient C/C++ speech recognition
FlashlightFast, flexible machine learning library written entirely in C++.
sherpa-onnxOffline speech recognition for many platforms and languages
KaldiC++ toolkit designed for speech recognition researchers.
FunASRVersatile speech recognition toolkit for real-world audio
faster-whisperFaster Whisper transcription with reduced memory use
SpeechBrainAll-in-one conversational AI toolkit based on PyTorch
HandyOffline speech-to-text application
ESPnetEnd-to-End speech processing toolkit
PocketSphinxLightweight speech recognition for embedded and mobile use
deepspeech.pytorchImplementation of DeepSpeech2 using Baidu Warp-CTC.
EpicenterTranscription application with global speech-to-text functionality
JuliusTwo-pass large vocabulary continuous speech recognition engine
Speech NoteWrite notes, transcribe speech, translate text and read content aloud
aTrainPrivate local transcription of recorded speech through a polished GUI
hyprwhsprNative speech-to-text designed for Arch / Omarchy
VocalinuxLocal voice dictation through a clean desktop interface
osttOpen Speech-to-Text
SimonFlexible speech recognition software

Read our verdict in the software roundup.


Best Free and Open Source Software Explore our comprehensive directory of recommended free and open source software. Our carefully curated collection spans every major software category.

This directory is part of our ongoing series of informative articles for Linux enthusiasts. It features hundreds of detailed reviews, along with open source alternatives to proprietary solutions from major corporations such as Google, Microsoft, Apple, Adobe, IBM, Cisco, Oracle, and Autodesk.

You’ll also find interesting projects to try, hardware coverage, free programming books and tutorials, and much more.

Discovered a useful open source Linux program that we haven’t covered yet? Let us know by completing this form.
Subscribe

Please read our Comment Policy before commenting.

Notify of
guest
0 Comments
Oldest
Newest Most Voted