| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Bump version to v0.10.0-alpha.3
This is the 0.9.3 release of Deep Speech, an open speech-to-text engine. In accord with semantic versioning, this version is not backwards compatible with earlier versions. However, models exported for 0.7.X and 0.8.X should work with this release. This is a bugfix release and retains compatibility with the 0.9.0, 0.9.1 and 0.9.2 models. All model files included here are identical to the ones in the 0.9.0 release. As with previous releases, this release includes the source code:
Under the MPL-2.0 license. And the acoustic models:
deepspeech-0.9.3-models.pbmm
deepspeech-0.9.3-models.tflite
In addition we're releasing experimental Mandarin Chinese acoustic models trained on an internal corpus composed of 2000h of read speech:
deepspeech-0.9.3-models-zh-CN.pbmm
deepspeech-0.9.3-models-zh-CN.tflite
all under the MPL-2.0 license.
The model files with the ".pbmm" extension are memory mapped and thus memory efficient and fast to load. The model files with the ".tflite" extension are converted to use TensorFlow Lite, has post-training quantization enabled, and are more suitable for resource constrained environments.
The acoustic models were trained on American English with synthetic noise augmentation and the .pbmm model achieves an 7.06% word error rate on the LibriSpeech clean test corpus.
Note that the model currently performs best in low-noise environments with clear recordings and has a bias towards US male accents. This does not mean the model cannot be used outside of these conditions, but that accuracy may be lower. Some users may need to train the model further to meet their intended use-case.
In addition we release the scorer:
deepspeech-0.9.3-models.scorer
which takes the place of the language model and trie in older releases and which is also under the MPL-2.0 license.
There is also a corresponding scorer for the Mandarin Chinese model:
deepspeech-0.9.3-models-zh-CN.scorer
We also include example audio files:
which can be used to test the engine, and checkpoint files for both the English and Mandarin models:
deepspeech-0.9.3-checkpoint.tar.gz
deepspeech-0.9.3-checkpoint-zh-CN.tar.gz
which are under the MPL-2.0 license and can be used as the basis for further fine-tuning.
The hyperparameters used to train the model are useful for fine tuning. Thus, we document them here along with the training regimen, hardware used (a server with 8 Quadro RTX 6000 GPUs each with 24GB of VRAM), and our use of cuDNN RNN.
In contrast to some previous releases, training for this release occurred as a fine tuning of the previous 0.8.2 checkpoint, with data augmentation options enabled. The following hyperparameters were used for the fine tuning. See the 0.8.2 release notes for the hyperparameters used for the base model.
The weights with the best validation loss were selected at the end of 200 epochs using --noearly_stop.
The optimal lm_alpha and lm_beta values with respect to the LibriSpeech clean dev corpus remain unchanged from the previous release:
For the Mandarin Chinese model, the following values are recommended:
This release also includes a Python based command line tool deepspeech, installed through
pip install deepspeech
Alternatively, quicker inference can be performed using a supported NVIDIA GPU on Linux. (See below to find which GPU's are supported.) This is done by instead installing the GPU specific package:
pip install deepspeech-gpuOn Linux, macOS and Windows, the DeepSpeech package does not use TFLite by default. A TFLite version of the package on those platforms is available as:
pip install deepspeech-tfliteAlso, it exposes bindings for the following languages
Python (Versions 3.5, 3.6, 3.7, 3.8 and 3.9) installed via
pip install deepspeechAlternatively, quicker inference can be performed using a supported NVIDIA GPU on Linux. (See below to find which GPU's are supported.) This is done by instead installing the GPU specific package:
pip install deepspeech-gpuOn Linux (AMD64), macOS and Windows, the DeepSpeech package does not use TFLite by default. A TFLite version of the package on those platforms is available as:
pip install deepspeech-tfliteNodeJS (Versions 10.x, 11.x, 12.x, 13.x, 14.x and 15.x) installed via
npm install deepspeech
Alternatively, quicker inference can be performed using a supported NVIDIA GPU on Linux. (See below to find which GPU's are supported.) This is done by instead installing the GPU specific package:
npm install deepspeech-gpu
On Linux (AMD64), macOS and Windows, the DeepSpeech package does not use TFLite by default. A TFLite version of the package on those platforms is available as:
npm install deepspeech-tfliteElectronJS versions 5.0, 6.0, 6.1, 7.0, 7.1, 8.0, 9.0, 9.1, 9.2, 10.0, 10.1, and 11.0 are also supported
C which requires the appropriate shared objects are installed from native_client.tar.xz (See the section in the main README which describes native_client.tar.xz installation.)
.NET which is installed by following the instructions on the NuGet package page.
In addition there are third party bindings that are supported by external developers, for example
This is the 0.9.2 release of Deep Speech, an open speech-to-text engine. In accord with semantic versioning, this version is not completely backwards compatible with earlier versions. However, models exported for 0.7.X and 0.8.X should work with this release. This is a bugfix release and retains compatibility with the 0.9.0 and 0.9.1 models. All model files included here are identical to the ones in the 0.9.0 release. As with previous releases, this release includes the source code:
Under the MPL-2.0 license. And the acoustic models:
deepspeech-0.9.2-models.pbmm
deepspeech-0.9.2-models.tflite
In addition we're releasing experimental Mandarin Chinese acoustic models trained on an internal corpus composed of 2000h of read speech:
deepspeech-0.9.2-models-zh-CN.pbmm
deepspeech-0.9.2-models-zh-CN.tflite
all under the MPL-2.0 license.
The model files with the ".pbmm" extension are memory mapped and thus memory efficient and fast to load. The model files with the ".tflite" extension are converted to use TensorFlow Lite, has post-training quantization enabled, and are more suitable for resource constrained environments.
The acoustic models were trained on American English with synthetic noise augmentation and the .pbmm model achieves an 7.06% word error rate on the LibriSpeech clean test corpus.
Note that the model currently performs best in low-noise environments with clear recordings and has a bias towards US male accents. This does not mean the model cannot be used outside of these conditions, but that accuracy may be lower. Some users may need to train the model further to meet their intended use-case.
In addition we release the scorer:
deepspeech-0.9.2-models.scorer
which takes the place of the language model and trie in older releases and which is also under the MPL-2.0 license.
There is also a corresponding scorer for the Mandarin Chinese model:
deepspeech-0.9.2-models-zh-CN.scorer
We also include example audio files:
which can be used to test the engine, and checkpoint files for both the English and Mandarin models:
deepspeech-0.9.2-checkpoint.tar.gz
deepspeech-0.9.2-checkpoint-zh-CN.tar.gz
which are under the MPL-2.0 license and can be used as the basis for further fine-tuning.
The hyperparameters used to train the model are useful for fine tuning. Thus, we document them here along with the training regimen, hardware used (a server with 8 Quadro RTX 6000 GPUs each with 24GB of VRAM), and our use of cuDNN RNN.
In contrast to some previous releases, training for this release occurred as a fine tuning of the previous 0.8.2 checkpoint, with data augmentation options enabled. The following hyperparameters were used for the fine tuning. See the 0.8.2 release notes for the hyperparameters used for the base model.
The weights with the best validation loss were selected at the end of 200 epochs using --noearly_stop.
The optimal lm_alpha and lm_beta values with respect to the LibriSpeech clean dev corpus remain unchanged from the previous release:
For the Mandarin Chinese model, the following values are recommended:
This release also includes a Python based command line tool deepspeech, installed through
pip install deepspeech
Alternatively, quicker inference can be performed using a supported NVIDIA GPU on Linux. (See below to find which GPU's are supported.) This is done by instead installing the GPU specific package:
pip install deepspeech-gpuOn Linux, macOS and Windows, the DeepSpeech package does not use TFLite by default. A TFLite version of the package on those platforms is available as:
pip install deepspeech-tfliteAlso, it exposes bindings for the following languages
Python (Versions 3.5, 3.6, 3.7, 3.8 and 3.9) installed via
pip install deepspeechAlternatively, quicker inference can be performed using a supported NVIDIA GPU on Linux. (See below to find which GPU's are supported.) This is done by instead installing the GPU specific package:
pip install deepspeech-gpuOn Linux (AMD64), macOS and Windows, the DeepSpeech package does not use TFLite by default. A TFLite version of the package on those platforms is available as:
pip install deepspeech-tfliteNodeJS (Versions 10.x, 11.x, 12.x, 13.x, 14.x and 15.x) installed via
npm install deepspeech
Alternatively, quicker inference can be performed using a supported NVIDIA GPU on Linux. (See below to find which GPU's are supported.) This is done by instead installing the GPU specific package:
npm install deepspeech-gpu
On Linux (AMD64), macOS and Windows, the DeepSpeech package does not use TFLite by default. A TFLite version of the package on those platforms is available as:
npm install deepspeech-tfliteElectronJS versions 5.0, 6.0, 6.1, 7.0, 7.1, 8.0, 9.0, 9.1, 9.2, 10.0, 10.1, and 11.0 are also supported
C which requires the appropriate shared objects are installed from native_client.tar.xz (See the section in the main README which describes native_client.tar.xz installation.)
.NET which is installed by following the instructions on the NuGet package page.
In addition there are third party bindings that are supported by external developers, for example
This is the 0.9.1 release of Deep Speech, an open speech-to-text engine. In accord with semantic versioning, this version is not completely backwards compatible with earlier versions. However, models exported for 0.7.X and 0.8.X should work with this release. This is a bugfix release and retains compatibility with the 0.9.0 models. All model files included here are identical to the ones in the 0.9.0 release. As with previous releases, this release includes the source code:
Under the MPL-2.0 license. And the acoustic models:
deepspeech-0.9.1-models.pbmm
deepspeech-0.9.1-models.tflite
In addition we're releasing experimental Mandarin Chinese acoustic models trained on an internal corpus composed of 2000h of read speech:
deepspeech-0.9.1-models-zh-CN.pbmm
deepspeech-0.9.1-models-zh-CN.tflite
all under the MPL-2.0 license.
The model files with the ".pbmm" extension are memory mapped and thus memory efficient and fast to load. The model files with the ".tflite" extension are converted to use TensorFlow Lite, has post-training quantization enabled, and are more suitable for resource constrained environments.
The acoustic models were trained on American English with synthetic noise augmentation and the .pbmm model achieves an 7.06% word error rate on the LibriSpeech clean test corpus.
Note that the model currently performs best in low-noise environments with clear recordings and has a bias towards US male accents. This does not mean the model cannot be used outside of these conditions, but that accuracy may be lower. Some users may need to train the model further to meet their intended use-case.
In addition we release the scorer:
deepspeech-0.9.1-models.scorer
which takes the place of the language model and trie in older releases and which is also under the MPL-2.0 license.
There is also a corresponding scorer for the Mandarin Chinese model:
deepspeech-0.9.1-models-zh-CN.scorer
We also include example audio files:
which can be used to test the engine, and checkpoint files for both the English and Mandarin models:
deepspeech-0.9.1-checkpoint.tar.gz
deepspeech-0.9.1-checkpoint-zh-CN.tar.gz
which are under the MPL-2.0 license and can be used as the basis for further fine-tuning.
The hyperparameters used to train the model are useful for fine tuning. Thus, we document them here along with the training regimen, hardware used (a server with 8 Quadro RTX 6000 GPUs each with 24GB of VRAM), and our use of cuDNN RNN.
In contrast to some previous releases, training for this release occurred as a fine tuning of the previous 0.8.2 checkpoint, with data augmentation options enabled. The following hyperparameters were used for the fine tuning. See the 0.8.2 release notes for the hyperparameters used for the base model.
The weights with the best validation loss were selected at the end of 200 epochs using --noearly_stop.
The optimal lm_alpha and lm_beta values with respect to the LibriSpeech clean dev corpus remain unchanged from the previous release:
For the Mandarin Chinese model, the following values are recommended:
This release also includes a Python based command line tool deepspeech, installed through
pip install deepspeech
Alternatively, quicker inference can be performed using a supported NVIDIA GPU on Linux. (See below to find which GPU's are supported.) This is done by instead installing the GPU specific package:
pip install deepspeech-gpuOn Linux, macOS and Windows, the DeepSpeech package does not use TFLite by default. A TFLite version of the package on those platforms is available as:
pip install deepspeech-tfliteAlso, it exposes bindings for the following languages
Python (Versions 3.5, 3.6, 3.7 and 3.8) installed via
pip install deepspeechAlternatively, quicker inference can be performed using a supported NVIDIA GPU on Linux. (See below to find which GPU's are supported.) This is done by instead installing the GPU specific package:
pip install deepspeech-gpuOn Linux (AMD64), macOS and Windows, the DeepSpeech package does not use TFLite by default. A TFLite version of the package on those platforms is available as:
pip install deepspeech-tfliteNodeJS (Versions 10.x, 11.x, 12.x, 13.x, 14.x and 15.x) installed via
npm install deepspeech
Alternatively, quicker inference can be performed using a supported NVIDIA GPU on Linux. (See below to find which GPU's are supported.) This is done by instead installing the GPU specific package:
npm install deepspeech-gpu
On Linux (AMD64), macOS and Windows, the DeepSpeech package does not use TFLite by default. A TFLite version of the package on those platforms is available as:
npm install deepspeech-tfliteElectronJS versions 5.0, 6.0, 6.1, 7.0, 7.1, 8.0, 9.0, 9.1, 9.2, 10.0 and 10.1 are also supported
C which requires the appropriate shared objects are installed from native_client.tar.xz (See the section in the main README which describes native_client.tar.xz installation.)
.NET which is installed by following the instructions on the NuGet package page.
In addition there are third party bindings that are supported by external developers, for example
Windows 8.1, 10, and Server 2012 R2 64-bits (at least AVX support, requires Redistribuable Visual C++ 2015 Update 3 (64-bits) for runtime).
OS X 10.10, 10.11, 10.12, 10.13, 10.14, and 10.15
Linux x86 64 bit with a modern CPU (at least AVX/FMA)
Linux x86 64 bit with a modern CPU (at least AVX/FMA) + NVIDIA GPU (Compute Capability at least 3.0, see NVIDIA docs)
Raspbian Buster on Raspberry Pi 3, Pi 4
Linux/ARM64 built against Debian/ARMbian Buster and tested on LePotato boards
Java Android (7.0-11.0) bindings (+ demo app). Tested on Google Pixel 2 ; Sony Xperia Z Premium ; Nokia 1.3, TF Lite model only.
iOS with Swift bindings (experimental). Tested on iPhone Xs.
TFLite Delegation API is here...
This is the 0.9.0 release of Deep Speech, an open speech-to-text engine. In accord with semantic versioning, this version is not completely backwards compatible with earlier versions. However, models exported for 0.7.X and 0.8.X should work with this release. As with previous releases, this release includes the source code:
Under the MPL-2.0 license. And the acoustic models:
deepspeech-0.9.0-models.pbmm
deepspeech-0.9.0-models.tflite
In addition we're releasing experimental Mandarin Chinese acoustic models trained on an internal corpus composed of 2000h of read speech:
deepspeech-0.9.0-models-zh-CN.pbmm
deepspeech-0.9.0-models-zh-CN.tflite
all under the MPL-2.0 license.
The model files with the ".pbmm" extension are memory mapped and thus memory efficient and fast to load. The model files with the ".tflite" extension are converted to use TensorFlow Lite, has post-training quantization enabled, and are more suitable for resource constrained environments.
The acoustic models were trained on American English with synthetic noise augmentation and the .pbmm model achieves an 7.06% word error rate on the LibriSpeech clean test corpus.
Note that the model currently performs best in low-noise environments with clear recordings and has a bias towards US male accents. This does not mean the model cannot be used outside of these conditions, but that accuracy may be lower. Some users may need to train the model further to meet their intended use-case.
In addition we release the scorer:
deepspeech-0.9.0-models.scorer
which takes the place of the language model and trie in older releases and which is also under the MPL-2.0 license.
There is also a corresponding scorer for the Mandarin Chinese model:
deepspeech-0.9.0-models-zh-CN.scorer
We also include example audio files:
which can be used to test the engine, and checkpoint files for both the English and Mandarin models:
deepspeech-0.9.0-checkpoint.tar.gz
deepspeech-0.9.0-checkpoint-zh-CN.tar.gz
which are under the MPL-2.0 license and can be used as the basis for further fine-tuning.
The hyperparameters used to train the model are useful for fine tuning. Thus, we document them here along with the training regimen, hardware used (a server with 8 Quadro RTX 6000 GPUs each with 24GB of VRAM), and our use of cuDNN RNN.
In contrast to some previous releases, training for this release occurred as a fine tuning of the previous 0.8.2 checkpoint, with data augmentation options enabled. The following hyperparameters were used for the fine tuning. See the 0.8.2 release notes for the hyperparameters used for the base model.
The weights with the best validation loss were selected at the end of 200 epochs using --noearly_stop.
The optimal lm_alpha and lm_beta values with respect to the LibriSpeech clean dev corpus remain unchanged from the previous release:
For the Mandarin Chinese model, the following values are recommended:
This release also includes a Python based command line tool deepspeech, installed through
pip install deepspeech
Alternatively, quicker inference can be performed using a supported NVIDIA GPU on Linux. (See below to find which GPU's are supported.) This is done by instead installing the GPU specific package:
pip install deepspeech-gpuOn Linux, macOS and Windows, the DeepSpeech package does not use TFLite by default. A TFLite version of the package on those platforms is available as:
pip install deepspeech-tfliteAlso, it exposes bindings for the following languages
Python (Versions 3.5, 3.6, 3.7 and 3.8) installed via
pip install deepspeechAlternatively, quicker inference can be performed using a supported NVIDIA GPU on Linux. (See below to find which GPU's are supported.) This is done by instead installing the GPU specific package:
pip install deepspeech-gpuOn Linux (AMD64), macOS and Windows, the DeepSpeech package does not use TFLite by default. A TFLite version of the package on those platforms is available as:
pip install deepspeech-tfliteNodeJS (Versions 10.x, 11.x, 12.x, 13.x, 14.x and 15.x) installed via
npm install deepspeech
Alternatively, quicker inference can be performed using a supported NVIDIA GPU on Linux. (See below to find which GPU's are supported.) This is done by instead installing the GPU specific package:
npm install deepspeech-gpu
On Linux (AMD64), macOS and Windows, the DeepSpeech package does not use TFLite by default. A TFLite version of the package on those platforms is available as:
npm install deepspeech-tfliteElectronJS versions 5.0, 6.0, 6.1, 7.0, 7.1, 8.0, 9.0, 9.1 and 9.2 are also supported
C which requires the appropriate shared objects are installed from native_client.tar.xz (See the section in the main README which describes native_client.tar.xz installation.)
.NET which is installed by following the instructions on the NuGet package page.
In addition there are third party bindings that are supported by external developers, for example
Bump VERSION to 0.9.0-alpha.12
Bump VERSION to 0.9.0-alpha.11
Merge pull request #3337 from lissyx/bump-0.9.0a10 Bump VERSION to 0.9.0-alpha.10
Bump VERSION to 0.9.0-alpha.9
Merge pull request #3315 from lissyx/bump-v0.9.0-alpha.8 Bump VERSION to v0.9.0-alpha.8
| Back | FazBrowse Home | New Git URL |