Comments (4)
I added the following to augmentor, and used the following snippet from online augmentation tutorial
rir_data_path = f'{data_dir}/dataset'
!python {NEMO_ROOT}/scripts/dataset_processing/get_openslr_rir_data.py --data_root {rir_data_path}
rir_manifest_path = os.path.join(rir_data_path, 'processed', 'rir.json')
!head -n 3 {rir_manifest_path}
Then to use the augmentation I applied the following
audio_augmentations = dict(
speed = dict(
sr=16000,
prob=0.3,
resample_type='kaiser_fast',
min_speed_rate=0.95,
max_speed_rate=1.05,
),
noise = dict(
manifest_path=rir_manifest_path,
prob=0.5,
min_snr_db=0,
max_snr_db=15,
),
)
finetune_config.model.train_ds.augmentor = audio_augmentations
Am I correct and thanks @okuchaiev
from nemo.
Yes, code looks fine to me. But for impulse you should use impulse pertubation not noise pertubation.
Sample can be found here:
from nemo.
@nithinraok that's what I thought, However in Titanet-Large they use noise instead of impulse, and it says we are using impulse perturbation. So, does that mean in their training they made an error using RIR corpora for noise instead of pulse perturbation.
NeMo/examples/speaker_tasks/recognition/conf/titanet-large.yaml
Lines 14 to 26 in 6442bb6
The paper statement:
(just realized you are the first author x.x)
Thank you @nithinraok
from nemo.
I don;t remember details exactly but as far I remember RIR corpora also has noise samples as well along with impulse responses, and I have not added impulse section to this config file but was added to titanet-small config.
from nemo.
Related Issues (20)
- training config used for training stt_en_quartznet15x5 HOT 2
- llama2 training hangs when pp_size > 1 HOT 2
- Integration of Turn-Taking Models into Nemo Framework for Enhanced Realistic Conversations
- FileNotFoundError: Model stt_fa_fastconformer_hybrid_large was not found. HOT 6
- [Feature] Add Support on Multiple Metrics Reporting during Training Progress for Validation
- checkpoints not saved due to wrong loss comparison?
- when "write_predictions_to_file" is true,generate will fail。 HOT 2
- "RuntimeError: start (4) + length (1) exceeds dimension size (4)." when running cache aware streaming inference
- slow validation process HOT 2
- Optimizing Learning Rate Parameters in Model Fine-tuning
- AUDIO FILE SIZE for fine tuning STT En FastConformer Hybrid Transducer-CTC Large Streaming Multi HOT 1
- `EncDecCTCModel.transcribe(audio=...)` changed to `EncDecCTCModel.transcribe(paths2audio_files=...)` HOT 7
- Enormous number of `.nemo` checkpoints produced in training HOT 4
- [Conversion] How to convert Finetuned T5 checkpoint ended with `.ckpt` to `.nemo` checkpoint with NeMo toolkit?
- Can't launch NeMo containers with CUDA support
- Latest huggingface transformers version breaking nlp modules HOT 6
- Any tts models in nemo that can simulated human laughter and other human sounds?
- setuptools 70.0.0 results in ImportError: cannot import name 'packaging' from 'pkg_resources' HOT 3
- Question about the settings in speech_data_simulator HOT 4
- The training hangs in the middle on multiple nodes, showing low power consumption and 100% GPU utilization.
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from nemo.