Skip to main content
zenodoopen

Machine Learning for Bird Song Learning (ML4BL) dataset

<p><strong>General description</strong></p> <p>This dataset contains Zebra Finch decisions about perceptual similarity on song units. All the data and files are used for reproducing the results of the paper &#39;Bird song comparison using deep learning trained from avian perceptual judgments&#39; by the same authors.&nbsp;</p> <p><strong>Git&nbsp;repo on Zenodo:</strong>&nbsp;<a href="https://doi.org/10.5281/zenodo.5545932">https://doi.org/10.5281/zenodo.5545932</a><br> <strong>Git&nbsp;repo access:&nbsp;</strong><a href="https://github.com/veronicamorfi/ml4bl/tree/v1.0.0">https://github.com/veronicamorfi/ml4bl/tree/v1.0.0</a></p> <p><strong>Directory organisation:</strong><br> ML4BL_ZF<br> |_files<br> &nbsp; &nbsp; |_Final_probes_20200816.csv - all trials and decisions of the birds (aviary 1 cycle 1 data are removed from experiments)<br> &nbsp; &nbsp; |_luscinia_triplets_filtered.csv - triplets to use for training<br> &nbsp; &nbsp; |_mean_std_luscinia_pretraining.pckl - mean and std of luscinia triplets used for trianing<br> &nbsp; &nbsp; |_*_cons_* - % side consistency on triplets (train/test) - train set contains both train and val splits<br> &nbsp; &nbsp; |_*_gt_* - cycle accuracy for triplets of the specific bird&nbsp;(train/test)&nbsp;- train set contains both train&nbsp;and val splits<br> &nbsp; &nbsp; |_*_trials_* - number of decisions made for a triplet&nbsp;(train/test)&nbsp;- train set contains both train and val splits<br> &nbsp; &nbsp; |_*_triplets_* - triplet information (aviary_cycle-acc_birdID, POS, NEG, ANC)&nbsp;(train/test)&nbsp;- train set contains both train and val splits<br> &nbsp; &nbsp; |_*_low*_ - low-margin (ambiguous) triplets (train/val/test)<br> &nbsp; &nbsp; |_*_high_ -&nbsp;high-margin (unambiguous) triplets (train/val/test)<br> &nbsp; &nbsp; |_*_cycle_bird_keys_* - unique aviary_cycle-acc_birdID&nbsp;keys (train/test) -&nbsp;train set contains both train and val splits<br> &nbsp; &nbsp; |_TunedLusciniaV1e.csv - pairwise distance of two recordings computed by Luscinia<br> &nbsp; &nbsp; |_training_setup_1_ordered_acc_single_cons_50_70_trials.pckl - dictionary containing everything needed for training the model&nbsp;(keys: &#39;train_keys&#39;, &#39;train_triplets&#39;, &#39;val_keys&#39;, &#39;vali_triplets&#39;, &#39;test_triplets&#39;, &#39;test_keys&#39;, &#39;train_mean&#39;, &#39;train_std&#39;)<br> |_melspecs - *.pckl - melspectrograms of recordings<br> |_wavs - *wav - recordings<br> |_README.txt</p> <p><strong>Recordings</strong></p> <p>887 syllables extracted from zebra finch song recordings, with a sampling rate of 48kHz and high pass filtered (100Hz), with a 20ms intro/outro fade.&nbsp;</p> <p><strong>Decisions</strong></p> <p>Triplets were created from the recordings and the birds made side based decisions about their similarity (see &#39;Bird song comparison using deep learning trained from avian perceptual judgments&#39; for further information).</p> <p><strong>Training dictionary Information</strong></p> <p>Dictionary keys:<br> &nbsp;&nbsp; &nbsp;&#39;train_keys&#39;, &#39;train_triplets&#39;, &#39;val_keys&#39;, &#39;vali_triplets&#39;, &#39;test_triplets&#39;, &#39;test_keys&#39;, &#39;train_mean&#39;, &#39;train_std&#39;</p> <p>train_triplets/vali_triplets/test_triplets:&nbsp;<br> &nbsp;&nbsp; &nbsp;Aviary_Cycle_birdID, POS, NEG, ANC, Decisions, Cycle_ACC(%), Consistency(%)<br> &nbsp;<br> train_keys/val_keys/test_keys:<br> &nbsp;&nbsp; &nbsp;Aviary_Cycle_birdID</p> <p>train_mean/train_std:<br> &nbsp;&nbsp; &nbsp;shape: (1, mel_bins)</p> <p>&nbsp;</p> <p><strong>Open Access</strong></p> <p>This dataset is available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license.</p> <p><br> <strong>Contact info</strong></p> <p>Please send any questions about the recordings to:<br> Lies Zandberg:&nbsp;Elisabeth.Zandberg@rhul.ac.uk</p> <p>Please send any feedback or questions about the code and the rest of the data&nbsp;to:<br> Veronica Morfi: g.v.morfi@qmul.ac.uk</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
0
Engagement
8

Topics