Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,185
datasets available to search
ShareScore release 0.9.0
Dataset results
7,185 results for “learning”
Video classification using deep learning
<p>Material accompanying the paper "Frame-by-frame annotation of video recordings using deep neural networks". Contains a selection of videos used in the paper, manual annotations, and results. Code is contained in the associated GitHub repository. See the readme for details.</p>
Developing a deep Learning network to retrieve ocean hydrographic profiles in the North Atlantic from combined satellite and in situ measurements: test datasets.
<p>We provide here the datasets used for the test and assessment of a deep learning algorithm which is presently candidate for the development of a daily 3D ocean product covering the North Atlantic at 1/10° resolution, over the 2010-2018 period, as part of the European Space Agency World Ocean Circulation project (ESA-WOC). The method is based on a stacked Long Short-Term Memory neural network, coupled to a Monte-Carlo dropout approach, and allows to project satellite-derived sea surface temperature, sea surface salinity and absolute dynamic topography data at depth after training with sparse co-located in situ vertical hydrographic profiles (Buongiorno Nardelli, 2020, doi:<a href="https://www.researchgate.net/deref/http%3A%2F%2Fdx.doi.org%2F10.3390%2Frs12193151?_sg%5B0%5D=0xE-347r7Hvb80klJcEo811AhUiXq-twG_E6l4yB-BfIKkVtW-lVLGcO02mTFkUczvozYYI0WCPyUBFR3kzWNGGZKg.ftvLheFrzHIJriO4qW2bdxalvR_TWt3MpwUfvto3EemhRgvDRGwJ9Mdy4Xr0IcGCfICivf4j-VqTgKxVvXRogA">10.3390/rs12193151</a>). </p> <p>The test dataset presented here includes different sets of co-located temperature and salinity vertical profiles: </p> <ul> <li>in situ observations extracted from the quality controlled Argo and CTD profiles produced by Copernicus Marine Environment Monitoring Service CORA 5.2 (<a href="http://marine.copernicus.eu/services-portfolio/access-to-products/">http://marine.copernicus.eu/services-portfolio/access-to-products/</a>, product_id: INSITU_GLO_TS_REP_OBSERVATIONS_013_001_b, doi: 10.17882/46219TS1, Szekely et al., 2019) and interpolated through a spline on a regularly spaced vertical grid (with 10 m intervals);</li> <li>climatological profiles extracted from World Ocean Atlas 2013 optimally interpolated monthly fields (Locarnini et al., 2013; Zweng et al., 2013), interpolated through a spline on a regularly spaced vertical grid (with 10 m intervals), upsized to a 1/10° horizontal grid through a cubic spline and linearly interpolated in time between the central day of each month;</li> <li>synthetic profiles obtained through three different techniques: multivariate EOF reconstruction, a 2 layer feed-forward network (with 1000 units in each hidden layer) and a stacked LSTM network (with 2 LSTM layers and 35 hidden units)</li> </ul> <p><em>References:</em></p> <p>Buongiorno Nardelli, B.: A Deep Learning network to retrieve ocean hydrographic profiles from combined satellite and in situ measurements, 2020, <em>submitted</em>.</p> <p>Locarnini, R. A., Mishonov, A. V., Antonov, J. I., Boyer, T. P., Garcia, H. E., Baranova, O. K., Zweng, M. M., Paver, C. R., Reagan, J. R., Johnson, D. R., Hamilton, M. and Seidov, D.: World Ocean Atlas 2013. Vol. 1: Temperature., S. Levitus, Ed.; A. Mishonov, Tech. Ed.; NOAA Atlas NESDIS, 73(September), 40, doi:10.1182/blood-2011-06-357442, 2013.</p> <p>Szekely, T., Gourrion, J., Pouliquen, S. and Reverdin, G.: The CORA 5.2 dataset for global in situ temperature and salinity measurements: Data description and validation, Ocean Sci., 15(6), 1601–1614, doi:10.5194/os-15-1601-2019, 2019.</p> <p>Zweng, M. M., Reagan, J. R., Antonov, J. I., Mishonov, A. V., Boyer, T. P., Garcia, H. E., Baranova, O. K., Johnson, D. R., Seidov, D. and Bidlle, M. M.: World Ocean Atlas 2013, Volume 2: Salinity, NOAA Atlas NESDIS, 119(1), 227–237, doi:10.1182/blood-2011-06-357442, 2013.</p> <p> </p>
MEG dataset nonlinguistic auditory statistical learning
<p>MEG data of 24 healthy adults with an auditory nonlinguistic statistical learning paradigm plus data from two subsequent behavioral tasks. For closer description of data see data description file. </p>
DeepLabCut: markerless pose estimation of user-defined body parts with deep learning
<p>This data entry contains <strong>annotated mouse data from the <a href="https://www.nature.com/articles/s41593-018-0209-y">DeepLabCut Nature Neuroscience paper</a></strong>.</p> <p>This data entry contains a public release of annotated mouse data from the DeepLabCut paper. The trail-tracking behavior is part of an investigation into odor guided navigation, where one or multiple wildtype (C57BL/6J) mice are running on a paper spool and following odor trails. These experiments were carried out by Alexander Mathis & Mackenzie Mathis in the Murthy lab at Harvard University. </p> <p>Data was recorded by two different cameras (640×480 pixels with Point Grey Firefly (FMVU-03MTM-CS), and at approximately 1,700×1,200 pixels with Grasshopper 3 4.1MP Mono USB3 Vision (CMOSIS CMV4000-3E12)) at 30 Hz. The latter images were cropped around mice to generate images that are approximately 800×800. </p> <p>Here we share 1066, frames from multiple experimental sessions observing 7 different mice. Pranav Mamidanna labeled the snout, the tip of the left and right ear as well as the base of the tail in the example images. The data is organized in <a href="https://www.nature.com/articles/s41596-019-0176-0">DeepLabCut 2.0 project structure</a> with images and annotations in the labeled-data folder. The names are pseudocodes indicating mouse id and session id, e.g. m4s1 = mouse 4 session 1.</p> <p>Code for loading, visualizing & training deep neural networks available at <a href="http://https://github.com/DeepLabCut/DeepLabCut"> https://github.com/DeepLabCut/DeepLabCut</a>.</p>
Dataset for "Machine Learning Stability and Bandgaps of Lead-Free Perovskites for Photovoltaics"
<p>Datasets used in the publication "Machine Learning Stability and Bandgaps of Lead-Free Perovskites for Photovoltaics" [doi:10.1002/adts.201900178].</p> <p>All structures were relaxed with the following parameters using Quantumwise QATK 2017:</p> <p>- SG15-GGA norm-conserving (Vanderbilt) pseudopotentials employed in a LCAO-approach (200 Hartree cutoff)<br> - 2x1x2-cubic-perovskite-supercells, relaxed from cubic 11.4Åx5.7Åx11.4Å-structures (forces < 0.01eV/Å)<br> - 300K Fermi-Dirac-smearing<br> - a 6x12x6 k-point grid (Monkhorst-Pack)</p> <p><br> Specifically, the included files are:</p> <p><strong>db_2.data: </strong>the actual database used for model building (json-format)<br> <strong>lead_set.data:</strong> the "external" test set used to test predictive power with out of sample compounds (json-format)<br> <strong>load_stanley_c.py:</strong> a python script to parse the .json-files to a python-dictionary including the structures (relaxed and unrelaxed) as <a href="https://gitlab.com/ase/ase">ASE</a>-atoms</p> <p>The format of the datafiles is as follows (-1 generally denote values not parsed from the raw data):<br> {<br> "<idstring>" : {<br> "trajectory" : n/a,<br> "energy" : total DFT energy in eV,<br> "rstruc" : relaxed structure, 3-tuple: (cell-vectors, scaled_positions, elements),<br> "gaps" : { "opt_gap", "ind_gap } - both direct and indirect gap,<br> "effective_mass" : n/a,<br> "iterations" : number of relaxation steps,<br> "calc" : some calculation metadata,<br> "ustruc" : unrelaxed input structure,<br> <br> }<br> }<br> Missing ids relate to structures filtered out, because the calculation didn't converge.</p> <p>Some code which works with a different representation of this data can be found at https://github.com/jstanai/Machine-Learning-Perovskite-Properties-for-Photovoltaics</p> <p> </p> <p> </p>
GAP-20 machine learning force field for phosphorus
<p>This dataset contains the force-field parameter files and reference database described in the manuscript "A general-purpose machine-learning force field for bulk and nanostructured phosphorus" (to be published).</p>
Data for: Machine learning identifies robust matrisome markers and regulatory mechanisms in cancer
<p>The expression and regulation of matrisome genes - the ensemble of extracellular matrix, ECM, ECM-associated proteins and regulators as well as cytokines, chemokines and growth factors - is of paramount importance for the many biological processes and signals within the tumor microenvironment. The availability of large and diverse multi-omics data enables mapping and understanding the regulatory circuitry governing the tumor matrisome to an unprecedented level, though such a volume of information requires robust approaches to data analysis and integration. In this study, we show that combining Pan-Cancer expression data from The Cancer Genome Atlas (TCGA) with genomics, epigenomics and microenvironmental features from TCGA and other sources enables the identification of “landmark” matrisome genes and machine learning-based reconstruction of their regulatory networks in 74 clinical and molecular subtypes of human cancers and approx. 6700 patients. These results, enriched for prognostic genes and cross-validated markers at the protein level, unravel the role of genetic and epigenetic programs in governing the tumor matrisome and allow the prioritization of tumor-specific matrisome genes (and their regulators) for the development of novel therapeutic approaches.</p>
Systematic Data Analysis and Diagnostic Machine Learning Reveal Differences between Compounds with Single- and Multitarget Activity
<p>The deposited files contain balanced data sets of multi-target (MT) and single-target (ST) compounds (CPDs) used for machine learning studies (https://dx.doi.org/10.1021/acs.molpharmaceut.0c00901). The first file (st_mt_data.tsv) contains 15,142 MT- and 15,081 ST-CPDs and the second (st_dt_data.tsv) 1828 DT- and 1776 ST-CPDs. For each CPD, a nonstereo_aromatic_SMILES representation, the original ChEMBL_cid, UniProt (target) IDs, and CPD category (CPD_CAT) (i.e. DT/MT/ST) is provided. DT stands for 'diverse-target' and denotes a subset of MT-CPDs (as detailed in the publication). In addition, a CPD is tagged “Y” if it continued to be present in the data set after removal of 50% randomly selected CPDs or 50% CPD nearest neighbors (NN), respectively.</p>
Adele 3D seismic survey segy format used in the FORCE 2020 machine learning competition for fault identification
<p>Adele seismic 3D survey segy format used in the FORCE 2020 machine learning competition for fault identification.</p> <p>Dataset is courtesy of GEOSCIENCE Australia who need to be acknowledged in each publication</p> <p> </p>
A Deep Learning Dataset for Tomato Pest Leafminer TUTA ABSOLUTA
<p>The images of tomato leafminer (<em>Tuta absoluta</em>) were taken in in-house plots between August 2018 and May 2019 in Arusha, Tanzania. Under net-house that were controlled from other others. <em>T.absoluta</em> larvae were inoculated on the commonly grown tomato varieties at the early growth stage (herein, on the second day after transplanting). The images were taken for the first 2 weeks after inoculation. Images captured the canopy of the plants. </p> <p><strong>File Description</strong><br> All Images are in the <strong>.zip</strong> files; "dataset_1_H.zip" has 1926 Images, dataset_1_NH.zip has 325 Images, dataset_2 .zip has 3482 Images and the files labels are in "file_labels.csv" the image file name in column "FileName" and respective label in column "Label", labels meaning "1" refer to healthy (plants not inoculated with <em>T.absoluta</em> larvae and "2" refer to <em>T.absoluta</em> affected plants. A total of 4341 image files are labelled. </p> <p> </p>
Airborne Radar Quality Control with Machine Learning
<p>This repository contains the radar data collected by ELDORA required to train and test the random forest model discussed in "Airborne radar quality control with machine learning" by Alexander DesRosiers and Michael M. Bell at the Colorado State University Department of Atmospheric Science. The model used in the manuscript is also contained in a '.pkl' file. Upon publication, a link to the paper will be provided here. Finer points of the methodology were discussed in the manuscript and the python script (make_radarQC_rf_model.py) is commented to guide users through the process of creating the model.</p>
Multi-fidelity Generative Deep Learning Turbulent Flows
<p>Data sets for the two numerical examples in the paper <a href="https://arxiv.org/abs/2006.04731">Multi-fidelity Generative Deep Learning Turbulent Flows</a> as well as two pre-trained models. In this work, a novel multi-fidelity deep generative model is introduced for the surrogate modeling of high-fidelity turbulent flow fields given the solution of a computationally inexpensive but inaccurate low-fidelity solver. The resulting surrogate is able to generate physically accurate turbulent realizations at a computational cost magnitudes lower than that of a high-fidelity simulation. The deep generative model developed is a conditional invertible neural network, built with normalizing flows, with recurrent LSTM connections that allow for stable training of transient systems with high predictive accuracy. Data is provided from OpenFOAM LES simulations for turbulent flow over backwards step and flow around an array of cylinders.</p> <p>Data-set Files:</p> <ul> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/backward_step_testing.tar.gz?versionId=320a523f-0015-4ba3-8c6e-66733ab5a1af">backward_step_testing.tar.gz</a> - Backward step testing data.</li> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/backward_step_training.tar.gz?versionId=ac34ac15-973d-4fb8-8881-faa17eced69f">backward_step_training.tar.gz</a> - Backward step training data.</li> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/cylinder_array_testing.tar.gz?versionId=ebbea725-c8b6-4338-978b-dc73f943552e">cylinder_array_testing.tar.gz</a> - Cylinder array testing data.</li> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/cylinder_array_training.tar.gz?versionId=b0409c34-fb19-45c4-bbba-6968fdcfc4d8">cylinder_array_training.tar.gz</a> - Cylinder array training data.</li> </ul> <p>Pre-trained Models:</p> <ul> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/bstepWorkspace400.zip">bstepWorkspace400.zip</a> - Backward step pre-trained model.</li> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/cylinderWorkspace400.zip">cylinderWorkspace400.zip</a> - Cylinder array pre-trained model.</li> </ul> <p> </p>
Combinatorial and machine learning approaches for the analysis of Cu2ZnGeSe4: influence of the off-stoichiometry on defect formation and solar cell performance
<p>Dataset of the results published in the <a href="https://zenodo.org/record/4742379#.YMzExOgzYmJ">J. Mater. Chem. A, 2021, 9, 10466</a>. The files represent: i) the measured compositional and optoelectronic data of each solar cell, as well as the data generated from the Raman spectra analysis; ii) Raman spectra of the representative cells; iii) Machine Learning discriminants.</p> <p>The elemental composition of the different cells of the combinatorial sample was determined by X-ray fluorescence (XRF) using a Fischerscope XDV system with a 1 mm spot diameter, a 50 kV acceleration voltage, a Ni10 lter and a 45 s acquisition time. Raman analysis with blue (442 nm) and green (532 nm) excitation wavelengths were performed on the bare absorber, while measurements with NIR (785 nm) were performed in complete devices using Horiba Jobin Yvon FHR640 and iHR320 monochromators coupled with CCD detectors. The first monochromator is optimized for the UV and visible spectral ranges and was used with 442 nm (He–Cd gas laser) and 532 nm (solid state laser) excitation wavelengths. The second monochromator is optimized for the NIR range and was used with a 785 nm (solid state laser) excitation wavelength. The power density of the lasers was kept below 150 W cm<sup>2</sup> and the spot size was ~70 <span class="math-tex">\(\mu\)</span>m. The measurements were performed in a backscattering configuration through a specific probe designed at IREC. The J–V characteristics of the devices were obtained under simulated AM1.5 illumination (1000 W m2 intensity at room temperature) using a pre-calibrated Class AAA solar simulator (Abet Technologies Sun 3000).</p>
Data and scripts related to: Rapid coordination of effective learning by the human hippocampus
<p>This data set contains intracranial EEG data (ASCII format), eye-tracking data from an EyeLink 1000 remote system (edf format), behavioral data, and MATLAB code to reproduce the analyses reported in the manuscript, “Rapid coordination of effective learning by the human hippocampus” published in <em>Science Advances.</em></p> <p>The file <strong>KragelEtal21_SciAdv.zip</strong> contains the raw data divided into folders according to content type, for each of the six participants in the study, and the MATLAB code necessary to reproduce all analyses. MATLAB live scripts provide examples of how to reproduce the main analyses reported in the manuscript.</p> <p>External datasets:</p> <p>In addition to the dataset provided here, three open-access datasets are analyzed in the manuscript.</p> <p> - The <a href="http://figrim.mit.edu/">FIGRIM Dataset</a> contains eye-tracking data during a continuous recognition task.</p> <p> - Two additional eye-tracking datasets during free viewing of repeated scenes are provided in “<a href="https://datadryad.org/stash/dataset/doi:10.5061/dryad.9pf75">An extensive dataset of eye movements during viewing of complex images</a>,” namely the Memory I and Memory II datasets.</p> <p>To reproduce region of interest analyses outside of the hippocampus, both the seven-network cortical parcellation developed by <a href="https://surfer.nmr.mgh.harvard.edu/fswiki/CorticalParcellation_Yeo2011">Yeo, Krienen et al.</a>, and the <a href="https://identifiers.org/neurovault.image:1702">Harvard-Oxford cortical atlas</a> are required.</p> <p>Stimuli:</p> <p>The scenes used in this study are part of <a href="https://cocodataset.org">Microsoft COCO</a>. Scenes were selected from the 2017 Train images. Image identifiers are maintained.</p> <p>Salience model:</p> <p>To reproduce analyses that consider the visual salience of each scene, DeepGaze II model predictions for each stimulus are required. Tensorflow models and a Jupyter notebook demonstrating their use are available for <a href="https://deepgaze.bethgelab.org/">download</a>.</p> <p>Software dependencies:</p> <p>The code in this project was developed using MATLAB r2017b. The following external packages are required for code execution. Some external packages are included in the repository.</p> <p>- fieldtrip (<a href="https://github.com/fieldtrip/fieldtrip">https://github.com/fieldtrip/fieldtrip</a>)<br> - spm12 (<a href="https://github.com/spm/spm12">https://github.com/spm/spm12</a>)<br> - BOSC (<a href="https://doi.org/10.1016/j.neuroimage.2010.08.064">https://doi.org/10.1016/j.neuroimage.2010.08.064</a>)<br> - Edf2Mat (<a href="https://github.com/uzh/edf-converter">https://github.com/uzh/edf-converter</a>)<br> - boundedline (<a href="https://github.com/kakearney/boundedline-pkg">https://github.com/kakearney/boundedline-pkg</a>)<br> - export_fig (https://github.com/altmany/export_fig)</p> <p>License:</p> <p>The included code is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, either version 3 of the License, or any later version. See the file COPYING for more details. The release of this software includes functions from other toolboxes that are covered under their respective licenses.</p>
Galaxy Zoo DECaLS: Detailed Visual Morphology Measurements from Volunteers and Deep Learning for 314,000 Galaxies
<p>This repository contains the data released in the paper "Galaxy Zoo DECaLS: Detailed Visual Morphology Measurements from Volunteers and Deep Learning for 314,000 Galaxies" <em>(DOI to follow on publication).</em></p> <p>We release detailed morphology catalogues, both volunteer and automated, for Galaxy Zoo DECaLS.</p> <p>- gz_decals_volunteers_1_and_2 contains volunteer classifications for galaxies classified during the GZD-1 and GZD-2 campaigns.</p> <p>- gz_decals_volunteers_5 similarly contains classifications from the GZD-5 campaign. Note that GZD-5 used a modified schema designed to better detect mergers and weak bars, and includes many galaxies with only approx. five volunteer responses.</p> <p>- gz_decals_auto_posteriors contains the predicted posteriors for volunteer responses to all galaxies used in any campaign. The full posteriors are recorded as Dirichlet distribution concentrations. gz_decals_auto_posteriors also summarises these posteriors as the automated equivalent of previous Galaxy Zoo data releases;<strong> the expected vote fractions (mean posteriors)</strong>. Note that not all posteriors/vote fractions are relevant for every galaxy; we suggest assessing relevance using the estimated fraction of volunteers that would have been asked each question.</p> <p>We include a schema document, schema.md, to define the column names in each catalogue.</p> <p>We also release the galaxy images shown to volunteers on www.galaxyzoo.org during GZD-5. The images on which the automated classifier was trained may be derived from these volunteer-facing images. These images are split into four zip files, each of which contains images named by iauname inside a subfolder named by the first four characters in their iauname. Not all images were labelled during GZD-5 - refer to the catalog for training labels. We are working with the Zenodo team to add these large files to this repository - meanwhile, you can download them from The University of Manchester <a href="https://docs.google.com/document/d/1YgpnxiSJ7ffOW6FY8pX0pw93LTu8rLIdPL2PYhxW1fo/edit?usp=sharing">here</a>.</p> <p>The .csv and .parquet files contain identical data. Parquet is a fast column-oriented binary format which can be read with pd.read_parquet(loc, columns=[some columns]).</p> <p>You may also be interested in the <a href="https://github.com/mwalmsley/zoobot">github repository</a> which contains code to reproduce the model and to fine-tune it for new tasks (including pretrained weights).</p> <p>We will release updates if needed via Zenodo versioning. We recommend using the latest version of this repository. You can check the version you are currently viewing on the right-hand sidebar.</p> <p>Please cite the paper (DOI to follow on publication) when using the data in this repository.</p> <p>---</p> <p>History</p> <p>v0.0.1 (submission) provides the catalog files.</p> <p>v0.0.2 (first revision) renames the catalog files, adds flags for poorly sized galaxies, and includes the galaxy images via the University of Manchester</p>
Dataset: Reinforcing Cybersecurity Hands-on Training With Adaptive Learning
<p>This repository contains supplementary materials for the following conference paper:<br> <br> Pavel Seda, Jan Vykopal, Valdemar Švábenský, Pavel Čeleda.<em><br> Reinforcing Cybersecurity Hands-on Training With Adaptive Learning. </em><br> In Proceedings of the 51st IEEE Frontiers in Education Conference (FIE 2021).<br> <a href="https://doi.org/10.1109/FIE49875.2021.9637252">https://doi.org/10.1109/FIE49875.2021.9637252</a><br> <br> Preprint available at: <a href="https://arxiv.org/abs/2201.01574">https://arxiv.org/abs/2201.01574</a></p> <p><strong>How to cite</strong></p> <p>If you use or build upon the materials, please use the BibTeX entry below to cite the original paper (not only this web link).</p> <p>Some of the linked repositories have their separate citation entry; please use that one as well, if possible.</p> <pre><code>@inproceedings{Seda2021reinforcing, author = {Seda, Pavel and Vykopal, Jan and \v{S}v\'{a}bensk\'{y}, Valdemar and \v{C}eleda, Pavel}, title = {{Reinforcing Cybersecurity Hands-on Training With Adaptive Learning}}, booktitle = {Proceedings of the 51st IEEE Frontiers in Education Conference}, series = {FIE '21}, location = {Lincoln, NE, USA}, publisher = {IEEE}, address = {New York, NY, USA}, month = {10}, year = {2021}, pages = {1--9}, numpages = {9}, isbn = {978-1-6654-3851-3}, url = {https://doi.org/10.1109/FIE49875.2021.9637252}, doi = {10.1109/FIE49875.2021.9637252}, }</code></pre> <p> </p>
Results from the RDM Survey - LEARN project (December 2016)
<p>Data obtained from the open survey developed by the LEARN project (http://www.learn-rdm.eu/) as a self-assessment tool to assist institutions discover how ready they are for managing research data. This dataset replaces the first one published at http://doi.org/10.5281/zenodo.61903. The survey is based on the issues posed to institutions by the LERU Roadmap for Research Data published at the end of 2013, and available at: http://www.learn-rdm.eu/material/leru_roadmap_for_research_data<br> The survey has thirteen questions addressing the main elements to be taken into account in developing an institutional strategy for research data management. Each question has three possible answers representing green, yellow or red light. The more ‘green light’ responses recorded, the readier an institution probably is for managing its research data.</p> <p>The survey is available in English at http://learn-rdm.eu/en/rdm-readiness-survey/ and in Spanish at http://learn-rdm.eu/encuesta-rdm/</p>
Negative Sampling Improves Hypernymy Extraction Based on Projection Learning
<p>We present a new approach to extraction of hypernyms based on projection learning and word embeddings. In contrast to classification-based approaches, projection-based methods require no candidate hyponym-hypernym pairs. While it is natural to use both positive and negative training examples in supervised relation extraction, the impact of negative examples on hypernym prediction was not studied so far. In this paper, we show that explicit negative examples used for regularization of the model significantly improve performance compared to the state-of-the-art approach on three datasets from different languages.</p> <p>The <strong>russian</strong> model.</p> <p>$ python -V; pip show tensorflow numpy scipy scikit-learn gensim | egrep -i '(name|version)'<br> Python 3.5.2 :: Continuum Analytics, Inc.<br> Name: tensorflow<br> Version: 0.12.1<br> Name: numpy<br> Version: 1.12.0<br> Name: scipy<br> Version: 0.18.1<br> Name: scikit-learn<br> Version: 0.18.1<br> Name: gensim<br> Version: 0.13.4.1</p> <p>The <strong>english</strong><strong>-combined</strong> model has been trained using the well-known word embeddings dataset based on Google News: GoogleNews-vectors-negative300.bin on EVALution, BLESS, K&H+N, ROOT09 combined. The <strong>english</strong><strong>-</strong><strong>evalution</strong> model is traned on EVALution only.</p> <p>$ python -V; pip show tensorflow numpy scipy scikit-learn gensim | egrep -i '(name|version)'<br> Python 3.5.2 :: Anaconda custom (64-bit)<br> Name: tensorflow<br> Version: 0.12.1<br> Name: numpy<br> Version: 1.11.3<br> Name: scipy<br> Version: 0.18.1<br> Name: scikit-learn<br> Version: 0.18.1<br> Name: gensim<br> Version: 0.13.4.1</p>
Dataset of experimental measurements for "Demonstration of quantum advantage in machine learning"
<p>Dataset of experimental measurements for "Demonstration of quantum advantage in machine learning", <em>npj Quantum Information</em><strong> 3</strong>, Article number: 16 (2017).</p>
Dataset supporting "Using Machine Learning to decide when to Precondition Cylindrical Algebraic Decomposition with Groebner Bases"
<p>Dataset supporting the paper:</p> <p>Z. Huang, M. England, J.H. Davenport and L.C. Paulson<br> Using Machine Learning to decide when to Precondition Cylindrical Algebraic Decomposition with Groebner Bases.<br> Proceedings of the 18th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC '16), pp. 45--52. IEEE, 2016. Digital Object Identifier: 10.1109/SYNASC.2016.020 </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.