Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Weekly supervised Multilingual Data Set to train Named Entity Recognition for Symptom Extraction
<p>Data Sets were generated using the Weakly Supervised NER pipeline (https://github.com/HUMADEX/Weekly-Supervised-NER-pipline) to train the symptom extraction NER models. </p> <p><strong>Supported Languages and dataset locations for the specific language:</strong></p> <p> English (base language): https://huggingface.co/HUMADEX/english_medical_ner<br> German: https://huggingface.co/HUMADEX/german_medical_ner<br> Italian: https://huggingface.co/HUMADEX/italian_medical_ner<br> Spanish: https://huggingface.co/HUMADEX/spanish_medical_ner<br> Greek: https://huggingface.co/HUMADEX/german_medical_ner<br> Slovenian: https://huggingface.co/HUMADEX/slovenian_medical_ner<br> Polish: https://huggingface.co/HUMADEX/polish_medical_ner<br> Portuguese: https://huggingface.co/HUMADEX/portugese_medical_ner</p> <p> </p> <p><strong>Dataset Building </strong></p> <ul> <li>Data Integration and Preprocessing</li> <li>Data Cleaning</li> <li>Annotation with Stanza's i2b2 Clinical Model </li> <li>Translation into the targeted language</li> <li>Word Alignment </li> <li>Data Augmentation </li> </ul> <p><strong>Acknowledgement</strong><br>This dataset had been created as part of joint research of HUMADEX research group (https://www.linkedin.com/company/101563689/) and has received funding by the European Union Horizon Europe Research and Innovation Program project SMILE (grant number 101080923) and Marie Skłodowska-Curie Actions (MSCA) Doctoral Networks, project BosomShield ((rant number 101073222). Responsibility for the information and views expressed herein lies entirely with the authors.</p> <p><strong>Authors:</strong><br>dr. Izidor Mlakar, Rigona Sallauka, dr. Umut Arioz, dr. Matej Rojc</p> <p><strong>Please cite as:</strong></p> <p><span>Article title: Weakly-Supervised Multilingual Medical NER For Symptom Extraction For Low-Resource Languages</span><br><span>Doi: 10.20944/preprints202504.1356.v1</span><br><span>Website: </span><a title="https://www.preprints.org/manuscript/202504.1356/v1" href="https://www.preprints.org/manuscript/202504.1356/v1">https://www.preprints.org/manuscript/202504.1356/v1</a></p>
CO2 fluxes measurement data set for the lower segment of the River Noce, Italy
<p>This data set reports on data recorded during four measurement campaigns in the lower segment of the River Noce, Italy.<br>The measurements were aimed at characterising CO2 fluxes along the river, and were conducted using a combination of loggers (time-resolved data) and sampling (space-resolved data).<br>The data were collected in the Santa Giustina reservoir (upstream reservoir), along the residual flow reaches that extend between the Santa Giustina reservoir and the outlet of the hydropower diversion system in Mezzocorona, and downstream of the diversion outlet.</p> <p>The data is provided as .csv and Matlab .mat files.<br>For more information about the data set and data acquisition, please contact the author or refer to the following article:</p> <p><br>Dolcetti, G., Piccolroaz, S., Bruno, M. C., Calamita, E., Larsen, S., Zolezzi, G. & Siviglia, A. Quantification of carbopeaking and CO2 fluxes in a regulated Alpine river.</p>
Mastodon example data set - 10 tracked phallusia mammillata embryos
<ul> <li>This dataset originates from the Phallusia mammillata embryonic development dataset, which was published on figshare. <ul> <li><a href="https://figshare.com/projects/Phallusiamammillata_embryonic_development/64301" rel="nofollow">https://figshare.com/projects/Phallusiamammillata_embryonic_development/64301</a>.</li> </ul> </li> <li>The dataset is described in the following publication: <ul> <li>Léo Guignard et al., Contact area–dependent cell communication and the morphological invariance of ascidian embryogenesis. Science369, eaar5663(2020). DOI:<a href="https://doi.org/10.1126/science.aar5663" rel="nofollow">10.1126/science.aar5663</a></li> </ul> </li> <li>The dataset has been converted to Mastodon file format using these scripts: <ul> <li>Astec file format to CSV: <a href="https://github.com/GuignardLab/LineageTree/blob/master/src/LineageTree/legacy/export_csv.py">https://github.com/GuignardLab/LineageTree/blob/master/src/LineageTree/legacy/export_csv.py</a></li> <li>CSV to Mastodon file format: <a href="https://github.com/mastodon-sc/mastodon-deep-lineage/blob/master/src/test/java/org/mastodon/mamut/astec/AstecReader.java">https://github.com/mastodon-sc/mastodon-deep-lineage/blob/master/src/test/java/org/mastodon/mamut/astec/AstecReader.java</a></li> </ul> </li> <li>Limitations resulting from the conversion: <ul> <li>Image data is not included in the projects. This would require additional effort. However, it is probably not necessary for the intended use cases (tests of lineage clustering, lineage registration, and Blender visualization).</li> <li>Only datasets Pm01 – Pm10 have been converted.</li> <li>The cell names were missing for Pm11-Pm13. To handle this, the conversion scripts would need to be adapted. Thus, these datasets are not included.</li> <li>The uploaded data for Pm14 is faulty on figshare. Thus it is not included.</li> <li>Half-Pm1, U0126PM1 and U0126PM2 do not contain lineages or coordinates. Thus, these datasets are not included.</li> <li>When opening the Mastodon projects, a message appears stating that the image data is missing. It can be chosen, to open the project without image data ("dummy image data"). This must be confirmed. After that, it should work.</li> </ul> </li> </ul>
Cubic insulin data set collected from multiple lattices on i03 at Diamond Light Source
<p>Dataset collected as part of routine commissioning work, found to have more than one crystal present at the point where data were collected, allowing multiple lattices to be processed. </p> <p> </p> <p>While three lattices are present one is substantially weaker than the other two.</p> <p> </p> <p>Uploading to enable methods development and also to use for tutorials on how to use dials software.</p> <p> </p> <p>Processing data with the usual dials scripts (which will point to this deposition) result in statistics shown below. Tutorial to be uploaded to https://github.com/graeme-winter/dials_tutorials when available.</p> <p> </p> <p><code> -------------Summary of merging statistics-------------- </code></p> <p><code> Suggested Low High Overall</code><br><code>High resolution limit 1.51 4.10 1.51 1.48</code><br><code>Low resolution limit 54.89 54.93 1.54 54.89</code><br><code>Completeness 98.9 100.0 83.7 95.3</code><br><code>Multiplicity 54.7 78.1 4.6 53.5</code><br><code>I/sigma 18.9 89.2 0.3 18.4</code><br><code>Rmerge(I) 0.139 0.060 1.423 0.139</code><br><code>Rmerge(I+/-) 0.138 0.060 1.316 0.138</code><br><code>Rmeas(I) 0.140 0.061 1.590 0.140</code><br><code>Rmeas(I+/-) 0.140 0.060 1.609 0.140</code><br><code>Rpim(I) 0.016 0.007 0.683 0.016</code><br><code>Rpim(I+/-) 0.023 0.009 0.894 0.023</code><br><code>CC half 1.000 1.000 0.271 1.000</code><br><code>Anomalous completeness 97.8 100.0 67.8 92.3</code><br><code>Anomalous multiplicity 28.7 43.6 2.6 28.3</code><br><code>Anomalous correlation 0.038 0.276 -0.050 0.049</code><br><code>Anomalous slope 0.667 </code><br><code>dF/F 0.060 </code><br><code>dI/s(dI) 0.629 </code><br><code>Total observations 669615 52255 2342 670277</code><br><code>Total unique 12231 669 507 12529</code></p>
Data set of study Examining the impact of cognitive control and discrimination on mental health outcomes in diverse Pakistani and Afghan communities
<div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <p>This dataset is part of a study examining the impact of cognitive control and discrimination on mental health outcomes among diverse Pakistani and Afghan communities. It includes demographic variables such as age, gender, and ethnicity, as well as psychological measures assessing cognitive control, discrimination experiences, and various mental health outcomes.</p> </div> </div> </div> </div> <div> <div> <div> <div> </div> <div><span>4o</span></div> <span></span></div> </div> </div> <div> </div> <div> <div> </div> </div> </div> <div> <div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> <div> <div> <div> <div> <div> <div> </div> <div> <div> </div> </div> <div> <div> <div> <div> <div> <div> <div> <div> </div> </div> </div> </div> </div> <div> <div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div>
Data set of study The role of REM sleep in consolidating executive function: a prefrontal cortex perspective
Open the record for dataset details and reuse information.
Data sets for the MSCA project ElectronsOnRun
<p>Data created during MSCA project ElectronsOnRun. </p>
DWUG DE Sense: A data set of historical word sense annotations in German
<p>This data collection contains a subset of <a href="https://zenodo.org/record/5543723">DWUG DE</a> word usage data annotated with classical word sense definitions (<em>DWUG DE Sense</em>, see <code>data/*/judgments_senses.csv</code>). From these annotations aggregated and cleaned sense labels were derived (<code>labels/*/labels_senses.csv</code>). From these labels we derived additional binary semantic proximity labels between use pairs ('0' for different sense, '1' for same sense, <code>labels/*/labels_proximity.csv</code>) and change labels reflecting sense changes between the two time periods from which word usages were sampled (<code>stats/*/stats_groupings.csv</code>).</p> <p>The sense labels were derived from the sense annotation by removing instances where not at least 2/3 annotators agree on the label (<code>maj_2</code>/<code>maj_3</code>). Note that the binary proximity labels were <em>derived</em> from the sense annotation, and not directly judged by humans (in contrast to other <a href="https://www.ims.uni-stuttgart.de/data/wugs">WUG data sets</a>). Note that consequently also the change scores EARLIER, LATER and COMPARE were not calculated directly from human judgments, but from the inferred binary proximity labels. Please find the code aggregating and cleaning the data, deriving proximity labels and deriving change labels in the <a href="https://github.com/Garrafao/WUGs">WUG repository</a>.</p> <p>Please find more information on the provided data in the paper referenced below.</p> <p>Version: 1.0.1, 01.11.2024. Correct or remove some normalization and lemmatization errors in the uses. Updated references.</p> <h3>Reference</h3> <p>Dominik Schlechtweg, Frank D. Zamora-Reina, Felipe Bravo-Marquez, Nikolay Arefyev. 2024. <a href="https://doi.org/10.1007/s10579-024-09771-7">Sense Through Time: Diachronic Word Sense Annotations for Word Sense Induction and Lexical Semantic Change Detection</a>. Language Resources and Evaluation.</p> <p>Dominik Schlechtweg. 2023. <a href="http://dx.doi.org/10.18419/opus-12833">Human and Computational Measurement of Lexical Semantic Change</a>. PhD thesis. University of Stuttgart.</p>
Data set to: Spawning is accompanied by increased thermal performance in blue mussels
<p>Data for the manuscript "Spawning is accompanied by increased thermal performance in blue mussels".</p>
Data set for Site Specific Lewy Pathology Circuits
<p>Processed dataset and example code related to manuscript titled: Site-Specific Seeding of Lewy Pathology Induces Distinct Pre-Motor Cellular and Dendritic Vulnerabilities in the Cortex</p>
Data sets of the analysis of 1sg and 2sg subject expression with the verbs 'creer' and 'saber' in a corpus of spoken Spanish
<p>These are the two data sets used for the quantitative analysis in the paper "“Perspectival factors and <em>pro</em>-drop – A corpus study of speaker/addressee pronouns with <em>creer </em>‘think/believe’ and <em>saber </em>‘know’ in spoken Spanish” (Peter Herbeck;<em> to appear </em>in <em>Glossa – a journal of general linguistics</em>). It contains the values of the annotation of null and overt speaker/addressee pronouns in the Madrid and Alcalá samples of the corpus PRESEEA (2014-). The first data set was used for the quantitative analysis of subject expression (null/overt) with the verbs <em>creer </em>and <em>saber </em>according to person (1sg vs. 2sg) and polarity (negative vs. positive verb forms). The second data set was used for the analysis of 1sg subject expression according to the complement type of <em>creer </em>and <em>saber</em>.</p>
D4.3 Assessment of DRALOD drying plant performance data set dryer march 2021
<p>Data and meta-data for dryer operation March 2021</p>
Data set for the article "Recently photoassimilated Carbon and fungus-delivered Nitrogen are spatially correlated at the cellular scale in the ectomycorrhizal tissue of Fagus sylvatica"
<p>This dataset contains data that support the manuscript</p> <p>Mayerhofer et al (2021) "Recently photoassimilated Carbon and fungus-delivered Nitrogen are spatially correlated at the cellular scale in the ectomycorrhizal tissue of<em> Fagus sylvatica", </em>The New Phytologist, DOI:10.1111/nph.17591</p> <p>It contains the following data:</p> <p>(1) NanoSIMS imaging data, which was used for Fig. 4-7, is provided in NanoSIMS_control_root_tip.zip and NanoSIMS_labelled_root_tip.zip. Each zip-files contains:</p> <ul> <li>the original NanoSIMS images (.im)</li> <li>their related checkfiles (.chk_im)</li> <li>ROIs description (.rois.zip)</li> </ul> <p>of the unlabelled control and the labelled root tip section, respectively. ".im" and ".chk_im" are the original image data aquisition files from the NanoSIMS instrument. "rois.zip" files describe selected regions of interests and were created utilizing the OpenMIMS plugin (Center for Nano Imaging, https://nano.bwh.harvard.edu/MIMSsoftware) for the image analysis software ImageJ (National Institutes of Health, Bethesda, MD, USA).</p> <p>(2) Means and standard deviations of all measured elements and isotopes of each region of interest, as obtained via the .rois.zip files from the NanoSIMS images, are reported in NanoSIMS_ROI_data.csv (used for Fig.7).</p> <p>(3) Linescan_data.csv contains data used for Fig. 8.</p> <p>(4) IRMS_roots_data.csv contains data of root segments and mycorrhizal root tips analysed with isotope-ratio mass spectrometry (EA-IRMS) (used for Fig. 2)</p> <p>Description of column meanings from csv data files can be found in the according "_description" files.</p>
Coherent Electric Field Manipulation of Fe3+-spins in PbTiO3. Open data set
<p>Data supporting figures 3 and 4 of the related publication.</p>
DNA Stack Data Structure: Data Set of Polyacrylamide Gels
<p>Data set to accompany paper: A Last-In First-Out Stack Data Structure Implemented in DNA, Lopiccolo and Shirt-Ediss et al., Nature Comms 12, 4861 (2021). <a href="https://doi.org/10.1038/s41467-021-25023-6">doi:10.1038/s41467-021-25023-6</a>.</p> <p>The data set contains 32 polyacrylamide gel experiments (reaction sequences and gel images) used for creating the electrophoretic mobility graph in Figure S13 of the SI of the above paper.</p>
Simulation and analysis data set for apo-conformational kinetics and gated ligand binding to HIV-1 protease
<p>The data set provided here accompanies a study described in the manuscript:</p> <p>S. Kashif Sadiq, Abraham Muñiz Chicharro, Patrick Friedrich, Rebecca Wade, A multiscale approach for computing gated ligand binding from molecular dynamics and Brownian dynamics simulations. (2021) Preprint available: https://doi.org/10.1101/2021.06.22.449380</p> <p>This study combines molecular dynamics MD simulations and associated conformational analyses and Markov state models (MSMs) with Brownian dynamics (BD) simulations to compute conformation gated ligand association kinetics to HIV-1 protease.</p> <p>To download the data, go to a directory where you would like to download the files. Then for each of the provided tar files enter the following command:</p> <p>tar xvf $X.tar</p> <p>where $X is the name prefix of the corresponding tar file.</p> <p>The unpacked data set creates a ./data sub-directory which itself contains two further sub-directories: MD and BD. Please see README.txt files within these sub directories for further instructions on the software tools and scripts that have been provided therein for using and reproducing the data set. The MD README.txt is found within: data_MD_MSM_analysis.tar, the BD README.txt is found within: data_BD_examples.tar.</p> <p>Please note, the python Jupyter notebook and associated module for further analysis of the MSM from the pre-defined feature set calculated in the study as well as other analyses can also be found at:</p> <p>https://github.com/kashifsadiq/hiv1pr-msm/</p> <p>MD trajectory files are provided for further analysis but are not required to reproduce the MSM and conformational analyses reported in the study. To facilitate overview, MSM analysis has been stored in several object files. To exactly reproduce the reported MD/MSM analyses, untar only the 1) data_MD_MSM_analysis.tar and 2) data_MD_MSMobj.tar files and work through the python Jupyter notebook.</p> <p> </p>
Data set for methoprene and fenoxycarb's effects on FCM (unpublished)
<p>Data set for different concentrations of methoprene and fenoxycarb's effects on FCM (unpublished), measuring developmental and performance parameters.</p>
SPLReePlan Data set
<p>SPLReePlan Evaluation data set and documentation. </p>
Magnetotelluric data set of the northern Songliao Block, NE China
<p>The study area covers 120,000 square kilometers in NE China. Field collection of MT data was carried out with MTU-5 instruments, manufactured by Phoenix Geophysics of Canada. This study used a total of 138 broadband MT sites. The signal acquisition time was less than one day at each site. The signal acquisition time was about 20 hours at each site. Each MT site measured 5 components, 3 magnetic field components and 2 mutually orthogonal horizontal electric field components, approximately 10 km site spacing along several lines.</p>
Data set related to the manuscript "Mesoscopic simulations of the in situ NMR spectra of porous carbon based supercapacitors: Electronic structure and adsorbent reorganisation effects"
<p>Graphical files in the agr format for all the figures in the manuscript entitled "Mesoscopic simulations of the in situ NMR spectra of porous carbon based supercapacitors: Electronic structure and adsorbent reorganisation effects". Examples of input files for the lattice simulations are also provided.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.