Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11,174
datasets available to search
ShareScore release 0.7.1
Dataset results
11,174 results for “identifiers”
Calculated Parameters of Thyroid Homeostasis: de-identified Data
<p>De-identified data of patients with different thyroid conditions. TI: Thyrotropic insufficiency; TT: Thyrotoxicosis; PH: Primary hypothyroidism; ST: Secondary hyperthyroidism. Parameters are reported with units of measurement and reference interval.</p>
List of (tentatively) identified non-target structures
<p>This is a list of candidate structures of organic contaminants (tentatively) identified in a riverbank filtration system in The Netherlands by non-target screening of high-resolution mass spectrometry data (associated with ACS Publication <a href="https://doi.org/10.1021/acs.est.9b01750">10.1021/acs.est.9b01750</a>).</p>
Dataset of Horizon scanning to identify invasion risk of ornamental plants marketed in Spain
<p>Full dataset for the research entitled "Horizon scanning to identify invasion risk of ornamental plants marketed in Spain". We classified non-native species into six different lists based on their invasion status in Spain and elsewhere, their climatic suitability in Spain, and their potential environmental and socioeconomic impacts.</p>
Raw data supporting Identifying invertebrates from pitfall, flight interception traps and hand collecting
<p>Raw data supporting identifying invertebrates from pitfall, flight interception traps and hand collecting: comparing metabarcoding with traditional methods.</p> <p>Two step PCRs were performed on each sample replicate using modified primers mICOIintF and jgHCO2198 followed by the Nextera XT index kit v2 Set A (Illumina).</p> <p>The pool was loaded onto an illumina MiSeq using a MiSeq Reagent Kit v2 500 cycle kit (Illumina), with 10% Phi-X to generate 250-bp paired-end reads.</p>
French Word Sense Disambiguation with Princeton WordNet Identifiers
<p>This is a dataset for the Word Sense Disambiguation of French using Princeton WordNet identifiers. It contains two training corpora : the SemCor and the WordNet Gloss Corpus, both automatically translated from their original English version, and with sense tags automatically aligned. It contains also a test corpus : the task 12 of SemEval 2013, originally sense annotated with BabelNet identifiers, converted into Princeton WordNet 3.0.</p>
Identifiers with Images (EOL v2): identifiers_with_images.csv.gz
For questions or use cases calling for large, multi-use aggregate data files, please visit the EOL Services forum at <p></p>http://discuss.eol.org/c/eol-services
Identifying Harmful Marine Dinoflagellates: Harmful Dinoflags
Faust, Maria A. and Rose A. Gulledge. Identifying Harmful Marine Dinoflagellates. Smithsonian Contributions from the United States National Herbarium, volume 42: 1-144 (including 48 plates, 1 figure and 1 table).<p></p>Faust, Maria A. and Rose A. Gulledge. Identifying Harmful Marine Dinoflagellates. Smithsonian Contributions from the United States National Herbarium, volume 42: 1-144 (including 48 plates, 1 figure and 1 table).
UnientrezDB: Large-scale Gene Ontology Annotation Dataset and Evaluation Benchmarks with Unified Entrez Gene Identifiers
<p>Our work focuses on providing a comprehensive dataset and benchmarks for evaluating gene ontology annotations using a unified system of Entrez Gene Identifiers.</p>
Integration of expression datasets to identify biomarkers for accurate Gleason scoring in Prostate Cancer -- Supplementary data
<p>This dataset contains expression data from multiple sources used to identify biomarker candidates for prostate cancer aggressiveness. The data includes transcriptional expression levels, patient metadata, and other relevant features utilized in our machine-learning models. The training dataset was extracted from the repository described at Matos-Filipe, et al. (2022) [1].</p> <p>Raw ML metrics from models gaussian Naïve Bayes classifiers using this dataset are available in </p> <p> </p> <p>[1] <span><span><span>Matos-Filipe, P, et al. "</span></span></span>The usage of transcriptomics datasets as sources of Real-World Data for clinical trialling". <span>bioRxiv (</span><span>2022). </span><span><span>doi:</span> https://doi.org/10.1101/2022.11.10.515995</span></p>
Supporting Data for "Identifying the Origins of Nanoplastics in the Abyssal South Atlantic Using Backtracking Lagrangian Simulations with Fragmentation"
<p>Supporting Data for "Identifying the Origins of Nanoplastics in the Abyssal South Atlantic Using Backtracking Lagrangian Simulations with Fragmentation", published in the Ocean and Coastal Research Journal. The repository consists of the code used for the simulations and analysis, and the supplementary information document associated with the main manuscript, the Lagrangian simulation outputs and aditional data used for the analysis.</p>
Dataset and models of TMLR 2024 Paper "Identifying and Clustering Counter Relationships of Team Compositions in PvP Games for Efficient Balance Analysis"
<p>This is a part of dataset and models of the paper published in TMLR 2024 (Transactions on Machine Learning Research, <a href="https://jmlr.org/tmlr/" target="_blank" rel="noopener">https://jmlr.org/tmlr/</a>).</p> <p>Including training datasets, testing datasets, and models.</p> <p>The example program for using this file will be put on the author's github repo branch: <a href="https://github.com/DSobscure/cgi_drl_platform/tree/game_balance_measures_tmlr" target="_blank" rel="noopener">https://github.com/DSobscure/cgi_drl_platform/tree/game_balance_measures_tmlr</a></p> <p> </p>
Satellite-based precipitation estimates using a dense rain gauge network over the Southwestern Brazilian Amazon: Implication for identifying trends in dry season rainfall
<h1>Satellite-based precipitation estimates using a dense rain gauge network over the Southwestern Brazilian Amazon.</h1>
EOL full taxon identifier map
<p>A mapping of taxon identifiers from EOL resources, of the form: node_id, resource_pk, resource_id, page_id, preferred_canonical_for_page</p> <ul> <li>node_id: internal to EOL; useful for some API calls</li> <li>resource_pk: identifier according to the classification provider </li> <li>resource_id: identifies the classification provider (see below) </li> <li>page_id: EOL taxon concept identifier; official, for sharing </li> <li>preferred_canonical_for_page: canonical name preferred by EOL for this taxon concept </li> </ul> <p>commas within entries are "escaped, by, quoting", and quotes-within-quotes are ""double quoted""</p> <p>resource_ids, their names and descriptions, are available at: <a href="https://eol.org/resources.json">https://eol.org/resources.json </a></p> <p>To view one at a time: https://eol.org/resources/[enter ID here] </p> <p>There is also a summary file with resources ids, links, and resource names here: <a href="https://github.com/KatjaSchulz/eolResources">https://github.com/KatjaSchulz/eolResources</a></p> <p>And there is a smaller mapping file just for the major EOL classification sources: <a href="../doi/10.5281/zenodo.13769681">EOL taxon identifier map</a></p>
Source Data for Manuscript: Identifying genomic data use with the Data Citation Explorer
<p>This page contains the source data for the manuscript describing the Data Citation Explorer, currently in review for publication. The preprint version can be found on this page.</p> <p>Files:</p> <p><strong>DCE_manual_eval_sample.xlsx:</strong></p> <p>This file was used to manually evaluate hits generated by the Data Citation Explorer. There are two separate sheets: one with publications returned by searches in PubMed and PubMed Central and another with publications returned by searches in Dimensions. Column descriptions can be found in the file itself. Each row in each evaluation sheet refers to a pair between a JAMO record and a linked publication.</p> <p><strong>DCE_citation_report.csv</strong></p> <p>Contains JAMO record IDs and PubMed IDs from the initial 2020 DCE trial run. There are 238,994 unique JAMO IDs and 30,641 unique PubMed IDs. 78,104 JAMO records are linked with publications.</p> <p>Columns:</p> <ul> <li>jamo_id - unique JAMO record ID</li> <li>sample_group - Sample strata from which manually evaluated records were pulled</li> <li>citation_count - Number of citations associated with each record</li> <li>citations - comma-delimited PubMed IDs for linked publications</li> <li>sampled - True/False, denoting which records were included in the initial evaluation sample</li> <li>notes - descriptions for why certain sampled records were excluded from manual evaluation</li> <li>unprocessed - True/False. These 7,890 records contained anomalous fields that caused them to be rejected for processing. They are represented as zero-length files in the archive.</li> </ul> <p><strong>DCE_source_files.zip:</strong></p> <p>This folder contains 3 files for each JAMO record in DCE_citation_report.tsv. For each JAMO record listed in the citation report, three files are provided:</p> <ol> <li>JAMO_ID_source.yaml - The fields extracted from the JAMO record that were relevant to the citation search, including any previously known PMIDs (manually curated).</li> <li>JAMO_ID_expand.yaml - The source record augmented with additional metadata discovered in other resources, including the citations that were discovered based on querying PubMed Central for the values in those metadata fields.</li> <li>JAMO_ID_audit.json - The audit path as a directed acyclic graph, in JSON.</li> </ol>
Fuτure - dataset for studies, development, and training of algorithms for reconstructing and identifying hadronically decaying tau leptons
<h1> Data description</h1> <h2>MC Simulation</h2> <p><br>The <strong>Fuτure</strong> dataset is intended for studies, development, and training of algorithms for reconstructing and identifying hadronically decaying tau leptons. The dataset is generated with Pythia 8, with the full detector simulation being performed by Geant4 with the CLIC-like detector setup CLICdet (CLIC_o3_v14) setup. Events are reconstructed using the Marlin reconstruction framework and interfaced with Key4HEP. Particle candidates in the reconstructed events are reconstructed using the PandoraPF algorithm.</p> <p>In this version of the dataset no γγ -> hadrons background is included.</p> <h2>Samples</h2> <p><br>This dataset contains e+e- samples with Z->ττ, ZH,H->ττ and Z->qq events, with approximately 2 million events simulated in each category.</p> <p>The following processes e+e- were simulated with Pythia 8 at sqrt(s) = 380 GeV:</p> <ul> <li>p8_ee_qq_ecm380 [Z -> qq events]</li> <li>p8_ee_ZH_Htautau [ZH -> Ztautau]</li> <li>p8_ee_Z_Ztautau_ecm380 [ZH -> Ztautau]</li> </ul> <p>The .root files from the MC simulation chain are eventually processed by the software found in <a href="https://github.com/HEP-KBFI/ml-tau-en-reg">Github</a> in order to create flat ntuples as the final product.</p> <h2><br>Features</h2> <p><br>The basis of the ntuples are the particle flow (PF) candidates from PandoraPF. Each PF candidate has four momenta, charge and particle label (electron / muon / photon / charged hadron / neutral hadron). The PF candidates in a given event are clustered into jets using generalized kt algorithm for ee collisions, with parameters p=-1 and R=0.4. The minimum pT is set to be 0 GeV for both generator level jets and reconstructed jets. The dataset contains the four momenta of the jets, with the PF candidates in the jets with the above listed properties.</p> <p>Additionally, a set of variables describing the tau lifetime are calculated using the software in <a href="https://github.com/HEP-KBFI/ml-tau-en-reg">Github</a>. As tau lifetime is very short, these variables are sensitive to true tau decays. In the calculation of these lifetime variables, we use a linear approximation.</p> <p>In summary, the features found in the flat ntuples are:</p> <p> </p> <table> <tbody> <tr> <td><strong>Name</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>reco_cand_p4s</td> <td>4-momenta per particle in the reco jet.</td> </tr> <tr> <td>reco_cand_charge</td> <td>Charge per particle in the jet.</td> </tr> <tr> <td>reco_cand_pdg</td> <td>PDGid per particle in the jet.</td> </tr> <tr> <td>reco_jet_p4s</td> <td>RecoJet 4-momenta.</td> </tr> <tr> <td>reco_cand_dz</td> <td>Longitudinal impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dz_err</td> <td>Uncertainty of the longitudinal impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dxy</td> <td>Transverse impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dxy_err</td> <td>Uncertainty of the transverse impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>gen_jet_p4s</td> <td>GenJet 4-momenta. Matched with RecoJet within a cone of radius dR < 0.3.</td> </tr> <tr> <td>gen_jet_tau_decaymode</td> <td>Decay mode of the associated genTau. Jets that have associated leptonically decaying taus are removed, so there are no DM=16 jets. If no GenTau can be matched to GenJet within dR < 0.4, a fill value is used.</td> </tr> <tr> <td>gen_jet_tau_p4s</td> <td>Visible 4-momenta of the genTau. If no GenTau can be matched to GenJet within dR<0.4, a fill value is used.</td> </tr> </tbody> </table> <p>The ground truth is based on stable particles at the generator level, before detector simulation. These particles are clustered into generator-level jets and are matched to generator-level τ leptons as well as reconstructed jets. In order for a generator-level jet to be matched to generator-level τ lepton, the τ lepton needs to be inside a cone of dR = 0.4. The same applies for the reconstructed jet, with the requirement on dR being set to dR = 0.3. For each reconstructed jet, we define three target values related to τ lepton reconstruction:</p> <ul> <li> a binary flag <strong>isTau</strong> if it was matched to a generator-level hadronically decaying τ lepton. <strong>gen_jet_tau_decaymode</strong> of value -1 indicates no match to generator-level hadronically decaying τ.</li> <li> the categorical decay mode of the τ <strong>gen_jet_tau_decaymode</strong> in terms of the number of generator level charged and neutral hadrons. Possible <strong>gen_jet_tau_decaymode</strong> are {0, 1, . . . , 15}.</li> <li> if matched, the visible (neglecting neutrinos), reconstructable pT of the τ lepton. This is inferred from the <strong>gen_jet_tau_p4s</strong></li> </ul> <h2>Contents:</h2> <ul> <li>qq_test.parquet</li> <li>qq_train.parquet</li> <li>zh_test.parquet</li> <li>zh_train.parquet</li> <li>z_test.parquet</li> <li> z_train.parquet</li> <li>data_intro.ipynb</li> </ul> <h2>Dataset characteristics</h2> <p> </p> <table> <tbody> <tr> <td><strong>File</strong></td> <td><strong># Jets</strong></td> <td><strong>Size</strong></td> </tr> <tr> <td>z_test.parquet</td> <td> <pre>870 843</pre> </td> <td>171 MB</td> </tr> <tr> <td>z_train.parquet</td> <td> <pre>3 483 369</pre> </td> <td>681 MB</td> </tr> <tr> <td>zh_test.parquet</td> <td> <pre>1 068 606</pre> </td> <td>213 MB</td> </tr> <tr> <td>zh_train.parquet</td> <td> <pre>4 274 423</pre> </td> <td>851 MB</td> </tr> <tr> <td>qq_test.parquet</td> <td> <pre>6 366 715</pre> </td> <td>1.4 GB</td> </tr> <tr> <td>qq_train.parquet</td> <td> <pre>25 466 858</pre> </td> <td>5.6 GB</td> </tr> </tbody> </table> <p>The dataset consists of 6 files of 8.9 GB in total.</p> <h2>How can you use these data?</h2> <p>The .parquet files can be directly loaded with the Awkward Array Python library.<br>An example how one might use the dataset and the features is given in <strong>data_intro.ipynb</strong></p>
Natural history specimens collected and/or identified and deposited.
Natural history specimen data collected and/or identified by John Howieson, <a href="https://orcid.org/0000-0003-1378-293X">https://orcid.org/0000-0003-1378-293X</a>. Claims or attributions were made on Bionomia, <a href="http://bionomia.net">https://bionomia.net</a> using specimen data from the Global Biodiversity Information Facility, <a href="https://gbif.org">https://gbif.org</a>.
The largest variable UCHII identified in GLOSTAR and CORNISH surveys
<p>The GLOSTAR and CORNISH 5 GHz images for the 38 variable sources identified in this work. The beam size is 1.5\arcsec for each image in the two surveys. Each figure is centered at the position of the identified \hii\ region. </p>
GWAS to single cell: Intersecting single-cell transcriptomics and genome wide association studies identifies crucial cell-populations and candidate genes for atherosclerosis.
<p><strong>Background</strong></p> <p>Genome-wide association studies (GWAS) have discovered hundreds of common genetic variants for atherosclerotic disease and cardiovascular risk factors. The translation of susceptibility loci into biological mechanisms and targets for drug discovery remains challenging. Intersecting genetic and gene expression data has led to identification of candidate genes. However, the assayed tissues are often non-diseased and heterogeneous in cell composition confounding the candidate prioritization. We collected single-cell transcriptomics (scRNA-seq) from atherosclerotic plaques and aimed to identify cell-type-specific expression of disease-associated genes. </p> <p> </p> <p><strong>Methods and Results</strong></p> <p>To identify disease-associated candidate genes, we applied gene-based analyses using GWAS summary statistics from 46 atherosclerotic, cardiometabolic, and other traits. Next we intersected these candidates with single-cell transcriptomics (scRNA-seq) to identify those genes that are specifically expressed in individual cell (sub)populations of atherosclerotic plaques. We derive an enrichment score and show that loci that associated with coronary artery disease demonstrated a prominent substrate in plaque smooth muscle cells (<em>SKI</em>, <em>KANK2</em>, <em>SORT1</em>), endothelial cells (<em>SLC44A1</em>, <em>ATP2B1</em>), and macrophages (<em>APOE</em>, <em>HNRNPUL1</em>). Further sub clustering of SMC-subtypes revealed genes in risk loci for coronary calcification specifically enriched in a synthetic cluster of SMCs. To verify the robustness of our approach, we used liver-derived scRNAseq-data and showed enrichment of circulating lipids-associated loci in hepatocytes.</p> <p><br> <strong>Conclusion</strong></p> <p>We confirm known gene-cell pairs relevant for atherosclerotic disease, and discovered novel pairs pointing to new biological mechanisms amenable for therapy. We present an intuitive single-cell transcriptomics driven workflow rooted in human large-scale genetic studies to identify putative candidate genes and affected cells associated with cardiovascular traits.</p> <p> </p>
Data and R-scripts for "Land-use trajectories for sustainable land system transformations: identifying leverage points in a global biodiversity hotspot" (V2)
<p>Sustainable land system transformations are necessary to avert biodiversity and climate collapse. However, it remains unclear where entry points for transformations exist in complex land systems. Here, we conceptualize land systems along land-use trajectories, which allows us to identify and evaluate leverage points; i.e., entry points on the trajectory where targeted interventions have particular leverage to influence land-use decisions. We apply this framework in the biodiversity hotspot Madagascar. In the Northeast, smallholder agriculture results in a land-use trajectory originating in old-growth forests, spanning forest fragments, and reaching shifting hill rice cultivation and vanilla agroforests. Integrating interdisciplinary empirical data on seven taxa, five ecosystem services, and three measures of agricultural productivity, we assess trade-offs and co-benefits of land-use decisions at three leverage points along the trajectory. These trade-offs and co-benefits differ between leverage points: two leverage points are situated at the conversion of old-growth forests and forest fragments to shifting cultivation and agroforestry, resulting in considerable trade-offs, especially between endemic biodiversity and agricultural productivity. Here, interventions enabling smallholders to conserve forests are necessary. This is urgent since ongoing forest loss threatens to eliminate these leverage points due to path-dependency. The third leverage point allows for the restoration of land under shifting cultivation through vanilla agroforests and offers co-benefits between restoration goals and agricultural productivity. The co-occurring leverage points highlight that conservation and restoration are simultaneously necessary. Methodologically, the framework shows how leverage points can be identified, evaluated, and harnessed for land system transformations under the consideration of path-dependency along trajectories.</p>
Data for `Identifying New Pulsating Variables and Eclipsing Binaries Using TESS Data'
<p>This study presents a series of surveys using TESS data to identify new δ Scuti and γ Doradus stars, as well as eclipsing binaries with pulsating components. Preliminary catalogs of newly discovered variables are being made publicly available to encourage community use as the project progresses. Please visit this website for updates. </p> <p> 1. New δ Scuti stars, γ Doradus stars, and eclipsing binaries from a subset of 709,000 selected AF-type stars observed by TESS. Version 4.5 is an update to Version 4.0 (the first release for Part IV), while Part III was provided in Version 3.0.</p> <p> The attached CSV file <strong>NewVar_AF50w_R2.csv,</strong> contains the catalog of newly identified variables from Phases I and II of the survey (i.e. R2 includes R1 from Version 3.0). The accompanying file <strong>ms2RNAAS_2025AF50w_IV.pdf</strong> is the initial draft describing the Phase-II identifications. The published paper can be found in 2025 <em>Research Notes of the AAS</em>, <strong>Vol. 9, No. 1, 23</strong> (Zhou 2025, https://iopscience.iop.org/article/10.3847/2515-5172/adaf8c (ADS code: 2025RNAAS...9...23Z). These unpublished discoveries have been compiled into catalogs of δ Scuti and γ Doradus stars available at Zenodo https://zenodo.org/records/17096360 (DOI: <a href="https://doi.org/10.5281/zenodo.17096360">10.5281/zenodo.17096360</a>) -- which currently include <strong>118,410</strong> and <strong>41,622</strong> entries, respectively, as of the release date.</p> <p> 2. New Pulsating Variable Stars and Eclipsing Binaries around BL Cam (added in Version 2.0)</p> <p> The attached CSV file NewVar_BLCam_R2D.csv provides a catalog of the new variables with the following columns: TIC ID, Simbad main ID, Gaia DR3 ID, RA_deg Dec_deg(J2000), Tmag, Teff, SpType, Lum, logg, mass, VarType_Notes</p> <p> 3. New Pulsating Variable Stars and Eclipsing Binaries near NGC 6302 (Version 1.0)</p> <p> Check corresponding files containing the string "NGC 6302".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.