Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

14

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

14 results for “version control”

Learn how ShareScore rates datasets ↗
zenodo48/100

MarFERReT: an open-source, version-controlled reference library of marine microbial eukaryote functional genes

<p>Metatranscriptomics generates large volumes of sequence data about transcribed genes in natural environments. Taxonomic annotation of these datasets depends on availability of curated reference sequences. For marine microbial eukaryotes, current reference libraries are limited by gaps in sequenced organism diversity and barriers to updating libraries with new sequence data, resulting in taxonomic annotation of only about half of eukaryotic environmental transcripts. Here, we introduce version 1.0 of the Marine Functional EukaRyotic Reference Taxa (MarFERReT), an updated marine microbial eukaryotic sequence library with a version-controlled framework designed for taxonomic annotation of eukaryotic metatranscriptomes. We gathered 902 marine eukaryote genomes and transcriptomes from multiple sources and assessed these candidate entries for sequence quality and cross-contamination issues, selecting 800 validated entries for inclusion in the library. MarFERReT v1 contains reference sequences from 800 marine eukaryotic genomes and transcriptomes, covering 453 species- and strain-level taxa, totaling nearly 28 million protein sequences with associated NCBI and PR2 Taxonomy identifiers and Pfam functional annotations. An accompanying MarFERReT project repository hosts containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT.<br><br>MarFERReT is linked to a code repository hosting containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT here:&nbsp;<a href="https://github.com/armbrustlab/marferret">https://github.com/armbrustlab/marferret</a></p> <p>The raw source data for the 902 candidate entries considered for MarFERReT v1.1.1, including the 800 accepted entries, are available for download from their respective online locations. The source URL for each of the entries is listed here in MarFERReT.v1.1.1.entry_curation.csv, and detailed instructions and code for downloading the raw sequence data from source are available in the MarFERReT code repository (<a href="https://github.com/armbrustlab/marferret/blob/main/docs/process_clean_marmicrodb.log.sh">link</a>). &nbsp;&nbsp;</p> <p>This repository release contains MarFERReT database files from the v1.1.1 MarFERReT release using the following MarFERReT library build scripts: <strong>assemble_marferret.sh</strong>, <strong>pfam_annotate.sh</strong>, and <strong>build_diamond_db.sh</strong><br><br>The following MarFERReT data products are available in this repository:</p> <p><strong>MarFERReT.v1.1.1.metadata.csv</strong><br>This CSV file contains descriptors of each of the 902 database entries, including data source, taxonomy, and sequence descriptors. Data fields are as follows:</p> <ol> <li><strong>entry_id</strong>: Unique MarFERReT sequence entry identifier.</li> <li><strong>accepted:&nbsp;</strong>Acceptance into the final MarFERReT build (Y/N). The Y/N values can be adjusted to customize the final build output according to user-specific needs.</li> <li><strong>marferret_name</strong>: A human and machine friendly string derived from the NCBI Taxonomy organism name; maintaining strain-level designation wherever possible.</li> <li><strong>tax_id</strong>: The NCBI Taxonomy ID (taxID).</li> <li><strong>pr2_accession</strong>: Best-matching PR2 accession ID associated with entry</li> <li><strong>pr2_rank</strong>: The lowest shared rank between the entry and the pr2_accession</li> <li><strong>pr2_taxonomy</strong>: PR2&nbsp;Taxonomy classification scheme of the pr2_accession</li> <li><strong>data_type</strong>: Type of sequence data; transcriptome shotgun assemblies (TSA), gene models from assembled genomes (genome), and single-cell amplified genomes (SAG) or transcriptomes (SAT).</li> <li><strong>data_source</strong>: Online location of sequence data; the Zenodo data repository (<a href="../">Zenodo</a>), the datadryad.org repository (<a href="http://datadryad.org/">datadryad.org</a>), MMETSP re-assemblies on Zenodo (MMETSP)17, NCBI GenBank (<a href="https://www.ncbi.nlm.nih.gov/genbank/">NCBI</a>), JGI Phycocosm (<a href="https://phycocosm.jgi.doe.gov/phycocosm/home">JGI-Phycocosm</a>), the TARA Oceans portal on Genoscope (<a href="http://www.genoscope.cns.fr/tara/">TARA</a>), or entries from the Roscoff Culture Collection through the METdb database repository (<a href="https://metdb.sb-roscoff.fr/metdb/">METdb</a>).</li> <li><strong>source_link</strong>: URL where the original sequence data and/or metadata was collected.</li> <li><strong>pub_year</strong>: Year of data release or publication of linked reference.</li> <li><strong>ref_link</strong>: Pubmed URL directs to the published reference for entry, if available.</li> <li><strong>ref_doi</strong>: DOI of entry data from source, if available.</li> <li><strong>source_filename</strong>: Name of the original sequence file name from the data source.</li> <li><strong>seq_type</strong>: Entry sequence data retrieved in nucleotide (nt) or amino acid (aa) alphabets.</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file.</li> <li><strong>source_name:</strong> Full organism name from entry source</li> <li><strong>original_taxID</strong>: Original NCBI taxID from entry data source metadata, if available</li> <li><strong>alias:</strong> Additional identifiers for the entry, if available</li> </ol> <p><br><strong>MarFERReT.v1.1.1.curation.csv</strong><br>This CSV file contains curation and quality-control information on the 902 candidate entries considered for incorporation into MarFERReT v1, including curated NCBI Taxonomy IDs and entry validation statistics. Data fields are as follows:</p> <ol> <li><strong>entry_id:</strong> Unique MarFERReT sequence entry identifier</li> <li><strong>marferret_name:&nbsp;</strong>Organism name in human and machine friendly format, including additional NCBI taxonomy strain identifiers if available.</li> <li><strong>tax_id</strong>: Verified NCBI taxID used in MarFERReT</li> <li><strong>taxID_status</strong>: Status of the final NCBI taxID (Assigned, Updated, or Unchanged)</li> <li><strong>taxID_notes</strong>: Notes on the original_taxID</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file</li> <li><strong>n_pfams</strong>: Number of Pfam domains identified in protein sequences</li> <li><strong>qc_flag</strong>: Early validation quality control flags for the following: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.</li> <li><strong>flag_Lasek</strong>: Flag notes from Lasek-Nesselquist and Johnson (2019); contains the flag 'FLAG_LASEK' indicating ciliate samples reported as contaminated in this study.</li> <li><strong>VV_contam_pct</strong>: Estimated contamination reported for MMETSP entries in Van Vlierberghe et al., (2021).</li> <li><strong>flag_VanVlierberghe:&nbsp;</strong>Flag for a high level of estimated contamination, from 'flag_VanVlierberghe' &nbsp;values over 50%: FLAG_VV.</li> <li><strong>rp63_npfams</strong>: Number of ribosomal protein Pfam domains out of 63 total.</li> <li><strong>rp63_contam_pct</strong>: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, as described in the Technical Validation section from analysis of 63 Pfam ribosomal protein domains.</li> <li><strong>flag_rp63</strong>: Flag for a high level of estimated contamination, from 'rp63_contam_pct' &nbsp;values over 50%: FLAG_RP63.</li> <li><strong>flag_sum:&nbsp;</strong>Count of the number of flag columns (`qc_flag`, `flag_Lasek`, `flag_VanVlierberghe`, and `flag_rp63`). All entries with one or more flag are nominally rejected ('accepted' = N); entries without any flags are validated and accepted ('accepted' = Y).</li> <li><strong>accepted:&nbsp;</strong>Acceptance into the final MarFERReT build (Y or N).</li> </ol> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.proteins.faa.gz</strong><br>This Gzip-compressed FASTA file contains the 27,951,013 final translated and clustered protein sequences for all 800 accepted MarFERReT entries. The sequence defline contains the unique identifier for the sequence and its reference (mftX, where 'X' is a ten-digit integer value).&nbsp;</p> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.taxonomies.tab.gz</strong><br>This Gzip-compressed tab-separated file is formatted for interoperability with the DIAMOND protein alignment tool commonly used for downstream analyses and contains some columns without any data. Each row contains an entry for one of the MarFERReT protein sequences in MarFERReT.v1.proteins.faa.gz. Note that 'accession.version' and 'taxid' are populated columns while 'accession' and 'gi' have NA values; the latter columns are required for back-compatibility as input for the DIAMOND alignment software and LCA analysis.&nbsp;</p> <p>The columns in this file contain the following information:</p> <ol> <li><strong>accession</strong>: (NA)</li> <li><strong>accession.version</strong>: The unique MarFERReT sequence identifier ('mftX').</li> <li><strong>taxid</strong>: The NCBI Taxonomy ID associated with this reference sequence.</li> <li><strong>gi</strong>: (NA).</li> </ol> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.proteins_info.tab.gz</strong><br>This Gzip-compressed tab-separated file contains a row for each final MarFERReT protein sequence with the following columns:</p> <ol> <li><strong>aa_id</strong>: the unique identifier for each MarFERReT protein sequence.</li> <li><strong>entry_id</strong>: The unique numeric identifier for each MarFERReT entry.</li> <li><strong>source_defline</strong>: The original, unformatted sequence identifier</li> </ol> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.best_pfam_annotations.csv.gz<br></strong>This Gzip-compressed CSV file contains the best-scoring Pfam annotation for intra-species clustered protein sequences from the 800 validated MarFERReT entries; derived from the hmmsearch annotations against Pfam 34.0&nbsp; functional domains. This file contains the following fields:</p> <ol> <li><strong>aa_id</strong>: The unique MarFERReT protein sequence ID ('mftX').</li> <li><strong>pfam_name</strong>: The shorthand Pfam protein family name.</li> <li><strong>pfam_id</strong>: The Pfam identifier.</li> <li><strong>pfam_eval</strong>: hmm profile match e-value score</li> <li><strong>pfam_score:</strong> hmm profile match bitscore</li> </ol> <p><br><strong>MarFERReT.v1.1.1.dmnd</strong><br>This binary file is the indexed database of the MarFERReT protein library with embedded NCBI taxonomic information generated by the DIAMOND makedb tool using the build_diamond_db.sh script from the MarFERReT /scripts/ library. This can be used as the reference DIAMOND database for annotating environment sequences from eukaryotic metatranscriptomes.&nbsp;<br><br></p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

caseysaenger/ForamMgCa_PSM: files and scripts for revised version of manuscript "Calibration and validation of environmental controls on planktic foraminifera Mg/Ca using global core-top data".

<p>files and scripts for revised version of manuscript &quot;Calibration and validation of environmental controls on planktic foraminifera Mg/Ca using global core-top data&quot;. Saenger, C. and M. N. Evans. Resubmitted to Paleoceanography and Paleoclimatology, May 3, 2019.</p>

openother-openOct 2018View details →
zenodo36/100

Results from a 2015 survey on Git/Distributed Version Control at Imperial College London

<p>These are the - anonymised - results from a survey run at Imperial College London in November-December 2015. The survey was aimed at user of distributed version control systems, in particular Git. Before publishing the results I deleted all comments to avoid individuals being identified. The survey was designed to inform internal decision making but the general results may be of value to others.</p>

opencc-by-4.0Aug 2016View details →
zenodo36/100

Datasets for "Peatland evaporation across hemispheres: contrasting controls and sensitivity to climate warming driven by plant functional types" - Version 2

<p>Version 2 of datasets used for analyses in the paper titled "Peatland evaporation across hemispheres: contrasting controls and sensitivity to climate warming driven by plant functional types" submitted to Biogeosciences. There are two datasets - one from Kopuatai bog, Aotearoa New Zealand, and one from Mer Bleue bog, Canada - which contain gap-filled and filtered data that were used to produce the results of our study.</p> <p>Due to improvements made to our methodology following the paper peer review process, the data in this version slightly differs from that of the previous version. Information on these revisions can be found in the README file below or the Discussion/Peer Review tab of our paper.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2024View details →
zenodo32/100

AWI-CM3 version 3.0 pre-industrial control data

<p>Pre-industrial control simulation output from AWI-CM3 version 3.0</p> <p>Years: 1850-2014</p> <p>Frequency: Annual means</p> <p>Variables:</p> <ul> <li>T2M: near surface (2m) air temperature</li> <li>PRECIP: precipitation</li> <li>a_ice: sea ice concentraion</li> <li>m_ice: sea ice mass per unit area</li> <li>salt: salinity</li> <li>w: vertical velocity</li> <li>MLD2: Mixed layer depth according to Large et al. (https://doi.org/10.1175/1520-0485(1997)027%3C2418:STSFAB%3E2.0.CO;2)</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Replication Package for "Using Reinforcement Learning to Sustain the Performance of Version Control Repositories"

<p>The replication package is organized into two containers, each of which is responsible for reproducing figures and analyses for each RQ.</p> <p>## Decompress Package</p> <p>```<br>$ tar -xvJf rl4monorepos.tar.xz<br>```</p> <p>## RQ1</p> <p>1. Import the container</p> <p>```<br>$ docker import rq1.tar rq1<br>```</p> <p>2. Regenerate figures</p> <p>```<br>$ docker container run -v &lt;outputdir&gt;:/out -e R_SCRIPT=figures.R rq1&nbsp;<br>```</p> <p>3. Re-execute statistical tests</p> <p>```<br>$ docker container run -e R_SCRIPT=stats-test.R rq1<br>```</p> <p>## RQ2</p> <p>1. Import the container</p> <p>```<br>$ docker import rq2.tar rq2<br>```</p> <p>2. Regenerate figures and print AUC values</p> <p>```<br>$ docker container run -v &lt;outputdir&gt;:/out -e R_SCRIPT=figures.R rq2&nbsp;<br>```</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Supplementary material 1 from: Seehausen ML, Branco M, Afonso C, Kenis M (2023) Testing a modified version of the EPPO decision-support scheme for release of classical biological control agents of plant pests using Ganaspis cf. brasiliensis and Cleruchoides noackae as case studies. NeoBiota 87: 121-141. https://doi.org/10.3897/neobiota.87.103187

Decision-support scheme for release of classical biological control agents of plant pests – Environmental impact assessment (EIA)

opencc-zeroAug 2023View details →
zenodo28/100

Supplementary material 2 from: Seehausen ML, Branco M, Afonso C, Kenis M (2023) Testing a modified version of the EPPO decision-support scheme for release of classical biological control agents of plant pests using Ganaspis cf. brasiliensis and Cleruchoides noackae as case studies. NeoBiota 87: 121-141. https://doi.org/10.3897/neobiota.87.103187

Assessment of Ganaspis cf. brasiliensis

opencc-zeroAug 2023View details →
zenodo28/100

Supplementary material 3 from: Seehausen ML, Branco M, Afonso C, Kenis M (2023) Testing a modified version of the EPPO decision-support scheme for release of classical biological control agents of plant pests using Ganaspis cf. brasiliensis and Cleruchoides noackae as case studies. NeoBiota 87: 121-141. https://doi.org/10.3897/neobiota.87.103187

Assessment of Cleruchoides noackae

opencc-zeroAug 2023View details →
ClinicalTrials.gov24/100

Global Controlled Trial on Effects of an Online Self-Help Program for of Ambitious Altruists on Their Mental Health, Wellbeing, and Productivity: Comparing Versions With IFS vs. CBT, Buddy- vs. Group-

ClinicalTrials.gov study NCT06442072. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov24/100

Open Randomized Controlled Trial to Evaluate the Efficacy and Safety of Remifentanil Versus Nitrous Oxide in External Cephalic Version at Term in Singleton Pregnancy in Breech Presentation

ClinicalTrials.gov study NCT01735669. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov24/100

Cross-Cultural of the Validity, Reliability and Interpretability of Thai-version of Urticaria Control Test

ClinicalTrials.gov study NCT02285049. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov24/100

A Prospective Randomnised Controlled Trial Comparing Overall Patient Compliance in a Bariatric Surgical Pathway Using the Standard Versus a More Intensified and Interactive Version of the "Get Ready"

ClinicalTrials.gov study NCT07297342. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
geo20/100

Differential gene expression in human THP-1 monocytes expressing the RdRP transgene (WT version) compared to THP-1 empty vector control cells

GEO Series GSE62753. Homo sapiens. 4 samples. Type: Expression profiling by array.

openGEO-OpenNov 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record