Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

11 results for “manual curation”

Learn how ShareScore rates datasets ↗
zenodo52/100

BioASQ-QA: A manually curated corpus for Biomedical Question Answering

<p>The BioASQ question answering (QA) benchmark dataset contains questions in English, along with golden standard (reference) answers and related material. The dataset has been designed to reflect real information needs of biomedical experts and is therefore more realistic and challenging than most existing datasets. Furthermore, unlike most previous QA benchmarks that contain only exact answers, the BioASQ-QA dataset also includes ideal answers (in effect summaries), which are particularly useful for research on multi-document summarization. The dataset combines structured and unstructured data. The material linked with each question comprise documents and snippets, which are useful for Information Retrieval and Passage Retrieval experiments, as well as concepts that are useful in concept-to-text Natural Language Generation. Researchers working on paraphrasing and textual entailment can also measure the degree to which their methods improve the performance of biomedical QA systems. Last but not least, the dataset is continuously extended, as the BioASQ challenge is running and new data are generated.</p>

opencc-by-2.5Dec 2022View details →
zenodo48/100

A pangenome-guided manually curated library of transposable elements for Zymoseptoria tritici

<p>A manually-curated TE consensus library generated using a panel of 19 reference genomes for&nbsp;<em>Zymoseptoria tritici</em><sup>1-3</sup>&nbsp;along with reference genome assemblies for the sister species&nbsp;<em>Z. ardabiliae</em>,&nbsp;<em>Z. brevis</em>,&nbsp;<em>Z. pseudotritici</em>, and&nbsp;<em>Z. passerinii<sup>4</sup></em>.&nbsp;</p> <p>&nbsp;</p> <p><strong>Methods</strong></p> <p>Putative TE consensus sequences were first obtained by annotating all 23 genome assemblies<sup>1&ndash;4</sup>&nbsp;with Earl Grey with default settings (v3.0;&nbsp;<a href="https://github.com/TobyBaril/EarlGrey">https://github.com/TobyBaril/EarlGrey</a>)<sup>5,6</sup>. Consensus sequences generated from each reference genome were clustered using CD-Hit-Est (v4.8.1)<sup>7,8</sup>&nbsp;to group sequences with 90% similarity across 80% of the longer sequence length (<em>-n 8 -d 0 -aL 0.8 -c 0.90 -G 0 -g 1 -b 500 -r 1</em>)&nbsp;to reduce redundancy whilst preventing the collapsing of chimeric sequences. Consensus sequences &lt;100bp were removed, as these are unlikely to represent true TE sequences. Each consensus sequence was then subject to manual curation as described by Goubert et al. (2022)<sup>9</sup>. Briefly, genomic copies of each TE were obtained using a &ldquo;BLAST, Extract, Extend&rdquo; process to recover genomic copies from each of the 23 reference genome assemblies with 1,000 flanking bases at either end<sup>9,10</sup>. For families with &gt;100 BLASTN hits, the 25 longest hits were selected, along with 75 random hits. Multiple alignments were generated for each putative TE family using MAFFT (v7.505) with the --auto flag<sup>11</sup>. Columns composed of &gt;=80% gaps were removed with T-COFFEE (v13.45.0.4846264)<sup>12</sup>. Subsequently, all sequence alignments were manually curated to define TE boundaries and remove regions of low conservation and rare insertions. Following manual curation, new majority-rule consensus sequences were generated with EMBOSS (v6.6.0.0) cons<sup>13</sup>. TE-Aid (<a href="https://github.com/clemgoub/TE-Aid/">https://github.com/clemgoub/TE-Aid/</a>) was used to aid visual inspection and to identify diagnostic features for classification of extended consensus sequences. Following this, TIRs were recorded if present, and nhmmscan (HMMER v3.3.2)<sup>14</sup>&nbsp;was used to identify homology to known curated elements in Dfam (v3.7). Combining this information, each TE consensus sequence was manually classified using available information following the naming convention &lsquo;&gt;ZymTri_2023_family_[n]#[Classification]/[Family]&rsquo;. Consensus sequences classified with low confidence have a &lsquo;?&rsquo; added to the name, as well as the string &lsquo;_LowConf&rsquo;. To reduce redundancy in the final TE library, sequences were clustered to the family-level using the 80-80-80 rule implemented in CD-hit-est<sup>9,15 </sup>(<em>-d 0 -aS 0.8 -c 0.8 -G 0 -g 1 -b 500 -r 1</em>). The representative sequence for each cluster was manually selected to select the sequence with the highest classification confidence, also defined as the &lsquo;most intact consensus&rsquo;. Chimeric sequences erroneously clustered were manually separated to retain sequences for the chimeric TE and the individual elements that generated the chimer.</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>1.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Badet, T., Oggenfuss, U., Abraham, L., McDonald, B. A. &amp; Croll, D. A 19-isolate reference-quality global pangenome for the fungal wheat pathogen Zymoseptoria tritici.&nbsp;<em>BMC Biol.</em>&nbsp;<strong>18</strong>, 12 (2020).</p> <p>2.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Goodwin, S. B.&nbsp;<em>et al.</em>&nbsp;Finished genome of the fungal wheat pathogen Mycosphaerella graminicola reveals dispensome structure, chromosome plasticity, and stealth pathogenesis.&nbsp;<em>PLoS Genet.</em>&nbsp;<strong>7</strong>, e1002070 (2011).</p> <p>3.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Plissonneau, C., Hartmann, F. E. &amp; Croll, D. Pangenome analyses of the wheat pathogen Zymoseptoria tritici reveal the structural basis of a highly plastic eukaryotic genome.&nbsp;<em>BMC Biol.</em>&nbsp;<strong>16</strong>, 5 (2018).</p> <p>4.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Feurtey, A.&nbsp;<em>et al.</em>&nbsp;Genome compartmentalization predates species divergence in the plant pathogen genus Zymoseptoria.&nbsp;<em>BMC Genomics</em>&nbsp;<strong>21</strong>, 588 (2020).</p> <p>5.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Baril, T., Imrie, R. M. &amp; Hayward, A. Earl Grey: a fully automated user-friendly transposable element annotation and analysis pipeline. (2022) doi:10.21203/rs.3.rs-1812599/v1.</p> <p>6.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Baril, T., Galbraith, J. &amp; Hayward, A.&nbsp;<em>Earl Grey</em>. (Zenodo, 2023). doi:10.5281/ZENODO.8116025.</p> <p>7.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Li, W. &amp; Godzik, A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences.&nbsp;<em>Bioinformatics</em>&nbsp;<strong>22</strong>, 1658&ndash;1659 (2006).</p> <p>8.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Fu, L., Niu, B., Zhu, Z., Wu, S. &amp; Li, W. CD-HIT: accelerated for clustering the next-generation sequencing data.&nbsp;<em>Bioinformatics</em>&nbsp;<strong>28</strong>, 3150&ndash;3152 (2012).</p> <p>9.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Goubert, C.&nbsp;<em>et al.</em>&nbsp;A beginner&rsquo;s guide to manual curation of transposable elements.&nbsp;<em>Mob. DNA</em>&nbsp;<strong>13</strong>, 7 (2022).</p> <p>10.&nbsp;&nbsp;&nbsp;Camacho, C.&nbsp;<em>et al.</em>&nbsp;BLAST+: Architecture and applications.&nbsp;<em>BMC Bioinformatics</em>&nbsp;<strong>10</strong>, 1&ndash;9 (2009).</p> <p>11.&nbsp;&nbsp;&nbsp;Katoh, K. &amp; Standley, D. M. MAFFT multiple sequence alignment software version 7: Improvements in performance and usability.&nbsp;<em>Mol. Biol. Evol.</em>&nbsp;<strong>30</strong>, 772&ndash;780 (2013).</p> <p>12.&nbsp;&nbsp;&nbsp;Notredame, C., Higgins, D. G. &amp; Heringa, J. T-coffee: a novel method for fast and accurate multiple sequence alignment.&nbsp;<em>J. Mol. Biol.</em>&nbsp;<strong>302</strong>, 205&ndash;217 (2000).</p> <p>13.&nbsp;&nbsp;&nbsp;Rice, P., Longden, L. &amp; Bleasby, A. EMBOSS: The European Molecular Biology Open Software Suite.&nbsp;<em>Trends Genet.</em>&nbsp;<strong>16</strong>, 276&ndash;277 (2000).</p> <p>14.&nbsp;&nbsp;&nbsp;Wheeler, T. J. &amp; Eddy, S. R. nhmmer: DNA homology search with profile HMMs.&nbsp;<em>Bioinformatics</em>&nbsp;<strong>29</strong>, 2487&ndash;2489 (2013).</p> <p>15.&nbsp;&nbsp;&nbsp;Wicker, T.&nbsp;<em>et al.</em>&nbsp;A unified classification system for eukaryotic transposable elements.&nbsp;<em>Nat. Rev. Genet.</em>&nbsp;<strong>8</strong>, 973&ndash;982 (2007).</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials

<p>Toxicogenomics (TGx) approaches are increasingly applied to gain insight into the possible toxicity mechanisms of engineered nanomaterials (ENMs). Omics data can be valuable to elucidate the mechanism of action of chemicals and develop predictive models in toxicology. While vast amounts of transcriptomics data from ENM exposures have already been accumulated, a unified, easily accessible and reusable collection of transcriptomics data for ENMs is currently lacking. In an attempt to improve the FAIRness of already existing transcriptomics data for nanomaterials, we curated a collection of homogenized transcriptomics data from human, mouse and rat ENM exposures <em>in vitro</em> and <em>in vivo</em>.</p>

opencc-by-4.0Jul 2020View details →
zenodo44/100

EukRibo: a manually curated eukaryotic 18S rDNA reference database

<p>EukRibo is a manually curated database of reference small-subunit ribosomal RNA gene (18S rDNA) sequences of eukaryotes, specifically aimed at taxonomic annotation of high-throughput metabarcoding datasets. Unlike other reference databases of ribosomal genes, it is not meant to exhaustively capture all publicly available 18S rDNA sequences from the INSDC repositories, but to represent a subset of highly trustable sequences covering the whole known diversity of eukaryotes, with a focus on protists, manually verified taxonomic identifications, and relatively low genetic redundancy.</p> <p>EukRibo is part of a suite of public resources generated by the UniEuk project (www.unieuk.org), which are all designed to follow a common taxonomic framework for maximal interoperability. The high level of taxonomic accuracy of EukRibo, together with a newly designed, phylogenetically-informed annotation approach, allow high confidence in the taxonomic annotation of environmental metabarcodes, as well as identification of new eukaryotic diversity at various taxonomic levels using a connected components approach.</p> <p>*&nbsp;&nbsp; *&nbsp;&nbsp; *</p> <p>Accompanying preprint available at <a href="https://doi.org/10.1101/2022.11.03.515105">https://doi.org/10.1101/2022.11.03.515105</a>.</p> <p>*&nbsp;&nbsp; *&nbsp;&nbsp; *</p> <p><strong>EukRibo ReadMe file, versions 1 and 2</strong></p> <p>Each EukRibo release consists of <strong>4 files</strong>:<br> - a <strong>tsv table </strong>containing the taxonomic and other information about the 18S rDNA sequences included in the release<br> - a <strong>fasta file </strong>containing the <strong>full sequences </strong>as retrieved from the INSDC repositories (NCBI, EMBL-EBI/ENA, DDBJ)<br> - a <strong>fasta file </strong>containing the <strong>variable region V4 </strong>extracted from all these sequences (based on the fragment amplified with the Tara-Oceans V4 primers)<br> - a <strong>fasta file </strong>containing the <strong>variable region V9 </strong>extracted from the subset of sequences where it is present (based on the fragment amplified with the Tara-Oceans V9 primers)</p> <p>The primary goal of EukRibo was to be used to annotate the EukBank meta-dataset of available V4 metabarcoding datasets, and therefore all sequences included in EukRibo contain the variable region V4.<br> Only a subset of these sequences (about 75%) also contain the variable region V9; this is because many 18S rDNA sequences in the INSDC repositories stop before the V9 fragment.</p> <p>Sequences with slightly incomplete V4 or V9 fragments were kept if phylogenetically useful - i.e. if they are the only available representatives of a certain taxonomic lineage.<br> <strong>V4&nbsp;&nbsp; &nbsp;</strong>We allowed up to 50 missing positions in the relatively conserved area at the 5&#39; end of the V4 fragment (for an average fragment length of about 380 bp); no sequence incomplete at the 3&#39; end of the V4 fragment is included.<br> <strong>V9&nbsp;&nbsp; &nbsp;</strong>We allowed up to 30 missing positions in the relatively conserved area at the 3&#39; end of the V9 fragment (for an average length of about 135 bp); no sequence incomplete at the 5&#39; end of the V9 fragment is included.<br> We allowed a higher proportion of missing positions for the V9 region because being more conservative would imply losing too many sequences, including entire taxonomic lineages.</p> <p><strong>Version 1 of EukRibo</strong><br> This is the starting version of EukRibo that was used for the taxonomic annotation of the EukBank dataset, with taxonomy strings that were fixed as of October 2020.<br> - Contains 46,345 sequences with a sufficiently complete V4 region; 46,299 with the actual complete V4 region and 46 (about 0.1%) with missing positions at the 5&#39; end.<br> - Of these, 34,438 also include a sufficiently complete V9 region; 23,226 with the actual complete V9 region and 11,206 (about 33%) with missing positions at the 3&#39; end.</p> <p><strong>Version 2 of EukRibo</strong><br> This is a version of EukRibo that was made taxonomically compatible with version 3 of the EukProt database (<a href="https://doi.org/10.1101/2020.06.30.180687">https://doi.org/10.1101/2020.06.30.180687</a>), with taxonomic revisions as of July 2022 as well as additional information on the included selection of sequences that was not provided in the tsv file of version 1.<br> - Contains the exact same selection of sequences as in version 1, with the addition of genus <em>Meteora</em>, the last remaining known supergroup-level eukaryotic lineage for which an 18S rDNA was not previously available. (The <em>Meteora </em>sequence contains the full V4 fragment but does not include a sufficiently complete V9 fragment.)<br> - Only 34,432 sequences with a sufficiently complete V9 region are now retained because of 6 previously unrecognised chimeric sequences where the V9 fragment does not originate from the same organism as the V4 fragment.</p> <p><strong>Files in EukRibo version 1</strong>:<br> 46345_EukRibo.tsv.gz<br> 46345_EukRibo_full_seqs.fas.gz<br> 46345_EukRibo_V4.fas.gz<br> 34438_EukRibo_V9.fas.gz</p> <p>The tsv file contains 6 columns:<br> <strong>gb_accession </strong>- INSDC accession number of the sequence<br> <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2 </strong>- binning of the taxa into strictly monophyletic clades of evolutionary and/or ecological significance<br> <strong>UniEuk_taxonomy_string </strong>- full UniEuk-compatible taxonomic annotation of the sequence<br> - an unlimited number of levels is allowed (going down to strain for isolated organisms or to clone for environmental sequences)<br> - informal names are used for phylogenetically supported clades without formal name<br> <strong>V9 </strong>- presence (&#39;Y&#39;) or absence (&#39;N&#39;) of a sufficiently complete V9 fragment in the sequence</p> <p><strong>Files in EukRibo version 2</strong>:<br> 46346_EukRibo-02.tsv.gz<br> 46346_EukRibo-02_full_seqs.fas.gz<br> 46346_EukRibo-02_V4.fas.gz<br> 34432_EukRibo-02_V9.fas.gz</p> <p>The tsv file now contains 12 columns:<br> <strong>gb_accession</strong>, <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2</strong>, <strong>UniEuk_taxonomy_string</strong><br> &nbsp;&nbsp; &nbsp;- same columns as in version 1<br> <strong>alternative_strain_names </strong>(new) - provides alternative strain/isolate names when known to help cross-linking genetic data coming from the same organism<br> <strong>V4 </strong>(new) - indicates whether the V4 fragment is complete (&#39;yes - complete&#39;) or missing positions at the 5&#39; end (&#39;yes - partial&#39;)<br> <strong>V9 </strong>(emended content) - now contains more precise information than in version 1 about whether it is complete (&#39;yes - complete&#39;), missing positions at the 3&#39; end (&#39;yes - partial&#39;), or was excluded, and the 6 possible reasons why (&#39;no - missing&#39;, &#39;no - too incomplete&#39;, &#39;no - chimera&#39;, &#39;no - bad quality&#39;, &#39;no - deletion in V9&#39;, &#39;no - Ns in V9&#39;)<br> <strong>EukProt_ID_same_strain </strong>(new) - accession of EukProt datasets from the same isolate<br> <strong>EukProt_ID_different_strain </strong>(new) - accession of EukProt datasets from a different isolate of the same species<br> <strong>columns_modified_since_previous_version </strong>(new) - lists all of the 6 pre-existing columns that have a modified content compared to version 1<br> <strong>remarks </strong>(new) - additional information such as presence of an intron in the V9 fragment, taxonomic identity of the two parts of chimeric sequences, or the presence of Ns or a deletion in the V4 or the V9 fragment (but insufficient to warrant exclusion)</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

RDF version of the data from Saarimaki et al. Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials (Version 1.0.0) [Zenodo Dataset] (2020)

<p>This is an RDFied version of the dataset published by&nbsp;Saarimaki et al. Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials (Version 1.0.0) [Zebodo Dataset] (2020)</p> <p>The original dataset publication DOI:&nbsp;<a href="http://doi.org/10.5281/zenodo.4146981">http://doi.org/10.5281/zenodo.4146981</a></p> <p>The Original publication authors:&nbsp;Saarimaki, Laura Aliisa, Federico, Antonio, Lynch, Iseult, Papadiamantis, Anastasios G., Tsoumanis, Andreas, Melagraki, Georgia, Afantitis, Antreas, Serra, Angela, &amp; Greco, Dario</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

TF-Marker: A comprehensive manually curated database for transcription factors and related markers in specific cell and tissue types in human.

<p>Here, we developed the TF-Marker database (TF-Marker, http://bio.liclab.net/TF-Marker/) which is committed to a comprehensive manual curation of TFs and related markers with experimental evidence in specific cell and tissue types in human. Currently, through reviewing <strong>2,091</strong> published literature, we have manually classified TFs and related markers into five types according to their functions: 1) <strong>TF</strong>: TFs, which regulate the expression of markers; 2) <strong>T Marker</strong>: markers, which are regulated by TFs (TF and T Marker pairs can identify cell types more specifically); 3) <strong>I Marker</strong>: markers, which influence the activity of TFs (I Markers can also influence the development of specific cells and tissues); 4) <strong>TFMarker</strong>: TFs, which play roles as markers (TFMarkers are cell/tissue-specific TFs used as cell markers in biology experiments); and 5) <strong>TF Pmarker</strong>: TFs, which play roles as potential markers. By curating thousands of published literature, <strong>5,905</strong> entries including <strong>1,316</strong> TFs, <strong>1,092</strong> T Markers, <strong>473</strong> I Markers, <strong>1,600</strong> TFMarkers and <strong>1,424</strong> TF Pmarkers, were annotated in <strong>383</strong> cell types and <strong>95</strong> tissue types in human. Moreover, TF-Marker divided markers into disease markers and tissue/cell-specific markers. TF-Marker is an elaborate database, which provides TFs and related markers supported by experimental evidence. We believe TF-Marker will provide strong support for research into cell/tissue-specific TFs and related markers.</p>

opencc-by-4.0Oct 2021View details →
dryad32/100

Supplementary tables S5, S7, S9, S10, original protein models fasta files used for alignments, aligned and manually curated protein modes files used for phylogenies (PHYLIP format), and phylogenetic trees of plant cell wall decomposition gene families from 44 basidiomycete genomes (.tre files)

<p><span><span><span><span><span><span><span><span><span><span><span>Litter-decomposing Agaricales play key role in terrestrial carbon cycling, but little is known about their decomposition mechanisms. We assembled datasets of 42 gene families involved in plant-cell-wall decomposition from seven newly sequenced litter decomposers and 35 other Agaricomycotina members, mostly white-rot and brown-rot species. Using sequence similarity and phylogenetics, we split the families into phylogroups and compared their gene composition across nutritional strategies. Subsequently, we used Raman spectroscopy to examine the ability of litter decomposers, white-rot fungi, and brown-rot fungi to decompose crystalline cellulose. Both litter decomposers and white-rot fungi share the enzymatic cellulose decomposition, whereas brown-rot fungi possess a distinct mechanism that disrupts cellulose crystallinity. However, litter decomposers and white-rot fungi differ with respect to hemicellulose and lignin degradation phylogroups, suggesting adaptation of the former group to the litter environment. Litter decomposers show high phylogroup diversity, which is indicative of high functional versatility within the group, whereas a set of white-rot species shows adaptation to bulk-wood decomposition. In both groups, we detected species that have unique characteristics associated with hitherto unknown adaptations to diverse wood and litter substrates. Our results suggest that the terms white-rot fungi and litter decomposers mask a much larger functional diversity.</span></span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroJun 2020View details →
zenodo32/100

Manually Curated Library of Transposable Elements (TEs) and TE Annotations for Drosophila amaguana

<p>This data collection provides a manually curated library of transposable elements (TEs) for <em>Drosophila amaguana</em>, including consensus sequences, genome-wide TE annotations, and individual TE copy sequences. The library was built using <em>de novo</em> generated by EDTA (Extensive <em>de novo</em> TE Annotator) and RepeatModeler, curated with MCHelper, and further processed for genome annotation using RepeatMasker and OneCodeToFindThemAll.&nbsp;</p> <p>Below is a description of the included files:</p> <ul> <li><strong><code>Dama_curated_TE_library.fasta</code>:</strong>&nbsp;Contains 737 consensus TE sequences manually curated for&nbsp;<em>D. amaguana</em>. Sequence identifiers include classification and origin (e.g., new families or similarity to known elements).&nbsp;The identifier for each sequence in the FASTA file includes information about its classification:<br><br> <ul> <li><strong>For sequences corresponding to potentially new families</strong>: The identifier consists of a three-letter abbreviation for <em>D. amaguana</em> (Dama), followed by the new family identifier and the superfamily name, all separated by underscores. <em>Example: </em>Dama_NF_BELPAO_1.</li> <li><strong>For consensus sequences that show similarity to TE sequences previously reported in other species</strong>: The identifier includes the abbreviation&nbsp;Dama, followed by the superfamily name and an abbreviation for the species in which the TE was previously reported, all separated by underscores. Example:&nbsp;Dama_Helitron-1_DVir.<br><br></li> </ul> </li> <li><strong><code>Dama_TE_annotations.out</code>:</strong> Genome-wide annotation file of TE insertions produced with RepeatMasker and post-processed using OneCodeToFindThemAll to merge fragmented elements.</li> <li><strong><code>Dama_TE_copies.fasta</code>: </strong>FASTA file containing the extracted sequences of all annotated TE copies from the <em>D. amaguana</em> genome.</li> <li><strong><code>TEcopies_sequences.sh</code>:</strong> Shell script used to extract TE copy sequences from the genome using the annotation coordinates.</li> <li><strong><code>Dynamics_Dama.ipynb</code>: </strong>Jupyter Notebook for the analysis of transposable element (TE) dynamics in<strong>&nbsp;</strong><em>D. amaguana.</em></li> </ul>

opencc-by-4.0Nov 2024View details →
dryad32/100

Supplementary tables S5, S7, S9, S10, original protein models fasta files used for alignments, aligned and manually curated protein modes files used for phylogenies (PHYLIP format), and phylogenetic trees of plant cell wall decomposition gene families from 44 basidiomycete genomes (.tre files)

Open the record for dataset details and reuse information.

publicJun 2020View details →
geo20/100

Manual curation for improved genome annotation of the functionally extinct northern white rhinoceros (Ceratotherium simum cottoni)

GEO Series GSE300824. Ceratotherium simum simum. 8 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2026View details →
zenodo16/100

[DEPRECATED] A manually-curated categorisation of Java Maven libraries along Python PyPI Topics (dataset)

<p>This dataset has been superseded by <strong><a title="https://zenodo.org/records/10480832" href="../records/10480832">https://zenodo.org/records/10480832</a></strong></p> <p>&nbsp;</p> <p>-- Test edit (new version?)</p>

restrictedJan 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record