Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

35

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

35 results for “Read Mapping”

Learn how ShareScore rates datasets ↗
zenodo44/100

Explore maps as you read a comic book

<p>This repository contains three documents accompanying the article 'Explore maps as you read a comic book.</p> <p>The first document compiles several types of breaks encountered in various pan-scalar maps. Each type is also assigned a progression note after an analysis of their effects on the user.</p> <p>The next table illustrates an example of good generalisation for each of the hydrographic patterns observed in the article. This generalisation is based on three key concepts of progressivity, promoting a continuity of meaning, coherence, and rhythm.</p> <p>The final document represents a scale master inspired by the methodology of (Brewer and Buttenfield, 2007). In the article, we demonstrate how to adapt it to pan-scalar maps to better visualize sequences, as well as two types of rhythms: map cadence and abstraction cadence.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Data for Bovine breed-specific augmented reference graphs facilitate accurate sequence read mapping and unbiased variant discovery

<p><strong>Description of the datasets</strong></p> <p>Data are organized as folders and compressed with tar.gz.</p> <p>There are two compressed data folder: <strong>data </strong>which used for cattle genome graphs experiment and&nbsp;<strong>data_human</strong> which we used for human genome graphs experiment.&nbsp;</p> <p><strong>Cattle genome graphs experiments</strong></p> <p>First you need to unzip the file using command <em>tar -xvzf data.tar.gz</em>. After unzipping, the data folder is organized as follows:</p> <ul> <li>Utilities: contain bovine ARS-UCD 1.2 fasta reference with the accompanying index.</li> <li>Bin: contain the softwares used in the paper (vg, liftover, vcf2diploid)</li> <li>Part1: data for analysis in variant prioritization section, further subdivided into: <ul> <li>vcf_sim: variant files from four animal in each breed used to simulate reads</li> <li>reads_sim: simulated short reads used for read mapping</li> <li>vcf_freq: variants augmented to graphs filtered based on allele frequency</li> </ul> </li> <li>Part2: data used for analysis in the section of graph mapping with breeds-filtered variants, further subdivided into: <ul> <li>vcf_breed: variant files used to graphs construction.</li> </ul> </li> <li>Part3: data used for analysis in the section of consensus genome, further subdivided into: <ul> <li>read_sims: simulated reads as in the part1, but the coordinates are liftovered to the new consensus genomes.</li> <li>reference: contain the original reference and consensus references.</li> <li>vcf_consensus: contain major allele variants to construct consensus genomes.</li> </ul> </li> <li>Part4: data analysis in the section of whole genome graph construction and variant genotyping. <ul> <li>vcf_construct: variants from chromosome 1-29 from 82 Brown Swiss used to construct BSW whole genome graph.</li> <li>BSW_graph: whole genome Brown Swiss graph with the three accompanying indexes (xg,gcsa, and gbwt).</li> </ul> </li> </ul> <p><strong>Human genome graphs experiments</strong></p> <p>First you need to unzip the <em>data_human</em> file using command <em>tar -xvzf data</em><em>_hum.tar.gz</em>. After unzipping, the data folder is organized as follows:</p> <ul> <li>reference: the g1k_v37 reference used as a graph backbone</li> <li>vcf_sim: variant files from four individuals in each population used to simulate reads</li> <li>reads_sim: simulated short reads used for read mapping</li> <li>vcf_freq: variants augmented to graphs filtered based on allele frequency</li> </ul>

opencc-by-4.0Dec 2019View details →
zenodo40/100

Sistematic Mapping review of literature about public policies innovation and reading competences

<p>The systematic mapping method makes possible to identify what is known about a topic, what has been researched, what aspects remain unknown, and to obtain information about current trends and future challenges regarding this topic.&nbsp;In this Excel , the following articles were included: studies on innovative public policies related to reading skills, those published only in journals that can be found in two databases: Web of Science (WOS) and Scopus, published from January 2015 to December 2019, and related to the educational sector (social sciences, psychology, multidisciplinary, educational research, medical education, environmental education).</p>

opencc-by-4.0Oct 2020View details →
dryad40/100

Mapped read data and files and scripts from: Vicariance followed by secondary gene flow in a young gazelle species complex

<p>Grant's gazelles have recently been proposed to be a species complex comprising three highly divergent mtDNA lineages (<em>Nanger granti</em>, <em>N. notata</em> and <em>N. petersii</em>). The three lineages have non-overlapping distributions in East Africa, but without any obvious geographical divisions, making them an interesting model for studying the early stage evolutionary dynamics of allopatric speciation in detail. Here we use genomic data obtained by restriction site-associated (RAD) sequencing of 106 gazelle individuals to shed light on the evolutionary processes underlying Grant's gazelle divergence, to characterize their genetic structure and to assess the presence of gene flow between the main lineages in the species complex. We date the species divergence to 134,000 years ago, which is recent in evolutionary terms. We find population subdivision within <em>N. granti</em>, which coincides with the previously suggested two subspecies, <em>N.g. granti</em> and <em>N.g. robertsii</em>. Moreover, these two lineages seem to have hybridized in Masai Mara. Perhaps more surprisingly given their extreme genetic differentiation, <em>N. granti</em> and <em>N. petersii</em> also show signs of prolonged admixture in Mkomazi, which we identified as a hybrid population most likely founded by allopatric lineages coming into secondary contact. Despite the admixed composition of this population, elevated X-chromosomal differentiation suggests that selection may be shaping the outcome of hybridization in this population. Our results therefore provide detailed insights into the processes of allopatric speciation and secondary contact in a recently radiated species complex.</p>

opencc-zeroOct 2020View details →
zenodo40/100

Read-mapping for next-generation sequencing data (Drosophila melanogaster)

<p>Code, logs and quality-control data for whole-genome resequencing of Sussex-LH<sub>M</sub> and RG <em>Drosophila melanogaster</em>.</p> <p>Mapping code is in the archive lhm_mapping_scripts.zip</p> <p>Mapping logs are in the in the archive lhm_mapping_logs.zip</p> <p>Other zip archives contain the quality control data.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>

opencc-by-4.0Oct 2016View details →
zenodo40/100

Read-mapping for next-generation sequencing data (Wolbachia)

<p>Code, log files and QC data. NCBI SRA accession number SRP091004. Note that sequencing Wolbachia was not a central aim of the project, and was undertaken in order to maximise the amount of information that could be extracted from the raw genome sequence data targetted at the fruit-fly host.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>

opencc-by-4.0Oct 2016View details →
zenodo40/100

Extension of Partial Gene Transcripts by Iterative Mapping of RNA-Seq Raw Reads

<p>Trinity assembled transcriptome of <em>Drosophila&nbsp;melanogaster</em> and <em>Osmia </em><em>bicornis</em></p>

opencc-by-4.0Jun 2018View details →
zenodo40/100

Data related to research article: Towards mouse genetic-specific RNA-sequencing read mapping

<p>This dataset contains data related to the research article: &quot;Towards mouse genetic-specific RNA-sequencing read mapping&quot;.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Baseline assemblies for "ntLink: a toolkit for de novo genome assembly scaffolding and mapping using long reads" protocol

<p>ntLink is a flexible&nbsp;<em>de novo</em> genome scaffolding toolkit which can be run in various modes depending on the desired user output, with multiple new functionalities recently introduced. Here, we provide the baseline assembly datasets used in the ntLink protocol paper &quot;ntLink: a toolkit for <em>de novo </em>genome assembly scaffolding and mapping using long reads&quot;. The provided assemblies are ABySS (short-read) and Flye (long-read) assemblies of&nbsp;<em>Caenorhabditis elegans&nbsp;</em>genome sequencing data. The ABySS (v2.1.4) assembly utilized paired-end short reads (accession&nbsp;DRR008444), and was run with the following parameters:&nbsp;k=64&nbsp;l=40 s=1000&nbsp;q=15 B=10G j=8&nbsp;kc=3&nbsp;H=4 S=1000-10000 N=9.The&nbsp;<em>C. elegans</em>&nbsp;Flye (v2.5) assembly was run using Oxford Nanopore long reads (accession SRR10028109) and the following parameters:&nbsp;--nano-raw SRR10028109.fastq&nbsp;-g100m -t48.</p>

opencc-by-4.0Jan 2023View details →
dryad40/100

Variant calling in the Goldilocks Zone: how reference genome choice and read mapping stringency impact heterozygosity estimates and phylogenetic analyses

Open the record for dataset details and reuse information.

publicNov 2025View details →
dryad40/100

Mapped read data and files and scripts from: Vicariance followed by secondary gene flow in a young gazelle species complex

Open the record for dataset details and reuse information.

publicOct 2020View details →
zenodo36/100

Next-generation sequencing of newborn screening genes: The accuracy of short-read mapping

<p>We examine the effect of high homology genomic regions on the mapping of genes related to newborn screening while taking different read lengths and patient&#39;s ethnic background into consideration.</p>

opencc-by-4.0Aug 2020View details →
zenodo36/100

DoTA-seq processed datasets (barcode: read mappings)

<p>Processed DoTA-seq sequencing reads to generate barcode: read mapping tables. This data is further processed to generate figures 1-3.</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Test data for RonaQC - mapped SARS-CoV-2 reads

<p>This dataset includes test data for <a href="https://ronaqc.netlify.app/">RonaQC</a></p> <p>RonaQC accepts mapped SARS-CoV-2 reads (BAM format), generated from the SARS-CoV-2 bioinformatic pipelines like ARTIC, and any control samples from the respective sequencing run (negative/positive) as input. It will then assess the levels of cross contamination and primer contamination in the samples, and determine if the samples are reliable for detecting SARS-CoV-2, phylogenetic analysis, and/or submission to public databases.</p> <p><br> The dataset includes SARS-CoV-2 sequenced reads compiled by <a href="https://github.com/CDCgov/datasets-sars-cov-2">CDCgov/datasets-sars-cov-2</a>&nbsp;[1].&nbsp;</p> <p>These were reads were processed using the <a href="https://github.com/connor-lab/ncov2019-artic-nf">ncov2019-artic-nf pipelines</a>, which is a Nextflow pipeline for running the <a href="https://github.com/artic-network/fieldbioinformatics">ARTIC network&#39;s fieldbioinformatics tools</a>,&nbsp;with a focus on ncov2019.&nbsp;</p> <p><br> This dataset includes:&nbsp;</p> <ul> <li><strong>FailedQC </strong>- A cohort of 24 samples failed basic QC metrics, covering 8 possible failure scenarios, Illumina platform, amplicon-based approach&nbsp;&nbsp; &nbsp;</li> <li><strong>VOCRepresentatives </strong>- A cohort of 16 samples from 10 representative CDC defined VOI/VOC lineages as of 06/15/2021, Illumina platform, amplicon-based approach&nbsp;&nbsp; &nbsp;</li> <li><strong>Test </strong>- Smaller test samples, including sequenced negative controls of varying quality</li> </ul> <p>[1] &nbsp;Timme, Ruth E., et al. &quot;Benchmark datasets for phylogenomic pipeline validation, applications for foodborne pathogen surveillance.&quot; PeerJ 5 (2017): e3893.&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Mapped reads for the 24h labeling experiment of Herzog et al., Nature Methods 2017

<p>This are the mapped reads for samples GSM2666816,GSM2666817,GSM2666818,GSM2731767,GSM2731768,GSM2731769 against the Mouse genome (mm10, gtf from Ensembl 90). Reads were mapped (after rRNA removel) using STAR:</p> <p>STAR --runMode alignReads --runThreadN 8 --genomeDir ens90.STAR-index --readFilesIn X.fastq --outSAMmode NoQS&nbsp; --outSAMattributes nM MD&nbsp;</p> <p>These are the test data for GRAND-SLAM (https://github.com/erhard-lab/gedi/wiki/GRAND-SLAM)</p>

opencc-by-4.0Mar 2019View details →
zenodo36/100

Inputs & results for "Identifying circular DNA using short-read mapping"

<div>* `inputs`: Contains genome assemblies, annotations, and sample sheets for the Nextflow pipeline</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; * `parasite_species`: Contains the sample sheet and input data for the parasite/related species dataset</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; * `parasitoid_wasps`: Contains the sample sheet and input data for the parasitoid wasp dataset</div> <div>&nbsp;</div> <div>* `results`: Contains filtered BAM and coverage files, figures, and `GenomeInfo`-filtered example files for the parasitoid wasp and parasite/related species datasets</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; * `parasite_species`:</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; * `bam_files`: Contains BAM files with mapped distances &gt;= 1 kb</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;* `coverage_files`: Contains coverage depth files filtered by the BAM files</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;* `insert_filtered_results`: Contains examples of filtered file outputs from the `GenomeInfo` class. The complete set of outputs can be found on Zenodo.</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;* `fig`: Figures generated for the pub.</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; * `parasitoid_wasps`:</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;* `bam_files`: Contains BAM files with mapped distances &gt;= 1 kb</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;* `coverage_files`: Contains coverage depth files filtered by the BAM files</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;* `fig`: Figures generated for the pub.</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;* `hyposoter_didymator_blastx_results`: Contains the manual BLASTx results from the _Hyposoter didymator_ search.</div>

opencc-by-4.0Aug 2024View details →
dryad36/100

How to select appropriate hue ranges for sequential color schemes on choropleth maps? A quantitative evaluation using map reading experiments

<p>We propose map reading experiments to quantitatively evaluate the selection of hue ranges for sequential color schemes on choropleth maps. In these experiments, 60 sequential color schemes with six base hues and ten hue ranges were employed as experimental color schemes, and a total of 414 college students were invited to complete identification, comparison, and ranking tasks. Both controlled and real-map experiments were performed, each involving a web-based survey and an eye-tracking experiment. In the controlled experiments, the shapes of the map objects were relatively regular, and attribute data were randomized. In contrast, the shapes were complex in real-map experiments, and real data were employed. Our findings show that widely used color schemes with a hue range of 0º yield poor performance in all tasks; 15º hue ranges yield good performance in the comparison and ranking tasks but poor performance in the identification task. For large hue ranges of 120-360º, participants showed good performance in the identification task but poor performance in the comparison and ranking tasks. For 30-60º hue ranges, participants achieved excellent performance in the comparison and ranking tasks and acceptable performance in the identification task. We also found that the ratings of 0-60º ranges were high.</p>

opencc-zeroAug 2023View details →
zenodo36/100

Public database of multilingual map reading test

<p>The Excel file contains the filtered data records of the map-reading study of the Research Group on Experimental Cartography at the E&ouml;tv&ouml;s Lor&aacute;nd University (ktk.elte.hu). The data collection started in the autumn of 2015 and lasted until April 2022.&nbsp;The file contains three sheets: demographic_questions; correct_answers; map_reading_database. The first two sheets contain the questions asked, the&nbsp;answer codes, and the correct answers. The third one has 511 records, which is the result of a filtering of&nbsp;the original 805 fills. The filtering excluded the unfinished tests, and the ones with fill time below 2.5 minutes and above 15 minutes.</p>

opencc-by-4.0Sep 2023View details →
dryad36/100

How to select appropriate hue ranges for sequential color schemes on choropleth maps? A quantitative evaluation using map reading experiments

Open the record for dataset details and reuse information.

publicAug 2023View details →
dryad32/100

Hooded crow genome assembly (PacBio long reads + BioNano optical maps + DoveTail HiC maps)

<p>Structural variation (SV) is an important component of mutations providing the raw material for evolution. Here, we uncover the genome-wide spectrum of intra- and interspecific SV segregating in natural populations of seven songbird species in the genus Corvus. Combining short-read (N = 127) and long-read re-sequencing (N = 31), as well as optical mapping (N = 16), we apply both assembly- and read mapping approaches to detect SV and characterize a total of 220,452 insertions, deletions and inversions. We exploit sampling across wide phylogenetic timescales to validate SV genotypes and assess the contribution of SV to evolutionary processes in an avian model of incipient speciation. We reveal an evolutionary young (~530,000 years) cis-acting 2.25-kb retrotransposon insertion reducing expression of the NDP gene with consequences for premating isolation. Our results attest to the wealth and evolutionary significance of SV segregating in natural populations and highlight the need for reliable SV genotyping.</p>

opencc-zeroJun 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record