Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
355
datasets available to search
ShareScore release 0.9.0
Dataset results
355 results for “data extraction”
Results of integrating asynchronous data provenance methods, metadata extraction, and similarity detection into a standards-based RDMS
<p>Related to the thesis: "Asynchronous Tracking and Description of Research Data Changes in Distributed Systems with Interoperable Metadata"</p>
"CausalXtract: a flexible pipeline to extract causal effects from live-cell time-lapse imaging data" datasets
<p>Datasets from the article:</p> <p><strong>CausalXtract: a flexible pipeline to extract causal effects from live-cell time-lapse imaging data </strong></p> <p>by Franck Simon, Maria Colomba Comes, Tiziana Tocci, Louise Dupuis, Vincent Cabeli, Nikita Lagrange, Arianna Mencattini, Maria Carla Parrini, Eugenio Martinelli, Hervé Isambert.</p> <p> </p> <p>The <strong>original videos</strong> are uploaded as: "20161230.zip", "20170105.rar", "Video_2017_0517.zip".</p> <p><strong>Details </strong>for each <strong>experiment </strong>can be found in: "Experiments' details.zip".</p> <p>The <strong>ROIs </strong>(ROI: Region of Interest) are uploaded as .tif files in: "EXTRACTED ROIs.zip".</p> <p>The <strong>MATLAB data</strong> is uploaded in "MATLAB_DATA.rar" and includes: the cancer cells' trajectories (subfolder: "TUMOR TRAJECTORIES"), the immune cells' trajectories (subfolder: "IMMUNE TRAJECTORIES"), the ROIs videos as .mat files for the detection of cells (subfolder: "ROI MAT"), the ROIs further cropped for the extraction of shape descriptors (folder: "ROI_TU MAT"). The ROIs videos .mat files included in the last two subfolders are stopped after their apoptosis has been detected.</p>
Data from: Synthesis of [100]-only LiFePO<sub>4</sub> nanosheets for efficient electrochemical lithium extraction from low-grade brines
Open the record for dataset details and reuse information.
Data from: Exhaustive extraction of cyclopeptides from Amanita phalloides: guidelines for working with complex mixtures of secondary metabolites
Open the record for dataset details and reuse information.
Data from: Pioneer habitats drive high plant beta diversity and conservation value of mineral extraction sites
Open the record for dataset details and reuse information.
Data from: Extraction of alpha-tomatine from green tomatoes by subcritical water
Open the record for dataset details and reuse information.
Data from: Optimisation of biogenic synthesis of silver nanoparticles from flavonoid-rich Clinacanthus nutans leaf and stem aqueous extracts
Open the record for dataset details and reuse information.
Regeneration data of bryophyte fragments extracted from feces of Chloephaga picta and Attagis malouinus in Navarino Island, sub-Antarctic Chile
Open the record for dataset details and reuse information.
Extracted data from primary literature examining impacts of recreational activities on freshwater ecosystems
Open the record for dataset details and reuse information.
Labeled data for citation field extraction
Open the record for dataset details and reuse information.
Data for: A new threshold selection method for species distribution models with presence-only data: extracting the mutation point of the P/E curve by threshold regression
Open the record for dataset details and reuse information.
Spermatozoa RRBS data: Post-Bismark aligned files extracted for the methylation call for every single C in the CpG context of sperm cells of Murrah buffalo bulls under heat stress
Open the record for dataset details and reuse information.
Data from: Improving properties of chitosan/PVA films using cashew nut Testa extract – potential applications in food packaging
Open the record for dataset details and reuse information.
Data from: Minimally destructive hDNA extraction method for retrospective genetics of pinned historical Lepidoptera specimens
Open the record for dataset details and reuse information.
A STP-HSI index method for urban built-up area extraction based on multi-source remote sensing data
Open the record for dataset details and reuse information.
Data for: Environmental DNA storage and extraction method affects detectability for multiple aquatic invasive species
Open the record for dataset details and reuse information.
Data from: The effects of lipid extraction on δ13C and δ15N values and use of lipid-correction models across tissues, taxa, and trophic groups
Open the record for dataset details and reuse information.
Sample extraction and SNP sequencing data for: Identification of sex-linked SNP markers in wild populations of monomorphic birds
Open the record for dataset details and reuse information.
MitoFinder: efficient automated large-scale extraction of mitogenomic data in target enrichment phylogenomics
<p><strong>MitoFinder: efficient automated large-scale extraction of mitogenomic data in target enrichment phylogenomics</strong></p> <p>Rémi Allio<sup>1</sup>, Alex Schomaker-Bastos<sup>2,†</sup>, Jonathan Romiguier<sup>1</sup>, Francisco Prosdocimi<sup>2</sup>, Benoit Nabholz<sup>1</sup>, and Frédéric Delsuc<sup>1</sup></p> <p><sup>1</sup><em>Institut des Sciences de l’Evolution de Montpellier (ISEM), CNRS, EPHE, IRD, Université de Montpellier, Montpellier, France.</em></p> <p><sup>2</sup><em>Laboratório Multidisciplinar para Análise de Dados (LAMPADA), Instituto de Bioquímica Médica Leopoldo de Meis, Universidade Federal do Rio de Janeiro, Rio de Janeiro, Brasil.</em></p> <p><sup>†</sup><em> In Memoriam (08/01/2015) </em></p> <p> </p> <p><em><strong>Correspondence</strong></em></p> <p>Rémi Allio</p> <p>Email: <a href="mailto:remi.allio@umontpelier.fr">remi.allio@umontpellier.fr</a></p> <p>Frédéric Delsuc</p> <p>Email: <a href="mailto:frederic.delsuc@umontpellier.fr">frederic.delsuc@umontpellier.fr</a></p> <p> </p> <p><strong><em>Running head</em></strong></p> <p>Mitochondrial signal from UCE capture data</p> <p> </p> <p><strong>Abstract</strong><strong> </strong></p> <p>Thanks to the development of high-throughput sequencing technologies, target enrichment sequencing of nuclear ultraconserved DNA elements (UCEs) now allows routinely inferring phylogenetic relationships from thousands of genomic markers. Recently, it has been shown that mitochondrial DNA (mtDNA) is frequently sequenced alongside the targeted loci in such capture experiments. Despite its broad evolutionary interest, mtDNA is rarely assembled and used in conjunction with nuclear markers in capture-based studies. Here, we developed MitoFinder, a user-friendly bioinformatic pipeline, to efficiently assemble and annotate mitogenomic data from hundreds of UCE libraries. As a case study, we used ants (Formicidae) for which 501 UCE libraries have been sequenced whereas only 29 mitogenomes are available. We compared the efficiency of four different assemblers (IDBA-UD, MEGAHIT, MetaSPAdes, and Trinity) for assembling both UCE and mtDNA loci. Using MitoFinder, we show that metagenomic assemblers, in particular MetaSPAdes, are well suited to assemble both UCEs and mtDNA. Mitogenomic signal was successfully extracted from all 501 UCE libraries allowing confirming species identification using COI barcoding. Moreover, our automated procedure retrieved 296 cases in which the mitochondrial genome was assembled in a single contig, thus increasing the number of available ant mitogenomes by an order of magnitude. By leveraging the power of metagenomic assemblers, MitoFinder provides an efficient tool to extract complementary mitogenomic data from UCE libraries, allowing testing for potential mito-nuclear discordance. Our approach is potentially applicable to other sequence capture methods, transcriptomic data, and whole genome shotgun sequencing in diverse taxa.</p> <p> </p> <p><strong><em>Figures & Tables</em></strong></p> <p><strong>Figure 1.</strong> Conceptualization of the pipeline used to assemble and extract UCE and mitochondrial signal from ultraconserved element sequencing data.</p> <p><strong>Figure 2</strong>. Comparison of the efficiency of the assemblers in terms of: A) computational time, B) number of potentially mitochondrial contigs identified, and C) number of mitochondrial genes annotated. Violin plots reflect the data distribution with a horizontal line indicating the median. Note that for the three metagenomic assemblers, 5 CPUs were used compared to 35 CPUs for Trinity. Plots were obtained using PlotsOfData (Postma & Goedhart 2019).</p> <p><strong>Figure 3.</strong> Phylogenomic relationships of ants (Formicidae). AA) Mito-nuclear phylogenetic differences among subfamily relationships based on the UCE and mtDNA supermatrices obtained with the assembler MetaSPAdes assembler. Clades corresponding to subfamilies were collapsed. Inter-subfamily relationships with UFBS < 95% were collapsed. Non-maximal node support values are reported. B) The topology obtained reflects the results of phylogenetic analyses based on the amino acid mitochondrial supermatrix (using MetaSPAdes as assembler). Histograms reflect the percent of UCEs (light grey) and mitochondrial genes (dark grey) recovered for each species. Illustrative pictures (*): <em>Diacamma sp</em>. (Ponerinae; top left), <em>Formica sp</em>. (Formicinae; top right), and <em>Messor barbarus </em>(Myrmicinae; bottom right).</p> <p><strong>Table 1. </strong>Summary statistics on assembly results according to the assembler used. The values are averages over the 501 assemblies, except for the assembly time, which is a median value. The two tables report specific statistics for A) ultraconserved elements data, and B) mitochondrial data. Note that 35 CPUs were used for Trinity whereas 5 CPUs were used for other assemblers.</p> <p><strong>Table 2.</strong> Statistical comparison between the performances of the different assemblers. Statistical significance was estimated with a paired non parametric test (paired wilcoxon test). *** = <em>p</em><0.001; ** = <em>p</em><0.01; * = <em>p</em><0.05; NS = <em>p</em>>0.05; and (+)/(-) is the result of the comparison between the row and the column.</p> <p> </p> <p><strong><em>Appendices</em></strong></p> <p><strong>Appendix S1.</strong> List of the 501 UCE libraries (SRA accessions) and associated metadata.</p> <p><strong>Appendix S2.</strong> Summary statistics on mitochondrial signal recovered per species and depending on the assembler used. The table provides the number of contigs and genes recovered with MitoFinder and the size of each annotated gene.</p> <p><strong>Appendix S3.</strong> Summary statistics of barcoding analyses. Detailed results for both BOLDsystem and Megablast analyses are provided for each CO1 recovered with MitoFinder using MetaSPAdes.</p> <p><strong>Appendix S4.</strong> Detailed results of tree distance analyses realized with Dquad (Ranwez, Criscuolo, & Douzery 2010). Trees obtained with each assembler with mitochondrial amino acid supermatrix, mitochondrial nucleotide supermatrix, and UCE nucleotide supermatrix were compared with each others.</p> <p><strong>Appendix S5</strong>. List of Genbank accession numbers for newly generated mitchondrial contigs.</p> <p> </p> <p><strong><em>Zenodo supplementary files</em></strong></p> <p><strong>Assembly_results.tar.gz</strong> Contains all contigs obtained for each species with the different assemblers implemented in MitoFinder.</p> <p><strong>MitoFinder_annotations.tar.gz</strong> Contains MitoFinder annotations for each species. (based on the contigs obtained with MetaSPAdes)</p> <p><strong>UCE_results.tar.gz</strong> Contains all annotated UCE obtained for each species after UCE identification with PHYLUCE. (MetaSPAdes)</p> <p><strong>Final_mtDNA_alignments.tar.gz</strong> Contains the final mitochondrial gene alignments. (MetaSPAdes)</p> <p><strong>Final_UCE_alignments.tar.gz</strong> Contains the final UCE alignments. (MetaSPAdes)</p> <p><strong>Final_mtDNA_matrices.tar.gz</strong> Contains the final mi tochondrial supermatrices (AA and NT) used for the phylogenetic analyses. (MetaSPAdes)</p> <p><strong>Metaspades_final_UCE_matrix.phy</strong> The final UCE supermatrix used for the phylogenetic analyses. (MetaSPAdes)</p>
Fine-grained automated visual analysis of herbarium specimens for phenological data extraction: an annotated dataset of reproductive organs in Strepanthus herbarium specimens
<p>This dataset contains annotations of 31 herbarium specimens of <em>Streptanhus tortuosus Kellogg</em> for which we have we carefully and manually drew and annotated the contours of four reproductive organs: “bud”, “flower”, “immature fruit” and “mature fruit”.</p> <p>The dataset can be used to assess the ability of automated methods to count and detect precisely the shapes of these reproductive organs, with a view to conducting phenological studies.</p> <p>The annotations are formatted in accordance with the COCO data format, a usual format for object detection tasks in the field of Computer Vision. The annotations are divided into two files:</p> <ul> <li>train_21_full_masks.json contains the mask coordinates and labels of 21 herbarium sheets that can be used for training models</li> <li>test_10_full_masks.json contains the mask coordinates and labels of 10 other herbarium that can be used as a groundtruth file for evaluating the predictions, typically with the COCO evaluation scripts (<a href="https://github.com/cocodataset/cocoapi">https://github.com/cocodataset/cocoapi</a>)</li> </ul> <p>Please refer to the following publication for a first assessment of this dataset with a Mask-RCNN approach:</p> <p><em>H. Goëau, A. Mora-Fallas, J. Champ, N. Love, S. Mazer, E. Mata-Montero, A. Joly, P. Bonnet. </em>2020. New fine-grained method for automated visual analysis of herbarium specimens: a case study for phenological data extraction. <em>Applications in Plant Sciences </em></p> <p> </p> <p> </p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.