Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
15
datasets available to search
ShareScore release 0.9.0
Dataset results
15 results for “SAMPL challenge”
Challenges of sampling and how phylogenetic comparative methods help: Supplementary data
<p>Supplementary data and results files for the paper:</p> <p>Macklin-Cordes, Jayden L. & Erich R. Round (2022). Challenges of sampling and how phylogenetic comparative methods help: With a case study of the Pama-Nyungan laminal contrast. <em>Linguistic Typology</em> (advance online publication). <a href="https://doi.org/10.1515/lingty-2021-0025">https://doi.org/10.1515/lingty-2021-0025</a></p>
Data from: Treated like dirt: Robust forensic and ecological inferences from soil eDNA after challenging sample storage
<p>We investigated the effect of storage duration and conditions on the assessment of the soil biota with eDNA metabarcoding. We extracted eDNA from freshly collected soil samples and again from the same samples after storage under contrasting temperature conditions and contrasting exposure (open/closed tubes). We used four different primer sets targeting bacteria, fungi, protists (cercozoans), and general eukaryotes. <span>W</span>e quantified differences in richness, evenness, and community composition. Subsequently, we tested whether we could correctly infer habitat type and original sample identity after storage using a large reference dataset.</p> <p>This repository contains the un-demultiplexed fastq sequences.</p>
Data for manuscript, "An optimized workflow for MS-based quantitative proteomics of challenging clinical bronchoalveolar lavage fluid (BALF) samples"
<p>Clinical BALF samples are rich in biomolecules, including proteins, and useful for molecular studies of lung health and disease. However, MS based proteomic analysis of BALF is impeded by the dynamic range of protein abundance, and potential for interfering contaminants. We have developed a workflow that eliminates these challenges. By combining high abundance protein depletion, protein trapping, clean-up, and in-situ tryptic digestion, our workflow is compatible with both qualitative and quantitative MS-based proteomic analysis. The workflow includes collection of endogenous peptides for peptidomic analysis of BALF, if desired, as well as amenability to offline semi-preparative or microscale fractionation of peptide mixtures prior to LC-MS/MS analysis, for increased depth of analysis. We show the effectiveness of this workflow on BALF samples from COPD patients. Overall, our workflow should allow MS-based proteomics to be applied to a wide variety of studies focused on BALF clinical samples. </p> <p>Note: Due to the nature of some of the files, file <em>wendt005_ostr0103_18260_20210831_BALF_FAIMS_MS2_TMT16.msf, wendt005_ostr0103_18976_20230202_quantReport.msf, cmsptc_higgi022_18988_20230203_18976DW_EnF_hcdlT_1R.raw, cmsptc_higgi022_18988_20230203_18976DW_EnF_hcdlT_2R.raw, cmsptc_higgi022_18988_20230203_18976DW_EnF_hcdlT_3R.raw and cmsptc_higgi022_18988_20230203_18976DW_Eclipse_noFAIMS_quantReport.msf</em> were zipped into compressed folders before uploading.</p>
Data from: Treated like dirt: Robust forensic and ecological inferences from soil eDNA after challenging sample storage
Open the record for dataset details and reuse information.
Data from: Overcoming the challenge of small effective sample sizes in home-range estimation
Technological advances have steadily increased the detail of animal tracking datasets, yet fundamental data limitations exist for many species that cause substantial biases in home‐range estimation. Specifically, the effective sample size of a range estimate is proportional to the number of observed range crossings, not the number of sampled locations. Currently, the most accurate home‐range estimators condition on an autocorrelation model, for which the standard estimation frame‐works are based on likelihood functions, even though these methods are known to underestimate variance—and therefore ranging area—when effective sample sizes are small. Residual maximum likelihood (REML) is a widely used method for reducing bias in maximum‐likelihood (ML) variance estimation at small sample sizes. Unfortunately, we find that REML is too unstable for practical application to continuous‐time movement models. When the effective sample size N is decreased to N ≤ urn:x-wiley:2041210X:media:mee313270:mee313270-math-0001(10), which is common in tracking applications, REML undergoes a sudden divergence in variance estimation. To avoid this issue, while retaining REML's first‐order bias correction, we derive a family of estimators that leverage REML to make a perturbative correction to ML. We also derive AIC values for REML and our estimators, including cases where model structures differ, which is not generally understood to be possible. Using both simulated data and GPS data from lowland tapir (Tapirus terrestris), we show how our perturbative estimators are more accurate than traditional ML and REML methods. Specifically, when urn:x-wiley:2041210X:media:mee313270:mee313270-math-0002(5) home‐range crossings are observed, REML is unreliable by orders of magnitude, ML home ranges are ~30% underestimated, and our perturbative estimators yield home ranges that are only ~10% underestimated. A parametric bootstrap can then reduce the ML and perturbative home‐range underestimation to ~10% and ~3%, respectively. Home‐range estimation is one of the primary reasons for collecting animal tracking data, and small effective sample sizes are a more common problem than is currently realized. The methods introduced here allow for more accurate movement‐model and home‐range estimation at small effective sample sizes, and thus fill an important role for animal movement analysis. Given REML's widespread use, our methods may also be useful in other contexts where effective sample sizes are small.
Data from: Overcoming the challenge of small effective sample sizes in home-range estimation
Open the record for dataset details and reuse information.
Raw data of sequencing results of our study: Bovine milk microbiota: Evaluation of different DNA extraction protocols in challenging samples
<p>Clean reads of the repeated milk samples with used Primer Pairs V1V2 and V3V4</p> <p>Raw data of sequencing results (amplicon single variants)</p>
CAMI2 Challenge - Human Microbiome Project Toy Database - sample 19 - regenerated using recent RefSeq representative genomes
Open the record for dataset details and reuse information.
Spatially resolved transcriptomic profiling of degraded and challenging fresh frozen samples
GEO Series GSE221571. Mus musculus. 14 samples. Type: Other.
Mechanisms and pathways of bone metastasis: challenges and pitfalls of performing molecular research on patient samples
GEO Series GSE14776. Homo sapiens. 14 samples. Type: Expression profiling by array.
The Challenge of Stability in High-Throughput Gene Expression Analysis: Comprehensive Selection and Evaluation of Reference Genes for BALB/c Mice Spleen Samples in the Leishmania infantum Infection Mo
GEO Series GSE80709. Mus musculus. 47 samples. Type: Expression profiling by RT-PCR.
Performance of Immunochip 1.0 with hemolymph samples collected at 3 and 48 hours from Vibrio-challenged mussels
GEO Series GSE23535. Mytilus galloprovincialis. 8 samples. Type: Expression profiling by array.
Single cell RNA sequencing analysis of bronchoalveolar lavage samples from SARS-CoV-2 challenged Rhesus macaques
GEO Series GSE190913. Macaca mulatta. 31 samples. Type: Expression profiling by high throughput sequencing.
Dementia Care Challenges Among a Sample of Informal Caregivers in Egypt
ClinicalTrials.gov study NCT05572086. IPD Sharing: Not stated. Countries: 0. Publications: 0.
Blood samples from COPD subjects and Healthy Volunteers taken before and after LPS challenge
GEO Series GSE112811. Homo sapiens. 64 samples. Type: Expression profiling by array.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.