Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.9.0
Dataset results
4 results for “match statistics”
Data set for ODI cricket matches from 1987 to 2023 (extracted from ESPN Cricinfo) and code (R) used for a statistical study
<p>Here I present the data and code that has been used to study the statistical evolution of ODI cricket. The preprint for this research is available at: </p> <div> <div> <div> <table> <tbody> <tr> <td><a href="https://doi.org/10.48550/arXiv.2406.11652">https://doi.org/10.48550/arXiv.2406.11652</a> <div><span>Focus to learn more</span></div> </td> </tr> </tbody> </table> </div> </div> </div> <div> </div>
Minimizer collision statistics (BLEND: A Fast, Memory-Efficient, and Accurate Mechanism to Find Fuzzy Seed Matches in Genome Analysis)
<p>This dataset includes the statistics for the minimizers that generate the same hash value (i.e., collisions). The hash values are generated using a low-collision hash function and the SimHash technique in BLEND.</p> <p> </p> <p>*collision_stats.txt files include the overall collision statistics for a tool and configuration of the tool (i.e., the number of colliding minimizer pairs with a certain edit distance and their ratio to all number of collisions). For example blend_n3_collision_stats.txt shows the statistics for BLEND where the number of neighbors is set to 3 when running BLEND.</p> <p>_sim.csv files include all minimizer pairs with the same hash value and the edit distance between them.</p> <p> </p>
K-mer collision statistics (BLEND: A Fast, Memory-Efficient, and Accurate Mechanism to Find Fuzzy Seed Matches in Genome Analysis)
<p>This dataset contains 1,077 FASTA files and CSV files. Each FASTA file includes 25-character long sequences similar to each other.</p> <p>We have a CSV file for each tool (i.e., minimap2 and BLEND) and configuration (i.e., different number of neighbors in BLEND). CSV files include the non-identical k-mer pairs (16-mers) that generate the same hash value (i.e., collisions). These k-mers are extracted from sequences that are similar to each other. In each line, we show the hash value of the k-mers, the actual sequene pairs that the k-mers are extracted from, k-mer pairs that generate the same hash value, and the edit distance between these k-mers.</p> <p> </p>
Ultra-high throughput sequencing-based small RNA discovery and discrete statistical biomarker analysis in a collection of cervical tumors and matched controls
GEO Series GSE20592. Homo sapiens. 60 samples. Type: Non-coding RNA profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.