Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.9.0
Dataset results
5 results for “Censored Data”
Simulated data used in "The importance of censoring in competing risks analysis of the subdistribution hazard"
<p>The simulated data used for analysis in "The importance of censoring in competing risks analysis of the subdistribution hazard". Simulated using the method described in Additional file 1.<br> <br> <strong>Warning: Large file.</strong> Contains 1000 datasets of 300 observations each, for each of 105 parameter combinations (31,500,000 rows). Some programs (e.g. Excel) will not be able to open it in full.</p> <p>csv file with columns:</p> <p><strong>p.comp:</strong> risk of the competing event in exposure group A for this scenario [0 to 0.30 in increments of 0.05]<br> <strong>lnb.cens:</strong> log(hazard ratio) for loss to follow-up in old versus young individuals for this scenario [0 to 1 in increments of 0.25]<br> <strong>lnb.evt:</strong> log(subdistribution hazard ratio) for the event of interest in exposure group B vs group A for this scenario [0, 0.5, 1]<br> <strong>sim:</strong> ID of the simulated dataset for this scenario [1-1000]<br> <strong>exposure:</strong> exposure group (0 = A, 1 = B) of this individual<br> <strong>age:</strong> age group (0 = young, 1 = old) of this individual<br> <strong>time:</strong> time-to-event or censoring for this individual<br> <strong>evtcode:</strong> event type (0 = censoring, 1 = event of interest, 2 = competing event)<br> <strong>censcode:</strong> type of censoring (1 = end-of-study, 2 = loss to follow-up)</p>
Data for publication: "Quantifying the relationship between observed variables that contain censored values using Bayesian error-in-variables regression"
<p>This archive contains the two datasets used in the publication: Vermeiren, Charles, Munoz: Quantifying the relationship between observed variables that contain censored values using Bayesian error-in-variables regression <br>Preprint: <a href="https://hal.science/hal-04764660" rel="nofollow">https://hal.science/hal-04764660</a><br><br>The first dataset is used to develop and test the model using cross-validation, the 2nd dataset is used as an independent, external dataset to test the model. For details, see the publication.</p> <p>The model code, combined with the data and outputs, are also available on GitHub: https://github.com/Peter-Vermeiren/EIVmodels </p>
Impact of Interval Censoring on Data Accuracy and Machine Learning Performance in Biological High-Throughput Screening
<div> <h2>Overview</h2> <div>Data and Results used in the publication entitled "Impact of Interval Censoring on Data Accuracy and Machine Learning Performance in Biological High-Throughput Screening"</div> </div> <div> <h3><strong>Data</strong></h3> <div>This folder contains the raw data used during this work.</div> <div>`EvoEF.csv` contains information on the library used (sequences, number of mutations, etc.) and the fitness (energy) used as continuous mean values. `mut.csv` contains the information about the combinatorial scaling (N vs N_norm), the number of mutations (m) and the probability of each variant using different distributions (uniform and binomial) at different $p_{WT}$.</div> <div>For further details on how the fitness values were calculated and how the combinatorial scale works, please refer to our prevoius [Paper](https://arxiv.org/abs/2405.05167).</div> <div> </div> <h3><strong>Results</strong></h3> <div> <div>This folder contains the results (outputs) of all scripts used. Such results are included in the form of `.npy` and `.npz` files. To load such files with numpy you should include the option `allow_pickle=True`.</div> </div> </div>
Data from: Accounting for heteroscedasticity and censoring in chromosome partitioning analyses
A fundamental assumption in quantitative genetics is that traits are controlled by many loci of small effect. Using genomic data, this assumption can be tested using chromosome partitioning analyses, where the proportion of genetic variance for a trait explained by each chromosome (h2c), is regressed on its size. However, as h2c-estimates are necessarily positive (censoring) and the variance increases with chromosome size (heteroscedasticity), two fundamental assumptions of ordinary least squares (OLS) regression are violated. Using simulated and empirical data we demonstrate that these violations lead to incorrect inference of genetic architecture. The degree of bias depend mainly on the number of chromosomes and their size distribution and are therefore specific to the species; using published data across many different species we estimate that not accounting for this effect overall resulted in 28% false positives. We introduce a new and computationally efficient resampling method that corrects for inflation caused by heteroscedasticity and censoring and that works under a large range of data set sizes and genetic architectures in empirical data sets. Our new method substantially improves the robustness of inferences from chromosome partitioning analyses.
Data from: Accounting for heteroscedasticity and censoring in chromosome partitioning analyses
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.