Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8
datasets available to search
ShareScore release 0.9.0
Dataset results
8 results for “C# source code”
Code and Source Data for "Knowledge-Guided Machine Learning can improve C cycle quantification in agroecosystems"
<p>Datasets for code and Source Data for the study "Knowledge-Guided Machine Learning can improve C cycle quantification in agroecosystems" https://doi.org/10.1038/s41467-023-43860-5. All files belong to Licheng Liu and Zhenong Jin at University of Minnesota. deposit_code_v2.zip contains packaged codes and sample runs for KGML-ag-Carbon training, validation and implementations. Source Data.zip contains data for generating the figures inside the study. </p> <p>Note: We used Pytorch 1.6.0 (<a href="https://pytorch.org/get-started/previous-versions/">https://pytorch.org/get-started/previous-versions/</a>, last access: 21 Oct 2023) and Python 3.7.11 (<a href="https://www.python.org/downloads/release/python-3711/">https://www.python.org/downloads/release/python-3711/</a>, last access: 21 Oct 2023) as the programming environment for model development. Statistical analysis, such as linear regression, was conducted using Statsmodels 0.14.0 (<a href="https://github.com/statsmodels/statsmodels/">https://github.com/statsmodels/statsmodels/</a>, last access: 21 Oct 2023) In order to use a GPU to speed-up the training process, we installed the CUDA Toolkit 10.1.243 (<a href="https://developer.nvidia.com/cuda-toolkit">https://developer.nvidia.com/cuda-toolkit</a>, last access: 21 Oct 2023). </p> <p><strong>To use the full kgml_lib function, please create a new environment with the same python and libs above.</strong></p>
Source data and code for: Existing fossil fuel extraction would warm the world beyond 1.5°C
<p>Source data and code for the study, "Existing fossil fuel extraction would warm the world beyond 1.5°C." Datasets 1-4 include mine-level data collected for China (Dataset 1), India (Dataset 2), and five other countries (Dataset 3) that are among the world's top nine coal producers - the United States, Indonesia, Australia, South Africa, and Poland. Dataset 4 includes global and country-level output data from the 1,000-run Monte Carlo simulation. <Committed_Reserves_Monte_Carlo_Input_Data.zip> includes data and code to replicate the Monte Carlo simulation.</p>
Third-order momentum advection on the quasi-hexagonal C-grid on the sphere: Data and source code
<p>This upload contains data and source code accompanying the paper submitted to JAMES (Journal of advances in modeling Earth systems) under the title 'Third-order momentum advection on the quasi-hexagonal C-grid on the sphere'</p> <p>See README files for further details.</p>
Preprocessed C# Source Codes for Machine Learning
<p>The dataset comes from the HackerRank site, 329,937 C# source codes of 22 tasks were collected and all verified by unit tests.</p> <p>During the download process, source codes received only a unique serial number instead of the user name who solved the task and stored inside the 'task_name/origin' folder. After collecting the data, a new database was created, which included cleaned-up versions of the source codes ('task_name/cleaned' folders contains). Finally, a third set of data was extracted from this cleaned-up version, where a delimiter was inserted before and after each elementary expression to support easy processing and analysis processes ('task_name/reduced' folders contains). Inside the 'task_name' folder three csv files, which contain the equality checking result. The compressed folder also contains a vector space (and related files) made from the reduced data set. These four files are directly in the main folder.</p>
Codes and source data files for: Proximity labeling identifies LOTUS domain proteins that promote the formation of perinuclear germ granules in C. elegans
<p>The germ line produces gametes that transmit genetic and epigenetic information to the next generation. Maintenance of germ cells and development of gametes require germ granules—well-conserved membraneless and RNA-rich organelles. The composition of germ granules is elusive owing to their dynamic nature and their exclusive expression in the germ line. Using <i>C. elegans</i> germ granule, called P granule, as a model system, we employed a proximity-based labeling method in combination with mass spectrometry to comprehensively define its protein components. This set of experiments identified over 200 proteins, many of which contain intrinsically disordered regions. An RNAi-based screen identified factors that are essential for P granule assembly, notably EGGD-1 and EGGD-2, two putative LOTUS-domain proteins. Loss of <i>eggd-1</i> and <i>eggd-2</i> results in separation of P granules from the nuclear envelope, germline atrophy and reduced fertility. We show that intrinsically disordered regions of EGGD-1 are required to anchor EGGD-1 to the nuclear periphery while its LOTUS domains are required to promote perinuclear localization of P granules. Together, our work expands the repertoire of P granule constituents and provides new insights into the role of LOTUS-domain proteins in germ granule organization.</p>
Codes and source data files for: Proximity labeling identifies LOTUS domain proteins that promote the formation of perinuclear germ granules in C. elegans
Open the record for dataset details and reuse information.
A Large Corpus of C Source Code based on Gentoo packages
<p>Corpus of C packages extracted from the Gentoo packages, created for the JSEP publication.</p>
Source code and data for Ou et al. (2021) Updates to Paris climate pledges improve chances of limiting global warming to well below 2°C
<p>There are two folders in this repository. The <strong>GCAM-model</strong> folder contains the version of GCAM5.3 used to estimate emission pathways for this analysis. The <strong>data</strong> folder contains source data for our main results. Please check readme.pdf and our original paper for details. </p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.