Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Data set for training a ML model to predict duration of MPI application phases (HPC system) - with previous phase info
<p>This is the data used to train a ML model predicting the duration of MPI application phases, in a HPC system.</p> <p>There are 10 different data sets corresponding to different HPC applications.</p> <p>These data sets contain information regarding the previous MPI call with same ID and type.</p>
Data set for training a ML model to predict duration of MPI application phases (HPC system) - without previous phases info
<p>This is the data used to train a ML model predicting the duration of MPI application phases, in a HPC system.</p> <p>There are 11 different data sets corresponding to different HPC applications.</p> <p>These data sets do not contain information regarding previous MPI calls</p>
Data set for "Experimental investigation and constitutive description of mechanical anisotropy in soft, porous rocks"
<p>Experimental results</p>
Example for USA metrological years 2020 and 2022 results data sets and post processing python code
<p>The files contain LCF model output data for 2020 and 2022 metrological years in the US as well as a python post-processing code to aggregate the reults of these files.</p>
Requirements data sets (user stories)
<p>A collection of 22 data set of 50+ requirements each, expressed as user stories. </p> <p>The dataset has been created by gathering data from web sources and we are not aware of license agreements or intellectual property rights on the requirements / user stories. The curator took utmost diligence in minimizing the risks of copyright infringement by using non-recent data that is less likely to be critical, by sampling a subset of the original requirements collection, and by qualitatively analyzing the requirements. In case of copyright infringement, please contact the dataset curator (Fabiano Dalpiaz, f.dalpiaz@uu.nl) to discuss the possibility of removal of that dataset [see <a href="https://support.zenodo.org/help/en-gb/13-policies/140-what-is-your-take-down-procedure" target="_blank" rel="noopener">Zenodo's policies</a>]</p> <p>The data sets have been originally used to conduct experiments about ambiguity detection with the REVV-Light tool: https://github.com/RELabUU/revv-light</p> <p>This collection has been originally published in Mendeley data: https://data.mendeley.com/datasets/7zbk8zsd8y/1</p> <h2>Overview of the datasets [data and links added in December 2024]</h2> <p>The following text provides a description of the datasets, including links to the systems and websites, when available. The datasets are organized by macro-category and then by identifier.</p> <h3>Public administration and transparency</h3> <p><code>g02-federalspending.txt</code> (2018) originates from early data in the Federal Spending Transparency project, which pertain to the website that is used to share publicly the spending data for the U.S. government. The website was created because of the Digital Accountability and Transparency Act of 2014 (DATA Act). The specific dataset pertains a system called DAIMS or Data Broker, which stands for DATA Act Information Model Schema. The sample that was gathered refers to a sub-project related to allowing the government to act as a data broker, thereby providing data to third parties. The data for the Data Broker project is currently not available online, although the backend seems to be hosted in <a href="https://github.com/fedspendingtransparency/data-act-broker-backend" target="_blank" rel="noopener">GitHub</a> under a CC0 1.0 Universal license. Current and recent snapshots of federal spending related websites, including many more projects than the one described in the shared collection, can be found <a href="https://federal-spending-transparency.atlassian.net/jira/" target="_blank" rel="noopener">here</a>.</p> <p><code>g03-loudoun.txt</code> (2018) is a set of extracted requirements from a document, by the Loudoun County Virginia, that describes the to-be user stories and use cases about a system for land management readiness assessment called Loudoun County LandMARC. The source document can be found <a href="https://www.loudoun.gov/DocumentCenter/View/131287/APPENDIX-A-3---Future-State-To-Be-Business-User-Stories-and-Use-Cases?bidId=" target="_blank" rel="noopener">here</a> and it is part of the <a href="https://www.loudoun.gov/bids.aspx?bidID=459" target="_blank" rel="noopener">Electronic Land Management System and EPlan Review Project - RFP RFQ</a> issued in March 2018. More information about the overall LandMARC system and services can be found <a href="https://www.loudoun.gov/5823/LandMARC-Land-Management-Applications-Re" target="_blank" rel="noopener">here</a>.</p> <p><code>g04-recycling.txt</code>(2017) concerns a web application where recycling and waste disposal facilities can be searched and located. The application operates through the visualization of a map that the user can interact with. The dataset has obtained from a <a href="https://github.com/rafaellichen/Recycling-System/wiki/User-Stories" target="_blank" rel="noopener">GitHub website</a> and it is at the basis of a students' project on web site design; the <a href="https://github.com/rafaellichen/Recycling-System" target="_blank" rel="noopener">code is available</a> (no license).</p> <p><code>g05-openspending.txt</code> (2018) is about the OpenSpending project (<a href="https://www.openspending.org/" target="_blank" rel="noopener">www</a>), a project of the Open Knowledge foundation which aims at transparency about how local governments spend money. At the time of the collection, the data was retrieved from a Trello board that is currently unavailable. The sample focuses on publishing, importing and editing datasets, and how the data should be presented. Currently, OpenSpending is managed via a <a href="https://github.com/os-data" target="_blank" rel="noopener">GitHub repository</a> which contains multiple sub-projects with unknown license.</p> <p><code>g11-nsf.txt</code> (2018) refers to a collection of user stories referring to the NSF Site Redesign & Content Discovery project, which originates from a publicly accessible <a href="https://github.com/nsf-open/nsf" target="_blank" rel="noopener">GitHub repository</a> (GPL 2.0 license). In particular, the user stories refer to an early version of the NSF's website. The user stories can be found as <a href="https://github.com/nsf-open/nsf/issues?q=is%3Aissue+is%3Aclosed" target="_blank" rel="noopener">closed Issues</a>. </p> <h3>(Research) data and meta-data management</h3> <p><code>g08-frictionless.txt</code> (2016) regards the Frictionless Data project, which offers an open source dataset for building data infrastructures, to be used by researchers, data scientists, and data engineers. Links to the many projects within the Frictionless Data project are on <a href="https://github.com/frictionlessdata" target="_blank" rel="noopener">GitHub</a> (with a mix of Unlicense and MIT license) and <a href="https://frictionlessdata.io/" target="_blank" rel="noopener">web</a>. The specific set of user stories has been collected in 2016 by GitHub user @danfowler and are stored in a <a href="https://trello.com/b/MGC4RpTZ/frictionless-data-user-stories" target="_blank" rel="noopener">Trello board</a>.</p> <p><code>g14-datahub.txt</code> (2013) concerns the open source project <a href="https://datahubproject.io/" target="_blank" rel="noopener">DataHub</a>, which is currently developed via a <a href="https://github.com/datahub-project/datahub" target="_blank" rel="noopener">GitHub repository</a> (the code has Apache License 2.0). DataHub is a data discovery platform which has been developed over multiple years. The specific data set is an <a href="https://datahub.io/docs/dms/datahub/developers/user-stories#stories" target="_blank" rel="noopener">initial set of user stories</a>, which we can date back to 2013 thanks to a comment therein.</p> <p><code>g16-mis.txt</code> (2015) is a collection of user stories that pertains a repository for researchers and archivists. The source of the dataset is a public <a href="https://trello.com/b/hrulGmdz/repository-user-stories" target="_blank" rel="noopener">Trello repository</a>. Although the user stories do not have explicit links to projects, it can be inferred that the stories originate from some project related to the library of Duke University.</p> <p><code>g17-cask.txt</code> (2016) refers to the Cask Data Application Platform (CDAP). CDAP is an open source application platform (<a href="https://github.com/cdapio/cdap" target="_blank" rel="noopener">GitHub</a>, under Apache License 2.0) that can be used to develop applications within the Apache Hadoop ecosystem, an open-source framework which can be used for distributed processing of large datasets. The user stories are extracted from a document that includes requirements regarding dataset management for Cask 4.0, which includes the scenarios, user stories and a design for the implementation of these user stories. The raw data is available in the following <a href="https://cdap.atlassian.net/wiki/spaces/CE/pages/1595181111/Dataset+Management+User+Stories" target="_blank" rel="noopener">environment</a>.</p> <p><code>g18-neurohub.txt</code> (2012) is concerned with the <a href="https://neurohub.ca/" target="_blank" rel="noopener">NeuroHub platform</a>, a neuroscience data management, analysis and collaboration platform for researchers in neuroscience to collect, store, and share data with colleagues or with the research community. The user stories were collected at a time NeuroHub was still a research project sponsored by the UK Joint Information Systems Committee (JISC). For information about the research project from which the requirements were collected, see the following <a href="https://ora.ox.ac.uk/objects/uuid:df65547e-184a-4a21-b701-9e2a5bc240f4" target="_blank" rel="noopener">record</a>.</p> <p><code>g22-rdadmp.txt</code> (2018) is a collection of user stories from the Research Data Alliance's <a href="https://www.rd-alliance.org/groups/dmp-common-standards-wg/members/all-members/" target="_blank" rel="noopener">working group on DMP Common Standards</a>. Their <a href="https://github.com/RDA-DMP-Common" target="_blank" rel="noopener">GitHub repository</a> contains a collection of user stories that were created by asking the community to suggest functionality that should part of a website that manages data management plans. Each user story is stored as an <a href="https://github.com/RDA-DMP-Common/user-stories/issues" target="_blank" rel="noopener">issue on the GitHub's page</a>. </p> <p><code>g23-archivesspace.txt</code> (2012-2013) refers to ArchivesSpace: an open source, web application for managing archives information. The application is designed to support core functions in archives administration such as accessioning; description and arrangement of processed materials including analog, hybrid, and<br>born digital content; management of authorities and rights; and reference service. The application supports collection management through collection management records, tracking of events, and a growing number of administrative reports. ArchivesSpace is open source and its development is hosted in <a href="https://github.com/archivesspace/archivesspace" target="_blank" rel="noopener">GitHub</a> (Educational Community License, Version 2.0), with existing issues in a <a href="https://archivesspace.atlassian.net/jira/dashboards/last-visited" target="_blank" rel="noopener">board</a>, but the dataset only includes older user stories (not available any more in the board) between August 28, 2012 (the starting date of the community) until February 28, 2013.</p> <p><code>g24-unibath.txt</code> (2013) concerns the development of an institutional data repository for the University of Bath. This need was driven by changes in funder and publisher policy, as well as responses from the recent Research360 data management survey sent out to all University of Bath researchers. The purpose of this would be to provide a long-term archive of our research data, with the following benefits: Ensure long-term availability of data to our researchers; fulfil funder and publisher requirements; enable and track increased impact of our research through data re-use and citation by the wider community; encourage new collaborations and deepen existing relationships with industry; enable new types of research, both within the university and the wider sector. The requirements were identified in 2013 and the original document can be found <a href="https://purehost.bath.ac.uk/ws/portalfiles/portal/230431280/IDR_User_Stories_v1.0.pdf" target="_blank" rel="noopener">online</a>.</p> <p><code>g25-duraspace.txt</code> (2012) is a collection that originates from the development of the Data Dictionary Supplement component of the Data asset management system (DAMS) by DuraSpace (<a href="https://wiki.lyrasis.org/display/FF/Design+-+Audit+Service?preview=%2F68060244%2F68354291%2Fuser-stories.pdf" target="_blank" rel="noopener">original document</a>). DuraSpace, or DSpace, is an active project that can be found <a href="https://dspace.org/" target="_blank" rel="noopener">online</a> and is developed as an <a href="https://github.com/DSpace/DSpace" target="_blank" rel="noopener">open source project</a> (BSD 3.0). DSpace stores, preserves and disseminates digital cultural heritage content by supporting ingestion of digital objects and their metadata; management and curation of digital objects; easy access to the digital objects, by both listing and searching; long-term preservation of the digital objects.</p> <p><code>g26-racdam.txt</code> (2015) refers to a collection of requirements that were found in an online <a href="https://trello.com/b/Ou3OzOjR/rac-dam-user-stories" target="_blank" rel="noopener">Trello board</a> for a so-called RAC DAM system. Very limited information exists about the system, but the stories are organized into multiple epics: rights management, asset management, use, discovery, user management, reporting, curation, description, and preservation. It is possible to infer that the collection refers to an archiving system to be used by archivists and researchers. The online identity of the issue contributors makes us hypothesize this has to do with the Rockefeller Archive Center in New York City.</p> <p><code>g27-culrepo.txt</code> (2014-2015) is an extract of a document constructed by RepoExec, a group who manages multiple repositories and the overall repositories policy of the Cornell University Library (CUL). The document was publicly displayed and was accessible on their <a href="https://confluence.cornell.edu/display/culpublic/Charge" target="_blank" rel="noopener">Confuence website</a> as a <a href="https://confluence.cornell.edu/display/culpublic/IR+User+Stories+Working+Group?preview=/326381875/327639271/Repository_User_Stories_20150114.pdf" target="_blank" rel="noopener">PDF attachment</a>. The user stories in the document focus on a subset of the CUL’s institutional repository (IR) systems. </p> <p><code>g28-zooniverse.txt</code> (2014) originates from MICO – Media In Context - an EU-funded research project to develop an integrated platform for cross-media analysis, metadata publishing, querying and recommendation. The set of US describe in particular two main showcases within the MICO project, with the first one being Zooniverse, a citizen science platform (US27-59). The second showcase that is described by the user stories is InsideOut10. InsideOut10 is a start-up and consulting firm from Italy with extensive experience on media delivery platforms and web publishing (US1-US26). The projects are based on volunteers who have access to the platform and contribute by classifying data such as images, audio and video by performing recognition tasks that cannot easily be performed by a computer. The data has been extracted from a deliverable of the MICO project that is currently available <a href="https://www.mico-project.eu/wp-content/uploads/2016/04/CompendiumUseCaseRequirementsAnalysisDeliverable.pdf" target="_blank" rel="noopener">online</a>.</p> <h3>Information systems for specific domains</h3> <p><code>g12-camperplus.txt</code> (2017) concerns a software application called Camper+ (<a href="https://github.com/Notabela/Camper-Plus" target="_blank" rel="noopener">GitHub</a>, MIT License), which aims to support camp administrators, camp counselors, and parents in the context of camps organized for children. The collection of user stories is retrieved from the project's <a href="https://github.com/Notabela/Camper-Plus/wiki" target="_blank" rel="noopener">wiki</a>, which contains the <a href="https://github.com/Notabela/Camper-Plus/wiki/5.-User-Stories" target="_blank" rel="noopener">user stories</a> and organizes the roles into <a href="https://github.com/Notabela/Camper-Plus/wiki/4.-Personas" target="_blank" rel="noopener">personas</a>.</p> <p><code>g19-alfred.txt</code> (2015) describes a set of requirements from a European project called ALFRED (<a href="https://cordis.europa.eu/docs/projects/cnect/8/611218/080/deliverables/001-D23UserStoriesReportv15.pdf" target="_blank" rel="noopener">Deliverable 2.3</a>), which is about a system that provides support for older people; a “Personal Interactive Assistant for Independent Living and Active Ageing”. One of the main system objectives is to support older people to actively participate in society and act independently. The outputs of this project led to a <a href="https://github.com/ALFREDProject" target="_blank" rel="noopener">GitHub repository</a> (unspecified license).</p> <p><code>g21-badcamp.txt</code> (~2017) is a set of user stories that originates from a <a href="https://github.com/badcamp" target="_blank" rel="noopener">GitHub repository</a> (GPL 2.0 and unspecified licenses) of the <a href="https://www.badcamp.org/" target="_blank" rel="noopener">BADCamp event's website</a>, an annual conference that celebrates Drupal open source websites. The website gives general information about the event. Moreover, it is used as a supportive tool for all attendees. All sorts of features have been added over time to support all attendees of the event. While the original user stories are not available, the <a href="https://github.com/badcamp/camp_distro/issues" target="_blank" rel="noopener">current backlog of the website is accessible</a>.</p> <h3>Examples: first-version websites</h3> <p><code>g10-scrumalliance.txt</code>(2004) is a collection taken from a <a href="https://www.mountaingoatsoftware.com/agile/scrum/scrum-tools/product-backlog/example" target="_blank" rel="noopener">product backlog example</a> that is currently published on the Mountain Goat Software website. As stated on that website, these stories were written to describe the functionality of an early version of the <a href="https://www.scrumalliance.org/" target="_blank" rel="noopener">Scrum Alliance website</a>.</p> <p><code>g13-planningpoker.txt</code> (2010, estimated via the Internet Archive Wayback Machine) is an example -- like <code>g10-scrumalliance.txt</code> -- of a product backlog that is available on the Mountain Goat Software <a href="https://www.mountaingoatsoftware.com/agile/scrum/scrum-tools/product-backlog/example" target="_blank" rel="noopener">website</a>. This refers to the first version of the <a href="https://www.planningpoker.com/" target="_blank" rel="noopener">Planning Poker</a> website, which allows estimators in different locations to estimate collaboratively.</p> <p> </p>
Mastodon example data set - a few timepoints of drosophila embryogenesis
<p>Contains a three-dimensional time-lapse data set with 31 time points. It corresponds to a cropped sub-region of the Drosophila SiMView recording presented in the main text of Amat et al., Nature Methods 2014</p> <p>The dataset has been converted to file formats that can be used by the FIJI plugin Mastodon.</p> <p>Amat, F., Lemon, W., Mossing, D. <em>et al.</em> Fast, accurate reconstruction of cell lineages from large-scale fluorescence microscopy data. <em>Nat Methods</em> <strong>11</strong>, 951–958 (2014). https://doi.org/10.1038/nmeth.3036</p>
Data set for "Novel surrogate measures for improving water distribution systems' resilience via pipe diameter uniformity enhancement"
<p>This dataset contains the optimization results using resilience surrogate measures in four chosen cases (i.e., HAN, FOS, PES, MOD). It also includes mechanical reliability calculation results of optimized network layouts obtained by the surrogate measures.</p>
K-State - ARPA-E SMARTFARM Grain Sorghum 2021 Kansas site comprehensive sensor modalities data set.
<p>Comprehensive Year 1 data of ARPA-E SMARTFARM Grain Sorghum project titled "Establishing Validation Sites for Field-Level Emissions Quantification from Grain Sorghum in Southern Great Plains". Data sets includes Eddy Caovariance measurements of GHGs (CO2, CH4 and N2O) along with sub acre level soil moisture, soil temperarature, soil N and carbon, plant biomass and yield. This data is from the Kansas site of the project. </p>
Data Set : Role of intramolecular hydrogen bonding on photoelectron circular dichroism: the diastereoisomers of 1-Amino-2-Indanol
Open the record for dataset details and reuse information.
Data set
<p>Title: <strong>Weather Shocks in The Gambia: Macroeconomic Effects and Household Vulnerability - Dataset</strong></p> <p><em><strong>Description:</strong></em> This dataset supports the analysis in *Weather Shocks in The Gambia: Macroeconomic Effects and Household Vulnerability* exploring the economic impacts of climate-induced weather shocks in The Gambia through macroeconomic and vulnerability modeling. The dataset includes files necessary for estimating and analyzing Dynamic Stochastic General Equilibrium (DSGE) and Vector Autoregression (VAR) models, as well as figures demonstrating model results. It is structured as follows:</p> <p><em><strong>1. DSGE Model.zip: </strong></em>This file contains data required for estimating the DSGE model used in the study. The DSGE model is applied to analyze the macroeconomic impacts of weather shocks and the resulting policy implications.</p> <p><em><strong>2. VAR Model.zip</strong></em>: This file includes data for the estimation of the VAR model, which assesses the relationship between various economic indicators and weather-related shocks in The Gambia. This model is essential for understanding short-term dynamic responses in the economy.</p> <p><em><strong>3</strong><strong><em>.</em> Weather Shocks in The Gambia_Macroeconomic Effects and Households Vulnerability.zip:</strong></em> This file contains figures and visualizations generated from the DSGE model, illustrating the macroeconomic effects of weather shocks and household vulnerability over time. These figures provide a visual summary of key findings.</p> <p>Each file in this dataset is integral to replicating and expanding upon the findings of the study. The data and figures offer insights into the resilience and adaptation measures for economies facing climate-induced shocks, with specific relevance to developing countries. </p>
Lozova&Lytvynenko (data set)
<p>Data set of the research of adolescents' personal formation</p>
Figure 7 from: Jouladeh Roudbar A, Eagderi S, Esmaeili HR, Coad BW, Bogutskaya N (2016) A molecular approach to the genus Alburnoides using COI sequences data set and the description of a new species, A. damghani, from the Damghan River system (the Dasht-e Kavir Basin, Iran) (Actinopterygii, Cyprinidae). ZooKeys 579: 157-181. https://doi.org/10.3897/zookeys.579.7665
Figure 7 - Two views of Cheshmeh Ali, Damghan, type locality of Alburnoides damghani sp. n.
Data set to train a natural language classifier able to differentiate between 15 topics relevant to biodiversity informatics
<p><strong>Scope and size</strong><br> This data set is used to train a natural language processing classifier. The classifier shall be able to differentiate between 15 topics relevant to biodiversity informatics. The list of relevant topics was adapted from Searls (2012).</p> <p>The data set was split into training data, testing data (for tweaking and unit-testing the classifier) and validation data. Each data set is stored as PDF files in a separate directory.</p> <ul> <li>Training data (5494 pages)</li> <li>Test data (977 pages)</li> <li>Validation data (215 pages)</li> </ul> <p><strong>Data sources and licenses</strong><br> Details about the licenses for each data set can be found in the corresponding directories.</p> <ul> <li>Training data was compiled from MIT OpenCourseWare resources provided by MIT under a Creative Commons BY-NC-SA License.</li> <li>Testing data was compiled from MIT OpenCourseWare exams, provided by MIT under a Creative Commons License BY-NC-SA.</li> <li>Validation data was compiled from Wikipedia, provided under a Creative Commons License by Wikipedia editors and contributors.</li> </ul> <p><strong>Topic references</strong><br> Each topic references one or more MIT OpenCourseWare courses:</p> <ul> <li>Algorithms (Demaine, and Devadas, 2011)</li> <li>Artificial Intelligence (Winston, 2010)</li> <li>Building Dynamic Websites (Abelson, and Greenspun, 2003)</li> <li>Computational Biology (Kellis, 2015)</li> <li>Computer Graphics (Matusik, and Durand, 2012)</li> <li>Computer Science and Programming (Bell, Grimson, and Guttag, 2016)</li> <li>Databases (Madden, Morris, Stonebraker, and Curino, 2010)</li> <li>Data Structures (Demaine, 2012)</li> <li>Digital Image Processing (Clifford, Fisher, Greenberg, and Wells, 2007; Golland, 2005)</li> <li>Machine Learning (Singh, Jaakkola, and Mohammad, 2006)</li> <li>Machine Structures (Morris, and Madden, 2009)</li> <li>Natural Language Processing (Berwick, 2003; Collins, and Barzilay, 2005)</li> <li>Parallel Computing (Edelman, 2011)</li> <li>Software Engineering (Jackson, and Devadas, 2005)</li> <li>Structure and Interpretation of Computer Programs (Miller, and Goldman, 2016).</li> </ul> <p><strong>References</strong></p> <p>Harold Abelson, and Philip Greenspun. 6.171 Software Engineering for Web Applications. Fall 2003. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Ana Bell, Eric Grimson, and John Guttag. 6.0001 Introduction to Computer Science and Programming in Python. Fall 2016. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Robert Berwick. 6.863J Natural Language and the Computer Representation of Knowledge. Spring 2003. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Gari Clifford, John Fisher, Julie Greenberg, and William Wells. HST.582J Biomedical Signal and Image Processing. Spring 2007. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Michael Collins, and Regina Barzilay. 6.864 Advanced Natural Language Processing. Fall 2005. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Erik Demaine. 6.851 Advanced Data Structures. Spring 2012. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Erik Demaine, and Srini Devadas. 6.006 Introduction to Algorithms. Fall 2011. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Alan Edelman. 18.337J Parallel Computing. Fall 2011. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Polina Golland. 6.881 Representation and Modeling for Image Analysis. Spring 2005. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Daniel Jackson, and Srini Devadas. 6.170 Laboratory in Software Engineering. Fall 2005. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Manolis Kellis. 6.047 Computational Biology. Fall 2015. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Robert Miller, and Max Goldman. 6.005 Software Construction. Spring 2016. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Samuel Madden, Robert Morris, Michael Stonebraker, and Carlo Curino. 6.830 Database Systems. Fall 2010. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Wojciech Matusik, and Frédo Durand. 6.837 Computer Graphics. Fall 2012. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Morris, Robert, and Madden, Samuel. 6.033 Computer System Engineering. Spring 2009. Massachusetts Institute of Technology: MIT OpenCourseWare, http://hdl.handle.net/1721.1/118791. License: Creative Commons BY-NC-SA.</p> <p>David B. Searls. An online bioinformatics curriculum. 2012. PLoS computational biology, 8(9), p.e1002632.</p> <p>Rohit Singh, Tommi Jaakkola, and Ali Mohammad. 6.867 Machine Learning. Fall 2006. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Patrick Winston. 6.034 Artificial Intelligence. Fall 2010. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p> </p>
Convective self-compression of cratons and the stabilization of old lithosphere - DATA SET
<p>Data set for the research work titled, "Convective self-compression of cratons and the stabilization of old lithosphere"</p>
Tracing low-CO2 fluxes in incubation and 13C labeling experiments; data set
<p>Data set containing data from feature tests (1-3) as well as photosynthesis and respiration measurements.</p> <p> </p>
single-cell RNAseq data (data set 13) in the publication scFASTCORMICS: A contextualization algorithm to reconstruct metabolic multi-cell population models from single-cell RNAseq data
<p>The present dataset (dataset13) was used as input to build scFASTCORMICS models. The files correspond to the clusters identified by Seurat in the single-cell data from pancreas donor11 downloaded from the GEO website (<strong>GSE114297). </strong></p> <p>see the protocol: scFASTCORMICS: A contextualization algorithm to reconstruct metabolic multi-cell population models from single-cell RNAseq data</p> <p>and github: <a href="https://github.com/sysbiolux/scFASTCORMICS">https://github.com/sysbiolux/scFASTCORMICS</a></p> <p>For more information, version updates of the scFASTCORMICS. </p>
single-cell RNAseq data (data set 10) in the publication scFASTCORMICS: A contextualization algorithm to reconstruct metabolic multi-cell population models from single-cell RNAseq data
<p>The present dataset (dataset10) was used as input to build scFASTCORMICS models. The files correspond to the clusters identified by Seurat in the single-cell data from pancreas donor8 downloaded from the GEO website (<strong>GSE114297). </strong></p> <p>see the protocol: scFASTCORMICS: A contextualization algorithm to reconstruct metabolic multi-cell population models from single-cell RNAseq data</p> <p>and github: https://github.com/sysbiolux/scFASTCORMICS</p> <p>For more information, version updates of the scFASTCORMICS. </p>
single-cell RNAseq data (data set 6) in the publication scFASTCORMICS: A contextualization algorithm to reconstruct metabolic multi-cell population models from single-cell RNAseq data
<p>The present dataset (dataset6) was used as input to build scFASTCORMICS models. The files correspond to the clusters identified by Seurat in the single-cell data from pancreas donor4 downloaded from the GEO website (<strong>GSE114297). </strong></p> <p>see the protocol: scFASTCORMICS: A contextualization algorithm to reconstruct metabolic multi-cell population models from single-cell RNAseq data</p> <p>and github: <a href="https://github.com/sysbiolux/scFASTCORMICS">https://github.com/sysbiolux/scFASTCORMICS</a></p> <p>For more information, version updates of the scFASTCORMICS. </p>
single-cell RNAseq data (data set 2) in the publication scFASTCORMICS: A contextualization algorithm to reconstruct metabolic multi-cell population models from single-cell RNAseq data
<p>The present dataset (dataset2) was used as input to build scFASTCORMICS models. The files correspond to the clusters identified by Seurat in the single-cell data from normal mucosa samples downloaded from the GEO website (<strong>GSE81861). </strong></p> <p>see the protocol: scFASTCORMICS: A contextualization algorithm to reconstruct metabolic multi-cell population models from single-cell RNAseq data</p> <p>and github: https://github.com/sysbiolux/scFASTCORMICS</p> <p>For more information, version updates of the scFASTCORMICS. </p>
Data set for EMPIR project MeDDII Report A2.1.5 D3
<p>Data set for EMPIR project MeDDII Report A2.1.5 D3</p> <p>Validation report on the primary standards developed for the in-line measurement of the dynamic viscosity of Newtonian liquids with a target uncertainty of 2 % (k=2)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.