Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5,526

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

5,526 results for “information”

Learn how ShareScore rates datasets ↗
zenodo44/100

A route to school informational intervention for air pollution exposure reduction

<p>iSCAPE Dataset Reference No. = DS_PD_020</p> <p>Following datasets are gathered during the implementation of route to school intervention study in Antwerp&nbsp;(Belgium)</p> <ol> <li>Introductory Questionnaire Responses</li> <li>Feedback Questionnaire Responses</li> </ol>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Information Based Behavioural Intervention Study (Hasselt case)

<p>iSCAPE Dataset Ref. No. =&nbsp;DS_PD_010,&nbsp;DS_MA_002,&nbsp;DS_PD_015</p> <p>Following datasets are acquired and generated during the implementation of behavioural intervention study in Hasselt:</p> <ol> <li>Participant&#39;s detailed activity-travel diary&nbsp;</li> <li>Pollutant concentration maps</li> <li>Introductory and Followup questionnaire responses&nbsp;</li> </ol> <p>&nbsp;</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Preliminary Supplementary Information for "Kinetics of Deoxyribose-1-Phosphate Decay in Aqueous Solution"

<p>This is the dataset for our upcoming publication tenatively titled &quot;Kinetics of Deoxyribose-1-Phosphate Decay in Aqueous Solution&quot; and may serve as a preliminary Supplementary Information.</p> <p>We employed high-throughput UV spectroscopy-based monitoring of the apparent conversion of deoxyribosyl nucleoside phosphorolysis to access the kinetics of deoxyribose-1-phosphate hydrolysis in aqueous solution at different pH values and temperatures.</p> <p>Please see the files below for a general description of this entry and the full dataset(s).</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Project files provided as supporting information to the manuscript "Ligand-protein interactions in lysozyme investigated through a dual-resolution model"

<p><strong>README file for the project files provided as supporting information to the manuscript &quot;Ligand-protein interactions in lysozyme investigated through a dual-resolution model&quot;</strong></p> <p>February 12, 2020</p> <p>Authors: Raffaele Fiorentini, Kurt Kremer and Raffaello Potestio</p> <p>================================</p> <p>Overview</p> <p>The dataset&nbsp;is organised in three (compressed) subfolders (see the tree diagrams in each section):</p> <p>- annihilation<br> - decoupling<br> - density</p> <p>The figure deltaG_binding_ann_dec_comparison.png shows the results of binding free energy calculations comparing the values obtained both for annihilation and decoupling.</p> <p>The figure deltaG_binding_annih_gromacs_espp.png displays the results for Binding FE, comparing the values obtained in GROMACS and ESPResSo++.</p> <p>The README.pdf file contains detailed information about these folders and their content.</p> <p>================================</p> <p>The &quot;annihilation&quot; folder contains all results concerning the calculation of binding free energy in case of annihilation and it is divided in two parts:&nbsp;</p> <p>- complex<br> - ligand</p> <p>In &quot;complex&quot; are reported the results of Ligand-Protein FE both in ESPResSo++ and GROMACS. All simulations are fully-atomistic.&nbsp;</p> <p>In &quot;ligand&quot; are reported the results of ligand solvation free energy both in ESPResSo++ and GROMACS. All simulations are fully-atomistic.&nbsp;</p> <p>====</p> <p>The &quot;decoupling&quot; folder contains all results concerning the calculation of binding free energy in case of decoupling and it is divided in three parts:&nbsp;</p> <p>- complex-DualRes<br> - complex-FullyAT<br> - ligand</p> <p>In &quot;complex-DualRes&quot; are reported the results of Ligand-Protein FE only in ESPResSo++ (GROMACS cannot do decoupling). The system is simulated in Dual-Resolution. It is possible to find the trajectory files in the sub-directories &quot;lambdaindex-0&quot; and &quot;lambdaindex-30&quot;.</p> <p>In &quot;complex-fullyAT&quot; are reported the results of Ligand-Protein FE only in ESPResSo++. The system simulated is fully-atomistic. It is possible to find the trajectory file in the sub-directories &quot;lambdaindex-0&quot; and &quot;lambdaindex-30&quot;.</p> <p>In &quot;ligand&quot; are reported the results of ligand solvation free energy only in ESPResSo++. All simulations are fully-atomistic. It is possible to find the trajectory file in the sub-directories &quot;lambdaindex-0&quot; and &quot;lambdaindex-20&quot;.</p> <p>====</p> <p>The &quot;density&quot; folder contains the data for the tuning of the c parameter of the steric repulsion among residues. This parameter is tuned so that the water density attains the value computed in all-atom simulations.</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

A Wi-Fi Channel State Information (CSI) and Received Signal Strength (RSS) data-set for human presence and movement detection

<p>This data-set consists of antenna-wise received signal strength (RSS) and channel state information (CSI) data. Both types of data have been captured using the <a href="https://dhalperi.github.io/linux-80211n-csitool/">Intel CSI Tools</a>. The RSS data have been used in our paper &quot;Detecting Human Movement from Ambient Wi-Fi Signal Strength&quot;.</p> <p>This release extends the README with a data dictionary for the annotations. We hope to add more information about the data acquisition process (e.g., data acquisition protocols).</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

3D scans of two types of railway ballast including shape analysis information

<p>This data set contains 3D scanner data of two types of railway ballast &ldquo;Calcite&rdquo; (stems from Croatia) and &ldquo;Kieselkalk&rdquo;, also known as Helvetic Siliceous Limestone, (stems from Switzerland).<br> From each type of ballast 25 stones are scanned. The files are provided in .ply format.<br> For the scanned meshes several shape descriptors are provided: elongation, flatness, sphericity, convexity index.<br> Additional to the 3D scans, both simplified and rounded versions of the meshes are included.<br> For these meshes information on three different angularity indices are available.<br> The scanned ballast types are the same, as&nbsp; previously investigated in uniaxial compression tests and direct shear tests:<br> Suhr, Bettina, &amp; Six, Klaus. (2018).<br> &quot;Compression tests and direct shear test of two types of railway ballast [Data set]&quot;<br> Zenodo. http://doi.org/10.5281/zenodo.1423742</p> <p>&nbsp;</p> <p>A detailed shape analysis of the results is conducted in:<br> Bettina Suhr, William A. Skipper, Roger Lewis, and Klaus Six<br> &quot;Shape analysis of railway ballast stones: curvature-based calculation of particle angularity&quot;<br> <em>Scientific Reports, </em><strong>2020</strong><em>, 10</em>, 6045<br> DOI: https://doi.org/10.1038/s41598-020-62827-w</p> <p>A summary of several shape descriptors can be found in:<br> B. Suhr and K. Six:<br> &quot;Simple particle shapes for DEM simulations of railway ballast --&nbsp; influence of shape descriptors on packing behaviour&quot;<br> Granular Matter, <strong>2020</strong><em>, 22</em><br> DOI: https://doi.org/10.1007/s10035-020-1009-0</p> <p><br> This data set is organised as follows:<br> 1_ScanMeshesCleaned<br> &nbsp;&nbsp;&nbsp; scanned meshes:<br> &nbsp;&nbsp;&nbsp; K_1.ply&nbsp; -&nbsp; K_25.ply Calcite (German: Kalzit)<br> &nbsp;&nbsp;&nbsp; KK_1.ply - KK_25.ply Kieselkalk<br> 2_CSE1 &nbsp;<br> &nbsp;&nbsp;&nbsp; simplifications of the scanned meshes, little simplifications, used in the detailed shape analysis<br> &nbsp;&nbsp;&nbsp; CSE1_AngInfo.csv: contains values of three different angularity indices used in the detailed shape analysis<br> 3_CSE2 &nbsp;<br> &nbsp;&nbsp;&nbsp; simplifications of the scanned meshes, more simplified, used in the detailed shape analysis<br> &nbsp;&nbsp;&nbsp; CSE2_AngInfo.csv: contains values of three different angularity indices used in the detailed shape analysis<br> 4_CSE3 &nbsp;<br> &nbsp;&nbsp;&nbsp; simplifications of the scanned meshes, even more simplified, used in the detailed shape analysis<br> &nbsp;&nbsp;&nbsp; CSE3_AngInfo.csv: contains values of three different angularity indices used in the detailed shape analysis<br> 5_CSE4 &nbsp;<br> &nbsp;&nbsp;&nbsp; simplifications of the scanned meshes, most simplified, used in the detailed shape analysis<br> &nbsp;&nbsp;&nbsp; CSE4_AngInfo.csv: contains values of three different angularity indices used in the detailed shape analysis<br> 6_RoundedMeshes<br> &nbsp;&nbsp;&nbsp; artificially rounded versions of the scanned ballast meshes, used in the detailed shape analysis<br> &nbsp;&nbsp;&nbsp; RoundedMeshes_AngInfo.csv: contains values of three different angularity indices used in the detailed shape analysis<br> 7_TestBodies &nbsp;<br> &nbsp;&nbsp;&nbsp; meshes of artificial test bodies, constructed for testing different angularity indices in the detailed shape analysis<br> &nbsp;&nbsp;&nbsp; TestBodies_AngInfo.csv: contains values of three different angularity indices used in the detailed shape analysis<br> scanMeshesInfo.csv: summary of several shape descriptors of the scanned meshes<br> README.txt &nbsp;</p> <p><br> Check the README.txt file for more information on the technical aspects of scanning.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

An agent-based model of the origins of modern linguistic complexity – supplementary information

<p>A central question in the evolution of human language is whether it emerged as a result of one specific event or from a mosaic-like constellation of different phenomena and their interactions. Three potential processes have been identified by recent research as the potential&nbsp;<em>primum mobile</em>&nbsp;for the origins of modern linguistic complexity:&nbsp;Self-domestication, characterized by a reduction in reactive aggression and often associated with a gracilization of the face; changes in early brain development manifested by&nbsp;globularization&nbsp;of the skull; and&nbsp;demographic expansion&nbsp;of&nbsp;H. sapiens&nbsp;during the Middle Pleistocene. We developed an agent-based model to investigate how these three factors influence transmission of information within a population. Our model shows that there is an optimal degree of both hostility and mental capacity at which the amount of transmitted information is the largest. It also shows that linguistic communi- ties formed within the population are strongest under circumstances where individuals have high levels of cognitive capacity available for information processing and there is at least a certain degree of hos- tility present. In contrast, we find no significant effects related to population size.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Questionnaire data to research small-scale farmers' information sharing for adapting to climate change in Mozambique (2019-2020)

<p>Data collected from individual questionnaires with local communities of 4 districts of Mozambique in November 2019 and July 2020. It contains as well data from nine individual questionnaires to institutions (government and NGOs) working with local communities for their development.</p> <p>Data are replies from interviews containing open and closed questions about a) climate change adaptation options necessary for Mozambican small scale farmers, about b) the most used and preferred information sources of farmers, about c) the main barriers for a better exchange of information, and about d) proposals for improving it. The questionnaire can be consulted in Appendix A (in English and Portuguese). The open questions had the purpose to understand the causes and explanations about the themes presented. The closed questions followed a 0-5 likert scale approach, where 5 meant a very important factor and 0 non important one. This format was pursued for developing statistical analysis and comparison between the different types of participants. We used the same questions and format for interviewing farmers and stakeholders, although the questionnaire for farmers included also personal aspects like gender, age, and education.</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Phylogenomics of manakins (Aves: Pipridae) using alternative locus filtering strategies based on informativeness

<p>Data&nbsp;used in phylogenomic analyses of manakin birds.&nbsp;</p> <p>Datasets number&nbsp;1 to 7 include sequence alignments for each locus analyzed, and datasets 4 to 7 also contain gene trees used as input for&nbsp;ASTRAL.</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Connecting U.S. Supreme Court Case Information and Opinion Authorship (SCDB) to Full Case Text Data (CAP), 1791-2011

<p>This dataset was constructed to connect the rich metadata created by the Supreme Court Database (SCDB) to the Caselaw Access Project (CAP) full-text court opinion data. Since the SCDB includes only substantive opinions, it is necessarily a subset of the full range of opinions available through CAP.</p> <p>There are two parts to this data: the map connecting each SCDB ID to its corresponding CAP case number, and a more advanced (but error-prone) version in which the authorship of each opinion text identified for the case in CAP is attributed to the Justice who wrote it. Each of these data products have been hand-corrected to the best of this author&#39;s ability.</p> <p><strong>SCDB-CAP map</strong></p> <p>The SCDB-&gt;CAP map began as a relatively straightforward automated matching process, based on the US Reports citation for each case as expressed in both SCDB and CAP. Slightly over 80% of SCDB entries found a single CAP data match this way. From there, the data was entirely hand-corrected, with non-matches or duplicate matches individually investigated and manually corrected.</p> <p>Some SCDB entries simply could not be matched to an appropriate CAP text. Initially, the entirety of US Reports volume 44 was missing, but with the help of CAP staff, the volume was located as having been filed in the New York jurisdiction rather that the United States jurisdiction. The case numbers were then added to the map, but until the volume is relocated to the United States jurisdiction, it may be necessary to also incorporate the New York jurisdiction in full text analysis so that the cases from volume 44 can be searched. 108 more missing cases are from US Reports volume 131, which was a &quot;catch up&quot; volume published in the 19th century. These catch-up cases, many heard by the Supreme Court decades prior, were numbered with lowercase roman numerals instead of the ordinary&nbsp; numbers, which is almost certainly why CAP&#39;s software dismissed the catch-up section as prefatory material. Many of the rest of the errors seem largely to be examples where the SCDB project recognized a separate court action that CAP did not. Perhaps most of these seem to have been later rehearings for a case previously decided, which in the 19th century particularly were commonly reported out at the end of the first decision text. While SCDB sometimes gave these subsequent but related actions a separate SCDB entry, CAP seems to have largely incorporated them as part of the text of the main case. Additionally, there were a few that simply could not be found, despite a careful look through each database as well as the original US Reports and sometimes adjacent volumes. Finally, the cases were only matched up through the 2011 court term. After the 2011 term, the mismatches between CAP and SCDB were extensive and frequently seemed impossible to resolve.</p> <p>Even so, with the manual correction, the overall error rate is low. Of 28,304 cases, only 191 do not have a match, and of those, 108 are contained within the vol. 131 &quot;catch up&quot; volume. Since most of the rest are extremely short subsequent actions that were separately noted by SCDB, the effect of these non-matched cases would seem to be small in most cases.</p> <p>The typical use case would be that the researcher would generate some kind of results based on searching in the CAP full text, then could use the CAP ID to look up the SCDB ID in the map. With the SCDB ID, of course, the rich metadata from the SCDB can then be connected to each result as needed.</p> <p><strong>Opinion authorship</strong></p> <p>Being able to use the rich metadata of SCDB in conjunction with a case&#39;s full text is exciting, but it immediately prompts a further question -- what if the texts could be attributed directly to the Justices who authored them? SCDB produces its data in two forms; one is &quot;case centered,&quot; where each record represents one case, and the other is &quot;justice centered,&quot; in which each record is the vote of one Justice in one case. CAP, in turn, breaks the total text of the case into distinct opinions, and tries to attribute those opinions to their authors by scraping a string of text from the raw input. Therefore, the challenge was to connect these two sources at the opinion level.</p> <p>Connecting the opinions, like connecting the cases, involved an initial match by machines, followed by manual correction and revision. In this case, the scope of the manual effort was much larger than that posed by the case-level connection, and more errors were noted in both SCDB and CAP.</p> <p>The matching process involved a number of steps. First a list of opinions was generated from the CAP data, then matched to SCDB using the SCDB-CAP connector data described above. (Thus, a case without a CAP match in the SCDB-CAP data will not appear in the opinion author data either.) CAP opinions were numbered in the order they were encountered in each CAP case JSON object, and these numbers are used to distinguish the opinions.</p> <p>Next, a round of automatic matching was performed. If there was only one opinion, and only one author listed in the SCDB data, then the majority opinion author (as listed in SCDB) was safely assumed to be the author. If there was no author listed in SCDB, &quot;percuriam&quot; was recorded as the author in this data. If there were exactly two opinions and two authors, the process was also straightforward, as the SCDB-identified majority opinion author was assigned to opinion 1, and the remaining author assigned opinion 2.</p> <p>Subsequently, cases with more than two opinions were processed. A potential match (i.e. a &quot;guess&quot;) for each opinion in a given case was created by listing each Justice identified by SCDB as having written an opinion in the case. These guesses were then parsed using a semi-automatic procedure with Levenshtein distance fuzzy name matching. With sufficiently conservative parameters, a successful fuzzy match meant that the non-successful guesses for that opinion could be deleted. These sorted guesses were then reviewed manually. Particular care was also taken for any opinion that contained authored opinions by Justices who had similar names (for example, Clark and Black differ by only a single letter). These sorts of cases, as well as instances of co-authorship, were identified and fixed manually.</p> <p>Those opinions whose authorship could not be matched then were fixed by hand. These included some where the CAP author strings were more complicated than SCDB&#39;s strict interpretation; others where the OCR in CAP which contained the Justice name was especially bad; and a number of others where &quot;Mr. Chief Justice&quot; couldn&#39;t be directly matched with an author name by the machine. After this light manual correction, almost 500 opinions with substantial errors remained to be individually investigated in depth, by examining the CAP record, the SCDB record, and images of the US Reports for that case. For these last tough customers, errors in the source data were commonly the cause of matching problems. Typically these were of three kinds: examples where CAP should have split the text but didn&#39;t (e.g. 2 opinions together in one opinion entry in CAP); examples where SCDB either did not identify or mis-identified an author (such as attributing it to Swayne when it was written by Miller); and examples of non-valid opinions (such as where CAP mistakenly split the opinion too early, leaving an opinion fragment).</p> <p>For these errors, a system of codes was created in the author field to signal the error type so that researchers can be suitably cautious. The error code is always at the beginning of the field and is followed by a comma and the names of each author, separated by a comma with no space to facilitate parsing. Note also that co-authors are listed as comma-separated names in this same field with no error code. Researchers will probably want to disaggregate this field to create duplicate records with each individual author for most purposes. The justice number field also contains information about all justices authoring the opinion but the error codes have been omitted here.</p> <ul> <li>!C -- error: multiple opinion texts combined (i.e. CAP splitting error)</li> <li>!X -- error: unattributed or misattributed opinion (not listed in SCDB as writer)</li> <li>!D -- error: extra opinion that should be deleted, i.e. not a valid opinion</li> <li>!W -- error: listed as Writers by SCDB, but should be co-authors</li> </ul> <p>&nbsp;</p> <p><strong>Data file structure</strong></p> <p>&quot;scdb_cap-051820.tsv&quot; is a Tab-separated data file containing 5 columns: SCDB ID, CAP ID, US Reports citation, case date, and case name (the latter three from the SCDB data).</p> <p>&quot;scdb-cap-opinion-authorship_051920.tsv&quot; is a Tab-separated data file containing seven columns: SCDB ID, CAP ID, US Reports citation, case name, opinion number in the case, opinion author, and SCDB justice ID. See above for caveats about disaggregating and error codes in fields six and seven.</p> <p><strong>Errors</strong></p> <p>It is likely that errors remain in this data, and it is also hoped that some of the errors beyond the author&#39;s immediate control might be fixed in the upstream data so that they can be corrected here. Authors would be grateful for error reports, and also reports of errors fixed, if any.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Excavator-generated information from Linux drivers (Decoder Use-Case A)

<p>This dataset is released as part of DECODER&#39;s D6.2 deliverable. It contains the information generated by the Excavator tool for easing the verification with Frama-C of the watchdog and ethernet Linux drivers that have been selected as Use-Case A of the project.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Supporting Data for: Information Retrieval Interfaces in Virtual Reality - A Scoping Review Focused on Current Generation Technology

<p>This is the full data set of all reviewed research items obtained from Google Scholar, Web of Science and Scopus for the Scoping Literature Review&nbsp;<em><a href="https://doi.org/10.1371/journal.pone.0246398">Information Retrieval Interfaces in Virtual Reality - A Scoping Review Focused on Current Generation VR technology</a>.</em></p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

WEST Diversity Panel GBS information (Ferguson et al., 2020)

<p>Genotyping by sequencing information (5,512,653 SNPs on 850 individuals [842 unique]) used in&nbsp;the manuscript &quot;Machine learning enabled phenotyping for GWAS and TWAS of WUE traits in 869 field-grown sorghum accessions&quot; (Ferguson et al., 2020) DOI:10.1101/2020.11.02.365213</p> <p>The information, data, or work presented herein was funded in part by the Advanced Research Projects Agency-Energy (ARPA-E), U.S. Department of Energy, under Award Number DE-DE-AR0000661. The views and opinions of the authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Dataset of 'Complete flow characterization from snapshot PIV, fast probes and physics-informed neural networks'

<p>Dataset of the article 'Complete flow characterization from snapshot PIV, fast probes and physics-informed neural networks' (https://doi.org/10.1016/j.cma.2023.116652). The codes processing data here are on https://github.com/AlvaroMS90/Complete-flow-characterization-from-snapshot-PIV-fast-probes-and-physics-informed-neural-networks.</p> <p>This project has received funding from the European Research Council (ERC) under the European Union&rsquo;s Horizon 2020 research and innovation program (grant agreement No 949085) and by MCIN/AEI /10.13039/501100011033 and the European Union &lsquo;NextGenerationEU/PRTR&rsquo; as part of the grant FJC2020-044342-I.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Survey of open data and information seeking in Kenya's Urban Slums and Rural Settlements

<p>This dataset provides survey responses from 240 people surveyed as part of the &quot;Investigating the Impact of Kenya&rsquo;s Open Data Initiative on Marginalized Communities: Case Study of Urban Slums and Rural Settlements&quot; project.</p> <p>The data, collected in mid-2013 looks at issues of where citizens look for data, and how successful they have been in getting government information from different sources, as well as their awareness of the Kenya open data portal, and their interest in getting information through different digital channels in future.</p> <p>Descriptive statistics have been analysed in the publication &quot;Open Government Data for Effective Public Participation: Findings of a Case Study Research Investigating The Kenya&#39;s Open Data Initiative in Urban Slums and Rural Settlements&quot;, but no further analysis has yet been carried out.</p> <p><strong>Data descriptions</strong></p> <p>The Codebook.csv file lists variable names and the questions asked to elicit each response.</p> <p>JHC-Data.csv contains the results from the questionnaires collected through structured in-person interview in the three locations.&nbsp; The questionnaires were administered at chiefs centres, community resource centres, constituency development fund office and religious centres). The questionnaires were filled in by every 2nd these centres.</p> <p><strong>More information</strong></p> <p>More information on the research project can be found at http://opendataresearch.org/project/2013/jhc</p>

opencc-by-sa-4.0Aug 2014View details →
zenodo44/100

Open-data release of aggregated Australian school-level information. Edition 2016.1

<p>The file set is a freely downloadable aggregation of information about Australian schools. The individual files represent a series of tables which, when considered together, form a relational database. The records cover the years 2008-2014 and include information on approximately 9500 primary and secondary school main-campuses and around 500 subcampuses. The records all relate to school-level data; no data about individuals is included. All the information has previously been published and is publicly available but it has not previously been released as a documented, useful aggregation. The information includes:<br /> (a) the names of schools<br /> (b) staffing levels, including full-time and part-time teaching and non-teaching staff<br /> (c) student enrolments, including the number of boys and girls<br /> (d) school financial information, including Commonwealth government, state government, and private funding<br /> (e) test data, potentially for school years 3, 5, 7 and 9, relating to an Australian national testing programme know by the trademark 'NAPLAN'<br /> <br /> Documentation of this Edition 2016.1 is incomplete but the organization of the data should be readily understandable to most people. If you are a researcher, the simplest way to study the data is to make use of the SQLite3 database called 'school-data-2016-1.db'. If you are unsure how to use an SQLite database, ask a guru.<br /> <br /> The database was constructed directly from the other included files by running the following command at a command-line prompt:<br />   <em>sqlite3 school-data-2016-1.db &lt; school-data-2016-1.sql</em><br /> Note that a few, non-consequential, errors will be reported if you run this command yourself. The reason for the errors is that the SQLite database is created by importing a series of '.csv' files. Each of the .csv files contains a header line with the names of the variable relevant to each column. The information is useful for many statistical packages but it is not what SQLite expects, so it complains about the header. Despite the complaint, the database will be created correctly.<br /> <br /> Briefly, the data are organized as follows.<br /> (a) The .csv files ('comma separated values') do not actually use a comma as the field delimiter. Instead, the vertical bar character '|' (ASCII Octal 174 Decimal 124 Hex 7C) is used. If you read the .csv files using Microsoft Excel, Open Office, or Libre Office, you will need to set the field-separator to be '|'. Check your software documentation to understand how to do this.<br /> (b) Each school-related record is indexed by an identifer called 'ageid'. The ageid uniquely identifies each school and consequently serves as the appropriate variable for JOIN-ing records in different data files. For example, the first school-related record after the header line in file 'students-headed-bar.csv' shows the ageid of the school as 40000. The relevant school name can be found by looking in the file 'ageidtoname-headed-bar.csv' to discover that the the ageid of 40000 corresponds to a school called 'Corpus Christi Catholic School'.<br /> (3) In addition to the variable 'ageid' each record is also identified by one or two 'year' variables. The most important purpose of a year identifier will be to indicate the year that is relevant to the record. For example, if one turn again to file 'students-headed-bar.csv', one sees that the first seven school-related records after the header line all relate to the school Corpus Christi Catholic School with ageid of 40000. The variable that identifies the important differences between these seven records is the variable 'studentyear'. 'studentyear' shows the year to which the student data refer. One can see, for example, that in 2008, there were a total of 410 students enrolled, of whom 185 were girls and 225 were boys (look at the variable names in the header line).<br /> (4) The variables relating to years are given different names in each of the different files ('studentsyear' in the file 'students-headed-bar.csv', 'financesummaryyear' in the file 'financesummary-headed-bar.csv'). Despite the different names, the year variables provide the second-level means for joining information acrosss files. For example, if you wanted to relate the enrolments at a school in each year to its financial state, you might wish to JOIN records using 'ageid' in the two files and, secondarily, matching 'studentsyear' with 'financialsummaryyear'.<br /> (5) The manipulation of the data is most readily done using the SQL language with the SQLite database but it can also be done in a variety of statistical packages.<br /> (6) It is our intention for Edition 2016-2 to create large 'flat' files suitable for use by non-researchers who want to view the data with spreadsheet software. The disadvantage of such 'flat' files is that they contain vast amounts of redundant information and might not display the data in the form that the user most wants it.<br /> (7) Geocoding of the schools is not available in this edition.<br /> (8) Some files, such as 'sector-headed-bar.csv' are not used in the creation of the database but are provided as a convenience for researchers who might wish to recode some of the data to remove redundancy.<br /> (9) A detailed example of a suitable SQLite query can be found in the file 'school-data-sqlite-example.sql'. The same query, used in the context of analyses done with the excellent, freely available R statistical package (http://www.r-project.org) can be seen in the file 'school-data-with-sqlite.R'.</p>

opencc-zeroDec 2015View details →
zenodo44/100

Data licences and organization type of contributors to the Global Biodiversity Information Facility as of 19 January 2016

<p>Data from the Global Biodiversity Information Facility were extracted using R (version 3.2.0) on 9 July 2015 using the rgbif package (version 0.9.0) (Chamberlain, S., Ram, K., Barve, V. &amp; Mcglinn, D. (2015) Package ‘rgbif’: Interface to the Global 'Biodiversity' Information Facility 'API' http://cran.r-project.org/web/packages/rgbif/rgbif.pdf). The ‘rights’ statements was extracted for all occurrence datasets with one or more observations. A total of 12,458  datasets were extracted, but only about 11% of the datasets have an explicit data-useage-rights statement at the dataset level. However, some datasets use the occurrence level ‘rights’ and ‘accessRights’ fields. To extract these data the rights information was obtained from the first record of each dataset where a rights statement was missing at the dataset level.</p> <p>The datasets were categorized into 13 different types depending on the origin of the observations.</p> <ol> <li>Biodiversity Information Facility or data centre</li> <li>Botanical Garden or Herbarium</li> <li>Citizen science</li> <li>Commercial</li> <li>Data publisher</li> <li>Educational</li> <li>Government</li> <li>Museum</li> <li>Network</li> <li>Parks Authority or Nature Reserve</li> <li>Research institution</li> <li>Society</li> <li>Foundations</li> </ol>

opencc-zeroJan 2016View details →
zenodo44/100

Dataset for "Development of allocentric representations using self-motion information"

<p>The present .csv file contains the raw data from a 2017 data collection. Specifically, children from six to 11 years old were tested with a non-visual spatial orientation task, in which they were required to i) observe animal-shaped landmarks located in each of the 4 angles of the experimental room; ii) be guided by the experimenter along a two-legged segment; iii) indicate the location of the four landmarks.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Household information for houses in Kasungu district participating in the Maladrone study, 2021

<p>Each row in the dataset contains information on each household participating in the Maladrone study, 2021. The information is as follows:</p><p>uniqueid: Unique ID assigned to the household, comprising of two letters corresponding to the community (ML, CK, CP) and the study house number.</p><p>under_5: Number of people under the age of 5 who live in the house at the time of asking.</p><p>over_5: &nbsp;Number of people over the age of 5 who live in the house at the time of asking.</p><p>under_5<i>_</i>rwt: Number of people under the age of 5 who usually sleep in the room where the CDC light trap was set.</p><p>over_5<i>_</i>rwt: Number of people over the age of 5 who usually sleep in the room where the CDC light trap was set.</p><p>number_mosquito<i>_</i>nets: Number of mosquito nets in the household.</p><p>mosquito_nets_rwt: Number of mosquito nets usually used in the room where the CDC light trap was set.</p><p>under_5<i>_</i>nets: Number of under 5s in the household who usually slept under a net.</p><p>over_5_nets: Numer of over 5s in the household who usually slept under a net.</p><p>unde_5_nets_rwt: Number of under 5s in the household who usually slept under a net in the room where the CDC light trap was set.</p><p>over_5_nets_rwt: Number of under 5s in the household who usually slept under a net in the room where the CDC light trap was set.</p><p>roof_type: Main material used for the roof of the house. Choices were iron sheet, thatched or tiled.</p><p>eaves: Whether the eaves of the house were open, partally open, or closed.</p><p>windows: Whether the windows of the house were open, partally open, or closed.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Supporting Data: ontophylo: Reconstructing the evolutionary dynamics of phenomes using new ontology-informed phylogenetic methods

<p>This dataset contains all scripts and data for reproducing the analyses of the paper. The README files contain additional information.</p>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record