Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
37
datasets available to search
ShareScore release 0.9.0
Dataset results
37 results for “Tabular data”
Tabular summary of wine data.
<p>These data are the results of a chemical analysis of wines made out of grapes grown in a region in Italy. But it is derived from three different cultivators. The analysis determined the quantities of 13 constituents found in each of the three types of wines.</p>
R code, spatial and tabular data to fully reproduce STEPS simulations of population change for common brushtail possum, grassland melomys and northern brown bandicoot in northern Australia
<ol> <li> <p>The development of effective fire management for biodiversity conservation is a global challenge. The highly dynamic nature of fire, the difficulty in replicating 'real-world' fire experiments, and the need to understand population changes at large spatiotemporal scales make computer simulations particularly useful for identifying optimal fire management regimes for biodiversity conservation. </p> </li> <li> <p>We aimed to develop a flexible modelling approach with which to investigate how the spatiotemporal application of fire (i.e. management scenarios) influences savanna biodiversity. We used existing data from a landscape-scale fire experiment to develop population simulations for the common brushtail possum (<i>Trichosurus vulpecula</i>), grassland melomys (<i>Melomys burtoni</i>) and northern brown bandicoot (<i>Isoodon macrourus</i>) across the Kapalga area of Kakadu National Park in northern Australia. We simulated how populations were expected to change between 1995 and 2015 in response to the fire patterns observed at Kapalga over this period, and under a hypothetical management scenario of extensive prescribed burning.</p> </li> <li> <p>Our models predicted a substantial decline in all three species in response to the observed fire regime at Kapalga, suggesting that the fire patterns observed at Kapalga, with the associated mechanisms and interactions with other ecological processes, were not conducive with the persistence of native mammal populations. </p> </li> <li> <p>Our prescribed burning scenario had little effect on the predicted population trajectory of the common brushtail possum and grassland melomys, but markedly improved the population trajectory of the northern brown bandicoot. These inconsistencies highlight the need for a nuanced approach to fire management across northern Australian savannas, that is tailored to local conditions and management objectives. </p> </li> <li> <p>Synthesis and applications. The modelling approach outlined here, provides a basis for identifying fire patterns that are beneficial for conserving biodiversity, thereby increasing our capacity to establish clear targets for prescribed fire management. Importantly, this approach is flexible and can be easily adapted to other taxa and fire-prone ecosystems.</p> </li> </ol>
Tabular Data for figures of Soutar et al. 2022
<p>Tabular data of analysed data presented in figures of the manuscript "Regulation of mitophagy by the NSL complex underlies genetic risk for Parkinson’s disease at Chr16q11.2 and on the MAPT H1 allele" by Soutar et al., 2022.</p>
Tough Tables: Carefully Evaluating Entity Linking for Tabular Data
<p>Tough Tables (2T) is a dataset designed to evaluate table annotation approaches in solving the CEA and CTA tasks.<br> The dataset is compliant with the data format used in <a href="https://www.cs.ox.ac.uk/isg/challenges/sem-tab/2019/index.html">SemTab 2019</a>, and it can be used as an additional dataset without any modification. The target knowledge graph is DBpedia 2016-10.<br> Check out the <a href="https://github.com/vcutrona/tough-tables">2T GitHub repository</a> for more details about the dataset generation.</p> <p><strong>New in v3.0:</strong> We release the updated version of 2T! The target knowledge graphs are DBpedia <a href="http://downloads.dbpedia.org/wiki-archive/downloads-2016-10.html">2016-10</a> and Wikidata <a href="https://zenodo.org/record/6643443">20220521</a>. Starting from this version, the dataset is split into valid and test sets.</p> <p>This work is based on the following paper:</p> <blockquote> <p>Cutrona, V., Bianchi, F., Jimenez-Ruiz, E. and Palmonari, M. (2020). Tough Tables: Carefully Evaluating Entity Linking for Tabular Data. ISWC 2020, LNCS 12507, pp. 1–16.</p> </blockquote> <p>Note on License: This dataset includes data from the following sources. Refer to each source for license details:<br> - Wikipedia https://www.wikipedia.org/<br> - DBpedia https://dbpedia.org/<br> - Wikidata https://www.wikidata.org/<br> - SemTab 2019 https://doi.org/10.5281/zenodo.3518539<br> - GeoDatos https://www.geodatos.net<br> - The Pudding https://pudding.cool/<br> - Offices.net https://offices.net/<br> - DATA.GOV https://www.data.gov/</p> <p>THIS DATA IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.<br> <br> <strong>Changelog:</strong></p> <p><em><strong>v3.0</strong></em></p> <ul> <li>Both datasets require SemTab2020 CEA format (tab_id, row_id, col_id, entity). <ul> </ul> </li> <li>Tables IDs and artificial noise values differ from previous versions.</li> <li>Datasets are split into Valid and Test sets of tables.</li> <li>New GT for ToughTables-WD (2T_WD) <ul> <li>Entities Q23772518 and Q7327323 have been removed because they are no longer represented in WD</li> <li>Updated ancestor/descendant hierarchies to evaluate CTA.</li> </ul> </li> <li>Evaluation scripts are provided with the data sets.</li> </ul> <p><em><strong>v2.0</strong></em></p> <ul> <li>New GT for 2T_WD <ul> <li>A few entities have been removed from the CEA GT, because they are no longer represented in WD (e.g., dbr:Devonté points to wd:Q21155080, which does not exist)</li> <li>Tables codes and values differ from the previous version, because of the random noise.</li> <li>Updated ancestor/descendant hierarchies to evaluate CTA.</li> </ul> </li> </ul> <p><em><strong>v1.0</strong></em></p> <ul> <li>New Wikidata version (2T_WD)</li> <li>Fix header for tables CTRL_DBP_MUS_rock_bands_labels.csv and CTRL_DBP_MUS_rock_bands_labels_NOISE2.csv (column 2 was reported with id 1 in target - NOTE: the affected column has been removed from the SemTab2020 evaluation)</li> <li>Remove duplicated entries in tables</li> <li>Remove rows with wrong values (e.g., the Kazakhstan entity has an empty name "''")</li> <li>Many rows and noised columns are shuffled/changed due to the random noise generator algorithm</li> <li>Remove row "Florida","Floorida","New York, NY" from TOUGH_WEB_MISSP_1000_us_cities.csv (and all its NOISE1 variants)</li> <li>Fix header of tables: <ul> <li>CTRL_WIKI_POL_List_of_current_monarchs_of_sovereign_states.csv</li> <li>CTRL_WIKI_POL_List_of_current_monarchs_of_sovereign_states_NOISE2.csv</li> <li>TOUGH_T2D_BUS_29414811_2_4773219892816395776___videogames_developers.csv</li> <li>TOUGH_T2D_BUS_29414811_2_4773219892816395776___videogames_developers_NOISE2.csv</li> </ul> </li> </ul> <p><em><strong>v0.1-pre</strong></em></p> <ul> <li>First submission. It contains only tables, without GT and Targets.</li> </ul>
AusCAT Simulation Tabular Data
<p>This data is part of the NCLC Radiomics dataset provided by The Cancer Imaging Archive (TCIA) and is</p> <p>released under the Creative Commons Attribution 3.0 Unported License.</p> <p> </p> <p>CAUTION: This data is provided for experimental purposes only. Data is partly synthetically generated.</p> <p> </p> <p>Citations & Data Usage Policy</p> <p>Users of this data must abide by the TCIA Data Usage Policy and the Creative Commons Attribution</p> <p>3.0 Unported License under which it has been published. Attribution should include references to</p> <p>the following citations:</p> <p> </p> <p>Dataset Citation</p> <p>Aerts, H. J. W. L., Wee, L., Rios Velazquez, E., Leijenaar, R. T. H., Parmar, C., Grossmann, P., Carvalho, S., Bussink, J., Monshouwer, R., Haibe-Kains, B., Rietveld, D., Hoebers, F., Rietbergen, M. M., Leemans, C. R., Dekker, A., Quackenbush, J., Gillies, R. J., Lambin, P. (2019). Data From NSCLC-Radiomics [Data set]. The Cancer Imaging Archive. <https://doi.org/10.7937/K9/TCIA.2015.PF0M9REI></p> <p> </p> <p>Publication Citation</p> <p>Aerts, H. J. W. L., Velazquez, E. R., Leijenaar, R. T. H., Parmar, C., Grossmann, P., Carvalho, S., Bussink, J., Monshouwer, R., Haibe-Kains, B., Rietveld, D., Hoebers, F., Rietbergen, M. M., Leemans, C. R., Dekker, A., Quackenbush, J., Gillies, R. J., Lambin, P. (2014, June 3). Decoding tumour phenotype by noninvasive imaging using a quantitative radiomics approach. Nature Communications. Nature Publishing Group. <https://doi.org/10.1038/ncomms5006></p> <p> </p> <p>TCIA Citation</p> <p>Clark K, Vendt B, Smith K, Freymann J, Kirby J, Koppel P, Moore S, Phillips S, Maffitt D, Pringle M, Tarbox L, Prior F. The Cancer Imaging Archive (TCIA): Maintaining and Operating a Public Information Repository, Journal of Digital Imaging, Volume 26, Number 6, December, 2013, pp 1045-1057. <https://doi.org/10.1007/s10278-013-9622-7></p> <p> </p>
Tabular data of inappropriate doses of antibiotics among pediatric patients
<p>This dataset contains information of inappropriate doses of antibiotics among pediatric patients</p>
R code, spatial and tabular data to fully reproduce STEPS simulations of population change for common brushtail possum, grassland melomys and northern brown bandicoot in northern Australia
Open the record for dataset details and reuse information.
Cardiovascular synthetic tabular data
<p>This dataset is focuses on cardiovascular diseases. It is generated using a hybrid machine learning model that combines diffusion models with Transformers, emphasizing data privacy. The dataset has been meticulously validated for quality and utility, yielding auspicious results.<br>Validation and Metrics:<br>The dataset has undergone rigorous validation processes to ensure quality, utility, and privacy. These validations involved:</p> <ol> <li>Distance to the Closest Record (DCR): The dataset achieved a DCR of 1.2879. The DCR is a metric that measures the distance of the generated data to the closest record in the original dataset. A higher DCR indicates that the synthetic data closely mirrors the real data in terms of statistical properties, making it reliable for further analysis and research.</li> <li>Membership Inference Attack Accuracy: The dataset scored 0.6780 in this metric. Membership inference attack accuracy measures the likelihood of correctly inferring whether a particular data point was part of the training dataset. An accuracy of 0.6780 suggests that the model maintains a strong level of privacy. It is important to note that a score of 0.5 would indicate random guessing, hence the achieved score demonstrates significantly better privacy protection than random predictions.</li> <li>Statistical Tests: Comprehensive statistical tests were conducted to compare the synthetic data with real data. These tests ensure that the synthetic data has similar statistical properties and distributions to the original data.</li> <li>Machine Learning Efficiency: The utility of the dataset was also validated using machine learning models to ensure that the synthetic data is effective for training and can produce reliable predictive models. The results showed that models trained on this dataset performed well, reinforcing the practical utility of the data.</li> </ol> <p>The high DCR value and the membership inference attack accuracy highlight the balance between data utility and privacy, making this dataset an invaluable resource for researchers and practitioners focusing on cardiovascular diseases and machine learning.</p>
Global Fire Emissions Indicators, Country-Level Tabular Data: 1997-2015
The Global Fire Emissions Indicators, Country-Level Tabular Data: 1997-2015 contains country tabulations from 1997 to 2015 for the total area burned (hectares) and total carbon content (tons). The annual total area burned is for all fire types per country. There are two groups of total carbon content (TCC), annual totals for all six fire types per country and annual totals for each of six fire types per country which include Agricultural, Boreal, Tropical Deforestation, Peat, Savanna, and Temperate forest fires.
Tabular summary of a Supermarket data
<p>Contains information about the customer invoices, sales and order details of a Hotel.</p>
GO NIMS TABULAR DATA FROM THE SL9 IMPACT WITH JUPITER V1.0
The Near Infrared Mapping Spectrometer (NIMS) on the Galileo spacecraft took unique data of Comet Shoemaker-Levy/9's impact with Jupiter. A preliminary analysis of this data is presented in this submission to the Planetary Data System (PDS). It consists of nine small tables with detached labels and documentation.
GO UVS TABULAR DATA FROM THE SL9 IMPACT WITH JUPITER V1.0
The UltraViolet Spectrometer (UVS) on the Galileo spacecraft took unique data of Comet Shoemaker-Levy/9's impact with Jupiter. A preliminary analysis of this data is presented in this submission to the Planetary Data System (PDS). It consists of two small tables with detached labels and documentation.
GO UVS TABULAR DATA FROM THE SL9-G IMPACT WITH JUPITER V1.0
The UltraViolet Spectrometer (UVS) on the Galileo spacecraft took unique data of Comet Shoemaker-Levy/9's impact with Jupiter. A preliminary analysis of this data is presented in this submission to the Planetary Data System (PDS). It consists of four small tables with detached labels and documentation.
Supplementary Data and Models of Melt-based Thermo-barometer for paper "'No Free Lunch' in Tabular Geochemical Data: An example of Shallow versus Deep Machine Learning Algorithms for Geothermobarometry"
Open the record for dataset details and reuse information.
GO NIMS TABULAR DATA FROM THE SL9 IMPACT WITH JUPITER V1.0
The Near Infrared Mapping Spectrometer (NIMS) on the Galileo spacecraft took unique data of Comet Shoemaker-Levy/9's impact with Jupiter. A preliminary analysis of this data is presented in this submission to the Planetary Data System (PDS). It consists of nine small tables with detached labels and documentation.
GO UVS TABULAR DATA FROM THE SL9 IMPACT WITH JUPITER V1.0
The UltraViolet Spectrometer (UVS) on the Galileo spacecraft took unique data of Comet Shoemaker-Levy/9's impact with Jupiter. A preliminary analysis of this data is presented in this submission to the Planetary Data System (PDS). It consists of two small tables with detached labels and documentation.
GO UVS TABULAR DATA FROM THE SL9-G IMPACT WITH JUPITER V1.0
The UltraViolet Spectrometer (UVS) on the Galileo spacecraft took unique data of Comet Shoemaker-Levy/9's impact with Jupiter. A preliminary analysis of this data is presented in this submission to the Planetary Data System (PDS). It consists of four small tables with detached labels and documentation.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.