Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,394
datasets available to search
ShareScore release 0.9.0
Dataset results
2,394 results for “Containers”
Polygons with small denominator containing a small number of lattice points
<p>The denominator of a rational polytope \(P\) is an integer \(r\) such that the dilated polytope \(rP\) has lattice point vertices. The size of a polytope is the number of lattice points it contains. This dataset contains polygons with denominator 2 and 3 with small size, classified using a growing algorithm as described in [HHK24].</p> <p>The data consists of files "denom_r_size_k_polygons.txt" which record the denominator \(r\) size \(k\) polygons \(P\). Each entry consists of the vertices and volume of \(rP\), the Ehrhart \(\delta\)-vector/\(h^*\)-vector of \(P\), and an ID number, which is unique among polygons of given size and denominator. Entries are ordered by their ID number. There are 50,564 entries in total.</p> <p><strong>Example entry:</strong></p> <p>ID=1<br>Vertices=[[ 1, 0 ], [ 0, 1 ], [ 3, 5 ]]<br>Volume=7<br>DeltaVec=[ 1, 0, 3, 7, 3, 0 ]</p> <p>If you make use of this data, please cite [HHK24] and the DOI for this data:</p> <p>doi:10.5281/zenodo.14230584</p> <p><strong>References:</strong></p> <p>[HHK24] Girtrude Hamm, Johannes Hofscheier, Alexander Kasprzyk, Classification and Ehrhart Theory of Denominator 2 Polygons. (preprint) arxiv:2411.19183</p>
JSON files containing parameters of training gene models for ab-initio prediction software
<p>These are the JSON files containing parameters of training gene models for ab-initio prediction software. These training datasets are Phytophthora specific and can be further utilized for the gene prediction and annotation of other related Phytophthora strains.</p>
Dataset containing DTS-data used in Karttunen et al. "Quantifying coastal urban surface layer structure using distributed temperature sensing in Helsinki, Finland"
<p>This record contains DTS-data used in the following study:</p> <p>Karttunen et al. (2021): Quantifying coastal urban surface layer structure using distributed temperature sensing in Helsinki, Finland, submitted to AMTD</p> <p> </p> <p>DTS_highfreq_SMEARIII_Karttunen_et_al.zip contains continuous high frequency potential temperature profiles measured along the SMEAR III 31-metre tall mast. See more information on the data in the netCDF-file attributes and on the measurement setup in the related manuscript.</p> <p>DTS_statistics_SMEARIII_Karttunen_et_al.nc contains profiles for the turbulence temperature statistics calculated from the continuous DTS potential temperature profiles.See more information in the netCDF-file attributes and the related manuscript.</p> <p> </p>
A dataset of environmental audio recordings containing chainsaw events
<p>This audio dataset contains audio recordings (.wav format, 8kHz sampling rate) of various durations acquired using eight Cornel University SWIFT Autonomous Recording Units (ARUs) and corresponding metadata (.textgrid files). The recordings were collected within the Rodopi Mountain-Range National Park, Greece, at different seasons between 2018 and 2019. They are part of longer audio recordings. This audio content was used for the evaluation of a chainsaw sound detection system that was developed for the needs of a research project funded via a Single RTDI State Aid Action ”Research – Create – Innovate” grant (T1EDK-04488), which is co-financed by Greece and the European Union (European Regional Development Fund) as part of the Operational Program ”Competitiveness, Entrepreneurship and Innovation” of the National Strategic Reference Framework (NSRF) 2014-2020.</p> <p>Portions of each uploaded recording contain chainsaw events. The temporal location of these events, originally located manually by human listeners, are marked within the corresponding .textgrid files. Each audio recording with the corresponding textgrid file can be opened using the Praat program (<a href="https://www.fon.hum.uva.nl/praat/">https://www.fon.hum.uva.nl/praat/</a>).</p> <p>If you find this dataset useful please cite the following paper:</p> <p>N. Stefanakis, K. Psaroulakis, N. Simou and C. Astaras, "An open-access system for long-range chainsaw sound detection", in Proceedings of EUSIPCO (2022).</p> <p>The chainsaw sound detection algorithm that was developed is also open access and is available in the form of python code from <a href="https://github.com/spl-icsforth/An-open-access-system-for-long-range-chainsaw-sound-detection">https://github.com/spl-icsforth/An-open-access-system-for-long-range-chainsaw-sound-detection</a></p>
Database containing harmonized datasets - FAIRWAY Project Deliverable 3.3
<p>A database has been developed and delivered during the <a href="https://www.fairway-project.eu/">FAIRWAY project</a>. This database was developed as a response to the need to harmonize datasets and assessment methods related to pressure and state indicators for water quality in the EU member states, in order to compare and assess indicators using a harmonized approach.</p> <p>The dataset that is made available here provides two files:</p> <ul> <li>a <em>public version*</em> of the <strong>Excel database</strong>, which contains <strong>all "tabular" (non-GIS)</strong> data related to the 13 case studies that was gathered for the purposes of FAIRWAY's Monitoring & Indicators research theme. It is structured as one "data sheet" and one "summary sheet" per case study. The data sheets contain various parameters (ideally time-dependent data series i.e. time series) that were used, wherever possible, to compute relevant Agri-drinking water quality indicators (ADWIs) such as "nitrogen budget" (a compound Pressure indicator) or "lag time" (a statistically-inferred Link indicator).</li> <li>a <strong>ZIP folder</strong> containing <strong>all GIS data</strong> gathered for the FAIRWAY's Monitoring & Indicators research theme. The GIS files are grouped in subfolders, by case study, and then by keywords describing the nature of the spatial data.</li> </ul> <p>The Excel database contains near 385,000 rows of data from the 13 case study sites, with more than 65 parameters and more than 500 sub-parameters. The dataset also contains spatial information in a GIS-data zipped folder. The spatial mapping information can be made visible using basic <a href="https://www.qgis.org/en/site/">QGIS</a> project files (.qgz), so that GIS data from each case study can be explored.</p> <p>The indicators database can be used in several ways. It may be used to explore data or to calculate additional indicators. Depending of the case studies’ interests, the most commonly available State indicators are about nitrate and pesticides concentrations in water.</p> <p>From a practical point of view based on its actual content, the database may notably be used to explore statistical relations (or Links) between related Pressure and State indicators. This database can also be used as a spatial mapping portal for other usages.</p> <p>For more information on the database, follow <a href="https://fairway-is.eu/index.php/farm-management/workpackages/harmonised-indicator-database">this link</a>.</p> <p>* Note that this is a <em>public version</em> of the database, which means that all confidential data was removed from the data sheets.</p>
TransProteus, Predicting 3D shapes, masks, and properties of materials, liquids, and objects inside transparent containers from images
<p>We present TransProteus, a dataset, for predicting the 3D structure and properties of materials, liquids, and objects inside transparent vessels from a single image without prior knowledge of the image source and camera parameters. Manipulating materials in transparent containers is essential in many fields and depends heavily on vision. This work supplies a new procedurally generated dataset consisting of 50k images of liquids and solid objects inside transparent containers. The image annotations include 3D models and material properties (color/transparency/roughness...) for the vessel and its content. The synthetic (CGI) part of the dataset was procedurally generated using 13k different objects, 500 different environments (HDRI), and 1450 material textures (PBR) combined with simulated liquids and procedurally generated vessels. In addition, we supply 104 real-world images of objects inside transparent vessels with depth maps of both the vessel and its content.</p> <p>Note that there are two files here:</p> <p><a href="https://zenodo.org/api/files/12b013ca-36be-4156-afd4-c93b5fa22093/Tansproteus_SimulatedLiquids2_New_No_Shift.7z">Transproteus_SimulatedLiquids2_New_No_Shift.7z</a></p> <p>and</p> <p><br> <a href="https://zenodo.org/api/files/2b833de0-4007-4682-ad5b-5e08bd63597e/TranProteus2.7z?versionId=f16e7126-8750-41f7-99e6-d35ca60399cc">TranProteus2.7z </a>, contain subset of the virtual CGI data set.</p> <p>https://zenodo.org/api/files/12b013ca-36be-4156-afd4-c93b5fa22093/Tansproteus_SimulatedLiquids2_New_No_Shift.7z</p> <p><a href="https://zenodo.org/api/files/2b833de0-4007-4682-ad5b-5e08bd63597e/TransProteus_RealSense_RealPhotos.7z">TransProteus_RealSense_RealPhotos.7z </a>: Contain real-world photos scanned with real sense with depth map of both the vessel and its content</p> <p>See ReadMe file in side the downloaded files for more details</p> <p>The full dataset (>100gb) can be found here:</p> <p><a href="https://e.pcloud.link/publink/show?code=kZfx55Zx1GOrl4aUwXDrifAHUPSt7QUAIfV">https://e.pcloud.link/publink/show?code=kZfx55Zx1GOrl4aUwXDrifAHUPSt7QUAIfV</a></p> <p>https://<a href="http://icedrive.net/1/6cZbP5dkNG">icedrive.net/1/6cZbP5dkNG</a></p> <p>See: <a href="https://arxiv.org/pdf/2109.07577.pdf"> https://arxiv.org/pdf/2109.07577.pdf</a> for more details</p> <p><strong><a href="https://zenodo.org/record/4736111#.YVOAx3tE1H4">**This dataset is complementary to LabPics dataset with 8k real images of materials in vessels in chemistry labs, medical labs, and other settings. The LabPics dataset can be downloaded from here:</a></strong></p> <p><strong><a href="https://zenodo.org/record/4736111#.YVOAx3tE1H4">https://zenodo.org/record/4736111#.YVOAx3tE1H4</a></strong></p> <p> </p> <p><strong>************************************************************************************</strong></p> <p><a href="https://zenodo.org/api/files/12b013ca-36be-4156-afd4-c93b5fa22093/Tansproteus_SimulatedLiquids2_New_No_Shift.7z">Transproteus_SimulatedLiquids2_New_No_Shift.7z </a>and <a href="https://zenodo.org/api/files/2b833de0-4007-4682-ad5b-5e08bd63597e/TranProteus2.7z?versionId=f16e7126-8750-41f7-99e6-d35ca60399cc">TranProteus2.7z</a></p> <p>The two folders contain relatively similar data styles.<br> The data in No_Shift contain images that were generated with no camera shift in the camera paramters. If you try to predict 3d model from an image as a depth map, this is easier to use (Otherwise, you need to adapt the image using the shift). For all other purposes, both folders are the same, and you can use either or both. In addition, a real image dataset for testing is given in the RealSense file.</p> <p> </p> <p> </p> <p> </p>
All data of the manuscript "A self-sustained charge neutrality lightning model containing the channel decay and reactivation process" submitted to Geophysical Research Letters
<p>The data supports the manuscript entitled "A self-sustained charge neutrality lightning model containing the channel decay and reactivation process”. Microsoft Notepad can open the *.txt files, they contain the channel information of two intracloud flashes (IC1 and IC2) and the channel elctrical parameters at the first fork of positive or negative leader channels. A normal video player software can open Movies S1.avi, and it shows the entire development process of IC1 discharge.</p> <p>The data can be used freely for scientific purposes with the appropriate citation.</p>
Twitter Dataset - Over 200,000 Tweets containing the word "Vaccine" for research porpuses
<p>This dataset contains 220,085 tweets containing the word vaccine between December 9th and December 18th 2021 at different times during each day, extracted using the Twitter API v2. Each tweet was extracted at least 3 days after its initial posting time in order to register 3 days of engagements, and it doesn't include retweets.</p> <p>Includes:</p> <ul> <li>Tweet ID</li> <li>Text</li> <li>Author ID</li> <li>Date</li> <li>Like count</li> <li>Retweet count</li> <li>Quote count</li> <li>Reply count</li> <li>User data (Followers, Following, Tweet count, Account creation date, Verified status)</li> </ul> <p>Usernames are hidden for privacy reasons</p>
ams-icdd-usecases: Use cases for employing ICDD containers for infrastructure asset management
<p>This repository provides two use cases for infrastructure asset management using information containers according to the <a href="https://www.iso.org/standard/74389.html">Information Container for linked Document Delivery standard (ISO 25197)</a>. Use Case 1 demonstrates the data preparation and result collection of bridge visual inspection with requirement- and delivery container. Use Case 2 demonstrates the pavement maintenance plan based on the existing condition data. Therefore, an additional connection to a relational database in the information container is provided in Use Case 2, which is registered within the container using an extension <a href="https://icdd.vm.rub.de/ontology/icdd/ExtendedDocument/">EXDOC:Extension for document types for the ISO 21597 ICDD Part 1 Container ontology</a>.</p> <p><strong>Full Changelog</strong>: <a href="https://github.com/RUB-Informatik-im-Bauwesen/ams-icdd-usecases/commits/v0.1">https://github.com/RUB-Informatik-im-Bauwesen/ams-icdd-usecases/commits/v0.1</a></p>
Dataset related to aticle "Additive Fabrication of a Vascular 3D Phantom for Stereotactic Radiosurgery of Arteriovenous Malformations"The database contains 3D models in STL file format of a patient-specific brain arteriovenous malformation phantom reconstructed from computed tomography scans.
<p><em>The database contains 3D models in STL file format of a patient-specific brain arteriovenous malformation phantom reconstructed from computed tomography scans.</em></p>
FIGURE 3 Morphological trait sampling for all bird families. AVONET contains 718,662 in AVONET: morphological, ecological and geographical data for all birds
FIGURE 3 Morphological trait sampling for all bird families. AVONET contains 718,662 individual trait measurements, all of which are used to calculate species averages. However, sampling per species varies across families depending on taxonomy. Upper phylogram shows sampling under BirdLife International (11,009 species in 243 families). Families where sampling completeness is below 75% indicated by lighter shading. Most families with lower sampling are species poor (numbers in black circles show species richness). Lower panels show that sampling improves under more conservative taxonomic treatments of eBird (10,661 species in 249 families) and BirdTree (9993 species in 194 families). Coloured bars indicate the proportion of species in each family measured to different levels of completeness. 'Complete set' means a full set of all 9 core morphological traits (not necessarily from the same individual). 'Individuals' means any individual bird with one or more traits measured
Figures 27–32. Some smaller copal pieces containing T in Extinct or extant? A new species of Termitodius Wasmann, 1894, (Coleoptera: Scarabaeidae: Aphodiinae: Rhyparini) with a short review of the genus
Figures 27–32. Some smaller copal pieces containing T. woodruffi paratypes, host termites and other inclusions. 27) CMNC. 28) REWC. 29–30) CEMT. Scale line = 1 mm. 31) IAvH-E. 32) FSCA (ex. RLBC), arrow indicates male with genitalia extracted, see Fig. 17. Photos for Figures 29–30 by Vinícius Costa-Silva (CEMT).
Oral cancer speech corpus for the paper "Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens"
<p>Dataset accompanying the paper "<em>Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens</em>"</p> <p>The zip file contains five folders:</p> <p>- <strong>Database:</strong> contains csv files for each speaker which contain the processed features</p> <p>- <strong>Recordings: </strong>the original recording from the YouTube Oral Cancer speech dataset, without further preprocessing</p> <p>- <strong>Recordings_Normalised:</strong> same as recordings but after minimal audio preprocessing (min-max scaling)</p> <p>- <strong>Textgrids: </strong>contains the textgrids which are annotated on the word-level and on phoneme-level</p> <p>- <strong>TIMIT selection: </strong>contains the textgrids for the TIMIT speakers. We unfortunately cannot share the audio date as it is not open source. More information can be found <a href="https://catalog.ldc.upenn.edu/LDC93s1">here.</a></p>
The Huanan Market Origin of SARS-CoV-2 is unlikely: The ancestral lineage containing specimen appears to have arisen from laboratory contamination
<p>• There is universal agreement that the lineage B/L is not the ancestral SARS-CoV-2; lineage A/S is the most ancestral lineage<br> • Until the Gao paper, no lineage A/S virus was identified at the Huanan Market, making the market an unlikely origin for the pandemic<br> • The Gao paper found one specimen, A20, with both lineage A/S and B/L<br> • The lineage B/L reads in A20 and the specimen Ct matched the expected findings<br> • The lineage A/S reads were anomalously high compared to the Ct and meet the definition of a statistical outlier<br> • The SARS-CoV-2 reads from the A20 sample also had two SNVs not seen in GISAID sequences until at least 60-90 days after the specimen was collected from the market<br> • This analysis supports a finding that the ancestral lineage A/S virus sequences in sample A20 were not present on the glove on January 1, 2020 when the sample was collected but instead arose by inadvertent laboratory contamination later, probably during metagenomic sequencing</p> <p>The absence of an unimpeachable ancestral lineage specimen at the Hunan Market makes it unlikely the market was the origin of the pandemic.</p>
Text-fig. 11. CGM 94-138, right mandible fragment containing p/4–m/3 of Libycochoerus massai from Moghara, Egypt. a: lingual view; b: stereo occlusal view; c: buccal view. in New Suoid Fossils (Mammalia, Artiodactyla) From The Miocene Of Moghara, Egypt, And Gebel Zelten, Libya: Biochronological Implications
Text-fig. 11. CGM 94-138, right mandible fragment containing p/4–m/3 of Libycochoerus massai from Moghara, Egypt. a: lingual view; b: stereo occlusal view; c: buccal view.
Database that contains all images (plus 180 more) employed in the article: "Image features for quality analysis of thick blood smears employed in malaria diagnosis"
<p>We share with you a bank of images obtained from microscopic fields of thick blood smears employed in the malaria diagnosis, and also the .csv file that contains the labels for each image.</p> <p>The images are saved with a unique name that is found in the first column of the .csv file. The second column contains the labels from each image, according to their unique names.</p> <p>The labeling process was done with the online toolbox Labelbox. Labelbox, "Labelbox," Online, 2020. [Online]. Available: https://labelbox.com </p> <p>If you are interested in using our database, cite our article as a way to recognize our work. We will be grateful for that. </p> <p>CITATION: Fong Amaris, W.M., Martinez, C., Cortés-Cortés, L.J. et al. Image features for quality analysis of thick blood smears employed in malaria diagnosis. Malar J 21, 74 (2022). https://doi.org/10.1186/s12936-022-04064-2</p> <p>URL of our paper: https://malariajournal.biomedcentral.com/articles/10.1186/s12936-022-04064-2</p> <p><strong>--- This is the link where you can find our images Bank: https://drive.google.com/drive/folders/1Qrv0e4bSEtkeqtPABz-klQp-6D6OjU-X?usp=sharing </strong></p> <p>It is important you to know that along with this .txt file, we are sharing the .csv file that contains 600 names of images (in the first column) with their respective labels (second column aside)</p> <p>This file corresponds to the instructions of an extended label file related to 600 images (and 600 new labels) in contrast to our previous label file with 420 labels from 420 images (https://www.researchgate.net/publication/359439520_Database420LabelsInstructionstxt ; https://www.researchgate.net/publication/359438904_Database420Labelscsv).</p> <p>Best Regards</p> <p> </p>
Container flows on road, rail and waterways along Rhine-Alpine corridor (Rhine section) at NUTS-2 level with cost-time-emissions estimates and accessibility-frequency-availability of modes
<p>The present dataset is used to estimate the heterogeneous mode choice preferences of shippers, that are presented in the following article :<br> "A Logit Mixture Model Estimating the Heterogeneous Mode Choice Preferences of Shippers Based on Aggregate Data"<br> (Nicolet, A., Negenborn, R. R. & Atasoy, B., A Logit Mixture Model Estimating the Heterogeneous Mode Choice Preferences of Shippers Based on Aggregate Data. IEEE Open Journal of Intelligent Transportation Systems, Vol. 3, 2022, pp. 650-661.)</p>
DCM containing analog series
<p>The file contains 1400 analog series that consist of DCM compounds from PubChem and their bioactive analogs from ChEMBL. For all 14,796 analogs in these series, analog series ID (AS_ID), ChEMBL_ID or PubChem compound identifier (PubChem_cid) and SMILES are provided. The targets corresponding to ChEMBL compounds in the series are also provided (ChEMBL_ID_targets). In addition, DCM compounds are indicated.</p>
Datasets containings simulation results
<p>These are the datasets containings the simulation results presented in "Approximate Bayes factors: flexible and likelihood-free model comparison with ABrox".</p>
X-Ray Structures of Target-Ligand Complexes Containing Compounds with Assay Interference Potential
<p>A total of 2755 crystallographic complexes with ligands containing PAINS-defining substructures were extracted from the Protein Data Bank (PDB). PDB identifiers of these structures are made available together with the the corresponding PDB_PAINS (component identifier, aromatic nonstereo SMILES, PAINS class). </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.