Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4,151

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

4,151 results for “discovery”

Learn how ShareScore rates datasets ↗
edi60/100

GRIME AI Water Segmentation Model for the USGS Monitoring Site Discovery Farms Waterway AO1 Near Antigo, WI, 2023-2024

Ground-based observations from fixed-mount cameras have the potential to fill an important role in environmental sensing, including direct measurement of water levels and qualitative observation of ecohydrological research sites. All of this is theoretically possible for anyone who can install a trail camera. Easy acquisition of ground-based imagery has resulted in millions of environmental images stored, some of which are public data, and many of which contain information that has yet to be used for scientific purposes. The goal of this project was to develop and document key image processing and machine learning workflows, primarily related to semi-automated image labeling, to increase the use and value of existing and emerging archives of imagery that is relevant to ecohydrological processes. This data package includes imagery, annotation files, water segmentation model and model performance plots, and model test results (overlay images and masks) for the USGS monitoring location Discovery Farms Waterway AO1 Near Antigo, WI (2023-2024). All imagery was acquired from the USGS Hydrologic Imagery Visualization and Information System (HIVIS; see https://apps.usgs.gov/hivis/camera/WI_AO1_STAFF for this specific data set) and/or the National Imagery Management System (NIMS) API. Water segmentation models were created by tuning the open-source Segment Anything Model 2 (SAM2, https://github.com/facebookresearch/sam2) using images that were annotated by team members on this project. The models were trained on the "water" annotations, but annotation files may include additional labels, such as "snow", "sky", and "unknown". Image annotation was done in Computer Vision Annotation Tool (CVAT) and exported in COCO format (.json). All model training and testing was completed in GaugeCam Remote Image Manager Educational Artificial Intelligence (GRIME AI, https://gaugecam.org/) software (Version: Beta 16). Model performance plots were automatically generated during this process.

openCC (other)Sep 2025View details →
edi60/100

Meteorological data from the Discovery Tree at the Andrews Experimental Forest, 2015 to present

The upper-canopy of forests is known to experience a very different leaf wetness than the rest of the forest: it is often simultaneously brighter, hotter, windier, and drier. The upper canopy also contains most of the leaf area, and because it absorbs most of the solar radiation, it accounts for the great majority of carbon and water exchanges in most forests. Critically, this is also the zone where most climate variations and stress likely manifest. The upper canopy is also the region of the forest that is sampled by satellite imagery. Intensive canopy microclimate monitoring provides connections to satellite-based imagery at varying temporal and spatial scales in order to scale across the Andrews landscape and improves our understanding of forest function and its response to climate change. A 50 meter old growth tree, called the Discovery Tree, was instrumented with various sensors. A thermal infrared camera was installed in March 2014, which collects surface temperatures of the old-growth forest and the adjacent secondary-growth forest. Since then, the scope of information being acquired in real-time has increased to include temperature, leaf wetness, relative humidity, soil temperature, soil moisture, wind direction and speed. This suite of data serves as a glimpse into the canopy and soil processes we are unaware of when our feet are planted firmly on the ground. These measurements complement and leverage ongoing, long-term climate measurements collected in the sub-canopy and at the climate stations located across the Andrews forest, and potentially link with Lidar data on canopy structure and planned soil moisture measurements. Canopy thermal imaging and microclimate measurements have been established for ecophysiological applications such as monitoring the response of forest tree canopies to climate variations, including heat and drought stress.

openCC (other)Apr 2023View details →
zenodo48/100

Data from: A FAIR and modular image-based workflow for knowledge discovery in the emerging field of imageomics

<p>Data and results from the Imageomics Workflow. These include data files from the Fish-AIR repository (https://fishair.org/) for purposes of reproducibility and outputs from the application-specific imageomics workflow contained in the Minnow_Segmented_Traits repository (https://github.com/hdr-bgnn/Minnow_Segmented_Traits).</p> <p>Fish-AIR:<br> This is the dataset downloaded from&nbsp;Fish-AIR, filtering for Cyprinidae and the Great Lakes Invasive Network (GLIN) from the Illinois Natural History Survey (INHS) dataset. These files contain information about fish images, fish image quality, and path for downloading the images. The data download ARK ID is dtspz368c00q. (2023-04-05). The following files are unaltered from the Fish-AIR download. We use the following files:</p> <p>extendedImageMetadata.csv: A CSV file containing information about each image file. It has the following columns: ARKID, fileNameAsDelivered, format, createDate, metadataDate, size, width, height, license, publisher, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>imageQualityMetadata.csv: A CSV file containing information about the quality of each image. It has the following columns: ARKID, license, publisher, ownerInstitutionCode, createDate, metadataDate, specimenQuantity, containsScaleBar, containsLabel, accessionNumberValidity, containsBarcode, containsColorBar, nonSpecimenObjects, partsOverlapping, specimenAngle, specimenView, specimenCurved, partsMissing, allPartsVisible, partsFolded, brightness,&nbsp;<br> uniformBackground, onFocus, colorIssue, quality, resourceCreationTechnique. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>multimedia.csv: A CSV file containing information about image downloads. It has the following columns: ARKID, parentARKID, accessURI, createDate, modifyDate, fileNameAsDelivered, format, scientificName, genus, family, batchARKID, batchName, license, source, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>meta.xml: A XML file with the metadata about the column indices and URIs for each file contained in the original downloaded zip file. This file is used in the fish-air.R script to extract the indices for column headers.</p> <p>The outputs from the Minnow_Segmented_Traits workflow are:</p> <p>sampling.df.seg.csv: Table with tallies of the sampling of image data per species during the data cleaning and data analysis. This is used in Table S1 in Balk et al.&nbsp;</p> <p>presence.absence.matrix.csv: The Presence-Absence matrix from segmentation, not cleaned. This is the result of the combined outputs from the presence.json files created by the rule &ldquo;create_morphological_analysis&rdquo;. The cleaned version of this matrix is shown as Table S3 in Balk et al.</p> <p>heatmap.avg.blob.png and heatmap.sd.blob.png: Heatmaps of average area of biggest blob per trait (heatmap.avg.blob.png) and standard deviation of area of biggest blob per trait (heatmap.sd.blob.png). These images are also in Figure S3 of Balk et al.</p> <p>minnow.filtered.from.iqm.csv: Filtered fish image data set after filtering (see methods in Balk et al. for filter categories).</p> <p>burress.minnow.sp.filtered.from.iqm.csv: Fish image data set after filtering and selecting species from&nbsp;Burress et al. 2017.</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

Stocktaking GO FAIR Discovery IN - Use cases, infrastructure

<p>In order to build a better ecosystem for data discovery tools the Data Discovery Implementation Group of GO Fair (https://www.go-fair.org/implementation-networks/overview/discovery) collected use cases between 2019 and 2020 from a variety of sources. We also detail the &lsquo;Actors&rsquo; for these use cases and the &lsquo;Source&rsquo; providing links, whenever possible. Since we found over a hundred individual use cases, we decided to cluster them to provide a better overview. The clustering, as well as the results of a small survey among data infrastructure specialists to find how they rate the importance of the clusters are&nbsp;detailed in the documentation to this dataset, a draft of which can currently be found <a href="https://docs.google.com/document/d/1sq78eCFYgmWcMFYcbNonA2KrkUO1qdGr7d49tHuRcRM/edit?usp=sharing">here</a>. The code and data to produce the figures in the&nbsp;documentation are available as R code in the GO_FAIR_Discovery_Use_case-master.zip file. The use cases themselves are available as Excel sheet and csv.&nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo48/100

Problem discovery and resolution activities in the Apache HTTP Server Project (March 2001- March 2013).

<p>This is a dynamic visualization of problem discovery and resolution activities&nbsp;observed in the in the development of the Apache HTTP Server Project during the period March 2001- March 2013. The nodes in the network represent problems (software bugs). Anthropomorphic icons represent participants (software developers). The network edges connect participants to problems.&nbsp; Numerical labels record the internal identification numbers or participants and problems.&nbsp; The visible clusters represent the software modules. The central node is the project core module. Participants move closer to problems that attract their attention. When a participant allocates attention to a problem, an edge emerges connecting the two. The edge is green when a participant opens a bug report (i.e., when he discovers a new problem), red when the participant closes the bug report (i.e., when she solves an existing problem), and yellow when any other action is recorded. Problems (white nodes) are green when they first appear. They turn red immediately before being closed, and are yellow when the corresponding bug report is being modified. &nbsp;The animation advances by 0.05 seconds every day of historical time.</p> <p>The animation is produced using the Gource server control visualization tool developed by Andrew Caldwell (<a href="https://gource.io/">https://gource.io/</a>)</p>

opencc-by-4.0Jan 2023View details →
zenodo48/100

Annotated Benchmark of Real-World Data for Approximate Functional Dependency Discovery

<p><strong>Annotated Benchmark of Real-World Data for Approximate Functional Dependency Discovery</strong></p> <p>This collection consists of ten open access relations commonly used by the data management community. In addition to the relations themselves (please take note of the references to the original sources below), we added three lists in this collection that describe approximate functional dependencies found in the relations. These lists are the result of a manual annotation process performed by two independent individuals by consulting the respective schemas of the relations and identifying column combinations where one column implies another based on its semantics. As an example, in the <em>claims.csv</em> file, the <em>AirportCode</em> implies <em>AirportName</em>, as each code should be unique for a given airport.</p> <p>The file <em>ground_truth.csv</em> is a comma separated file containing approximate functional dependencies. <em>table</em> describes the relation we refer to, <em>lhs</em> and <em>rhs</em> reference two columns of those relations where semantically we found that <em>lhs</em> implies <em>rhs</em>.</p> <p>The file <em>excluded_candidates.csv</em> and <em>included_candidates.csv</em> list all column combinations that were excluded or included in the manual annotation, respectively. We excluded a candidate if there was no tuple where both attributes had a value or if the <em>g3_prime</em> value was too small.</p> <p><strong>Dataset References</strong></p> <ul> <li><em>adult.csv</em>: Dua, D. and Graff, C. (2019). <a href="http://archive.ics.uci.edu/ml">UCI Machine Learning Repository</a>. Irvine, CA: University of California, School of Information and Computer Science.</li> <li><em>claims.csv</em>: TSA Claims Data 2002 to 2006, <a href="https://www.dhs.gov/tsa-claims-data">published by the U.S. Department of Homeland Security</a>.</li> <li><em>dblp10k.csv</em>: Frequency-aware Similarity Measures. Lange, Dustin; Naumann, Felix (2011). 243&ndash;248. <a href="https://hpi.de/naumann/projects/repeatability/datasets/dblp-dataset.html">Made available as DBLP Dataset 2</a>.</li> <li><em>hospital.csv</em>: Hospital dataset used in Johann Birnick, Thomas Bl&auml;sius, Tobias Friedrich, Felix Naumann, Thorsten Papenbrock, and Martin Schirneck. 2020. Hitting set enumeration with partial information for unique column combination discovery. Proc. VLDB Endow. 13, 12 (August 2020), 2270&ndash;2283. https://doi.org/10.14778/3407790.3407824. <a href="https://owncloud.hpi.de/s/j6Z0yvXC0qhtGCk/download">Made available as part the dataset collection to that paper.</a></li> <li><em>t_biocase_...</em> files: t_bioc_... files used in Johann Birnick, Thomas Bl&auml;sius, Tobias Friedrich, Felix Naumann, Thorsten Papenbrock, and Martin Schirneck. 2020. Hitting set enumeration with partial information for unique column combination discovery. Proc. VLDB Endow. 13, 12 (August 2020), 2270&ndash;2283. https://doi.org/10.14778/3407790.3407824. <a href="https://owncloud.hpi.de/s/j6Z0yvXC0qhtGCk/download">Made available as part the dataset collection to that paper.</a></li> <li><em>tax.csv</em>: Tax dataset used in Johann Birnick, Thomas Bl&auml;sius, Tobias Friedrich, Felix Naumann, Thorsten Papenbrock, and Martin Schirneck. 2020. Hitting set enumeration with partial information for unique column combination discovery. Proc. VLDB Endow. 13, 12 (August 2020), 2270&ndash;2283. https://doi.org/10.14778/3407790.3407824. <a href="https://owncloud.hpi.de/s/j6Z0yvXC0qhtGCk/download">Made available as part the dataset collection to that paper.</a></li> </ul>

opencc-by-4.0Jun 2023View details →
edi48/100

H.J. Andrews Forest Discovery Trail: An interpretation of place based on curriculum of interpretive learning trail and field trip support, 2016

The H.J. Andrews Experimental Forest (HJA) in the Oregon Cascades is one of 24 sites in the Long-Term Ecological Research (LTER) Network. It supports research on forests, streams, and watersheds, and fosters collaborations between ecosystem science, education, natural resource management, and the humanities. The site currently hosts 85 interdisciplinary research projects, as well as experiential training for undergraduate and graduate students. In addition, the HJA runs a vibrant professional development program for teachers. Because much of the HJA’s terrain is steep and occupied with sensitive research materials, middle and high school visits are limited to tours in designated areas. The Discovery Trail was developed in 2011 as a place for visitors (~1800 in 2014) to explore the forest and site research themes from HJA headquarters, but it is not yet amenable to unguided educational exploration. We have designed an interpretive learning trail and field trip support framework for the Discovery Trail. Our primary objective is to educate students about place while guiding them to reflect upon their own relationships with place and personal responsibility for stewardship behavior. Long-term place-based conservation research is woven with creative writing from the HJA writer’s residency program and paired with reflection and creative inquiry. Interactive trail stops enable students to engage the forest from multiple perspectives. The Discovery Trail is wired for intranet wifi and content and assessment will be delivered by digital media (i.e. iPads). We will evaluate conceptual learning according to the Framework for the Next Generation Science Standards, as well as observe affective changes in sense of place, empowerment, and expressions of care or empathy through analysis of student responses to the trail activities. Because conservation attitudes require not just knowledge about systems, but also emotional connections to the material, our learning experience will in

openCC (other)Oct 2016View details →
zenodo44/100

SNP and indel discovery and genotyping in next-generation sequencing data

<p>Code, logs and data for discovery and genotyping of SNPs and indels, in the the D.melanogaster genome, using GATK HaplotypeCaller. Code is in the zipped folder named code.zip. Run logs for this code as in the zipped folder named logs.zip. The unfiltered vcf genotypes file is named lhm_rg_HC_2015-09-15.vcf.gz. The filtered vcf genotypes file is named f1.lhm_rg_HC_raw.vcf.gz. The vcf submitted to NCBI dbSNP (filtered, and with indels &gt;50bp and variants with null alternate alleles both removed) is named dbSNP.lhm_rg_HC_raw.vcf.gz. The folder local_reference.zip contains the reference assembly files against which genotypes were called against, and includes the code used to format the data prior to use. Also included is genotypes data from the two in-house reference line samples sequenced (BDGP6+ISO1 mito/dm6, Bloomington <em>Drosophila</em> Stock Center no. 2057)</p> <p>Samples are 220 Sussex-LH<sub>M</sub> hemiclones, and 2 RG. The first run did not include chromosome 4 and the mitochondrial genome, so these were genotyped separately, and then added to the rest of the results.</p> <p>The link for the NCBI dbSNP record is currently https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461and the submitter handle is MORROW_EBE_SUSSEX.</p> <p>At the time of writting, the NCBI D.melanogaster build is still being updated, and therefore ss identifiers, but not rs identifers are available.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>

opencc-by-4.0Oct 2016View details →
zenodo44/100

Structural variant discovery and genotyping in next-generation sequencing data

<p>Code, logs, data, and summaries for detection and genotyping of genomic structural variants in the D.melanogaster Sussex LHM hemiclones (and one in-house reference line individual), using Genomestrip/2.0</p> <p>The unfiltered CNV pipleline results are lhm_gs.cnvs.raw.vcf.gz</p> <p>Filtered CNV results (including removal of bad samples) are filtered.goodS.lhm_gs.cnvs.raw.vcf.gz</p> <p>The file uploaded to NCBI dbVAR (which comprises of the filtered CNVs and indels &gt;50bp from the HaplotypeCaller method) is lhm_sx16.dbVAR.vcf.gz</p> <p>The NCBI dbVAR accession number is nstd134. Code, logs and summary data are in the zipped archives, named accordingly. The archive reference_data.zip contains additional input files required for Genomestrip, including a shell script for making some of them. The file gstrip_lhm_RG_bams.list is also an input for Genomestrip, indicating bam file names and paths.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>

opencc-by-4.0Oct 2016View details →
zenodo44/100

Convex inference for community discovery in signed networks (European Parliament Voting Dataset)

<p>This repository contains the necessary tools to reproduce the experiments of the paper</p> <ul> <li>G. Santatmaría, V. Gómez (2015)<br> Convex inference for community discovery in signed networks.<br> NIPS 2015 Workshop: Networks in the Social and Information Sciences</li> </ul> <p>The method first maps the MAP problem on the Potts model as a hinge-loss minimization problem (see the paper for details). To run the code you need to install psl (included here) and if you want to additionally compare with other inference methods, such as max prod belief propagation or junction tree, you need to install the libDAI library (also included here)</p> <p>The directory europeanCongressData/ (~500 Mb) contains the votings of the EU parlament, including 300 votings events from the actual term, from May 2014 to June 2015, obtained from http://www.votewatch.eu/</p> <ul> <li>data/ : json files with the european votes</li> <li>network.net : signed network built from the votes</li> <li>political_parties.txt : "ground truth" party</li> <li>community_results/ : results for different number of communities and initial vertices</li> <li>dataComputations.py : used to build the signed network</li> <li>dataProcessing.py : used to build the signed network</li> </ul> <p>We would appreciate if you cite the paper after using the data or the code.</p> <p>DEPENDENCIES</p> <p>The code has been tested in Linux Mint 18.1 Serena and Ubuntu 14.04</p> <p>- For PSL library, you need to have<br>     java 1.8<br>     you may need to export JAVAHOME='/usr/lib/jvm/YOURJAVA1.8FOLDER'<br>     maven 3.x</p> <p>- For libDAI you will need:<br>     make doxygen graphviz libboost-dev libboost-graph-dev libboost-program-options-dev libboost-test-dev libgmp-dev cimg-dev libgmp-dev</p> <p>CODE TO RUN THE FOLLOWING EXPERIMENTS:</p> <p>Compare the performance in terms of structural balance of max prod bp and our method against an exact inference method (junction tree), with different number of communities</p> <p>INSTALL</p> <p>To install the experiments you have to follow the next steps:</p> <p>1 Build the libdai library by doing: make -B on the folder (libdai)</p> <p>2 Generate the class path of the groovy project:<br> mvn clean install<br> mvn dependency:build-classpath-Dmdep.outputFile=classpath.out</p> <p>on the psl root folder (You need to have java 1.8 and maven 3.x installed)</p> <p>3 Grant exec permissions to the run.sh script</p> <p>Options</p> <p>The main python file to run the experiments is</p> <p>evaluatebalanceon_sn.py.</p> <p>It accepts the following parameters:</p> <p>1 (Int) Nodes of the graph. In order to run the junction tree we recommend to set this paremeter to 150 or less<br> 2 (Int) The number of underlying communities<br> 3 (Float) The maximum amount of unbalance for the experiments. We recommend 0.45<br> 4 (Bool) Whether to use an heuristic to find the initial node for each community or to use directly random nodes from the ground truth communities. This heuristic looks alternatively for the nodes with highest negative degree and highest positive degree. For the case when the number of communities is equal to 2 (Ising Model), the heuristic is used by default.</p> <p>An example of execution would be:</p> <p>python evaluate_balance_on_sn.py 120 3 0.45 True True</p> <p>The results of the experiments are save in the folder results/<br> Scripts</p> <p>The main script of the hinge-loss method can be found in the folder psl/psl-example/src/main/java/edu/umd/cs/example/PottsCommunities.groovy</p> <p>Authors:</p> <p>Guillermo Santamaria &amp; Vicenc Gomez<br> Mar 5, 2017</p> <p>For further questions, please contact vicen.gomez@upf.edu</p>

opencc-by-4.0Dec 2014View details →
zenodo44/100

Video, image, and supplemental files linked in Burge et al. (2023) "Depredation by Bottlenose Dolphins Tursiops truncatus from Antillean Z-traps at Discovery Bay, Jamaica"

<p>Video,&nbsp;image, and supplementary text files linked in Burge et al. (2023), Caribbean Naturalist, 95: 1–25.</p><p><strong>Depredation by Bottlenose Dolphins </strong><i><strong>Tursiops truncatus</strong></i><strong> from Antillean Z-traps at Discovery Bay, Jamaica</strong></p><p>All video and image files referred to in the main text, figures, and tables are available from this repository. See Table 1 and Table S1 for additional details.</p><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Dataset for manuscript titled "Exploring SureChEMBL from a drug discovery perspective".

<p>This is the data directory for running the code available on the GitHub repository for the manuscript titled "Exploring SureChEMBL from a drug discovery perspective<strong></strong>". The GitHub repository is available at <a href="https://github.com/Fraunhofer-ITMP/patent-clinical-candidate-characteristics">https://github.com/Fraunhofer-ITMP/patent-clinical-candidate-characteristics.</a></p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Discovery of South African plant-based biomarkers as potential flagships for SARS- CoV-2 receptor

<h1><span>Table S1: </span><span>Distinguished metabolites in the extracts of <em><span>Artemisia annua </span></em><span>and <em>Artemisia afra </em></span>using UPLC-MS/MS set in positive ionization mode.</span></h1> <h1><span>Table S2: Identified and docked compound-based biomarkers from&nbsp;<em><span>Artemisia annua </span></em><span>and <em>Artemisia afra </em></span>(ESI+ scan).</span></h1>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Dataset - SciVisContest - Materials Discovery Challenge - version 2025

<p>This the updated version of the dataset that is made available for the SciVisContest Materials Discovery Challenge.</p> <p>The data provided for this challenge was generated for the specific use case of developing a new Al-based alloy suitable for additive manufacturing by blending different available aluminum metal scrap such as automotive Al-Si piston alloys and other alloys from different sectors. Different alloy designs were initially generated based on mixing ratios between available scrap alloys. The CALPHAD method was used to perform equilibrium and non-equilibrium calculations to predict relevant thermo-physical and mechanical variables such as the content of volatile elements like Mg and Zn, phase formation and their fractions, solidification intervals, thermo-physical parameters, yield strength and hot crack sensitivity for each alloy composition.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Relocated Seismicity Catalogs on the Discovery Transform Fault, 4S on the East Pacific Rise

<p>Two relocated earthquake catalogs are provided for the Discovery Transform Fault located at 4&ordm;S on the East Pacific Rise. There is a microseismicity catalog representing one year of activity recorded during a 2008 ocean bottom seismometer deployment, which includes 12,635 events with local magnitudes, M<sub>L</sub>, between 0 and 4.1. The second catalog includes 24 years (1 January 1990 - 1 April 2013) of earthquakes obtained from the global Centroid Moment Tensor (CMT) catalog, a total of 15 events, with seismic moment magnitudes, M<sub>W</sub>, between 5.4 and 6.0.</p> <p>Microseismicity was relocated using the HypoDD relocation algorithm (Waldhauser, 2001), while the CMT events were relocated using a teleseismic surface-wave cross-correlation technique (McGuire, 2008). The 15 CMT events all relocated into one of five distinct rupture patches on the Discovery Transform Fault. In general, microseismicity was found to be reduced within these large, repeating rupture patches.</p> <p>A more detailed description of the methodology used to relocate both catalogs, as well as a discussion on the correlation between seismic behavior and fault structure on the Discovery Transform Fault is provided in:</p> <p>Wolfson-Schwehr, M., Boettcher, M. S., McGuire, J. J., &amp; Collins, J. A. (2014). The relationship between seismicity and fault structure on the Discovery transform fault, East Pacific Rise. <em>Geochemistry, Geophysics, Geosystems,&nbsp;</em>15(9), 3698&ndash;3712. <a href="https://doi.org/10.1002/2014GC005445">https://doi.org/10.1002/2014GC005445</a></p> <p>Seismic Catalogs:</p> <ul> <li>Discovery_CMT_relocated_seismicity_1990_2013.csv</li> <li>Discovery_relocated_microseismicity_2008.csv</li> </ul> <p>Additional References:</p> <p>1.&nbsp;McGuire, J. J. (2008). Seismic cycles and earthquake predictability on East Pacific Rise transform faults.&nbsp;<em>Bulletin of the Seismological Society of America</em>,&nbsp;98(3), 1067-1084.&nbsp;<a href="https://www.whoi.edu/cms/files/McGuire_BSSA_2008_48643.pdf">https://www.whoi.edu/cms/files/McGuire_BSSA_2008_48643.pdf</a></p> <p>2. Waldhauser, F. (2001). hypoDD--A program to compute double-difference hypocenter locations.&nbsp;<br> &nbsp; &nbsp; <a href="https://academiccommons.columbia.edu/doi/10.7916/D8SN072H">https://academiccommons.columbia.edu/doi/10.7916/D8SN072H</a></p>

opencc-by-4.0May 2022View details →
zenodo44/100

Supporting data for: Low-cost anti-mycobacterial drug discovery using engineered E. coli

<p>Supporting data for: Low-cost anti-mycobacterial drug discovery using engineered E. col</p> <p>This dataset pertains to our work developing TESEC Mtb ALR, a genetically engineered strain of E. coli expressing the enzyme ALR derived from Mtb. We used the TESEC Mtb ALR strain in a high-throughput drug screen and identified benazepril as targeted inhibitor of the ALR enzyme. We then performed additional experiments to characterize the activity of benazepril against E. coli, Mtb and purified enzymes. Finally, we tested the extensibility of the platform by constructing and screening against similar strains for additional targets: Asd, CysH, DapB, and TrpD.</p> <p>These files include growth measurements, biochemical assays and other forms of biological data. They are packaged together with scripts used to analyze the data and present them in figures. Our goal in creating this archive was to present our complete analysis pipeline in the spirit of open science. It is not intended to be a stand-alone resource. Consult the associated manuscript for protocols, units of measurement and other essential technical context.</p> <p>Our scripts were written for Python 3.8. The raw data is presented as human-readable .csv files intended to be imported as Pandas DataFrames. Some data is also packaged as Python dictionaries saved with the Pickle package.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

False discovery rate calculations for genome-wide association study of reproductive fitness in Drosophila melanogaster (Sussex LHM sample)

<p>R code and results of applying false discovery (FDR) rate calculations to establish statistical signficance in a genome-wide association study of reproductive fitness in Drosophila melanogaster. Phenotype values were generated on hemiclone female and male lines from an outbred, laboratory adapted population. Thus, GWAS were previously performed seperately on the phenotype values for each sex, and also using a bivariate GWAS implemented in the R package multiPhen.</p> <p>FDR calculations were performed using the R package 'fdrtool' on all SNPs, and on LD-independent SNPs, the latter of which was used to determine p-value thresholds for genome-wide significance when all SNPs were considered.</p> <p>This version differs from the original in that: i) Some gene positions/names have been reassigned for accuaracy in the input data. ii) A file containing the p-value thresholds corresponding to an FDR of 0.1 has been added. The 95% credible intervals for each SNP association have been added to the results data files.</p>

opencc-by-4.0Sep 2017View details →
zenodo44/100

Examining LGBTQ+-related Concepts in the Semantic Web: Link Discovery, Concept Drift, Ambiguity, and Multilingual Information Reuse

<div> <h1>Examining LGBTQ+-related Concepts in the Semantic Web</h1> </div> <div> <h2>Introduction</h2> </div> <p>Welcome to the project. We study the links between LGBTQ+ ontologies and structured vocabularies. More specifically, we focus on GSSO, Homosaurus, QLIT, and Wikidata. The code is free for use with the license GPL 3,0. You can resue/extend the code for free as long as you give credits to us in your publication/data. Citation information will be added after the corresponding paper gets accepted. The paper is under submission and will be included soon.&nbsp;</p> <p>If you would like to extend this work, you may want to contact the experts in the acknowledgement before releasing your data/code about legal and ethical issues. The DOI for this version is 10.5281/zenodo.12684870. The latest code can be found at https://github.com/Multilingual-LGBTQIA-Vocabularies/Examing_LGBTQ_Concepts.&nbsp;</p> <p>To reproduce the results or extend our work, you need to take the following steps.</p> <div> <h2>Step 1: Preparing the data</h2> </div> <p>In this project, the following datasets were used:</p> <ul> <li>QLIT: version 1.0</li> <li>Homosaurus: version 3.5 and version 2.3</li> <li>Wikidata: retrieved from the SPARQL Endpoint (<a href="https://query.wikidata.org/sparql" rel="nofollow">https://query.wikidata.org/sparql</a>) and processed between 5th May and 8th May, 2024.</li> <li>GSSO: we used gsso.owl (version 2.0.10) obtained from its Github (<a href="https://github.com/Superraptor/GSSO">https://github.com/Superraptor/GSSO</a>).</li> <li>LCSH was obtained from the official website:&nbsp;<a href="https://id.loc.gov/authorities/subjects.html" rel="nofollow">https://id.loc.gov/authorities/subjects.html</a>&nbsp;on 9th May, 2024. The LCSH data was converted to its HDT format.</li> </ul> <p>Please put the corresponding files in the following folders (and change its names where necessary) to make sure that the Python scripts can find your code.</p> <ul> <li>./data/GSSO/gsso.owl</li> <li>./data/Homosaurus/v2.ttl and ./data/Homosaurus/v3.ttl</li> <li>./data/LCSH/lcsh.hdt (we used its HDT format for fast query and analysis). The original file is also attached: subjects.skosrdf.nt.</li> <li>./data/QLIT/Qlit-v1.ttl</li> </ul> <p>The case of Wikidata is more complicated. The following scripts were used for the retrival of data. These scripts are all in the folder ./data/wikidata/</p> <ul> <li>We used the Wikidata SPARQL endpoint:&nbsp;<a href="https://query.wikidata.org/" rel="nofollow">https://query.wikidata.org/</a></li> </ul> <p>The following relations from Wikidata were used while extracting triples.</p> <ul> <li>Wikidata - GSSO:&nbsp;<a href="http://www.wikidata.org/prop/direct/P9827" rel="nofollow">http://www.wikidata.org/prop/direct/P9827</a></li> <li>Wikidata - Homosaurus 2:&nbsp;<a href="http://www.wikidata.org/prop/direct/P6417" rel="nofollow">http://www.wikidata.org/prop/direct/P6417</a></li> <li>Wikidata - Homosaurus 3:&nbsp;<a href="http://www.wikidata.org/prop/direct/P10192" rel="nofollow">http://www.wikidata.org/prop/direct/P10192</a></li> <li>Wikidata - LCSH:&nbsp;<a href="http://www.wikidata.org/prop/direct/P244" rel="nofollow">http://www.wikidata.org/prop/direct/P244</a></li> </ul> <p>The generated files are:</p> <ul> <li>'wikidata-homosaurus-v2-links.nt'</li> <li>'wikidata-homosaurus-v3-links.nt'</li> <li>'wikidata-gsso-links.nt'</li> <li>'wikidata-qlit-links.nt'</li> <li>'wikidata-lcsh-links-all.nt'</li> </ul> <p>Please note that the case of Wikdiata-LCSH is more complicated: there are so many links that are nothing to do with the entities in our scope. We restrict it to only entities in the scope of this paper. See below for more details.</p> <p>You can find all the scripts in the corresponding folder in the data folder.</p> <p>All the SPARQL queries used can be found in the folder ./SPARQL/</p> <p>Note! For GSSO, the following two mistakes were corrected while preprocessing:</p> <ul> <li><a href="https://www.wikidata.org/wiki/Q1823134" rel="nofollow">https://www.wikidata.org/wiki/Q1823134</a>&nbsp;should not be used as a relation. We have replaced it with&nbsp;<a href="http://www.wikidata.org/prop/direct/P244" rel="nofollow">http://www.wikidata.org/prop/direct/P244</a>.</li> <li>Instead of referring to the page, we refer to the entity. We use&nbsp;<a href="http://www.wikidata.org/entity/" rel="nofollow">http://www.wikidata.org/entity/</a>* instead of&nbsp;<a href="https://www.wikidata.org/wiki/" rel="nofollow">https://www.wikidata.org/wiki/</a>*</li> </ul> <p>The redirection test was conducted on 30th April, 2024, between 6PM and 8PM. The files can be found in the folder of ./data/Homosaurus/redirect/.</p> <div> <h2>Integrating the data</h2> </div> <p>In the folder ./integrated_data/, you can find all the scripts related to the integrated data. Unfortunately, due to the CC-BY-NC-ND license of GSSO and Homosaurus, the integrated data will not be made available. But you can generate it with the instructions above and by using the following scripts.</p> <p>The script ./integrated_data/integrate.py takes advantage of the data generated. It first integrates a list of files of links. Then we go through the links between Wikidata and LCSH. Only those that are in the scope of the study are included.</p> <ul> <li>If your steps are correct and using the same version as we did, you should be able to get four files:</li> <li>a) the integrated file as integrated.nt</li> <li>b) the links that are relevant for this study: wikidata-lcsh-links-selected.nt.</li> <li>c) a plot of the distribution of the size of WCCs</li> <li>d) a mapping of entities and their corresponding ID of WCCs.</li> </ul> <div> <h2>Weakly Connected Components</h2> </div> <p>The weakly connected components (WCCs) were computed for the following three purposes:</p> <p>a) Discovering missing links. See the section below for details.</p> <p>b) The WCCs can be used for manual examination. These are entities that form clusters about related concepts. The intuition is that the larger they are, the more likely there is concept drift/change, ambiguity, and mistakes.</p> <p>c) Multilingual information reuse. Smaller WCCs with exactly one entity from each dataset (e.g. Homosaurus and Wikidata) can then be used to suggest labels for the one with fewer labels for some given languages. See below for more details.</p> <p>As mentioned above, the distribution has been plotted. You can find this plot here: ./integrated_data/frequency.png</p> <p>In the folder ./integrated_data/weakly_connected_components/, you can find all the WCCs and their links.</p> <p>Two examples were given in the folder. The largest WCC about sex, gender, fucking, etc. The other is about BDSM and fetish.</p> <div> <h2>Discovering missing and outdated links</h2> </div> <p>Taking advantage of WCCs, we can further find missing and outdated links. The scripts are in the folder ./discover_missing_links.</p> <p>Three examples were given. The first two is about discovering missing links. The last one is about finding outdated links.</p> <ul> <li> <p>The script ./discover_missing_links/discover_H3_LCSH.py and ./discover_missing_links/discover_QLIT_LCSH.py are scripts that outputs links that could be missing in Homosaurus and QLIT respectively. This was computed by looking at the WCCs. If two entities are both involved in the same WCC, there could be a link between them. The csv files in the same folder are the corresponding links found.</p> </li> <li> <p>The script ./discover_missing_links/find_qlit_outdated_links/ is used to discover the outdated links between QLIT and Homosaurus v3. There was only one link found.</p> </li> <li> <p>The 105 potentially missing links were taken for further review by Swedish-speaking experts from the QLIT team, which showed that 78 (72.38%) suggested links should be included: 38 (36.19%) can be included using skos:exactMatch and another 38 (36.19%) using skos:closeMatch. 28 (26.67%) suggested links are incorrect. The manual annotation are included in the file ./discover_missing_links/Annotated_found_new_links_qlit-lcsh.xlsx.</p> </li> </ul> <div> <h2>Multilingual Information Reuse</h2> </div> <p>You can find two attempts in the folders about the use of GSSO and Wikidata for Homosaurus respectively.</p> <ul> <li>./WCC-based-gsso-multilingual_info_reuse/</li> <li>./WCC-based-wikidata-multilingual_info_reuse/</li> </ul> <p>Additionally, we provide also some code for the reuse of Wikidata multilingual info for QLIT. It's in the folder</p> <ul> <li>./WCC-based-QLIT-info-reuse-from-Wikidata/</li> </ul> <p>They follow very similar steps:</p> <ol> <li> <p>Compute the one-to-one mapping using the WCCs. The script is named compute-one-to-one-mapping.py</p> </li> <li> <p>Extract the multilingual labels from sources. The corresponding file is extract_multilingual_labels_from_one_to_one_mappings.py</p> </li> <li> <p>Provide the extracted multilingual as suggestions for targeting entities. The name of the corresponding files are like "*suggesting-labels.py", where the * is replaced by the actual source/target.</p> </li> </ol> <p>For GSSO, we use the following relations:</p> <ul> <li><a href="http://www.w3.org/2000/01/rdf-schema#label" rel="nofollow">http://www.w3.org/2000/01/rdf-schema#label</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasRelatedSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasRelatedSynonym</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasSynonym</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasExactSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasExactSynonym</a></li> <li><a href="http://purl.org/dc/terms/replaces" rel="nofollow">http://purl.org/dc/terms/replaces</a></li> <li><a href="https://www.wikidata.org/wiki/Property:P5191" rel="nofollow">https://www.wikidata.org/wiki/Property:P5191</a></li> <li><a href="https://www.wikidata.org/wiki/Property:P1813" rel="nofollow">https://www.wikidata.org/wiki/Property:P1813</a></li> <li><a href="https://schema.org/alternateName" rel="nofollow">https://schema.org/alternateName</a></li> <li><a href="http://www.w3.org/2002/07/owl#annotatedTarget" rel="nofollow">http://www.w3.org/2002/07/owl#annotatedTarget</a></li> </ul> <p>Additioinally, we found the relation to be studied in the future:&nbsp;<a href="http://www.geneontology.org/formats/oboInOwl#hasNarrowSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasNarrowSynonym</a></p> <p>For Wikidata, there are only two:</p> <ul> <li><a href="http://www.w3.org/2000/01/rdf-schema#label" rel="nofollow">http://www.w3.org/2000/01/rdf-schema#label</a></li> <li><a href="http://www.w3.org/2004/02/skos/core#altLabel" rel="nofollow">http://www.w3.org/2004/02/skos/core#altLabel</a></li> </ul> <div> <h2>Additional analysis</h2> </div> <p>Additionally, we perform an analysis using only redirection and replacement for GSSO and Homosaurus. The scripts are in the folder ./additional_test_gsso_multilingual_info_reuse. We consider also Homosaurus v2. This additional analysis shows the following:</p> <ul> <li> <p>For the Turkish language, in total there are 103 triples about labels about 23 entities. The average suggested labels per entity is 3.0.</p> </li> <li> <p>For the Spanish language, in total there are 205 triples about labels about 43 entities. The average suggested labels per entity is 2.12.</p> </li> <li> <p>For the French language, in total there are 277 triples about labels about 47 entities. The average suggested labels per entity is 2.19.</p> </li> <li> <p>For the Danish language, in total there are 115 triples about labels about 47 entities. The average suggested labels per entity is 2.70.</p> </li> </ul> <p>Some analysis about the replacement relations of Homosaurus is in the folder ./data/Homosaurus/replace_relations_homosaurus/.</p> <p>Finally, some additional analysis is included in the folder ./analysis_integrated_graph. Currently, there is only one that is about outdated entities in Homosaurus v3. Some more analysis will be added in the future.</p> <div> <h2>Acknowledgement</h2> </div> <p>The authors appreciate the help of the following researchers:</p> <ul> <li>Siska Humlesj&ouml;, QLIT, G&ouml;teborgs Universitet (<a href="mailto:siska.humlesjo@lir.gu.se">siska.humlesjo@lir.gu.se</a>)</li> <li>Olov Kristr&ouml;m, former member of QLIT</li> <li>Jack van der Wel, IHLIA (<a href="mailto:jack@ihlia.nl">jack@ihlia.nl</a>)</li> <li>Clair Kronk, GSSO (<a href="mailto:clair.kronk@mountsinai.org">clair.kronk@mountsinai.org</a>)</li> </ul> <div> <p>If you would like to extend this work, you may want to contact them before releasing your data/code about legal and ethical issues.</p> <h2>Contact</h2> </div> <ul> <li>Shuai Wang, Vrije Universiteit Amsterdam (<a href="mailto:shuai.wang@vu.nl">shuai.wang@vu.nl</a>)</li> <li>Maria Adamidou, Vrije Universiteit Amsterdam (<a href="mailto:m.adamidou@student.vu.nl">m.adamidou@student.vu.nl</a>)</li> </ul> <p>&nbsp;</p> <p>Thank you very much for your interest in our project!</p>

opengpl-3.0-or-laterJul 2024View details →
zenodo44/100

Daatset for testing Feature Discovery

<p>Modified open-source data for testing the automated feature discovery approach published in "<a href="https://github.com/delftdata/autofeat/blob/v2.1/ICDE_FeatureDiscovery.pdf" target="_blank" rel="noopener">AutoFeat: Transitive Feature Discovery over Join Paths</a>" ICDE 2024</p> <p>More details about the datasets <a href="https://github.com/delftdata/autofeat/tree/v2.1?tab=readme-ov-file#datasets" target="_blank" rel="noopener">on Github</a>.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Improving the drug discovery process by using multiple classifier systems

<p>High-quality dataset gathered from ChEMBL version 22 based on UniProt accession P34972. Regarding to activity data potential, duplicates were ignored, no activity or data validity comments were allowed, only data from binding assays and with a pCheMBL value were kept. This led to a dataset composed of 3925 chemical compounds (instances) represented using 2132 features. The first 2048 features epitomize different chemical structures fingerprints (represented using FCFP_6 notation), while the remaining 84 are associated with several physicochemical descriptors (such as Fractional Polar Surface Area, Rotatable Bonds&nbsp;or Molecular Weight). Finally, the set was transformed into a binary classification set where the activity cut-off was defined at a pChEMBL value &gt; 7 and written to a tab-delimited text file. The final set contained 1977 active compounds and 1948 inactive compounds. Table 3 shows the codification of each feature grouped by type.</p>

opencc-by-4.0Jun 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record