Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

76,402,788

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

76,402,788 results

Learn how ShareScore rates datasets ↗
zenodo52/100

2-Dimensional habitat files for 47 representative marine species

<p><strong>2D marine species habitats in NetCDF format on 0.5*0.5 degree global regular grid.</strong></p> <p>Based on Close et al. (2006) and&nbsp;converted from .CSV format.&nbsp;</p> <p>Filename is in format &#39;presence_speciesnumber.nc&#39; where species numbers are listed in the README.txt file (identifier for each species). The README.txt file further contains each species&#39; species_group which is the assigned depth group for each species (1=0-200m epipelagic, 2=200-1000m mesopelagic, 3=sea floor demersal) and species_name which is the&nbsp;Latin name of each species with underscore in between.</p> <p>The species occurs where the variable &#39;presence&#39; equals 1 (in the accompanying paper we assume this to be the 1995-2014 climatological mean distribution).</p> <p>In the NetCDF files, the variable &#39;presence&#39; has as an attribute &#39;species&#39; which contains the Latin species name without underscore.</p>

opencc-by-4.0May 2023View details →
zenodo52/100

ALLINTERACT_RawData1_v3.01

<p>This dataset is part of EC Horizon 2020 project ALLINTERACT Widening and diversifying citizen engagement in science (872396).<br> It contains the raw data obtained from the fieldwork, which consists of: 1) Literature Review, 2) Social Media Analytics, 3) Focus Groups, 4) Survey and 5) Social Media Communicative Observation.<br> 1) Literature Review<br> The objective of the literature review was to address the following topics in gender and education: a) How citizens&rsquo; benefit from scientific research, b) Citizen awareness of the impact of scientific research, c) Awareness-raising initiatives succeeding at engaging citizens in scientific participation, including the Open Access movement and citizen science initiatives, d) Awareness-raising actions that foster the recruitment of new talent in sciences and e) Policies that promote awareness-raising actions and citizen engagement in science.&nbsp;<br> In order to do so, the searches were carried out in the top scientific databases, namely Web of Science (mainly in those journals indexed in Journal Citation Reports) and Scopus. The articles were published between 2010-2021 in journals indexed Q1 or Q2 in JCR or in Q1 journals indexed in Scopus. Relevant reports from EU-funded research projects and official EU documents were also included.<br> We provide one word file with the following information of each topic (a-e) in gender and education.<br> -&nbsp;&nbsp; &nbsp;Keywords used<br> -&nbsp;&nbsp; &nbsp;Criteria of selection<br> -&nbsp;&nbsp; &nbsp;Identified sources<br> -&nbsp;&nbsp; &nbsp;Outcomes<br> -&nbsp;&nbsp; &nbsp;Annexes: Grids with the details of the identified socurces</p> <p>2) Social Media Analytic<br> It is the raw data obtained from social media interactions (Twitter, Facebook, Instagram and Reddit) among citizens about citizen participation in science and research with social impact related to two Sustainable Development Goals: Quality Education and Gender Equality.&nbsp;<br> The data collection followed a twofold strategy 1) Top-Down, in which researchers identified and selected relevant Twitter and Instagram hashtags and Facebook and Reddit pages and 2) Bottom-Up, in which Twitter hashtags were selected based on daily Trending Topics.<br> The data was collected between March 9th and March 16th 2021 and has been obtained, cleaned and anonymized following Allinteract - Social Media Analytics Protocol (Flecha &amp; Pulido, 2021).<br> We provide five Excel files (one for each social network explored). Each file contains the main information of the extracted messages, however the information extracted in each case is slightly different.&nbsp;<br> -Twitter: Tweet ID, Time, Tweet Type, Retweeted By, Number of Retweets, Hashtags<br> -Facebook: Post ID, Video, Type, Likes, Created Time, Updated Time, Comment ID, Comment Likes, Comment Time, Page Likes<br> -Instagram: Likes, comments, date<br> -Reddit: Row ID, sub_id, sub_title, sub_score, sub_date, comment_id, comment_score, comment_date</p> <p>3) Focus Groups<br> This data file contains the pseudonymized transcription of a total of 6 focus groups in gender and 6 in education, which were conducted between October 2021 and February 2022. These focus groups are the pre-test and therefore, the groups are distributed in control group or experimental group. The participants of the gender focus groups were women (including vulnerable women) from a women&rsquo;s group, members of an LGBTQI group and women (including young women) from a women&rsquo;s group. The participants of the education focus groups were parents, teachers and students.<br> We provide a word file with the literal transcriptions of the focus groups in the language in which the focus groups were conducted (English, Spanish or Portuguese).</p> <p>4) Survey<br> This data file contains the anonym answers of the survey conducted with participants from 12 countries, through a CATI/CAWI method. The survey was conducted between November 2021 and February 2022 and consists of 59 questions. The exploitation of this data has been carried out with the SPSS software.&nbsp;<br> We provide an excel file with the 59 questions and the answers of 7507 participants.</p> <p>5) Social Media Communicative Observation<br> The Social Media Communicative Observation aims to explore the effects of introducing scientific pieces of evidence in social media interactions as an initiative to increase participation through awareness. In order to do so, scientific evidence on gender and education were introduced in 10 Facebook groups (5 related to gender and 5 to education), 10 Reddit communities (5 related to gender and 5 to education) and 2 Social Impact Platforms (Sappho and Adhyayana).&nbsp;<br> We provide an excel file with the anonymized interactions among users around the introduced piece of evidence. This Excel file contains the following information: Group of documents, document name, code, start, final, weight, segment, changed by, changed, created, comment, area and percentage (%).</p> <p>6) Focus Group &ndash; Post test</p> <p>This data file contains the pseudonymized transcription of a total of 6 focus groups post test</p> <p>Funding: We acknowledge support of this work by the project &quot;ALLINTERACT Widening and diversifying citizen engagement in science&rdquo; (872396) from the European Commission Horizon 2020 programme.&nbsp;</p> <p>Contact information<br> Ram&oacute;n Flecha (PI): ramon.flecha@ub.edu<br> Marta Soler Gallart (KMC Coordinator): marta.soler@ub.edu<br> Pavel Oveiko (Ethics Chair): pavel.ovseiko@rdm.ox.ac.uk<br> ALLINTERACT Project: allinteract@ub.edu</p> <p>References<br> Flecha, R., &amp; Pulido, C. (2021). Allinteract - Social Media Analytics Protocol is licensed under a Creative Commons Attribution - NonCommercial - ShareAlike 4.0 International License is available in https://archive.org/details/@crea_research</p> <p>How to cite this dataset<br> Soler-Gallart, M. (2021). D1.1.Allinteract Raw Data is licensed under a Creative Commons Attribution - NonCommercial - ShareAlike 4.0 International License</p>

opencc-by-4.0Apr 2021View details →
zenodo52/100

Raw Particle Number Size-Distribution Data of twin-DMPS equipped with two CPCs for nanoparticle detection for SMEAR II station, Hyytiälä, Finland, Spring 2017

<p>Raw size-Distribution data from twin-DMPS system (Aalto et al., 2001), where the nano-DMA (measuring up to 40 nm, short Hauke type DMA) is quipped with two detectors:<br> a TSI 3776 and a modified Airmodus A20 (Kangasluoma et al., 2015)</p> <p>Data acquired during in March-May 2017 at the SMEAR II station in Hyyti&auml;l&auml;, Finland.<br> Data associated with the publication Stolzenburg, Laurila et al. (2023), Atmos. Meas. Techn., &quot;Improved counting statistics of an ultrafine differential mobility particle size spectrometer system&quot;</p> <p>Files DMYYDDMM_A20.Dat contain the raw DMPS data, with YYMMDD indicating the day of the measurement.<br> Data are provided alternating between data acquired with the nano-DMA and with the long-DMA, on a scan by scan basis.<br> First line of each scan cycle (for both DMAs) always indicates the start and end times of the voltage scan.<br> Second line gives the parameters related to the DMPS as given below:<br> (sheath flow in [l per min], aerosol flow in [l per min], DMA inner electrode diameter in [m], DMA outer electrode diameter in [m], DMA classification length in [m], other parameters)<br> Following lines give<br> (for long-DMA): set voltage at DMA [in V], concentration measured by TSI3772 in [per cm3]<br> (for nano_DMA): et voltage at DMA [in V], concentration measured by TSI 3776 in [per cm3], concentration measured by mod. Airmodus A20 in [per cm3]</p> <p>File dmps_data_format_specifier.text gives a conversion from voltage to diameter and indicates the measurement time at each voltage during the stepping of the DMPS.<br> Needs to be used to convert measured concentrations in counts per set-interval.</p> <p>Files GR_J_overview.xlsx gives size-distribution derived quantities during that campaign.<br> Header defines Date, Growth Rate and Formation Rate measured at different sizes [in nm] and by the two different CPCs connected to the nano-DMA.<br> Growth rates in [nm per h], formation rate in [per cm3 per s].</p> <p>Other data related to the campaign can be obtained from the corresponding author upon reasonable request.<br> juha.kangasluoma@helsinki.fi</p> <p>References:</p> <p>Stolzenburg, Laurila et al. &quot;Improved counting statistics of an ultrafine differential mobility particle size spectrometer system&quot;,<br> Atmos. Meas. Techn., in press, 2023</p> <p>Aalto et al., &quot;Physical characterization of aerosol particles during nucleation events&quot;,<br> Tellus B, vol. 53, pp. 344-358, 2001</p> <p>Kangasluoma et al., &quot;Sub-3 nm Particle Detection with Commercial TSI 3772 and Airmodus A20 Fine Condensation Particle Counters&quot;,<br> Aerosol Sci. Techn., vol. 49, pp. 674-681, 2015</p>

opencc-by-4.0May 2023View details →
zenodo52/100

Resources from: Disparate patterns of genetic divergence in three widespread corals across a pan-pacific environmental gradient highlights species-specific adaptation trajectories

<p>The following files are contained in this repository:</p> <p><br> README.Hume_et_al_2022.zenodov4.txt - This document.</p> <p>scripts.Hume_et_al_2022.zenodov4.pdf - Contains the scripts, or locations of the scripts, used to conduct the data analyses detailed in the associated manuscript.</p> <p>acknowledgements_local_authorities.Hume_et_al_2022.zenodov1.pdf - Acknowledgements of local authorities for the collection of samples used in the associated study.</p> <p>TaraPacific_SST_timeseries_mean_productsV2mai2021.Hume_et_al_2022.zenodov1.csv - The historical temperature data set used for the RDA, Mantel tests and gradient Forest analysis.</p> <p>Pocillopora_meandrina_v3_11Islands.raw.Hume_et_al_2022.zenodov2.vcf.genozip - The Pocillopora SNPs referred to as &#39;raw&#39; in the Methods of the associated manuscript. Compressed using genozip (https://genozip.readthedocs.io/index.html).</p> <p>Pocillopora_meandrina_v3_11Islands.raw.Hume_et_al_2022.zenodov2.vcf.genozip.md5 - md5 of the the Pocillopora raw SNPs.</p> <p>Pocillopora_meandrina_v3_11Islands_maf05_minQ30_biallelic_nomiss.linked.Hume_et_al_2022.zenodov2.vcf.gz - The Pocillopora SNPs referred to as &#39;linked&#39; in the Methods of the associated manuscript.</p> <p>Pocillopora_meandrina_v3_11Islands_maf05_minQ30_biallelic_nomiss.linked.Hume_et_al_2022.zenodov2.vcf.gz.md5 - md5 of the the Pocillopora linked SNPs.</p> <p>Pocillopora_meandrina_v3_11Islands_maf05_minQ30_biallelic_nomiss_LD02.unlinked.Hume_et_al_2022.zenodov2.vcf.gz - The Pocillopora SNPs referred to as &#39;unlinked&#39; in the Methods of the associated manuscript.</p> <p>Pocillopora_meandrina_v3_11Islands_maf05_minQ30_biallelic_nomiss_LD02.unlinked.Hume_et_al_2022.zenodov2.vcf.gz.md5 - md5 of the the Pocillopora unlinked SNPs.</p> <p>Porites_lobata_v3_11Islands.raw.Hume_et_al_2022.zenodov2.vcf.genozip - The Pocillopora SNPs referred to as &#39;raw&#39; in the Methods of the associated manuscript. Compressed using genozip (https://genozip.readthedocs.io/index.html).</p> <p>Porites_lobata_v3_11Islands.raw.Hume_et_al_2022.zenodov2.vcf.genozip.md5 - md5 of the the Pocillopora raw SNPs.</p> <p>Porites_lobata_v3_11Islands_maf05_minQ30_biallelic_nomiss.linked.Hume_et_al_2022.zenodov2.vcf.gz - The Pocillopora SNPs referred to as &#39;linked&#39; in the Methods of the associated manuscript.</p> <p>Porites_lobata_v3_11Islands_maf05_minQ30_biallelic_nomiss.linked.Hume_et_al_2022.zenodov2.vcf.gz.md5 - md5 of the the Pocillopora linked SNPs.</p> <p>Porites_lobata_v3_11Islands_maf05_minQ30_biallelic_nomiss_LD02.unlinked.Hume_et_al_2022.zenodov2.vcf.gz - The Pocillopora SNPs referred to as &#39;unlinked&#39; in the Methods of the associated manuscript.</p> <p>Porites_lobata_v3_11Islands_maf05_minQ30_biallelic_nomiss_LD02.unlinked.Hume_et_al_2022.zenodov2.vcf.gz.md5 - md5 of the the Pocillopora unlinked SNPs.</p> <p>PANAMA2021.raw.Hume_et_al_2022.zenodov2.vcf.gz - The Millepora SNPs referred to as &#39;raw&#39; in the Methods of the associated manuscript.</p> <p>PANAMA2021.raw.Hume_et_al_2022.zenodov2.vcf.gz.md5 - md5 of the the Millepora raw SNPs.</p> <p>Millepora_REF_orthologue_genes.Hume_et_al_2022.zenodov2.csv - The Millepora gene list referred to as &#39;target genes&#39; in the Methods of the associated manuscript.</p> <p>Mil_transcriptom.Hume_et_al_2022.zenodov2.fa.gz - The Millepora de novo assembled transcriptome.</p> <p>Mil_transcriptom.Hume_et_al_2022.zenodov2.fa.gz.md5 - md5 of the Millepora de novo assembled transcriptome.</p> <p>&nbsp;</p> <p>mtORF Phylogeny</p> <p>TP-Johnston_mtORF-Pocillo.fa = all sequences</p> <p>TP-Johnston_mtORF-Pocillo.mafft.fa = mafft alignment</p> <p>TP-Johnston_mtORF-Pocillo.mafft.ML.nwk = ML tree newick</p> <p>&nbsp;</p> <p>Hellberg genotype network Porites</p> <p>TP-Hellberg_MM32-Porites.nex = all aligned sequences for this locus with indels encoded</p> <p>TP-Hellberg_MM100-Porites.nex = all aligned sequences for this locus with indels encoded</p> <p>TP-Hellberg_ATPaseB.nex = all aligned sequences for this locus with indels encoded,</p> <p>TP-Hellberg_POFAD.nex = POFAD multilocus genotypic distance,</p> <p>TP-Hellberg_Splitstree.nex= Multilocus genotype network in nexus format</p> <p><br> Gradient Forest Analysis</p> <p>Poc_abund.csv - Pocillopora SSH Occurrences per Site er Island</p> <p>Por_abund.csv - Porites SSH Occurrences per Site er Island</p> <p>mean_depth_por.csv - per site per island mean depth among Porites colonies</p> <p>mean_depth_poc.csv - per site per island mean depth among Pocillopora colonies</p>

opencc-by-4.0Oct 2022View details →
zenodo52/100

Data set discussed in "Beyond Fortune 500: Women in a Global Network of Directors"

<p>Bipartite graph of directors and companies. Generated from information on the Financial Times website (<a href="https://markets.ft.com/data/equities/results">https://markets.ft.com/data/equities/results</a>), retrieved on 17 September 2016.</p> <p>Blank fields are used for missing data.</p> <p><strong>comp_nodes.csv:</strong></p> <ul> <li>id: unique identifier</li> <li>ft_country: name of the country</li> <li>ft_sector: segment of the economy in which a company operates</li> <li>ft_industry: specific business (i.e., subset of sector) in which a company operates</li> <li>ft_employees_num: number of company&#39;s employees. &quot;NA&quot; if the vertex represents a person or if the company&#39;s number of employees is unknown.</li> </ul> <p><strong>comp_people_edges.csv:</strong></p> <ul> <li>person_id:</li> <li>comp_id: company identifier. It matches the identifier in comp_nodes.csv</li> </ul> <p><strong>people_one_mode_edges.csv:</strong></p> <p>Edges in the one-mode projection, in which two directors are connected if and only if they sit together on at least one board. Numbers correspond to the identifiers in unique_people_nodes.csv.</p> <p><strong>unique_people_nodes.csv:</strong></p> <ul> <li>ID: unique identifier</li> <li>age: years of age</li> <li>gender_base: &quot;Male&quot; or &quot;Female&quot;</li> </ul>

opencc-by-4.0Nov 2019View details →
zenodo52/100

Wearable data and self reported fatigue scores from a remote observational study in Sjogren's disease, SLE and healthy participants

<p>Fatigue is a subjective, complex, and multi-faceted phenomenon, commonly&nbsp;experienced as tiredness. However, pathological fatigue is a major debilitating symptom&nbsp;associated with overwhelming feelings of physical and mental exhaustion.&nbsp;To date,&nbsp;there is no consensus about reliable quantitative assessments of fatigue.</p> <p>We collected observational data for a period of one month from 296 participants (healthy volunteers, Sjogren&rsquo;s Syndrome, and Systemic Lupus Erythematosus patients) in the United States. Data comprised continuous multimodal digital data from Fitbit, including heart rate, physical activity, and sleep daily features, and app-based daily and weekly questions (e.g., pain, mood, general physical activity, and fatigue). When matching both sensor data and PROs, and excluding missing data, the dataset contains data from 183 subjects and 3950 recording days.</p> <p>The analysis of the association of digital data to self-reported fatigue was published at <em><strong>Rao C., et. al. (2023), Association of digital measures and&nbsp;self-reported fatigue: a remote observational&nbsp;study in healthy participants and participants&nbsp;with chronic inflammatory rheumatic disease, Frontiers in Digital Health</strong></em>.</p> <p>Demographics, digital parameters, and other information on this dataset can be found in the aforementioned manuscript and related supplementary material. Details on the data files can be found under README.txt.</p>

opencc-by-4.0Dec 2022View details →
zenodo52/100

iPlacenta: hIPSC placenta-on-a-chip RNAseq data from 3D vs 2D, day 0 vs day 4 differentiation

<p>RNAseq data from hIPSC dervived trophoblasts seeded in 3D (OrganoPlate) or 2D surface at day 0 or day 4 differentiation.&nbsp;</p> <p>Description of file names found below</p> <table> <tbody> <tr> <td> <p><strong>SampleID/File name</strong></p> </td> <td> <p><strong>Condition- Differentiation day</strong></p> </td> </tr> <tr> <td> <p>iPSC-THB-2D-D0-1</p> </td> <td> <p>2D-Day0</p> </td> </tr> <tr> <td> <p>iPSC-THB-2D-D0-2</p> </td> <td> <p>2D-Day0</p> </td> </tr> <tr> <td> <p>iPSC-THB-2D-D0-3</p> </td> <td> <p>2D-Day0</p> </td> </tr> <tr> <td> <p>&nbsp;</p> </td> <td> <p>&nbsp;</p> </td> </tr> <tr> <td> <p>iPSC-THB-2D-D4-4</p> </td> <td> <p>2D-Day4</p> </td> </tr> <tr> <td> <p>iPSC-THB-2D-D4-5</p> </td> <td> <p>2D-Day4</p> </td> </tr> <tr> <td> <p>iPSC-THB-2D-D4-6</p> </td> <td> <p>2D-Day4</p> </td> </tr> <tr> <td> <p>&nbsp;</p> </td> <td> <p>&nbsp;</p> </td> </tr> <tr> <td> <p>iPSC-THB-3D-D0-7</p> </td> <td> <p>3D-Day0</p> </td> </tr> <tr> <td> <p>iPSC-THB-3D-D0-8</p> </td> <td> <p>3D-Day0</p> </td> </tr> <tr> <td> <p>iPSC-THB-3D-D0-9</p> </td> <td> <p>3D-Day0</p> </td> </tr> <tr> <td> <p>&nbsp;</p> </td> <td> <p>&nbsp;</p> </td> </tr> <tr> <td> <p>iPSC-THB-3D-D4-10</p> </td> <td> <p>3D-Day4</p> </td> </tr> <tr> <td> <p>iPSC-THB-3D-D4-11</p> </td> <td> <p>3D-Day4</p> </td> </tr> <tr> <td> <p>iPSC-THB-3D-D4-12</p> </td> <td> <p>3D-Day4</p> </td> </tr> <tr> <td> <p>iPSC-THB-3D-D4-13</p> </td> <td> <p>3D-Day4</p> </td> </tr> </tbody> </table>

opencc-by-4.0Dec 2022View details →
zenodo52/100

Single-pulsar search for eccentric SMBHBs using NANOGrav 12.5-year data of PSR J1909--3744: Posterior samples

<p>This repository contains posterior samples for a Bayesian single-pulsar search for nanohertz gravitational waves originating from eccentric supermassive binaries, done using the NANOGrav 12.5-year dataset for PSR J1909-3744. The analysis is presented in Susobhanan 2023 [https://arxiv.org/abs/2210.11454].</p>

opencc-by-4.0Jun 2023View details →
zenodo52/100

Niedertiefenbach: neolithic collective burial

<p>Spatialite database (SQLite) with data of the Neolithic collective grave Niedertiefenbach in Hesse (Germany). Data collected from the published images and copies of the corresponding lithographs in the archive (Wurm et al., Fundberichte aus Hessen Bd. 3, 1963, 56-78). Data collection from CRC 1266: &quot;Scales of Transformation - Human-Environmental Interaction in Prehistoric and Archaic Societies&quot;. Project: &quot;Regional and Local Patterns of 3rd Millennium Transformations of Social and Economic Practices in the Central German Mountain Range (D2)&quot;. Funded by the DFG. DFG project number: 2901391021</p>

opencc-by-4.0Jun 2023View details →
zenodo52/100

Monte da Contenda geomagnetic prospection 2018

<p>Monte da Contenda geomagnetic prospection</p> <p>Report on the archaeological prospection of Monte da Contenda, Nossa Senhora da Gra&ccedil;a dos Degolados, Concelho Campo Maior, Alto Alentejo, Portugal.<br> Geophysical prospection of the CRC 1266 &quot;Scales of Transformation&quot;, Subproject F1: Climate Constraints of Western Mediterranean Socio-environmental Transformation and Potential Implications for Central Europe (https://www.sfb1266.uni-kiel.de). &nbsp;<br> The prospection took place in context of the &ldquo;project from Antonio Valera at the Patrimino Cultural Portugal&rdquo;. The prospection was carried out by the collaborating partners of Kiel University, Department of Pre- and Protohistory. The scientific question was to define the wider extend of the site documented by a first geophysical prospection and further ditches visible in aerial photos (Valera et al., 2014, fig. 5).</p> <p>Ribeiro, A.S.P., Rinne, C., and Valera, A.C., 2019. Geomagnetic investigations at Monte da Contenda, Arronches, Portugal &ndash; Results from the 2018 campaign. Journal of Neolithic Archaeology, 21, 61&ndash;73. https://doi.org/10.12766/jna.2019.3</p>

opencc-by-4.0Jun 2023View details →
zenodo52/100

Megalithgräber im Haldensleber Forst

<p>Digital data from<br> C. Rinne, Die Megalithgr&auml;ber im Haldensleber Forst, Landkreis B&ouml;rde.<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Fr&uuml;he Monumentalit&auml;t und Soziale Differenzierung Bd. 17 (Bonn 2019).<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;https://d-nb.info/1180865057<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;<br> DFG project https://gepris.dfg.de/gepris/projekt/128675135</p> <p>The two databases of the publication:<br> &nbsp;1. one for the monument data, based on the structure of the general database<br> &nbsp;for excavation of the cultural heritage (GDB LDA LSA) and<br> &nbsp;2. one for the geo data of the project.</p> <p>Both databases are provided as SQL statements creating the tables and adding the data.<br> The codepage is UTF-8.<br> The geo data is provided in srid 31468.<br> More comments within the SQL files.</p>

opencc-by-4.0Jun 2023View details →
zenodo52/100

Snow Water Equivalent Dataset for the South Fork of the San Joaquin River (2018/2021) and Senales (2019/2021)

<p>The dataset is related to: Premier, V., Marin, C., Bertoldi, G., Barella, R., Notarnicola, C., &amp; Bruzzone, L. (2022). Exploring the Use of Multi-source High-Resolution Satellite Data for Snow Water Equivalent Reconstruction over Mountainous Catchments.&nbsp;<em>The Cryosphere Discussions</em>, 1-42.</p> <p>It contains three hydrological seasons - from 1st of October 2018 to 30th of September&nbsp;2021 - of snow water equivalent (SWE) for the South Fork of the San Joaquin river in California (USA) and two hydrological seasons -&nbsp;from 1st of October 2019&nbsp;to 30th of September&nbsp;2021 - for the Schnals/Senales basin in South Tyrol (Italy). The product is daily and with a spatial resolution of 25 m. SWE values are in mm. Snow cover area (SCA) can be derived from the same by thresholding pixel containing SWE greater than 0 mm. Further information about the reference system is contained in the attributes of the netcdf files. Please, contact the authors for further questions. Information about the methodology and the input data for producing these time series is contained in the related article.</p>

opencc-by-4.0Jun 2023View details →
zenodo52/100

From the collective to the individual: transformation processes at the transition from the 4th to the 3rd millennium BC in the German low mountain zone

<p>Data collected by Clara Drummer, Kiel 2022.</p> <p>Clara Drummer, Vom Kollektiv zum Individuum: Transformationsprozesse am &Uuml;bergang vom 4. zum 3. Jahrtausend v. Chr. in der Deutschen Mittelgebirgszone. Scales of transformation Bd. 13 (Leiden 2022).https://d-nb.info/1241580332</p> <p>CRC 1266: &quot;Scales of Transformation - Human-Environmental Interaction in Prehistoric and Archaic Societies.&quot;<br> &quot;Regional and Local Patterns of 3rd Millennium Transformations of Social and Economic Practic-es in the Central German Mountain Range (D2)&quot; Deutsche Forschungsgemeinschaft (DFG) - Projektnummer 128675135 https://gepris.dfg.de/gepris/projekt/316739879</p> <p>Data for the analyses of the decisive transformation in the Hessian-Westphalian area from the Wartberg society to the Corded Ware groups. The work discusses above all the social aspects of the change. This includes, on the one hand, a more detailed analysis of burial rituals and, on the other hand, the integration of, for example, available aDNA results into the overall analysis.</p>

opencc-by-4.0Jun 2023View details →
zenodo52/100

Initial Sample of HYPERNETS Hyperspectral Surface Reflectance Measurements for Satellite Validation from the Bare soil at Marquardt, Germany

<p>The HYPERNETS&nbsp;project (www.hypernets.eu) aims to ensure that high-quality in situ measurements are available to support the (VNIR/SWIR) optical Copernicus products. Therefore, it established a new autonomous&nbsp;hyperspectral spectroradiometer (HYPSTAR&reg; - www.hypstar.eu) dedicated to land and water surface reflectance validation&nbsp;with instrument-pointing capabilities.&nbsp;In the prototype phase, the instrument is being deployed at 24 sites covering a range of water and land types and a range of climatic and logistic conditions. This dataset provides the first published data for the ATB HYPERNETS site in Marquardt, Germany [52&deg;27&#39;59.40&quot;N, 12&deg;57&#39;35.16&quot;E] (ATGE). It is a subset of the complete data record, consisting of the measurements which could be used&nbsp;for satellite validation.&nbsp;</p> <p>The provided&nbsp;NetCDF files are the L2A hypernets products with surface reflectances, their associated uncertainties and error-correlation information. The reflectance in the L2A products is&nbsp;the Hemispherical-directional Reflectance Factor (HDRF) defined as HDRF = &pi; L / E where L is the directional upwelling radiance (with the field o, view of 5&nbsp;dgrees), and E is the (hemispherical)&nbsp;downwelling irradiance (i.e. including both direct solar and diffuse sky irradiance). These reflectances have dimensions of wavelength and series, where each series is a set of measurements for a given geometry (combination of viewing zenith and azimuth angle). In addition to variables for&nbsp;wavelength and bandwidth, the files also contain variables that provide for each series the acquisition time, viewing and solar angles, number of valid scans used, and quality flags (typically, no flags are set in the data provided in this dataset).&nbsp;These NetCDF files also contain further relevant metadata as attributes. See&nbsp;https://hypernets-processor.readthedocs.io/ for further info.</p> <p>The HYPSTAR&reg;-XR sensor was installed on 11 Oct 2022 at the top of a 5m mast on an extended 5 m horizontal boom to minimise interruption of the field of view.&nbsp;The boom faces South at the right angle towards bare soil. The mast is located at 52.466778&deg;N, 12.959778&deg;E. Data are collected every 30 minutes between 9:00 and 17:00 (UTC) from different zenith and azimuth angle.</p> <p>The HYPSTAR&reg;-XR (eXtended Range) instruments deployed at each land HYPERNETS site consist of&nbsp;a VNIR and a SWIR sensor and autonomously collect data between 380-1700 nm at various viewing&nbsp;geometries and send it to a central server for quality control and processing. The VNIR sensor spans&nbsp;1330 channels between 380 and 1000 nm with a FWHM of 3 nm, and the SWIR sensor has 220 channels&nbsp;between 1000 and 1700 nm with a FWHM of 10 nm. The hypernets_processor (Goyens et al. 2021; De Vis et al.&nbsp;in prep.)&nbsp;automatically processes all this data into various products, including the&nbsp;L2A surface&nbsp;reflectance product provided here. All products have associated uncertainties (divided into random and systematic uncertainties, including error-correlation information) which were propagated using the CoMet toolkit (www.comet-toolkit.org).&nbsp;</p> <p>To obtain this dataset, we start&nbsp;from the full ATGE data record and omit&nbsp;all the data that do not pass all quality checks performed as part of the hypernets_processor. In addition, an additional screening procedure was also developed to remove outliers and supply the best quality data suitable for satellite validation. To remove the outliers, a sigma-clipping method is used. First, reflectances are extracted in separate 2-hour windows throughout the day (to account for BRDF differences due to different solar positions) for four different wavelengths (500, 900, 1100 and 1600 nm).&nbsp;Outliers in these reflectances are then identified by iteratively calculating the mean reflectance trend&nbsp;with time&nbsp;(by binning the data per maximum of 30 data points), calculating the standard deviation from this trend, and masking any data that is more than three standard deviations away from the trend. This process is repeated on the unmasked data until the standard deviation does not vary by more than 5% between two iterations. The masks for the four different wavelengths&nbsp;are then combined (keeping only measurements for which none of the four wavelengths is an outlier). The reflectances and associated uncertainties for any masked series (i.e. a geometry that is masked either by the sigma-clipping procedure or from the masks of the hypernets_processor) are replaced by NaNs. Any sequence that has more than half of its series masked is removed entirely.&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo52/100

Dataset for: Mudrik, N., & Charles, A. S. (2022). Multi-Lingual DALL-E Storytime. arXiv preprint arXiv:2212.11985.

<p>This dataset represents the comprehensive collection of data generated during the study presented in the paper available at https://arxiv.org/abs/2212.11985.</p> <p>If your research incorporates this data and results in a publication - Please cite both the dataset and the paper.</p>

opencc-by-4.0Jun 2023View details →
zenodo52/100

Datasets for evaluating scalable supervised learning for synthesize-on-demand chemical libraries

<p>This repository contains datasets for the manuscript &quot;Evaluating scalable supervised learning for synthesize-on-demand chemical libraries&quot;:</p> <ul> <li><strong>ams_all_preds.csv.gz</strong>: The AMS dataset predictions when using an RF or baseline model trained on the training dataset. Includes the predicted score and rank from each model for each compound. We started with 8,434,707 AMS compounds and detected that 247,025 were in the LC or MLPCN training data. These were removed from the AMS list, leaving 8,187,682 compounds to score. The compound matching was done on the SMILES that we canonicalized in rdkit.</li> <li><strong>ams_order_results.csv.gz</strong>: Information about the 1,024 compounds purchased from the AMS library. Excludes the 4 AMS compounds that were incompletely dissolved. Includes the chemical feature representation, information from the vendor, RF and baseline model predictions, screening results, and clustering results.</li> <li><strong>baseline_weight.npy</strong>: The saved Similarity Baseline model, which consists of the active compounds in the training data. This model was used to score the AMS library. See the <a href="https://github.com/gitter-lab/pria-ams-enamine">GitHub repository</a>&nbsp;for code to load the model and make predictions on new compounds.</li> <li><strong>cdd_training_data.tar.gz</strong>: The LC1234 and MLPCN PriA-SSB screening data exported from CDD.</li> <li><strong>enamine_costs_clustered_v3_with_nneighbor.csv.gz</strong>: Contains 5,620 Enamine compounds that were selected based on the RF prediction score and availability. This file also contains the Taylor-Butina cluster ID when clustering the training compounds, 1,024 tested AMS compounds, and top-ranked Enamine compounds at a 0.4 threshold. The nearest neighbor compounds in the training and AMS sets are also included along with compound information from Enamine, RF model scores, and chemical feature representations.</li> <li><strong>enamine_dose_response_curve_plots.xlsx</strong>: Images of the dose response curves from all three runs on the 68 Enamine compounds. If a compound was tested multiple times, multiple curves are shown in the same plot. The compound structure images and SMILES are exported from CDD, not generated with RDKit.</li> <li><strong>enamine_dose_response_curves.tsv</strong>: The dose response curve summaries from all three runs on the 68 Enamine compounds. If a compound was tested multiple times, only the highest-quality dose response curve was used.</li> <li><strong>enamine_final_list.csv.gz</strong>: The final 100 filtered compounds from&nbsp;<code>enamine_top_10000.csv.gz</code>. Contains compound information from Enamine as well as RF model scores, chemical feature representations, and clustering results.</li> <li><strong>enamine_PriA-SSB_dose_response_data.tar.gz</strong>: The dose response screening data from all three runs on the 68 Enamine compounds. The 2021-06-16 run was originally screened on 2020-08-24. 2021-06-16 is the date the compound identities were corrected. This run contains two 1,536 well plates.</li> <li><strong>enamine_top_10000.csv.gz</strong>: Top 10,000 predictions from the Enamine REAL dataset using the selected RF model. Contains compound information from Enamine as well as RF model scores, chemical feature representations, and clustering results.</li> <li><strong>master_df.csv.gz</strong>: The output of preprocessing the files in&nbsp;<code>cdd_training_data.tar.gz</code>. Contains 441,900 rows.</li> <li><strong>random_forest_classification_139.pkl</strong>: The saved RF classification model with&nbsp;hyperparameter ID 139. This model was used to score the AMS and Enamine REAL libraries. See the <a href="https://github.com/gitter-lab/pria-ams-enamine">GitHub repository</a> directory for code to load the model and make predictions on new compounds.</li> <li><strong>train_ams_real_cluster.csv.gz</strong>: Contains cluster IDs for Taylor-Butina clustering at a 0.4 threshold applied to the training compounds, 1,024 tested AMS compounds, and top-ranked compounds from Enamine. Includes the chemical features, dataset to which the compound belongs, leader compound for each cluster, and whether the compound is a known hit.</li> <li><strong>training_df_single_fold.csv.gz</strong>: This is all ten folds in&nbsp;<code>training_folds.tar.gz</code>&nbsp;merged for convenience. Contains 427,300 compounds.</li> <li><strong>training_df_single_fold_with_ams_clustering.csv.gz</strong>: Contains cluster IDs for Taylor-Butina clustering applied to the 427,300 training compounds and the 1,024 tested AMS compounds. Different clustering results are shown at the 0.2, 0.3, and 0.4 thresholds. Includes the leader compound for each cluster. Although the training and AMS compounds were clustered jointly, only the training compounds&#39; clusters are shown. The AMS compounds&#39; clusters are in&nbsp;<code>ams_order_results.csv.gz</code>.</li> <li><strong>training_folds.tar.gz</strong>: The LC1234 and MLPCN training data split into ten folds. This dataset with 427,300 compounds was used for cross validation and model selection. This dataset is derived from&nbsp;<code>master_df.csv.gz.</code></li> </ul> <p>If you use&nbsp;these&nbsp;datasets in a publication, please cite:</p> <p>Moayad Alnammi, Shengchao Liu, Spencer S. Ericksen, Gene E. Ananiev, Andrew F. Voter, Song Guo, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter.&nbsp;<a href="https://doi.org/10.1021/acs.jcim.3c00912">Evaluating scalable supervised learning for synthesize-on-demand chemical libraries</a>.&nbsp;<em>Journal of Chemical Information and Modeling</em>&nbsp;2023.</p> <p>See&nbsp;PubChem AID&nbsp;<a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1272365">1272365</a>, AID&nbsp;<a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1918986">1918986</a>,&nbsp;and the associated publications for details about the PriA-SSB screening data. The screening datasets were compiled from three separate sources that should all be cited if the training dataset is used in a publication:</p> <ul> <li>Moayad Alnammi, Shengchao Liu, Spencer S. Ericksen, Gene E. Ananiev, Andrew F. Voter, Song Guo, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter.&nbsp;<a href="https://doi.org/10.1021/acs.jcim.3c00912">Evaluating scalable supervised learning for synthesize-on-demand chemical libraries</a>.&nbsp;<em>Journal of Chemical Information and Modeling</em>&nbsp;2023.</li> <li>Shengchao Liu<sup>+</sup>, Moayad Alnammi<sup>+</sup>, Spencer S. Ericksen, Andrew F. Voter, Gene E. Ananiev, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter.&nbsp;<a href="https://doi.org/10.1021/acs.jcim.8b00363">Practical model selection for prospective virtual screening</a>.&nbsp;<em>Journal of Chemical Information and Modeling</em>&nbsp;2018.</li> <li>Andrew F. Voter<sup>+</sup>, Michael P. Killoran<sup>+</sup>, Gene E. Ananiev, Scott A. Wildman, F. Michael Hoffmann, James L. Keck.&nbsp;<a href="https://doi.org/10.1177/2472555217712001">A high-throughput screening strategy to identify inhibitors of SSB protein&ndash;protein interactions in an academic screening facility</a>.&nbsp;<em>SLAS Discovery</em>&nbsp;2018.</li> </ul> <ul> </ul>

opencc-by-4.0Oct 2021View details →
zenodo52/100

S33 | SOLUTIONSMLOS | Chemicals used for Modelling in SOLUTIONS

<p>This is the collection associated with list S33 SOLUTIONSMLOS on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>S33 | SOLUTIONSMLOS | <strong>Chemicals used for Modelling in SOLUTIONS</strong></p> <p>SOLUTIONSMLOS contains the 6462 chemicals used for modelling in the SOLUTIONS project (<a href="http://www.solutions-project.eu/">www.solutions-project.eu/</a>), provided by Jaroslav Slobodnik (EI).</p> <p>Update 14 Nov 2019: added CSV file. 6 Feb. 2020 CSV with corrected SMILES entries for PubChem import. 6 Nov 2020 structure fix for CAS 111360-16-8 (reported by Leon, PubChem). 17 July 2022: more SMILES fixes, plus one InChIKey change. 18 June 2023: one more SMILES fix (YLMOTKLYENPQLK-VMPITWQZSA-N) in CSV only.</p>

opencc-by-4.0Oct 2018View details →
zenodo52/100

SALLO validation experiment

<p>dataset containing psychophysical raw data and psychometric curves&#39; points of subjective equality (PSE) obtained in the&nbsp;left-right discrimination&nbsp;and in the bisection tasks, repeatedly performed in the visual and in the acoustic domains, with the head turned at 45&deg; left (-45&deg;), center (0) and 45&deg; right (+45). The clean dataset also contains the values of guess rate and lapse rate used to fit each psychometric curve.</p>

opencc-by-4.0Oct 2022View details →
zenodo52/100

Data and code for: Little directional change in the timing of Arctic spring phenology over the past twenty-five years

<p>Data and code accompanying the publication:&nbsp;Little directional change in the timing of Arctic spring phenology over the past twenty-five years.</p> <p>This resource contains 1. R-scripts to calculate yearly phenologies from raw temporally explicit flowerin, arthropod observation and bird nesting data from Zackenberg. The raw data is openly accessible through the Greenland Ecosystem Monitoring database (https://data.g-e-m.dk/), as well as an R-script to carry out most of the analyses presented in the publication. To facilitate the use of the analysis script, pre-produced annual phenologies of focal&nbsp;taxa are included as csv-tables.</p>

opencc-by-4.0Jun 2023View details →
zenodo52/100

Soil moisture sensor network, design, location attributes and soil properties, Hainich, Germany, project AquaDiva

<p>This dataset contains information of the small scale highly resolved soil moisture measurement network that is part of the of the AquaDiva Critical Zone exploratory, Hainich National Park, Germany. The dataset contains information on soil measurement locations, as well as attributes to the location, the design type (random locations vs transects), as well as locations attributes like distance to the next tree and soil properties. Measurement design was first introduced by Metzger et al., (2017), and used in Fischer et al., 2023. See there for more information.</p> <p><strong>References</strong></p> <p>Fischer-Bedtke, C., Metzger, J. C., Demir, G., Wutzler, T., and Hildebrandt, A.: Throughfall spatial patterns translate into spatial patterns of soil moisture dynamics &ndash; empirical evidence, Hydrology and Earth System Sciences, https://doi.org/10.5194/hess-2022-418, 2023.</p> <p>Metzger, J. C., Wutzler, T., Dalla Valle, N., Filipzik, J., Grauer, C., Lehmann, R., Roggenbuck, M., Schelhorn, D., Weckm&uuml;ller, J., K&uuml;sel, K., Totsche, K. U., Trumbore, S., and Hildebrandt, A.: Vegetation impacts soil water content patterns by shaping canopy water fluxes and soil properties, Hydrological Processes, 31, 3783&ndash;3795, https://doi.org/10.1002/hyp.11274, 2017.</p>

opencc-by-4.0Jun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record