Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

591

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

591 results for “Regression”

Learn how ShareScore rates datasets ↗
edi56/100

Synthesized Dataset of Length-Weight Regression Coefficients for Delta Fish

This dataset is a compilation of length-weight regression coefficients for fish species commonly found in the freshwater tidal habitats of the San Francisco Estuary. This effort was born out of the Delta Smelt Resiliency Strategy Aquatic Weed Control Action study, which, in order to calculate fish biomass, needed to calculate individual fish weights from their measured lengths. The Aquatic Weed Control study was supported by Interagency Ecological Program through the Endangered Species Act and is included in the Interagency Ecological Program 2017-2019 workplan. Weight is estimated from length using the exponential function W=a\ L^b. These can be calculated using the linear regression of the log-transformed equation (log⁡(W)=log⁡(a)+b log(L)). This dataset provides the species-specific a and b parameters. Associated publication(s) and relevant metadata information are included. Data was obtained either via database (fishbase.us) or peer-reviewed scientific papers.

openCC0Dec 2025View details →
zenodo52/100

iEEG Data for "Functional Group Bridge for Simultaneous Regression and Support Estimation"

<p>The repository contains analysis scripts and data used in Wang Z, Magnotti J, Beauchamp MS, Li M. Functional Group Bridge for Simultaneous Regression and Support Estimation, 2020. The data contains high-gamma brain responses across 8 subjects from &quot;congruency&quot; audio-visual experiment.</p>

opencc-by-4.0Mar 2022View details →
zenodo52/100

AMOC reconstruction between 1981 and 2016 from hydrographic data using an empirical linear regression model from Worthington, E. L., Moat, B. I., Smeed, D. A., Mecking, J. V., Marsh, R., and McCarthy, G. D.: A 30-year reconstruction of the Atlantic meridional overturning circulation shows no decline, Ocean Sci., 17, 285–299, https://doi.org/10.5194/os-17-285-2021, 2021.

<p>Dataset used to create Figure 8 in Worthington et al., 2021 (https://doi.org/10.5194/os-17-285-2021). Details of the data and methods can be found in the journal article.<br> <br> Worthington, E. L., Moat, B. I., Smeed, D. A., Mecking, J. V., Marsh, R., and McCarthy, G. D.: A 30-year reconstruction of the Atlantic meridional overturning circulation shows no decline, Ocean Sci., 17, 285&ndash;299,&nbsp;<a href="https://doi.org/10.5194/os-17-285-2021">https://doi.org/10.5194/os-17-285-2021</a>, 2021.</p>

opencc-by-4.0Jul 2022View details →
zenodo52/100

Parsimonious Random-Forest-Based Land-Use Regression Model Using Particulate Matter Sensors in Berlin, Germany

<p>The dataset consists of particulate matter pollution concentration, measured in three localities - Hermsdorf, Charlottenburg and Adlershof, in Berlin, Germany.</p> <p><a href="../api/records/10076056/draft/files/pm25_summer_rd_30s.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer_rd_30s.geojson</a> shows the observed PM2.5 concentration in a 30 second interval.</p> <p><a href="../api/records/10076056/draft/files/pm25_summer.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer.geojson</a> shows the concentrations shown is the local concentration (observed concentration - background concentration) in a 30 second interval. The background concentration is calculated as the lowest 5 percentile of the measured concentration for each measurement round.&nbsp;</p> <p><a href="../api/records/10076056/draft/files/PM2.5_lc_max.geojson/content" target="_blank" rel="noopener noreferrer">PM2.5_lc_max.geojson</a> contains the information from&nbsp;<a href="../api/records/10076056/draft/files/pm25_summer.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer.geojson</a> in a 25m resolution. Additionally, it contains the land use information for each coordinate.</p> <p>The original publication providing all necessary background information on study sites, methodology and data processing is the following: Venkatraman Jagatha, J., T. Sauter, C. Schneider (2024): Parsimonious Random-Forest-Based Land-Use Regression Model Using Particulate Matter Sensors in Berlin, Germany. MDPI Sensors, 24(13), 4193, DOI: 10.3390/s24134193. The paper is fully open access and can be downloaded at&nbsp;<a href="https://doi.org/10.3390/s24134193">https://doi.org/10.3390/s24134193</a>.</p> <p>Information on working with geojson file can be found under <a href="https://geojson.readthedocs.io/en/latest/">GeoJSON</a> .</p>

opencc-by-4.0Jun 2024View details →
zenodo52/100

Annual time series of global VIIRS nighttime lights for 2000-2024 at 500-m spatial resolution extrapolated using logistic regression

<p>The <a href="https://eogdata.mines.edu/products/vnl/"><strong>Annual Visible Night Light (VNL) V2</strong></a> (VIIRS) images at 500-m spatial resolution for the period 2012 to 2024 (Elvidge et al., 2021) have been used to extrapolate the values backwards for years 2000&ndash;2011. This was done by fitting a logistic regression (per pixel) and then predicting the values for the previous years (see nightlights_stack_500m.R). After consistent time-series have been produced, I also derived the difference between year 2024 and year 2000 (nightlights.difference_viirs.v21_m_500m_s_2000_2024_go_epsg4326_v20230318.tif): this shows average rate of change for the 25 years period. Use with caution: extrapolation of values can lead to artifacts. For most of the land surface, however, it appears that the growth of night lights follows exponential growth function and hence nights in the past can be represented accurately by fitting decay / logistic regression function.</p> <p>Original values from the Annual VNL V2 product have been converted from 0&ndash;200 to 0&ndash;2000 scale and are available as Cloud-Optimized GeoTIFFs.</p> <p>Principal components (PC1, PC2, PC3, PC4) were derived using SAGA GIS (sums-of-squares-and-cross-products matrix) method. The first PC1 usually matches the long-term mean value, PC2 matches the 1st derivation in values. File "nightlights_dmsp.v10_m_1km_s_19920101_20241231_go_epsg4326_v20251006.tif" contains 33 years 1992 to 2024, but at 1 km resolution.</p> <p>To cite the Annual VNL V2, please use:</p> <ul> <li>Elvidge, C. D., Zhizhin, M., Ghosh, T., Hsu, F. C., &amp; Taneja, J. (2021). <a href="https://doi.org/10.3390/rs13050922">Annual time series of global VIIRS nighttime lights derived from monthly averages: 2012 to 2019</a>. Remote Sensing, 13(5), 922. https://doi.org/10.3390/rs13050922</li> </ul> <p>Historic night light images (1 km resolution) are also available from <a href="https://doi.org/10.6084/m9.figshare.9828827.v10">Figshare</a>:</p> <ul> <li>Li, X., Zhou, Y., Zhao, M., &amp; Zhao, X. (2020). <a href="https://doi.org/10.1038/s41597-020-0510-y">A harmonized global nighttime light dataset 1992&ndash;2018</a>. Scientific data, 7(1), 168. https://doi.org/10.1038/s41597-020-0510-y</li> </ul>

opencc-by-4.0Mar 2023View details →
OpenNeuro48/100

Effects of Phase Regression on High-Resolution Functional MRI of the Primary Visual Cortex

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
zenodo48/100

Genomic evidence for the parallel regression of melatonin synthesis and signaling pathways in placental mammals

<p><strong>Supplementary Material for:</strong></p> <p>Emerling C.A., Springer M.S., Gatesy J., Jones Z., Hamilton D., Xia-Zhu D., Collin M.A.,&nbsp;and Delsuc F. (2021).&nbsp;Genomic evidence for the parallel regression of melatonin synthesis and signaling pathways in placental mammals.<strong><em> Open Research Europe</em></strong> 1:75. doi:10.12688/openreseurope.13795.1.</p> <p>&nbsp;</p> <p><strong>Supplementary File Legends:</strong></p> <p><strong>- Supplementary_Figure_S1.pdf:</strong>&nbsp;<em>AANAT</em> PAML &lsquo;master model&rsquo; showing branch categories, corresponding to &ldquo;Model 1: 24 ratio&rdquo; in Supplementary Table S7.</p> <p><strong>- Supplementary_Figure_S2.pdf:</strong>&nbsp;<em>ASMT</em> PAML &lsquo;master model&rsquo; showing branch categories, corresponding to &ldquo;Model 2: 24 ratio&rdquo; in Supplementary Table S8.</p> <p><strong>- Supplementary_Figure_S3.pdf:</strong>&nbsp;<em>MTNR1A</em> PAML &lsquo;master model&rsquo; showing branch categories, corresponding to &ldquo;Model 1: 27 ratio&rdquo; in Supplementary Table S9.</p> <p><strong>- Supplementary_Figure_S4.pdf:</strong>&nbsp;<em>MTNR1B</em> PAML &lsquo;master model&rsquo; showing branch categories, corresponding to &ldquo;Model 1: 46 ratio&rdquo; in Supplementary Table S10.</p> <p><strong>- Supplementary_Figure_S5.pdf:</strong>&nbsp;RAxML <em>AANAT</em> gene tree. Numbers at nodes correspond to bootstrap support values.</p> <p><strong>- Supplementary_Figure_S6.pdf:&nbsp;</strong>RAxML <em>ASMT</em> gene tree. Numbers at nodes correspond to bootstrap support values.</p> <p><strong>- Supplementary_Figure_S7.pdf:&nbsp;</strong>RAxML <em>MTNR1A</em>+<em>MTNR1B</em>&nbsp;tree. Numbers at nodes correspond to bootstrap support values.</p> <p><strong>- Supplementary_Figure_S8.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>MTNR1A</em> exon 2 in cetaceans. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S9.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>ASMT</em> in spalacids and <em>Fukomys damarensis</em>. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S10.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>MTNR1A</em> in hyracoids and <em>Cyclopes didactylus</em>. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S11.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>MTNR1A</em> in sirenians. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S12.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>AANAT</em> in sirenians and a polymorphic premature stop codon in exon 5 of <em>ASMT</em> in <em>Trichechus manatus</em>. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S13.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>MTNR1A</em> in <em>Condylura cristata</em>. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S14.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>MTNR1A</em> in <em>Phataginus tricuspis</em>. Read Supplementary Table S14 for further details.</p> <p><strong>- Supplementary_Figure_S15.pdf:&nbsp;</strong>PAML <em>AANAT</em> results, Model 1: 24 ratio (see Supplementary Table S7).</p> <p><strong>- Supplementary_Figure_S16.pdf:&nbsp;</strong>PAML <em>ASMT</em> results, Model 2: 24 ratio (see Supplementary Table S8).</p> <p><strong>- Supplementary_Figure_S17.pdf:&nbsp;</strong>PAML <em>MTNR1A</em> results, Model 1: 27 ratio (see Supplementary Table S9).</p> <p><strong>- Supplementary_Figure_S18.pdf:&nbsp;</strong>PAML <em>MTNR1B</em> results, Model 1: 46 ratio (see Supplementary Table S10).</p> <p><strong>- Supplementary_Table_S1.xlsx:&nbsp;</strong>List of species examined in this study and the sources of the genes. Source key: WGS: Sequences derived from NCBI&#39;s Whole Genome Shotgun database, with accession prefix provided; Whole Genome Sequencing of Short Reads: whole genomes were sequenced using short-read technologies. The methodologies&nbsp;varied for the species, and will be or have been published with other projects, so please contact the author(s) for information on the specific methodology and samples used (Xenarthrans, <em>Proteles cristatus</em>, <em>Otocyon megalotis</em>: Fr&eacute;d&eacute;ric Delsuc, e-mail: Frederic.Delsuc@umontpellier.fr; Crocodylians: John Gatesy, e-mail: jgatesy@amnh.org; <em>Dugong dugon</em>: Mark Springer, e-mail: mark.springer@ucr.edu; SRA: sequences derived from NCBI&#39;s Sequence Read Archive; GenBank: sequences derived from NCBI&#39;s nucleotide collection; Bowhead Whale Genome Resource: sequences derived from http://www.bowhead-whale.org; Ensembl: sequences derived from Ensembl genome browser (www.ensembl.org)l; Discovar de novo: sequences derived genomes assembled via Discovar de novo&nbsp; (<a href="https://software.broadinstitute.org/software/discovar/blog/">https://software.broadinstitute.org/software/discovar/blog/</a>). Coverage: indicates coverage of the whole genome (reported in NCBI or other source) or individual genes (derived from short read mapping). Scaffold and contig N50: reported in NCBI or other source.</p> <p><strong>- Supplementary_Table_S2.xlsx:&nbsp;</strong>Accession numbers and functionality of <em>AANAT</em> in species examined. If Accession # indicated as &ldquo;New&rdquo;, sequence generated for this study and can be found in Supplementary Dataset S1. Parentheses after accession number indicates coordinates for sequence on the contig / scaffold. Exon colors code for the following: green = putatively functional; yellow = missing (e.g., negative BLAST results, negative mapping results); pink = one or more inactivating mutations found. Abbreviations for mutations are as follows: del = deletion; ins = insertion; start = start codon mutation; stop = premature stop codon; ? = ambiguity whether the mutation is shared among all members of the clade. Abbreviations in brackets following an inactivating mutation indicate shared inactivating mutation. Key for each abbreviation follows: Bacu =&nbsp;<em>Balaenoptera acutorostrata</em>; BALA = Balaenidae; BALAEN = Balaenopteridae; Bbon =&nbsp;<em>Balaenoptera bonaerensis</em>; CAB =&nbsp;<em>Cabassous</em>; Ccap =&nbsp;<em>Cebus capucinus</em>; CETA = Cetacea; CHLAM = Chlamyphoridae; CHOL =&nbsp;<em>Choloepus</em>; Cjac =&nbsp;<em>Callithrix jacchus</em>; CING = Cingulata; DASY = Dasypodidae; DELP = Delphinidae; DERM = Dermoptera; Erob =&nbsp;<em>Eschrichtius robustus</em>; INIA =&nbsp;<em>Inia</em>; FOLI = Folivora; GALE =&nbsp;<em>Galeopterus</em>; LIPO =&nbsp;<em>Lipotes</em>; Lobl =&nbsp;<em>Lagenorhynchus obliquidens</em>; MANI = Manidae; MONO = Monodontidae; MYRM = Myrmecophagidae; MYST = Mysticeti; NPP = Not present in&nbsp;<em>Platanista</em>&nbsp;or Physeteroidea, but present in other Odontocetes; NPZ = Not present in Ziphiidae, but present in other Odontocetes; Oorc =&nbsp;<em>Orcinus orca</em>; PEUT = Tolypeutinae; PHOC = Phocoenidae; PHOL = Pholidota; PHOR = Chlamyphorinae; PILO = Pilosa; PHYS = Physeteroidea; PONT =&nbsp;<em>Pontoporia</em>; Schi =&nbsp;<em>Sousa chinensis</em>; SIRE = Sirenia; Tadu =&nbsp;<em>Tursiops aduncus</em>; TOLY =&nbsp;<em>Tolypeutes</em>; VERM = Vermilingua; XEN = Xenarthra.</p> <p><br> <strong>- Supplementary_Table_S3.xlsx:&nbsp;</strong>Accession numbers and functionality of <em>ASMT</em> in species examined. See Table S2 caption for details.</p> <p><strong>- Supplementary_Table_S4.xlsx:&nbsp;</strong>Accession numbers and functionality of <em>MTNR1A</em> in species examined. See Table S2 caption for details.</p> <p><strong>- Supplementary_Table_S5.xlsx:&nbsp;</strong>Accession numbers and functionality of <em>MTNR1B</em> in species examined. See Table S2 caption for details.</p> <p><strong>- Supplementary_Table_S6.xlsx:&nbsp;</strong>Codon frequency model selection. These are the results from one ratio dN/dS analyses using different codon frequency models.&nbsp;AIC = Akaike Information Criterion.</p> <p><strong>- Supplementary_Table_S7.xlsx:&nbsp;</strong>Results of <em>AANAT</em> PAML dN/dS analyses for mammals. Model: BG = branch(es) grouped with background; fixed 1 = branch(es) fixed at 1. p&rsquo;-value: p-value after Holm-Bonferroni correction for multiple testing. Model Comparison: if model comparison yields statistically significant differences (p &lt; 0.05), model comparison bolded and given green background; if model comparison is still significant after Holm-Bonferroni correction, asterisk (*) added. For most models, w only shown for branch(es) of interest. Numbers in front of taxonomic names in first row correspond to numbers in the master model shown in Supplementary Figure S1.</p> <p><strong>- Supplementary_Table_S8.xlsx:&nbsp;</strong>Results of <em>ASMT</em> PAML dN/dS analyses for mammals. Refer to Table S7 caption for additional details. Numbers in front of taxonomic names in first row correspond to numbers in the master model shown in Supplementary Figure S2.</p> <p><strong>- Supplementary_Table_S9.xlsx:&nbsp;</strong>Results of <em>MTNR1A</em> PAML dN/dS analyses for mammals. Refer to Table S7 caption for additional details. Numbers in front of taxonomic names in first row correspond to numbers in the master model shown in Supplementary Figure S3.</p> <p><strong>- Supplementary_Table_S10.xlsx:&nbsp;</strong>Results of <em>MTNR1B</em> PAML dN/dS analyses for mammals. Refer to Table S7 caption for additional details. Numbers in front of taxonomic names in first row correspond to numbers in the master model shown in Supplementary Figure S4.</p> <p><strong>- Supplementary_Table_S11.xlsx:&nbsp;</strong>Results of PAML analyses for sauropsids.</p> <p><strong>- Supplementary_Table_S12.xlsx:&nbsp;</strong>Results of BLASTing and mapping short reads from&nbsp;<em>Alligator mississippiensis</em>&nbsp;RNA sequencing experiments.</p> <p><strong>- Supplementary_Table_S13.xlsx:&nbsp;</strong>Supporting data for validating putative inactivating mutations. Validating data came from four general sources of information: mutations shared by more than one species within a clade, mutations shared by two sources of sequencing data for the same species, mutations validated by coverage of mapped short reads and statistically elevated dN/dS ratio estimates. For additional details, see Supplementary Tables S2&ndash;S5 and S7&ndash;S10, as well as Figure 2 and Supplementary Figures S8&ndash;S18.</p> <p><strong>- Supplementary_Dataset_S1.txt:</strong><strong>&nbsp;</strong>Genomic alignments in fasta format used to determine the pseudogene/functional&nbsp;status of all four melatonin genes in different taxonomic groups.</p> <p><strong>- Supplementary_Dataset_S2.txt:</strong><strong>&nbsp;</strong>Alignment of <em>AANAT</em>&nbsp;in phylip format used in maximum likelihood phylogenetic reconstruction with RAxML.&nbsp;</p> <p><strong>- Supplementary_Dataset_S3.txt:&nbsp;</strong>Alignment of <em>ASMT</em> in phylip format used in maximum likelihood phylogenetic reconstruction with RAxML.&nbsp;</p> <p><strong>- Supplementary_Dataset_S4.txt:&nbsp;</strong>Alignment of <em>MTNR1A</em> and <em>MTNR1B</em> in phylip format used in maximum likelihood phylogenetic reconstruction with RAxML.&nbsp;</p> <p><strong>- Supplementary_Dataset_S5.txt:</strong><strong>&nbsp;</strong>Codon&nbsp;alignments of <em>AANAT</em> used in selection pressure analyses&nbsp;with PAML.&nbsp;</p> <p><strong>- Supplementary_Dataset_S6.txt:&nbsp;</strong>Codon&nbsp;alignments of <em>ASMT</em> used in selection pressure analyses&nbsp;with PAML.</p> <p><strong>- Supplementary_Dataset_S7.txt:</strong><strong>&nbsp;</strong>Codon&nbsp;alignments of <em>MTNR1A</em> used in selection pressure analyses&nbsp;with PAML.</p> <p><strong>- Supplementary_Dataset_S8.txt: </strong>Codon&nbsp;alignments of <em>MTNR1B</em> used in selection pressure analyses&nbsp;with PAML.</p> <p><strong>- Supplementary_Dataset_S9.txt:&nbsp;</strong>Tree topologies in newick format used in selection pressure analyses&nbsp;with PAML.</p>

opencc-by-4.0Jun 2021View details →
zenodo48/100

Data sets for Span-level SNR Regression in EONs

<ul> <li><strong>DS1: </strong>the symbol rate is fixed and equals 64 Gbaud, and the channel loading factor is selected from [25 &minus; 100];</li> <li><strong>DS2:</strong> the symbol rate and channel occupancy status is randomly selected (uniformly distributed) from {32, 64, 96 (GBaud) and {0, 1}, respectively;</li> <li>&nbsp;<strong>DS3</strong>: both symbol rate and the channel loading factor are fixed and equal to 64 GBaud and 25%, respectively.</li> </ul>

opencc-by-4.0May 2022View details →
zenodo48/100

Files from TCGA-KIRC Study for Body Part Regression Tutorial

<p>The data here are in whole based upon data generated by the TCGA Research Network:&nbsp;<a href="https://cancergenome.nih.gov/">http://cancergenome.nih.gov/</a>.</p> <p><br> The DICOM files from the <a href="https://wiki.cancerimagingarchive.net/display/Public/TCGA-KIRC#580038695f8cd691bda43dda71b4093c69c7318">TCGA-KIRC&nbsp; </a>study were converted to nifti files. Moreover, the nifti files with greater size than 35 MB and smaller size than 5 MB were removed (to reduce the size of the dataset and to remove the files with few slices). Furthermore, the metadata from the DICOM files is saved in a separate excel-file.</p>

opencc-by-3.0Jul 2021View details →
zenodo48/100

Downsampling of CT-Lymph-Node Dataset for Body Part Regression Tutorial

<p>Down sampling of the<a href="https://wiki.cancerimagingarchive.net/display/Public/CT+Lymph+Nodes#19726546f04e74ab3631480694fcb72cac2e5477"> CT Lymph Node</a> dataset from the TCIA.<br> The files were down sampled to a pixel spacing of 7 mm/pixel. Through zero padding and cropping, all images are provided in the size of 64px x 64 px. Moreover, the HU values were clipped between -1000 HU and 1500 HU and rescaled to -1 and 1. To avoid aliasing effects, an additional Gaussian smoothing filter was applied before down sampling.</p> <p>This dataset was created for a Body Part Regression tutorial.</p>

opencc-by-3.0Jul 2021View details →
edi48/100

Sensor and nutrient data associated with the article Harrison et al. 2020. Prediction of stream nitrogen and phosphorus concentrations from high-frequency sensors using Random Forests Regression

This document describes a dataset used to produce Random Forests Regression models of stream nitrogen and phosphorus concentrations from high-frequency sensor data, as reported in: Harrison, J.W., Lucius, M.A., Farrell, J.L., Eichler, L.W., and Relyea, R.A. 2020. Prediction of stream nitrogen and phosphorus concentrations from high-frequency sensors using Random Forests Regression. Science of the Total Environment: https://doi.org/10.1016/j.scitotenv.2020.143005. The dataset consists of paired values of stream nitrogen and phosphorus concentrations and various high-frequency sensor parameters (water temperature, specific conductance, pH, fluorescent dissolved organic matter, turbidity, hydrostatic pressure, soil moisture) collected during baseflow and storm events from 2018 to 2019 as part of routine monitoring of eleven tributaries of Lake George, New York. This dataset does not include raw data; two levels of processing were performed: (1) erroneous values (extreme or otherwise outlying values with no apparent environmental cause) were removed from the sensor data as part of the routine QA/QC process of the Jefferson Project, and (2) one-hour rolling medians of the raw sensor data were calculated at a 1-minute timestep to maximize pairing of sensor data with nutrient concentrations. The resultant dataset was used to train and test the models presented in Harrison et al. 2020.

openCC (other)Jan 2021View details →
zenodo44/100

Investigating terrestrial isopod abundance in sandplain grassland using a multiple linear regression

<p>Most North American species of terrestrial isopod (Isopoda) have been introduced from Europe. Sandplain grassland is a globally rare habitat that is abundant on Nantucket Island, Massachusetts and the abundance of terrestrial isopods in the habitat has never been studied. The objective of this project was to develop a model to explain isopod abundance based on vegetation characteristics within Sandplain grassland and use this model to test for land management effects (prescribed burning and mowing) on isopod abundance. I counted terrestrial isopods from 175 pitfall traps set for one week and used multiple linear regression with several selection algorithms to select the best model. The vegetation characteristics I used as regressors do not appear to explain terrestrial abundance well and the final model only contains the percent grass coverage as a regressor. The model suggests that terrestrial isopods decrease in abundance with increasing grass coverage and it explains 29 percent of the data. When management effects are incorporated, the model suggests that mowing significantly increases isopod abundance.</p> <p>Funding for this project came from the Nantucket Islands Land Bank, Nantucket Land Council, and the Nantucket Biodiversity Initiative.</p> <p>Associated vegetation data is in the published &quot;Effects of Sandplain Grassland Management on Spider Richness and Abundance on Nantucket Island&quot; dataset.&nbsp; Sampling methods are in the thesis linked from that dataset.</p> <p>allisopodData.csv - isopod counts by trap<br> dataDictionary.csv - descriptions of variables<br> mckenna-foster_2009.pdf - a report submitted to NBI and used as part of a statistics class at the University of Wisconsin-Green Bay</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2009View details →
zenodo44/100

Supplementary Data - MORTALITY RATE DUE TO PULMONERY FIBROSIS ASSOCIATED WITH SARS- COV-2 INFECTION: SCOPE OF BEST FIT REGRESSION

<p>The dataset contains number of infected pateints - Death Frquencies - Mortality rate globally due to pulmonary fibrosis associated with&nbsp;SARS-COV-2 infection with effect from 21st Jan to 28 th April ,2020 . Data analysis report by best fit regression software Curve Expert V.1.4 supported with Spreadsheet ( Excel , Office 2007 ) are included for computation of statistical significance .</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Dataset: The effects of class balance on the training energy consumption of logistic regression models

<p>Two synthetic datasets for binary classification, generated with the Random Radial Basis Function generator from WEKA. They are the same shape and size (104.952 instances, 185 attributes), but the "balanced" dataset has 52,13% of its instances belonging to class c0, while the "unbalanced" one only has 4,04% of its instances belonging to class c0. Therefore, this set of datasets is primarily meant to study how class balance influences the behaviour of a machine learning model.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Covid-19 CT dataset for Body Part Regression Tutorial

<p>The dataset is a subset of CT scans from the <a href="https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70226443">Covid-19-AR</a> dataset from the Cancer Image Archive. The data were converted from the DICOM file format to the NifTI&nbsp;file format for better and easier handling. The converted dataset was created for a tutorial of the <a href="https://github.com/MIC-DKFZ/BodyPartRegression">bpreg</a> python package.</p> <p>Acknowledgment:<br> The dataset was funded with federal funds from the National Center for Advancing Translational Sciences&nbsp;&nbsp;UL1 TR003107 and the National Cancer Institute, Contract No. 75N91019D00024, Subcontract 20X023F.&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

IML workshop challenge on jet mass regression

<p>This dataset is associated with the LPCC IML (Lhc Physics Center at Cern Inter-experimental Machine Learning) working group.&nbsp; It was produced for the second IML annual workshop (April 2018).</p> <p>This dataset is part of a machine learning &quot;challenge&quot; on jet mass regression at future circular collider (FCC) conditions.&nbsp; Further details can be found on the challenge page, here:</p> <p><a href="https://gitlab.cern.ch/IML-WG/IML_challenge_2018/wikis/home">https://gitlab.cern.ch/IML-WG/IML_challenge_2018/wikis/home</a></p>

opencc-zeroMar 2018View details →
zenodo44/100

SCIAMACHY NO regression fit MCMC samples

<p><strong>SCIAMACHY mesosphere NO data and regression model samples</strong></p> <p>SCIAMACHY mesosphere daily zonal mean NO data and Markov-Chain Monte-Carlo samples from the regression coefficient distributions, derived from and&nbsp;for use with the&nbsp;<a href="https://github.com/st-bender/sciapy"><code>sciapy</code></a>&nbsp;regression module.</p> <p>This data set contains the following files:</p> <ul> <li><code>NO_regress_output_pGM_Lya_ltcs_exp1dscan60d_km32_float32.nc</code>, <code>NO_regress_output_pGM_Lya_ltcs_exp1dscan60d_km32_float64.nc</code>&nbsp;- samples from the&nbsp;regression coefficient distributions (single and double precision)</li> <li><code>NO_regress_quantiles_pGM_Lya_ltcs_exp1dscan60d_km32.nc</code> -&nbsp;the 0.1, 2.5, 16, 50, 84, 97.5, and 99.9 percentiles of the sampled distributions</li> <li><code>scia_nom_dzmNO_2002-2012_v6.2.1_2.2_akm0.002_geomag10_nw.nc</code>&nbsp;- the SCIAMACHY daily zonal mean NO data</li> <li><code>sciapy_regress_tutorial.ipynb</code> - example ipython notebook</li> </ul> <p><strong>MCMC Samples</strong></p> <p>The files <code>NO_regress_output..._float32.nc</code> and <code>NO_regress_output..._float64.nc</code> contain MCMC samples of the model as single and double precision floats. The file <code>NO_regress_quantiles....nc</code> contains the 0.1, 2.5, 16, 50, 84, 97.5, and 99.9 percentiles of the sampled distributions and is provided for convenience. The files contain the following parameters:</p> <ul> <li><code>kernel:log_sigma</code>, <code>kernel:log_rho</code> - the &quot;strength&quot; and &quot;lengthscale&quot; of the Mat&eacute;rn-3/2 Gaussian Process kernel</li> <li><code>mean:offset:value</code> - the constant offset of the NO model in [<span class="math-tex">\(10^6\)</span>&nbsp;cm<span class="math-tex">\(^{-3}\)</span>]</li> <li><code>mean:Lya:amp</code> - the Lyman-<span class="math-tex">\(\alpha\)</span>&nbsp;coefficient of the mean model in [<span class="math-tex">\(10^6\)</span>&nbsp;cm<span class="math-tex">\(^{-3}\)</span>&nbsp;/ Lyman-<span class="math-tex">\(\alpha\)</span>]</li> <li><code>mean:GM:amp</code> - the geomagnetic coefficient (AE) in&nbsp;[<span class="math-tex">\(10^6\)</span>&nbsp;cm<span class="math-tex">\(^{-3}\)</span>&nbsp;/ nT]</li> <li><code>mean:GM:tau0</code> - the constant lifetime of the geomagnetic lifetime in [d]</li> <li><code>mean:GM:taucos1</code>, <code>mean:GM:tausin1</code> - cosine and sine amplitudes of the yearly geomagnetic lifetime variation in [d]</li> </ul> <p><strong>Daily zonal mean NO data</strong></p> <p>The model was trained on the&nbsp;<a href="http://doi.org/10.5281/zenodo.1009078">SCIAMACHY mesosphere NO dataset</a>, binned into 10&deg; geomagnetic latitude bins using the provided <code>gm_lat</code> variable and using the standard error of the mean as data uncertainties. The data are uploaded as&nbsp;<code>scia_nom_dzmNO_2002-2012_v6.2.1_2.2_akm0.002_geomag10_nw.nc</code>&nbsp;and&nbsp;were prepared by running (after installing&nbsp;<code><a href="https://github.com/st-bender/sciapy">sciapy</a>)</code>:</p> <pre><code class="language-bash">bash&gt; scia_daily_zonal_mean.py -g -b'-90:90:10' -o &lt;daily_zonal_mean_NO.nc&gt; &lt;/path/to/SCIAMACHY_NO_NOM_orbits_20??_v6.2.1.nc&gt;</code></pre> <p><strong>Regression sampling</strong></p> <p>The samples were generated by running the following command:</p> <pre><code class="language-bash">bash&gt; python -m sciapy.regress &lt;daily_zonal_mean_NO.nc&gt; --proxies Lya:&lt;Lyman-alpha_file.dat&gt;,GM:&lt;AE_file.dat&gt; -A &lt;altitude&gt; -L &lt;geomag_latitude_bin&gt; -w 14 -b 800 -p 1400 -F \"\" -I GM --fit_annlifetimes GM --positive_proxies GM --lifetime_scan=60 --lifetime_prior exp -k -K Mat32 -O0 -m "nom_pGM_Lya_ltcs_exp1dscan60d_km32" -P</code></pre> <p>&nbsp;</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

Fast and accurate large multiple sequence alignments with a root-to-leaf regressive method

<p>This dataset contains a GitHub repository containing all the data, analysis, Nextflow workflows and Jupyter notebooks to replicate the manuscript&nbsp;titled &quot;Fast and accurate large multiple sequence alignments with a root-to-leaf regressive method&quot;.</p> <p>It also contains the Multiple Sequence Alignments (MSAs) generated and well as the main figures and tables from the manuscript.</p> <p>The repository is also available at GitHub (https://github.com/cbcrg/dpa-analysis) release `v1.2`.</p> <p>For details on how to use the regressive alignment algorithm, see the T-Coffee software suite (https://github.com/cbcrg/tcoffee).</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

swmm-nrtestsuite: Regression Test Suite for OWA SWMM

<p>Open Water Analytics (OWA) Stormwater Management Model (SWMM)</p> <p>Project Link:&nbsp;<br> &nbsp; https://github.com/OpenWaterAnalytics/swmm-nrtestsuite</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Electron Energy Regression in High-Granularity Calorimeter Prototype

<p>The dataset consists of simulations of calibrated reconstructed hits produced by a positron passing through the HGCAL test beam prototype. For the simulations, Monte Carlo method is used to produce the positrons with energy ranging from 20 to 350 GeV. The dataset contains the coordinates of the calibrated reconstructed hits in the prototype along with the calibrated energy in units of MIP.&nbsp;The HDF5 files can be extracted from the gzip files.</p>

opencc-by-4.0Jan 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record