Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

16

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

16 results for “linear regression”

Learn how ShareScore rates datasets ↗
zenodo52/100

AMOC reconstruction between 1981 and 2016 from hydrographic data using an empirical linear regression model from Worthington, E. L., Moat, B. I., Smeed, D. A., Mecking, J. V., Marsh, R., and McCarthy, G. D.: A 30-year reconstruction of the Atlantic meridional overturning circulation shows no decline, Ocean Sci., 17, 285–299, https://doi.org/10.5194/os-17-285-2021, 2021.

<p>Dataset used to create Figure 8 in Worthington et al., 2021 (https://doi.org/10.5194/os-17-285-2021). Details of the data and methods can be found in the journal article.<br> <br> Worthington, E. L., Moat, B. I., Smeed, D. A., Mecking, J. V., Marsh, R., and McCarthy, G. D.: A 30-year reconstruction of the Atlantic meridional overturning circulation shows no decline, Ocean Sci., 17, 285&ndash;299,&nbsp;<a href="https://doi.org/10.5194/os-17-285-2021">https://doi.org/10.5194/os-17-285-2021</a>, 2021.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Investigating terrestrial isopod abundance in sandplain grassland using a multiple linear regression

<p>Most North American species of terrestrial isopod (Isopoda) have been introduced from Europe. Sandplain grassland is a globally rare habitat that is abundant on Nantucket Island, Massachusetts and the abundance of terrestrial isopods in the habitat has never been studied. The objective of this project was to develop a model to explain isopod abundance based on vegetation characteristics within Sandplain grassland and use this model to test for land management effects (prescribed burning and mowing) on isopod abundance. I counted terrestrial isopods from 175 pitfall traps set for one week and used multiple linear regression with several selection algorithms to select the best model. The vegetation characteristics I used as regressors do not appear to explain terrestrial abundance well and the final model only contains the percent grass coverage as a regressor. The model suggests that terrestrial isopods decrease in abundance with increasing grass coverage and it explains 29 percent of the data. When management effects are incorporated, the model suggests that mowing significantly increases isopod abundance.</p> <p>Funding for this project came from the Nantucket Islands Land Bank, Nantucket Land Council, and the Nantucket Biodiversity Initiative.</p> <p>Associated vegetation data is in the published &quot;Effects of Sandplain Grassland Management on Spider Richness and Abundance on Nantucket Island&quot; dataset.&nbsp; Sampling methods are in the thesis linked from that dataset.</p> <p>allisopodData.csv - isopod counts by trap<br> dataDictionary.csv - descriptions of variables<br> mckenna-foster_2009.pdf - a report submitted to NBI and used as part of a statistics class at the University of Wisconsin-Green Bay</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2009View details →
zenodo40/100

Fig. 2. Linear regression models showing the relationship between Aphis citricola and Harmonia axyridis abundance. A in Behavioral responses of Aphis citricola (Hemiptera: Aphididae) and its natural enemy Harmonia axyridis (Coleoptera: Coccinellidae) to non-host plant volatiles

Fig. 2. Linear regression models showing the relationship between Aphis citricola and Harmonia axyridis abundance. A: Catnip (Nepeta cataria) + French marigold (Tagetes patula), B: ageratum (Ageratum houstonianum) + French marigold, C: catnip + ageratum, and D: native vegetation.

opencc-by-4.0Jun 2017View details →
zenodo32/100

Datasets for linear regression on Swedish Motor Insurance

<p>In this project we have 3 datasets. Training Set and Test set consists of the input data from Swedish Motor Insurance dataset which is dividen in ratio 80%-20%. Third dataset consists of our predictions for Sum of payments using linear regression.</p>

opencc-by-4.0Apr 2020View details →
zenodo32/100

FIGURES 1–5. Trichoptera larvae and linear regression. Figure 1, Ithytrichia lamellaris Eaton 1873, fifth instar larva with case. Figure 2, Hydropsyche tabacarui Botosaneanu 1960 in Tools for instar determination of European caddisfly larvae (Insecta: Trichoptera)

FIGURES 1–5. Trichoptera larvae and linear regression. Figure 1, Ithytrichia lamellaris Eaton 1873, fifth instar larva with case. Figure 2, Hydropsyche tabacarui Botosaneanu 1960, fifth instar larva, abdominal segments VI–IX (white ovals = pupal gill buds). Figure 3, Athripsodes longispinosus paleochora (Malicky 1972), fifth instar larva, head, right anterolateral (arrow = subocular ecdysial line). Figure 4, Halesus rubricollis (Pictet 1834), first instar larva, pronotum, dorsal (white numerals refer to setal positions; arrows indicate pits). Scale bars: 0.5 mm in Figs. 1 and 3, 1 mm in Fig. 2, 0.1 mm in Fig. 4. Figure 5, linear regression (with 95% confidence bands) of forewing length versus final instar head width, based on pooled data of 451 European Trichoptera species across all families, showing the regression equation, the coefficient of determination, and the probability level for the relationship.

opennotspecifiedJan 2021View details →
zenodo32/100

FIGURE. Scatter plots (N=200) and linear regression lines of the length and diameter of termite coprolites from the Lower Cretaceous Huolinhe Formation in eastern Inner Mongolia, China. The grey shading represents the 95% confidence interval of linear relationship. Note scatter plots depicting a k-means clustering analysis reveals three groups, indicated by circles of different colours; stars of different colour mean the clusters centroids which are the average length and diameter. in Termite coprolites (Blattodea: Isoptera) from the Early Cretaceous of eastern Inner Mongolia, Northeast China

FIGURE. Scatter plots (N=200) and linear regression lines of the length and diameter of termite coprolites from the Lower Cretaceous Huolinhe Formation in eastern Inner Mongolia, China. The grey shading represents the 95% confidence interval of linear relationship. Note scatter plots depicting a k-means clustering analysis reveals three groups, indicated by circles of different colours; stars of different colour mean the clusters centroids which are the average length and diameter.

opennotspecifiedJan 2022View details →
zenodo32/100

Generating Optimal Robust Continuous Piecewise Linear Regression with Outliers Through Combinatorial Benders Decomposition - Data Sets

<p>Data Sets for the Paper: &quot;Generating Optimal Robust Continuous Piecewise Linear Regression with Outliers Through Combinatorial Benders Decomposition&quot;.</p>

opencc-by-4.0Jul 2022View details →
zenodo32/100

TEXT-FIGURE 3. Scatterplots of morphometric measurements of Thecidellina leipnitzae sp. nov. Abbreviations: L, length; W, width; LDV, length of dorsal valve; T (max), maximal thickness; Lint, length of the interarea, Wint, width of hinge line or interarea. Relationships between ratios L/W and width, LDV/W and width, T(max)/W and width, Lint/W and width and Wint/W and width. Linear regression and regression coefficient (R²) indicated. The regression coefficient (R²) is indicated. N is the number of specimens measured. in Recent thecideide brachiopods from a submarine cave in the Department of Mayotte (France), northern Mozambique Channel

TEXT-FIGURE 3. Scatterplots of morphometric measurements of Thecidellina leipnitzae sp. nov. Abbreviations: L, length; W, width; LDV, length of dorsal valve; T (max), maximal thickness; Lint, length of the interarea, Wint, width of hinge line or interarea. Relationships between ratios L/W and width, LDV/W and width, T(max)/W and width, Lint/W and width and Wint/W and width. Linear regression and regression coefficient (R²) indicated. The regression coefficient (R²) is indicated. N is the number of specimens measured.

opennotspecifiedJun 2019View details →
zenodo32/100

Synthetic data set to evaluate and benchmark the performance of multiple linear regression algorithms in Scikit-Learn and SANElib

<p>The datasets respresent different numbers of columns and rows to measure the scalability of linear regression algorihms in terms of columns and rows.</p>

opencc-by-4.0Aug 2021View details →
zenodo32/100

Summary ouput data - Wasteaware Cities Benchmark Indicators - WABI 2023 - Global data analytics - Machine learning vs. Non-linear Regression

<p>This is the output&nbsp;dataset for the research publication &quot;<em>Socio-economic development drives solid waste management performance in cities: A global analysis using machine learning</em>&quot;. It features&nbsp;</p> <ul> <li>Metadata info used by R codes</li> <li>Summary of results for two modelling approaches (machine learning:&nbsp;Conditional random-forest and non-linear regression)</li> </ul> <p>The independent variables dataset&nbsp;analysed here refer to specific indicators of the WABI methodology (<a href="https://www.sciencedirect.com/science/article/pii/S0956053X14004905">https://www.sciencedirect.com/science/article/pii/S0956053X14004905</a>) that generates solid waste management and resource recovery profiles for cities. It was&nbsp;applied here for 40 cities around the world. The data input are available here: 10.5281/zenodo.7570174</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Figure 4. a, linear regression illustrating the relationship between log10 stride length and log10 stride speed. b in Morphological and performance modifications in the world's only marine lizard, the Galápagos marine iguana, Amblyrhynchus cristatus

Figure 4. a, linear regression illustrating the relationship between log10 stride length and log10 stride speed. b, linear regression illustrating the relationship between and log10 stride frequency and log10 stride speed for iguanids.

opennotspecifiedDec 2020View details →
zenodo32/100

TCHES artefact dataset for paper : Efficient Regression-Based Linear Discriminant Analysis for Side-Channel Security Evaluations

<p>This dataset allows reproducing the results of the CHES 2023 paper : &quot;Efficient Regression-Based Linear Discriminant<br> Analysis for Side-Channel Security Evaluations&quot;.</p> <p>It contains the side-channel measurements necessary to do so.</p> <p>Scripts and readme is available at the CHES artefact site : &lt;link not yet alive&gt;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo32/100

Output - Results of Random Forest and Multiple Linear regression analysis.

<p><strong>Hybrid streamflow modelling using machine learning and multi-model combination.</strong></p> <p>&nbsp;</p> <p><strong>Structure:</strong></p> <p><strong>MLR_output:</strong></p> <ul> <li>Validate <ul> <li>Different setups</li> </ul> </li> </ul> <p><strong>RF_output:</strong></p> <ul> <li>tune <ul> <li><em>all_stations</em></li> </ul> </li> <li>train <ul> <li><em>Different setups</em></li> </ul> </li> <li>Validate <ul> <li><em>Different setups</em></li> </ul> </li> </ul>

opencc-by-4.0Jun 2023View details →
zenodo28/100

Linear Regression BAM - UCentral Monitor

<p>Experimental measurement of PM2.5 in Bogota using a low-cost sensor developed in Universidad Central vs a BAM in Ferias Station</p>

opencc-by-4.0Nov 2020View details →
nasa20/100

Distributed Monitoring of the R2 Statistic for Linear Regression

The problem of monitoring a multivariate linear regression model is relevant in studying the evolving relationship between a set of input variables (features) and one or more dependent target variables. This problem becomes challenging for large scale data in a distributed computing environment when only a subset of instances is available at individual nodes and the local data changes frequently. Data centralization and periodic model recomputation can add high overhead to tasks like anomaly detection in such dynamic settings. Therefore, the goal is to develop techniques for monitoring and updating the model over the union of all nodes' data in a communication-efficient fashion. Correctness guarantees on such techniques are also often highly desirable, especially in safety-critical application scenarios. In this paper we develop <i>DReMo</i> --- a distributed algorithm with very low resource overhead, for monitoring the quality of a regression model in terms of its coefficient of determination (<i>R2</i> statistic). When the nodes collectively determine that <i>R2</i> has dropped below a fixed threshold, the linear regression model is recomputed via a network-wide convergecast and the updated model is broadcast back to all nodes. We show empirically, using both synthetic and real data, that our proposed method is highly communication-efficient and scalable, and also provide theoretical guarantees on correctness.

restrictednotspecifiedMar 2025View details →
zenodo12/100

Bulimia Nervosa data - linear regression

<p>Data from a sample of people affected by bulimia nervosa.</p>

restrictedNov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record