Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
164
datasets available to search
ShareScore release 0.9.0
Dataset results
164 results for “prediction algorithms”
The Prediction of the Rate of the Dropout of the Primary Schools Students by Using the Genetic Algorithm-Figure 10. Drawing of autocorrelation function and partial correlation for Males and Females primary stage students
<p>The instability of the time series is recognized, and to be more accurate, we draw each (Autocorrelation Function) ACF, and (Partial Autocorrelation Function) PACF in a row to assure the stability according to the figure (10).</p>
The Prediction of the Rate of the Dropout of the Primary Schools Students by Using the Genetic Algorithm-Figure 14. Drawing of autocorrelation function and partial correlation of the residues for primary stage males students
<p>After diagnosing and evaluating the models, the accommodating and the sufficiency of the models must be checked for primary stage female students, through applying the compute (Ljung- Box Q) to check the model accommodation on the function level 0.05 so the Q value occurs of primary stage female students: Ljung-Box Q' = 0.966626, With p-value = P(Chi-square(1) > 0.966626) = 0.3255 However, the Tabulated value equals 3.841 whilst the Q value is less than Tabulated value, so it accepts the Null Hypothesis which indicates the emptiness of the evaluated model out of the contrast accordance trouble. It's possible to notice the two parameters functions (Autocorrelation and Partial Correlation Functions) of the residues for females primary stage students, in which the residues value is located within the confidence interval limits, which means the residues series is random and the evaluated model is good and convenient as it is shown.</p>
The Prediction of the Rate of the Dropout of the Primary Schools Students by Using the Genetic Algorithm-Figure 9. Drawing the time series for males and females primary stage students
<p>The Unit Radix Dickey-Fuller Test is used to ensure the series’ stability. The results are: Dickey-Fuller Test Estimated Value = 0.736458 , Statistic Test = 0.380545 , P-Value = 0.794 We notice from the values above P-Value = 0.794on the abstract level of 0.05 which leads to refusing the Null Hypothesis and accepting the Alternative Hypothesis ( The Nonexistence of a Radix Unit) implies that the time series is stable. Figure (9) represents the time series of females and males in the primary stage students.</p>
The Prediction of the Rate of the Dropout of the Primary Schools Students by Using the Genetic Algorithm-Figure 7. Drawing the time chain for Females primary stage students after the first difference
<p>We use the Unit Radix Dickey-Fuller Test to assert the series’ stability. The results are: Dickey-Fuller Test Estimated Value = 0.47242, Statistic Test = 1.22797, P-Value = 0.9445 We notice from the values above P-Value =0.9445 on the abstract level of 0.05 which leads to accepting the Null Hypothesis and refusing the Alternative Hypothesis (Existence of a Radix Unit) implies that the time series is instable. By taking the first difference, it is noticed that the stability of the time series has been accomplished. See Figure 7.</p>
The Prediction of the Rate of the Dropout of the Primary Schools Students by Using the Genetic Algorithm-Figure 6. Drawing of auto correlation function and partial correlation for females primary stage students
<p>The instability of the time series is noticed and to be more precise we draw each (Autocorrelation Function) ACF, and (Partial Autocorrelation Function) PACF in a row to affirm the stability according to the figure 6.</p>
The Prediction of the Rate of the Dropout of the Primary Schools Students by Using the Genetic Algorithm-Figure 4. Drawing of autocorrelation function and partial correlation for males primary stage students
<p>We get to notice the stability of the time series, and to be more accurate we draw each (Autocorrelation Function) ACF, and (Partial Autocorrelation Function) PACF in a row to ensure the stability according to the figure (4).</p>
The Prediction of the Rate of the Dropout of the Primary Schools Students by Using the Genetic Algorithm-Figure 5. Drawing the time series for primary stage female students
<p>We use the Unit Radix Dickey-Fuller Test to assure the series’ stability. The results are: Dickey-Fuller Test Estimated Value = 0.233403, Statistic Test = 0.125769, P-Value = 0.6405 The values above P-Value = 0.6405 is noted on the abstract level of 0.05 which leads to refusing the Null Hypothesis and accepting the Alternative Hypothesis (The Nonexistence of a Radix Unit) implies that the time series is stable. Figure 5 represents the Time series of Female Primary Stage Students.</p>
The Prediction of the Rate of the Dropout of the Primary Schools Students by Using the Genetic Algorithm-Figure 8. Drawing of autocorrelation function and partial correlation for Females primary stage students
<p>The stability of the time series is observed and to be more accurate we draw each (Autocorrelation Function) ACF, and (Partial Autocorrelation Function) PACF in a row to assure the stability according to the figure (8).</p>
The Prediction of the Rate of the Dropout of the Primary Schools Students by Using the Genetic Algorithm-Figure 3. Drawing the time series for Males primary stage students after getting the first difference
<p>The Unit Radix Dickey-Fuller Test is used to assure the series’ stability. The results are: Dickey-Fuller Test Estimated Value = 0.276151, Statistic Test = 0.87796, P-Value = 0.8984 We get to notice from the values above P-Value = 0.8984 on the abstract level of 0.05 which leads to accepting the Null Hypothesis and refusing the Alternative Hypothesis (Existence of a Radix Unit) implies that the time series is instable. By taking the first difference, it is observed that the stability of the Time Series has been accomplished . See figure 3.</p>
The Prediction of the Rate of the Dropout of the Primary Schools Students by Using the Genetic Algorithm-Figure 2. Drawing of autocorrelation function and partial correlation for males primary stage
<p>We get to notice the instability of the time series, and to be more precise we draw each (Autocorrelation Function) ACF, and (Partial Autocorrelation Function) PACF in a row to ensure the stability according to the figure 2.</p>
The Prediction of the Rate of the Dropout of the Primary Schools Students by Using the Genetic Algorithm-Figure 1. Drawing the time series for males primary stage
<p>After collecting all the students’ dropout proportion for both males and females in the<br> primary stage, the first step of the Box-Jenkins is to draw the time chain data to understand the<br> chain's attitude.</p>
Prediction of Humpback Whale Sighting Zones based on Environmental Factors using Tree-based Algorithms
<p>This is the datased used in the paper: Prediction of Humpback Whale Sighting Zones based on Environmental Factors using Tree-based Algorithms</p>
Supporting data for RCANE: A Deep Learning Algorithm for Whole-genome Pan-Cancer Somatic Copy Number Aberration Prediction using RNA-seq Data.
<p>This is the data repository for <em>RCANE: A Deep Learning Algorithm for Whole-genome Pan-Cancer Somatic Copy Number Aberration Prediction using RNA-seq Data</em>. To use this dataset, please refer to <a href="https://github.com/HowardGech/RCANE" target="_blank" rel="noopener">https://github.com/HowardGech/RCANE</a>.</p>
Biometeorological Dataset for 'Novel algorithms for high resolution prediction of canopy evapotranspiration in grapevine'
<p>A head trained <strong><em>Vitis vinifera</em></strong> L. cv. Zinfandel vine was grafted on St. George rootstock (<em>V. rupestris</em>) then planted in a 1.1 m<sup>3</sup> plastic container filled with Yolo County, CA sourced sandy loam.<br> <br> To estimate evapotranspiration, we measured the wind speed, air temperature and relative humidity in vine canopies by mounting each vine with a suite of research grade sensors. We measured wind speed (units m ᐧ s<sup>-1</sup>) inside the vine canopy using a single needle anemometer (<em>East 30 Sensors</em>; Pullman, WA) that took instantaneous wind speed measurements every 10 seconds and recorded the average of the previous 12 instantaneous measurements for every 2-minute interval.</p> <p>We measured temperature (units <sup>o</sup>C) and relative humidity (units %) using HMP60L sensors (Campbell Scientific; Logan, UT) mounted both inside and outside of each vine canopy and recorded instantaneous measurements at each 2-minute interval. We filtered all biometeorological data using a 3-hour moving average to remove noise without causing any significant over or under-approximation of daily maxima and minima.</p> <p>We automated all data collection using two CR1000 data loggers (<em>Campbell Scientific</em>; Logan, UT), with 1 or 2 vines and associated sensors per logger, using custom CR1 programs. A single 30W solar cell and 12V lead acid battery powered the entire vine-sensor system.</p> <p>This dataset represents all sensor data from a single vine, as measured in August 2020. Columns are named accordingly and include units.</p> <p><strong>Please Note</strong>: The column named 'load_cell_kg' is not named accurately. The values given are in units of millivolts, and need to be translated from millivolts to kilograms. The 2020 calibration coefficient is 0.00330693663 millivolts per kilogram.</p>
Quasar Factor Analysis – An Unsupervised and Probabilistic Quasar Continuum Prediction Algorithm with Latent Factor Analysis
<p>Dataset used in <em>Quasar Factor Analysis – An Unsupervised and Probabilistic Quasar Continuum Prediction Algorithm with Latent Factor Analysis </em>[<a href="https://arxiv.org/abs/2211.11784">arXiv: <strong>2211.11784</strong></a>]. This dataset will be helpful to validate different continuum prediction model and study absorption systems.<br> <br> The descriptions of individual files can be found here:</p> <ul> <li><a href="/api/files/01c3b414-7572-4a24-9c67-fbab6ae05969/sdss-dr16.tar.gz?versionId=427c1aca-6f12-4e84-b91e-f9a3d678329e">sdss-dr16.tar.gz</a> : continuum prediction for ~100,000 quasar spectra from SDSS DR16, see Section 3.1 in <a href="https://arxiv.org/abs/2211.11784">arXiv:211.11784</a>;</li> <li><a href="https://zenodo.org/api/files/01c3b414-7572-4a24-9c67-fbab6ae05969/sdss-mock-with-dla-with-perturb.tar.gz?versionId=55bc466d-4cbf-43ba-9b83-209ffba4ec30">sdss-mock-with-dla-with-perturb.tar.gz </a> : ~150,000 mock quasar spectra to validate QFA performance with perturbation on quasar continuum from PCA template, see Section 3.2 in <a href="https://arxiv.org/abs/2211.11784">arXiv:211.11784</a>;</li> <li><a href="https://zenodo.org/api/files/01c3b414-7572-4a24-9c67-fbab6ae05969/sdss-mock-with-dla-without-perturb.tar.gz?versionId=efa5d05a-2a40-4a43-801e-c95b29a4efb2">sdss-mock-with-dla-without-perturb.tar.gz</a>: ~150,000 mock quasar spectra to validate QFA performance with quasar continuum directly from PCA template, see Section 3.2 in <a href="https://arxiv.org/abs/2211.11784">arXiv:211.11784</a>.<br> <br> <br> </li> </ul>
Different orthology inference algorithms generate similar predicted orthogroups among Brassicaceae species
Open the record for dataset details and reuse information.
Assessing predictive performance of supervised machine learning algorithms for a diamond pricing model
Open the record for dataset details and reuse information.
Seafloor Density Measurements, Prediction, and Associated Uncertainty for "Predicting global marine sediment density using the random forest regressor machine learning algorithm"
<p>Global seafloor density prediction results using the random forest regressor machine learning algorithm. </p> <p>Dataset S1. Seafloor density measurements. Columns are labeled with a header and include associated drilling project and measurement type for each sample. File format: CSV text file</p> <p>Dataset S2. Seafloor density prediction results from the random forest regressor machine learning algorithm at 5×5-arc minute resolution. Units are g/cm^3. File format: netCDF (.nc)</p> <p>Dataset S3. Seafloor density prediction standard deviation from the random forest regressor machine learning algorithm at 5×5-arc minute resolution. Units are g/cm^3. File format: netCDF (.nc)</p>
Subseasonal to Seasonal (S2S) Prediction Algorithms using Hybrid Machine Learning Techniques
<p>< S2S dataset.zip ></p><p>1.ECMWF observations/hindcast realizations</p><ul><li>hindcast-like-observations_2000-2019_biweekly_deterministic.zarr</li><li>forecast-like-observations_2020_biweekly_deterministic.zarr</li><li>ecmwf_hindcast-input_2000-2019_biweekly_deterministic.zarr</li><li>ecmwf_forecast-input_2020_biweekly_deterministic.zarr</li><li>hindcast-like-observations_2000-2019_biweekly_tercile-edges.nc</li></ul><p>2. External variables</p><ul><li>"nino" folder -> nino12.long.anom.data, nino34.long.anom.data : El Niño data</li><li>"Oscillation" folder<ul><li>-> ersst.v5.pdo.dat.text : PDO (Pacific Decadal Oscillation)</li><li>-> norm.nao.monthly.b5001.current.ascii.table.txt : NAO (North Atlantic Oscillation)</li><li>-> qbo.dat : QBO (Quasi Biennial Oscillation)</li></ul></li><li>"great_lake" folder -> N_seaice_extent_daily_v3.0 : Great lakes ice cover</li><li>observed-solar-cycle-indices.json : Sunspot cycles (two variables: original value and smoothed value)</li></ul><p>3. Region.txt : Region and its bound</p><p>4. Biweekly historical statistics data</p><ul><li>biw_stat_w34 folder -> data (mean, standard deviation, median, skewness, kurtosis) for Week 3-4</li><li>biw_stat_w56 folder -> data (mean, standard deviation, median, skewness, kurtosis) for Week 5-6</li></ul><p> </p><p>< ML_code.zip ></p><ul><li>ML codes for training, testing, and calculating RPSS based on Python3</li><li>Check run_val.sh and run_2020.sh </li></ul><p> </p>
Antioxidant capacity prediction of Nanoforms using supervised algorithms.
<p>This dataset contains p-chem properties of Nanoforms and experimental conditions for the antioxidant capacity measured with the DPPH assay.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.