Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Dataset - FetMRQC: an open-source machine learning framework for multi-centric fetal brain MRI quality control
<p>This dataset contains the data and model used in the paper</p> <blockquote> <p>Thomas Sanchez, Oscar Esteban, Yvan Gomez, Alexandre Pron, Mériam Koob, Vincent Dunet, Nadine Girard, Andras Jakab, Elisenda Eixarch, Guillaume Auzias, and Meritxell Bach Cuadra. "FetMRQC: an open-source machine learning framework for multi-centric fetal brain MRI quality control." <a href="https://arxiv.org/abs/2311.04780"><em>arXiv preprint arXiv:2311.04780</em></a> (2023).</p> </blockquote> <p>If you found this dataset useful or used it in your research, please cite this reference.</p> <p>This dataset contains manual quality annotations and image quality metrics (IQMs) obtained from 1647 stacks of T2-weighted (T2w) slices of fetal brain magnetic resonance (MR) images collected from 233 subjects at four different institutions Lausanne University Hospital (CHUV) in Switzerland, BCNatal at Hospital Sant Joan de Déu in Barcelona (Spain), University Children's Hospital Zürich (KISPI) in Switzerland and La Timone University Hospital in Marseille, France. The data were acquired on scanners from different vendors (Siemens at CHUV, BCNatal and Marseille, General Electrics at KISPI), MR sequences (Half Fourier Single-shot Turbo spin-Echo –HASTE– for Siemens scanners and Single-Short Fast Spin Echo –SS-FSE– for GE scanners), magnetic field strengths (1.5 T and 3 T), image resolutions, fields of view, repetition times and echo times, with both neurotypical and pathological cases.</p> <p>These data and the derived IQMs were used to train and evaluate models for quality assessment and quality control of fetal brain MR images. The code to reproduce the experiments is available on <a href="https://github.com/Medical-Image-Analysis-Laboratory/fetal_brain_qc">GitHub.</a></p> <p>Each entry describe the information for a single stack of T2w slices. It contains information regarding which subject it belongs to, its manual quality rating, scanner-related information and 332 IQMs, starting at the `centroid` column in the file. Further description of the data is available in the materials and methods section of the <a href="https://arxiv.org/abs/2311.04780">paper</a>.</p> <p>The model is a 2D nnUNet [1] segmentation network trained on the super-resolution reconstructed data and manual segmentations available as part of the<a href="https://www.synapse.org/#!Synapse:syn25649159/wiki/610007"> Fetal Tissue Annotation Challenge</a> (FeTA).</p> <p>Copyright (c) - All rights reserved. Medical Image Analysis Laboratory - Department of Radiology, Lausanne University Hospital (CHUV) and University of Lausanne (UNIL), Lausanne, Switzerland & CIBM Center for Biomedical Imaging. 2023.</p> <p>[1] Isensee, Fabian, et al. "nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation." <em>Nature methods</em> 18.2 (2021): 203-211.</p>
Sea-surface pCO2 maps for the Bay of Bengal based on machine learning algorithms
<p>The dataset contains two products, the first being sea-surface <em>p</em>CO<sub>2</sub> and the second being air-sea CO<sub>2</sub> flux for the Bay of Bengal region. The data is climatological data, with 12 months of the climatological year. Each of these data has a spatial resolution of 1/12<sup>o</sup>. The positive value of CO<sub>2</sub> flux indicates the outgassing of CO<sub>2</sub>, and the negative value shows the uptake of atmospheric CO<sub>2</sub>. This version contains an elaborate description in the attribute section of the NC files, which is missing from the previous versions.</p>
Hate Speech and Bias against Asians, Blacks, Jews, Latines, and Muslims: A Dataset for Machine Learning and Text Analytics
<h1>Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University Dataset on bias against Asians, Blacks, Jews, Latines, and Muslims </h1> <div> <h2> </h2> <h2>Description </h2> </div> <div> <p>The dataset is a product of a research project at Indiana University on biased messages on Twitter against ethnic and religious minorities. We scraped all live messages with the keywords "Asians, Blacks, Jews, Latinos, and Muslims" from the Twitter archive in 2020, 2021, and 2022.</p> <p>Random samples of 600 tweets were created for each keyword and year, including retweets. The samples were annotated in subsamples of 100 tweets by undergraduate students in Professor Gunther Jikeli's class 'Researching White Supremacism and Antisemitism on Social Media' in the fall of 2022 and 2023. A total of 120 students participated in 2022. They annotated datasets from 2020 and 2021. 134 students participated in 2023. They annotated datasets from the years 2021 and 2022. The annotation was done using the <a href="https://annotationportal.com/" target="_blank" rel="noreferrer noopener">Annotation Portal</a> (Jikeli, Soemer and Karali, 2024). The updated version of our portal, <a href="https://portal2.annotationportal.com/" target="_blank" rel="noreferrer noopener">AnnotHate</a>, is now publicly available. Each subsample was annotated by an average of 5.65 students per sample in 2022 and 8.32 students per sample in 2023, with a range of three to ten and three to thirteen students, respectively. Annotation included questions about bias and calling out bias. </p> </div> <div> <p>Annotators used a scale from 1 to 5 on the bias scale (confident not biased, probably not biased, don't know, probably biased, confident biased), using definitions of bias against each ethnic or religious group that can be found in the research reports from <a href="https://isca.indiana.edu/publication-research/social-media-project/Research-Report-BIAS-on-Twitter-against-Asians--Blacks-Jews-Latinos-Muslims-final-002.pdf" target="_blank" rel="noreferrer noopener">2022</a> and <a href="https://isca.indiana.edu/documents/BIAS%20Against%20Asian-Black-Hispanic-Jewish-and-%20Muslim-People%20on%20X-Twitter%20in%202021%20and%202022.pdf" target="_blank" rel="noreferrer noopener">2023</a>. If the annotators interpreted a message as biased according to the definition, they were instructed to choose the specific stereotype from the definition that was most applicable. Tweets that denounced bias against a minority were labeled as "calling out bias". </p> </div> <div> <p>The label was determined by a 75% majority vote. We classified “probably biased” and “confident biased” as biased, and “confident not biased,” “probably not biased,” and “don't know” as not biased. </p> </div> <div> <p>The stereotypes about the different minorities varied. About a third of all biased tweets were classified as general 'hate' towards the minority. The nature of specific stereotypes varied by group. Asians were blamed for the Covid-19 pandemic, alongside positive but harmful stereotypes about their perceived excessive privilege. Black people were associated with criminal activity and were subjected to views that portrayed them as inferior. Jews were depicted as wielding undue power and were collectively held accountable for the actions of the Israeli government. In addition, some tweets denied the Holocaust. Hispanic people/Latines faced accusations of being undocumented immigrants and "invaders," along with persistent stereotypes of them as lazy, unintelligent, or having too many children. Muslims were often collectively blamed for acts of terrorism and violence, particularly in discussions about Muslims in India. </p> </div> <div> <p>The annotation results from both cohorts (Class of 2022 and Class of 2023) will not be merged. They can be identified by the "cohort" column. While both cohorts (Class of 2022 and Class of 2023) annotated the same data from 2021,* their annotation results differ. The class of 2022 identified more tweets as biased for the keywords "Asians, Latinos, and Muslims" than the class of 2023, but nearly all of the tweets identified by the class of 2023 were also identified as biased by the class of 2022. The percentage of biased tweets with the keyword 'Blacks' remained nearly the same. </p> </div> <div> <p>*Due to a sampling error for the keyword "Jews" in 2021, the data are not identical between the two cohorts. The 2022 cohort annotated two samples for the keyword Jews, one from 2020 and the other from 2021, while the 2023 cohort annotated samples from 2021 and 2022.The 2021 sample for the keyword "Jews" that the 2022 cohort annotated was not representative. It has only 453 tweets from 2021 and 147 from the first eight months of 2022, and it includes some tweets from the query with the keyword "Israel". The 2021 sample for the keyword "Jews" that the 2023 cohort annotated was drawn proportionally for each trimester of 2021 for the keyword "Jews". </p> </div> <div> <h2> </h2> <h2>Content</h2> <h3>Cohort 2022 </h3> </div> <div> <p>This dataset contains 5880 tweets that cover a wide range of topics common in conversations about Asians, Blacks, Jews, Latines, and Muslims. 357 tweets (6.1 %) are labeled as biased and 5523 (93.9 %) are labeled as not biased. 1365 tweets (23.2 %) are labeled as calling out or denouncing bias. </p> </div> <div> <p>1180 out of 5880 tweets (20.1 %) contain the keyword "Asians," 590 were posted in 2020 and 590 in 2021. 39 tweets (3.3 %) are biased against Asian people. 370 tweets (31,4 %) call out bias against Asians. </p> </div> <div> <p>1160 out of 5880 tweets (19.7%) contain the keyword "Blacks," 578 were posted in 2020 and 582 in 2021. 101 tweets (8.7 %) are biased against Black people. 334 tweets (28.8 %) call out bias against Blacks. </p> </div> <div> <p>1189 out of 5880 tweets (20.2 %) contain the keyword "Jews," 592 were posted in 2020, 451 in 2021, and ––as mentioned above––146 tweets from 2022. 83 tweets (7 %) are biased against Jewish people. 220 tweets (18.5 %) call out bias against Jews. </p> </div> <div> <p>1169 out of 5880 tweets (19.9 %) contain the keyword "Latinos," 584 were posted in 2020 and 585 in 2021. 29 tweets (2.5 %) are biased against Latines. 181 tweets (15.5 %) call out bias against Latines. </p> </div> <div> <p>1182 out of 5880 tweets (20.1 %) contain the keyword "Muslims," 593 were posted in 2020 and 589 in 2021. 105 tweets (8.9 %) are biased against Muslims. 260 tweets (22 %) call out bias against Muslims. </p> </div> <div> <h3>Cohort 2023 </h3> </div> <div> <p>The dataset contains 5363 tweets with the keywords “Asians, Blacks, Jews, Latinos and Muslims” from 2021 and 2022. 261 tweets (4.9 %) are labeled as biased, and 5102 tweets (95.1 %) were labeled as not biased. 975 tweets (18.1 %) were labeled as calling out or denouncing bias. </p> </div> <div> <p>1068 out of 5363 tweets (19.9 %) contain the keyword "Asians," 559 were posted in 2021 and 509 in 2022. 42 tweets (3.9 %) are biased against Asian people. 280 tweets (26.2 %) call out bias against Asians. </p> </div> <div> <p>1130 out of 5363 tweets (21.1 %) contain the keyword "Blacks," 586 were posted in 2021 and 544 in 2022. 76 tweets (6.7 %) are biased against Black people. 146 tweets (12.9 %) call out bias against Blacks. </p> </div> <div> <p>971 out of 5363 tweets (18.1 %) contain the keyword "Jews," 460 were posted in 2021 and 511 in 2022. 49 tweets (5 %) are biased against Jewish people. 201 tweets (20.7 %) call out bias against Jews. </p> </div> <div> <p>1072 out of 5363 tweets (19.9 %) contain the keyword "Latinos," 583 were posted in 2021 and 489 in 2022. 32 tweets (2.9 %) are biased against Latines. 108 tweets (10.1 %) call out bias against Latines. </p> </div> <div> <p>1122 out of 5363 tweets (20.9 %) contain the keyword "Muslims," 576 were posted in 2021 and 546 in 2022. 62 tweets (5.5 %) are biased against Muslims. 240 tweets (21.3 %) call out bias against Muslims. </p> </div> <div> <h2> </h2> <h2>File Description</h2> </div> <div> <p>The dataset is provided in a csv file format, with each row representing a single message, including replies, quotes, and retweets. The file contains the following columns: </p> <p>'TweetID': Represents the tweet ID. </p> </div> <div> <p>'Username': Represents the username who published the tweet (if it is a retweet, it will be the user who retweetet the original tweet. </p> </div> <div> <p>'Text': Represents the full text of the tweet (not pre-processed). </p> </div> <div> <p>'CreateDate': Represents the date the tweet was created. </p> </div> <div> <p>'Biased': Represents the labeled by our annotators if the tweet is biased (1) or not (0). </p> </div> <div> <p>'Calling_Out': Represents the label by our annotators if the tweet is calling out bias against minority groups (1) or not (0). </p> </div> <div> <p>'Keyword': Represents the keyword that was used in the query. The keyword can be in the text, including mentioned names, or the username. </p> </div> <div> <p> ‘Cohort’: Represents the year the data was annotated (class of 2022 or class of 2023) </p> </div> <div> <h2> </h2> <h2>Acknowledgements </h2> </div> <div> <p>We are grateful for the technical collaboration with Indiana University's Observatory on Social Media (OSoMe). We thank all class participants for the annotations and contributions, including Kate Baba, Eleni Ballis, Garrett Banuelos, Savannah Benjamin, Luke Bianco, Zoe Bogan, Elisha S. Breton, Aidan Calderaro, Anaye Caldron, Olivia Cozzi, Daj Crisler, Jenna Eidson, Ella Fanning, Victoria Ford, Jess Gruettner, Ronan Hancock, Isabel Hawes, Brennan Hensler, Kyra Horton, Maxwell Idczak, Sanjana Iyer, Jacob Joffe, Katie Johnson, Allison Jones, Kassidy Keltner, Sophia Knoll, Jillian Kolesky, Emily Lowrey, Rachael Morara, Benjamin Nadolne, Rachel Neglia, Seungmin Oh, Kirsten Pecsenye, Sophia Perkovich, Joey Philpott, Katelin Ray, Kaleb Samuels, Chloe Sherman, Rachel Weber, Molly Winkeljohn, Ally Wolfgang, Rowan Wolke, Michael Wong, Jane Woods, Kaleb Woodworth, Aurora Young, Sydney Allen, Hundre Askie, Norah Bardol, Olivia Baren, Samuel Barth, Emma Bender, Noam Biron, Kendyl Bond, Graham Brumley, Kennedi Bruns, Leah Burger, Hannah Busche, Morgan Butrum-Griffith, Zoe Catlin, Angeli Cauley, Nathalya Chavez Medrano, Mia Cooper, Suhani Desai, Isabella Flick, Samantha Garcez, Isabella Grady, Macy Hutchinson, Sarah Kirkman, Ella Leitner, Elle Marquardt, Madison Moss, Ethan Nixdorf, Reya Patel, Mickey Racenstein, Kennedy Rehklau, Grace Roggeman, Jack Rossell, Madeline Rubin, Fernando Sanchez, Hayden Sawyer, Diego Scheker, Lily Schwecke, Brooke Scott, Megan Scott, Samantha Secchi, Jolie Segal, Katherine Smith, Constantine Stefanidis, Cami Stetler, Madisyn West, Alivia Yusefzadeh, Tayssir Aminou, Karen Fecht, Luciana Orrego-Hoyos, Hannah Pickett, and Sophia Tracy. </p> </div> <div> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296. </p> </div> <div> <p> </p> </div>
Calibration-free reaction yield quantification by HPLC with a machine-learning model of extinction coefficients
<p>This repository contains all the data and code associated with the manuscript "Calibration-free reaction yield quantification by HPLC with a machine-learning model of extinction coefficients"</p> <p>Mass spec and absorption chromatogram data are in reaction_set_1.zip, reaction_set_2.zip, and simulated_reaction_set.zip. The chemprop model trained on the Deep4Chem dataset is in Deep4Chem_chemprop.zip.</p>
Kinodata-3D: an in silico kinase-ligand complex dataset for kinase-focused machine learning.
<p><strong>Project Description</strong></p> <p>Drug discovery pipelines nowadays rely on machine learning models to explore and evaluate large chemical spaces. While the inclusion of 3D complex information is considered to be beneficial, structural ML for affinity prediction suffers from data scarcity. <br>We provide kinodata-3D, a dataset of <strong>~138 000</strong> docked complexes to enable more robust training of 3D-based ML models for kinase activity prediction (see <a href="https://github.com/volkamerlab/kinodata-3D-affinity-prediction">github.com/volkamerlab/kinodata-3D-affinity-prediction</a>).</p> <h2>Dataset</h2> <h3>1. Data</h3> <p>This data set consists of three-dimensional protein-ligand complexes that were generated using computational docking from the OpenEye toolkit. The modeled proteins cover the kinase family for which a fair amount of structural data, i.e. co-crystallized protein-ligand complexes in the PDB, enriched through KLIFS annotations, is available. This enables us to use template docking (OpenEye’s POSIT functionality) in which the ligand placement is guided according to a similar co-crystallized ligand pose. The kinase-ligand pairs to dock are sourced from binding assay data via the public ChEMBL archive, version 33. In particular, we use kinase activity data as curated through the <a href="https://github.com/openkinome/kinodata">OpenKinome kinodata</a> project. The final protein-ligand complexes are annotated with a predicted RMSD of the docked poses. The RMSD model is a simple neural network trained on a <a href="https://github.com/openkinome/kinase-docking-benchmark">kinase-docking benchmark</a> data set using ligand (fingerprint) similarity, docking score (ChemGauss 4), and Posit probability (see <a href="https://github.com/volkamerlab/kinodata-3D" target="_blank" rel="noopener">kinodata-3D repository</a>).</p> <p>The final data set contains in total <strong>138 286</strong> deduplicated kinase-ligand pairs, covering <strong>~98 000</strong> distinct compounds and ~<strong>271</strong> distinct kinase structures.</p> <h3>2. File structure</h3> <p>The archive <strong>kinodata_3d.zip </strong>uses the following file structure</p> <blockquote> <p>data/raw<br> | kinodata_docked_with_rmsd.sdf.gz<br> | pocket_sequences.csv<br> | mol2/pocket<br> | 1_pocket.mol2<br> | ...</p> </blockquote> <p>The file <strong>kinodata_docked_with_rmsd.sdf.gz</strong> contains the docked ligand poses and the information on the protein-ligand pair inherited from <em>kinodata</em>. The protein pockets located in <strong>mol2/pocket</strong> are stored according to the MOL2 file format.</p> <p>The pocket structures were sourced from KLIFS (<a href="https://klifs.net" target="_blank" rel="noopener">klifs.net)</a> and complete the poses in the aforementioned SDF file. The files are named <strong>{klifs_structure_id}_pocket.mol2</strong>. The structure ID is given in the SDF file along with the ligand poses.</p> <p>The file <strong>pocket_sequences.csv </strong>contains all KLIFS pocket sequences relevant to the kinodata-3D dataset.</p> <h3>3. Related code</h3> <p>The code used to create the poses can be found in the <a href="https://github.com/volkamerlab/kinodata-3D" target="_blank" rel="noopener">kinodata-3D repository</a>. The docking pipeline makes heavy use of the <a href="https://github.com/openkinome/kinoml" target="_blank" rel="noopener">kinoml</a> framework, which in turn uses <a href="https://www.eyesopen.com" target="_blank" rel="noopener">OpenEye's</a> Posit template docking implementation. The details of the original pipeline can also be found in the manuscript by <a href="https://www.biorxiv.org/content/10.1101/2023.09.11.557138v1">Schaller et al. (<strong>2023</strong>). Benchmarking Cross-Docking Strategies for Structure-Informed Machine Learning in Kinase Drug Discovery. <em>bioRxiv</em>.</a></p>
Replication Data for the paper "Predicting Food-Security Crises in the Horn of Africa Using Machine Learning"
<p>This folder contains all input data necessary to run the machine learning model as described in the paper "Predicting Food Security Crises in the Horn of Africa Using Machine Learning". </p><p>This model is developed at the Institute for Environmental Studies, Vrije Universiteit Amsterdam. </p><p>Questions or remarks can be send to tim.busker@vu.nl</p>
Data for: Morphological species delimitation in the Western Pond Turtle (Actinemys): Can machine learning methods aid in cryptic species identification?
<p>As the discovery of cryptic species has increased in frequency, there has been interest in whether geometric morphometric data can detect fine-scale patterns of variation that can be used to morphologically diagnose such species. We used a combination of geometric morphometric data and an ensemble of five supervised machine learning methods to investigate whether plastron shape can differentiate two putative cryptic turtle species, <em>Actinemys marmorata</em> and <em>Actinemys pallida</em>. <em>Actinemys</em> has been the focus of considerable research due to its biogeographic distribution and conservation status. Despite this work, reliable morphological diagnoses for its two species are still lacking. We validated our approach on two datasets, one consisting of eight morphologically disparate emydid species, and the other consisting of two subspecies of <em>Trachemys</em> (<em>T. scripta scripta</em>, <em>T. scripta elegans</em>). The validation tests returned near-perfect classification rates, demonstrating that plastron shape is an effective means for distinguishing taxonomic groups of emydids via machine learning methods. By contrast, the same methods did not return high classification rates for a set of alternative phylogeographic and morphological binning schemes in <em>Actinemys</em>. All classification hypotheses performed poorly relative to the validation datasets and no single hypothesis was unequivocally supported for <em>Actinemys</em>. Two hypotheses had machine learning performance that was marginally better than our remaining hypotheses. In both cases, those hypotheses favored a two-species split between <em>A. marmorata</em> and <em>A. pallida</em> specimens, lending tentative morphological support to the hypothesis of two <em>Actinemys</em> species. However, the machine learning results also underscore that <em>Actinemys</em> as a whole have lower levels of plastral variation than other turtles within Emydidae, but the reason for this morphological conservatism is unclear.</p>
Datasets for "Machine-Learning-Enhanced Symbolic Regression for Methane Storage Prediction in Covalent Organic Frameworks"
<p>This collection contains the datasets and associated files used in the research presented in the manuscript titled "Machine Learning-Enhanced Symbolic Regression for Methane Storage Prediction in Covalent Organic Frameworks". The datasets are critical for the development and validation of machine learning and symbolic regression models aiming to predict methane storage capacities in covalent organic frameworks (COFs).</p> <p><strong>Included Datasets:</strong></p> <ol> <li><code>COF_Data_for_ML.csv</code>: This dataset was utilized for the development of machine learning models.</li> <li><code>COF_Data_for_SISSO.csv</code>: This dataset was employed for the development of SISSO-based symbolic regression models.</li> <li><code>ML_vs_GCMC.xlsx</code>: This comparative dataset features GCMC-calculated results alongside machine learning predictions.</li> <li><code>Feature_Combination.xlsx</code>: This file contains data detailing all the feature combinations explored in the study.</li> <li><code>ML_SISSO_GCMC.xlsx</code>: This comparative dataset includes GCMC calculations, SISSO-based symbolic regression model predictions, and ML predictions.</li> <li><code>Crystallographic_Properties_of_535k_COFs.xlsx</code>: This consolidated dataset presents the crystallographic properties of 535,293 COFs.</li> </ol> <p><strong>Software Used:</strong></p> <ul> <li>Machine Learning Computations: Scikit-Learn (<a href="https://scikit-learn.org/stable/" target="_new">https://scikit-learn.org/stable/</a>)</li> <li>GCMC Simulations: RASPA2 (<a href="https://github.com/iRASPA/RASPA2" target="_new">https://github.com/iRASPA/RASPA2</a>)</li> <li>SISSO Calculations: SISSO toolkit (<a href="https://github.com/rouyang2017/SISSO" target="_new">https://github.com/rouyang2017/SISSO</a>)</li> <li>Crystallographic property calculations: Zeo++ (<a href="https://www.zeoplusplus.org/" target="_new">https://www.zeoplusplus.org/</a>)</li> </ul> <p>The datasets are provided to enable replication of the study's findings, encourage further research in the field, and facilitate the development of advanced predictive models by the scientific community. Researchers who use these datasets are requested to cite this Zenodo entry as well as the associated paper upon its publication.</p>
Data for Identifying Knot Types of Polymer Conformations by Machine Learning
<h1>Training Data and Generalizability Testsets</h1> <h2>For the publication PhysRevE.101.022502</h2> <pre>@article{PhysRevE.101.022502, title = {Identifying knot types of polymer conformations by machine learning}, author = {Vandans, Olafs and Yang, Kaiyuan and Wu, Zhongtao and Dai, Liang}, journal = {Phys. Rev. E}, volume = {101}, issue = {2}, pages = {022502}, numpages = {10}, year = {2020}, month = {Feb}, publisher = {American Physical Society}, doi = {10.1103/PhysRevE.101.022502}, url = {https://link.aps.org/doi/10.1103/PhysRevE.101.022502} }</pre> <h2>GitHub source code demo using this dataset: </h2> <p><a href="https://github.com/CompSoftMatterBiophysics-CityU-HK/Identify-Knot-Types-by-ML-PRE2020"><strong>🥨 https://github.com/CompSoftMatterBiophysics-CityU-HK/Identify-Knot-Types-by-ML-PRE2020</strong></a></p> <p><strong>The above GitHub repo provide <strong>a docker, training code, best model with weights, and two showcases of generalizability</strong>.</strong></p>
Dataset related to the article titled "Prediction of elastic modulus of basaltic rocks using machine learning methods."
<p>Dataset related to the article titled "Prediction of elastic modulus of basaltic rocks using machine learning methods."</p>
Small dataset machine-learning approach for efficient design space exploration: engineering ZnTe-based high-entropy alloys for water splitting
<p>Atomic structure data used in the research article entitled "Small Dataset Machine-Learning Approaches to Explore the Design Space of High-Entropy Alloys: Engineering ZnTe-based Multicomponent Alloys for the Photo-Splitting of Water"</p>
Nanofluid heat transfer and machine learning
<p>Table 1. Machine learning application for nanofluids in porous media.</p> <p>Table 2. Summary of machine learning application: Nanofluids in heat exchangers</p>
PERFORMANCE OF MACHINE LEARNING ALGORITHMS FOR LUNG CANCER PREDICTION: A COMPARATIVE STUDY
<p>This study compares the performance of five machine learning algorithms—logistic regression, support vector machines, random forests, gradient boosting, and neural networks—for lung cancer prediction using demographic, lifestyle, and medical data from the UCI Machine Learning Repository. Gradient boosting and random forests achieved the highest accuracy (89% and 87%, respectively) and AUC-ROC scores (0.93 and 0.92), while neural networks reached 90% accuracy but presented interpretability limitations. Key predictors included smoking history, chronic disease, and respiratory symptoms, aligning with established risk factors. Ensemble methods, particularly gradient boosting and random forests, provided an optimal balance of accuracy and interpretability, highlighting their potential for clinical applications in early lung cancer detection.</p>
Advancing shrub dendroecology: a cutting-edge machine learning method for measuring shrub rings
<p>This dataset contains the original data used for the manuscript "<span>Advancing shrub dendroecology: a cutting-edge machine learning method for measuring shrub rings". Image labels is structured as follows: </span></p> <p><span>Site code - Species - sample number </span></p> <p><span>Sites: </span></p> <ul> <li><span>F stands for Finse</span></li> <li><span>A stands for Abisko</span></li> </ul> <p><span>Species :</span></p> <ul> <li><span>DO stands fro Dryas octopetala </span></li> <li><span>EH for Empetrum hermaphroditum</span></li> </ul>
Accessibility Rank: A Machine Learning Approach for Prioritising Accessibility User Feedback
<p>This repository serves as a comprehensive collection of datasets, code scripts, and associated data used in my master's research conducted at the University of Auckland on accessibility-related reviews. The research findings and methodology are described in detail in our paper titled "Accessibility Rank: A Machine Learning Approach for Prioritising Accessibility User Feedback". By making these resources openly available, we aim to foster collaboration, reproducibility, and advancement in the field of accessibility research. Researchers and developers can leverage these datasets, associated data, and code scripts to gain insights, validate findings, and explore novel approaches to addressing accessibility challenges.</p> <p>We encourage users to refer to our paper for a comprehensive understanding of our research methodology, experimental setup, and results. Proper attribution and citation of our paper are appreciated when utilizing any part of this repository in further research or publications.</p>
Architectural Design Decisions for the Machine Learning Workflow: Dataset and Code
<p><strong>Title:</strong> Architectural Design Decisions for the Machine Learning Workflow: Dataset and Code</p> <p><strong>Authors:</strong> Stephen John Warnett; Uwe Zdun</p> <p><strong>About:</strong> This is the dataset and code artifact for the article entitled "Architectural Design Decisions for the Machine Learning Workflow".</p> <p><strong>Contents:</strong> The "_generated" directory contains the generated results, including latex files with tables for use in publications and the Architectural Design Decision model in textual and graphical form. "Generators" contains Python applications that can be run to generate the above. "Metamodels" contains a Python file with type definitions. "Sources_coding" contains our source codings and audit trail. "Add_models" contains the Python implementation of our model and source codings. Finally, "appendix" contains a detailed description of our research method.</p> <p><strong>Article Abstract: </strong>Bringing machine learning models to production is challenging as it is often fraught with uncertainty and confusion, partially due to the disparity between software engineering and machine learning practices, but also due to knowledge gaps on the level of the individual practitioner. We conducted a qualitative investigation into the architectural decisions faced by practitioners as documented in gray literature based on Straussian Grounded Theory and modeled current practices in machine learning. Our novel Architectural Design Decision model is based on current practitioner understanding of the topic and helps bridge the gap between science and practice, foster scientific understanding of the subject, and support practitioners via the integration and consolidation of the myriad decisions they face. We describe a subset of the Architectural Design Decisions that were modeled, discuss uses for the model, and outline areas in which further research may be pursued.</p> <p><strong>Objective:</strong> This article aims to study current practitioner understanding of architectural concepts associated with data processing, model building, and Automated Machine Learning (AutoML) within the context of the machine learning workflow.</p> <p><strong>Method:</strong> Applying Straussian Grounded Theory to gray literature sources containing practitioner views on machine learning practices, we studied methods and techniques currently applied by practitioners in the context of machine learning solution development and gained valuable insights into the software engineering and architectural state of the art as applied to ML.</p> <p><strong>Results:</strong> Our study resulted in a model of Architectural Design Decisions, practitioner practices, and decision drivers in the field of software engineering and software architecture for machine learning.</p> <p><strong>Conclusions:</strong> The resulting Architectural Design Decisions model can help researchers better understand practitioners' needs and the challenges they face, and guide their decisions based on existing practices. The study also opens new avenues for further research in the field, and the design guidance provided by our model can also help reduce design effort and risk. In future work, we plan on using our findings to provide automated design advice to machine learning engineers.</p>
Datasets for "Unexplored Antarctic meteorite collection sites revealed through machine learning"
<p>This archive provides datasets related to the following publication:</p> <p>V. Tollenaar, H. Zekollari, S. Lhermitte, D. Tax, V. Debaille, S. Goderis, P. Claeys, F. Pattyn, Unexplored Antarctic meteorite collection sites revealed through machine learning. Science Advances 8, eabj8138 (2022). <a href="https://doi.org/10.1126/sciadv.abj8138">DOI: 10.1126/sciadv.abj8138</a></p> <p>Contact: Veronica Tollenaar, Veronica.Tollenaar@ulb.be</p> <p>Users should cite the original publication when using all or part of the data. </p> <p>About the datasets: it includes a shapefile with the outline of the 613 Meteorite Stranding Zones (Fig. 7, "613MSZs.zip"), the observations used for classification, and the continent-wide probability to find meteorites (at 450-meter resolution, Fig. 5, "positive_classified.nc"). References to the literature are provided in the corresponding publication. Meteorite locations are based on the Meteoritical Bulletin Database (available at https://www.lpi.usra.edu/meteor/).</p> <p>- bias_above200m1kmbuff_expanded_dissolved: shapefile of polygons of unlabelled observations<br> - meteorite_locations_raw.csv: contains locations of meteorite finds as defined in the meteoritical bulletin consulted on 05/07/2019<br> - meteorite_types.csv: contains meteorite names and types as defined in the meteoritical bulletin consulted on 05/07/2019<br> - validation_neg.csv: contains locations of negative observations used for validation<br> - TEST_neg.csv: contains locations of negative test observations<br> - TEST_pos.csv: contains locations of positive test obesrvations<br> - MSZs_ranked: shapefile of ranked meteorite stranding zones<br> - Test_neg4326: shapefile of locations used as negative test data<br> - Cal_neg4326: shapefile of locations used as negative calibration/validation data<br> - TestMSZs_pos4326: shapefile of locations used as positive test data in MSZ-level assesment<br> - 613MSZs: shapefile of outlines of meteorite stranding zones<br> - positive_classified.nc: netcdf of positive classified observations with their estimated a posteriori probabilities</p>
Global mapping of lunar refractory elements: multivariate regression vs. machine learning
<p>The quantitative estimation of elemental concentrations at the spatial resolution of hyperspectral near-infrared (NIR) images<br> of the lunar surface is an important tool for understanding the processes relevant for the origin and evolution of the Moon. The NIR reflectance of the lunar regolith is an integrated response to the presence of refractory elements and soil alteration processes. Our approach was to define a combination of spectral parameters that are robust with respect to the effects of soil maturity.<br> We calibrated the spectral parameters with respect to elemental abundances measured by the Lunar Prospector Gamma Ray Spectrometer (LP GRS) and the Kaguya GRS (KGRS). For this purpose, we compared a classical multivariate linear regression (MLR) approach and the machine learning based support vector regression (SVR) technique applied to M3 global observations. The M 3 -based global elemental maps are consistent in distribution and range with the LP GRS and KGRS elemental maps<br> and do not show artifacts in immature areas such as small fresh craters. The results derived using MLR and SVR are compared to<br> sample-based ground truth data of the Apollo and Luna sample-return sites, where the root-mean-square deviations obtained by the<br> two regression models are similar. The main advantage of the proposed new algorithm is its ability to minimize artifacts due to space-weathering effects. The elemental maps of Mg and Ca provide additional information and reveal structures not always visible in the Fe map. The global elemental abundance maps derived for the fully calibrated M 3 observations might thus serve as important tools to investigate the lunar geology and evolution.</p>
Global estimates of marine gross primary production based on machine‐learning upscaling of field observations
<p>4 variables (excluding dimension variables):</p> <p>double GPP_LD_MLD_RF[Lon,Lat,Month] <br> units: mmol O2 m-2 d-1<br> fill value: NaN<br> long_name: Monthly mixed-layer integration of gross primary production trained from the<br> dataset determined by the light-dark bottle incubation using Random Forest<br> algorithm<br> coordinates: [Longitude, Latitude Month]</p> <p>double GPP_LD_ZEU_RF[Lon,Lat,Month] </p> <p> units: mmol O2 m-2 d-1<br> fill value: NaN<br> long_name: Monthly euphotic-zone integration of gross primary production trained from<br> the dataset determined by the light-dark bottle incubation using Random<br> Forest algorithm<br> coordinates: [Longitude, Latitude Month]</p> <p>double GPP_Triple_MLD_RF[Lon,Lat,Month] <br> units: mmol mmol O2 m-2 d-1<br> fillvalue: NaN<br> long_name: Monthly mixed-layer integration of gross primary production trained from<br> the dataset determined by the triple isotopes of dissolved oxygen using<br> Random Forest algorithm<br> coordinates: [Longitude, Latitude Month]<br> <br> double GPP_Triple_ZEU_RF[Lon,Lat,Month] <br> units: mmol O2 m-2 d-1<br> fill value: NaN<br> long_name: Monthly euphotic-zone integration of gross primary production trained from<br> the dataset determined by the triple isotopes of dissolved oxygen using<br> Random Forest algorithm</p> <p>3 dimensions:</p> <p> Lon Size:181<br> units: degree_north<br> long_name: Longitude</p> <p> Lat Size:91<br> units: degree_east<br> long_name: Latitude</p> <p> Month Size:13<br> units: Jan, Feb, Mar, Apr, May, Jun, Jul, Aug, Sep, Oct, Nov, Dec, Annuual_mean<br> long_name: Month</p> <p><br> Author: Yibin Huang & Nicolas Cassar<br> Correspond: nicolas.cassar@duke.edu<br> <br> Request_for_citation: If you use these data in publications or presentations, please cite: Huang,<br> Y., Nicholson, D., Huang, B., & Cassar, N. (2021). Global estimates of<br> marine gross primary production based on machine‐learning upscaling of<br> field observations. Global Biogeochemical Cycles, 35, e2020GB006718.<br> https://doi.org/10.1029/2020GB006718<br> <br> Creation date: Dec/6th/2021</p>
Machine Learning Dataset for Poultry Diseases Diagnostics - PCR annotated
<p>The dataset of poultry disease diagnostics was annotated using Polymerase Chain Reaction (PCR). Polymerase Chain Reaction (PCR) is a molecular biology technique for rapid diagnostics. We gathered both the fecal images and fecal samples from layers, cross and indigenous breeds of chicken from poultry farms in Arusha and Kilimanjaro regions in Tanzania between September 2020 and February 2021. Each fecal sample collected was coded to its corresponding image during data collection. PCR method is used for detection and identification of pathogens through amplification of DNA sequences unique to the pathogen. We used existing primers from literature to amplify the target DNA/RNA on the poultry fecal samples for PCR. The targets were Coccidiosis, Newcastle disease and Salmonella. We used the primers for PCR diagnostics at the molecular laboratory of the Nelson Mandela African Institution of Science and Technology (NM-AIST). The fecal samples were stored at -80 degrees celsius. The PCR diagnostics were conducted using reagents and kits from Zymo Research and the protocol is summarized in these five stages: 1. DNA sample loading 2. DNA extraction 3. Amplification; 4. Quantification and 5. Detection.</p> <p>All the PCR annotated fecal images are in the <strong><strong>.zip files</strong></strong>; “pcrcocci.zip” has 373 images, “pcrhealthy.zip” has 347 images, “pcrsalmo.zip” has 349 images, "pcrncd.zip" has 186 images. A total of 1,255 image files are labeled.</p> <p>The research project is funded by the Organization for Women in Science for the Developing World (OWSD) with Grant Award Number: 4500406715.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.