Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
27
datasets available to search
ShareScore release 0.9.0
Dataset results
27 results for “Learning analytics”
Attributes: A Curriculum Analytics System for measuring learning outcomes - Overview
<p><span><strong>Link to video </strong><a href="https://vimeo.com/1015456231?share=copy#t=0"><strong>https://vimeo.com/1015456231?share=copy - t=0</strong></a><br><br>The Curriculum Analytics System at the Instituto Tecnológico de Costa Rica, integrated into TEC Digital, supports faculty, coordinators, and students in assessing engineering learning outcomes during accreditation processes. The system offers two key user modules: one for coordinators to map and manage learning outcomes, and another for instructors to conduct assessments through the course portal. Coordinators oversee course and attribute mapping using visual representations of study plans, control points, and outcome visualizations. Instructors configure assignments and evaluate student submissions with standardized rating scales. The system tracks progress in real-time and generates</span> <span>comprehensive reports with performance metrics, facilitating continuous improvement in academic programs.</span></p> <p><strong><span>Key words: </span></strong><span>attributes, learning outcomes, TEC Digital, curriculum analytics, continuous improvement. </span></p>
Bayesian Symbolic Learning to Build Analytical Correlations from Rigorous Process Simulations: Application to CO2 Capture Technologies
<p>Dataset of process simulations results of the natural gas sweetening and flue gas treatment (first and second sheet, respectively as indicated by the sheet name in the .xlsx file). The dataset refers to the publication <em>Bayesian Symbolic Learning to Build Analytical Correlations from Rigorous Process Simulations: Application to CO<sub>2</sub> Capture Technologies </em>by V. Negri, Vàzquey D., Sales-Pardo, Marta, Guimerà, R. and Guillén-Gosàlbez, G. The training and testing dataset are used to generate the figures in the main manuscript and supplementary information. </p> <p> </p>
Antisemitism on Twitter: A Dataset for Machine Learning and Text Analytics
<h1><strong><span><span>Dataset from the Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University: </span></span></strong></h1> <p> </p> <div> <div> <p><span><span>The </span><span>Social Media</span><span> & Hate research lab at the Institute for the Study of Contemporary Antisemitism compiled this dataset using an annotation portal (Jikeli, Soemer, and Karali 2024), which was used to label tweets as either antisemitic or non-antisemitic, among other labels. Note that annotation was done on live data, including images and context, such as threads. All data was annotated by two experts, and all discrepancies were discussed</span><span> (Jikeli et al. 2023)</span><span>.</span></span><span> </span></p> </div> </div> <p><br><strong>Content: </strong></p> <p><span><span>This dataset </span><span>contains</span> <span>1</span><span>1</span><span>311</span><span> tweets </span><span>covering</span><span> a wide range of topics common in conversations about Jews, Israel, and antisemitism between January 2019 and </span><span>April 2023</span><span>. </span><span>The dataset consists of random samples of relevant keywords during this </span><span>time period</span><span>.</span><span> 1,</span><span>953</span><span> tweets (1</span><span>7</span><span>%) </span><span>are antisemitic </span><span>according to </span><span>the IHRA definition of antisemitism.</span><span> </span></span><span> </span></p> <div> <p><span><span>The distribution of tweets by year is as follows:</span><span> 1499 (</span><span>13</span><span>%) from 2019, 371</span><span>2</span><span> (</span><span>33</span><span>%) from 2020, </span><span>2591</span><span> (2</span><span>3</span><span>%) from 2021</span><span>, 2644 from 2022 </span><span>(23%)</span> <span>and 865 </span><span>(8%)</span> <span>f</span><span>rom 2023</span><span>. </span><span>6365</span><span> (</span><span>56</span><span>%) </span><span>contain</span><span> the keyword "Jews,"</span><span> 4134 </span><span>(</span><span>3</span><span>7</span><span>%) include "Israel," 529 (</span><span>5</span><span>%) feature the derogatory term "</span><span>ZioNazi</span><span>*," and 283 (</span><span>3</span><span>%) use the slur "K---s." Some tweets may </span><span>contain</span><span> multiple keywords. </span></span><span> </span></p> </div> <div> <p><span><span>725</span><span> out of the </span><span>6365</span><span> tweets with the keyword "Jews" (11%) and </span><span>664</span><span> out of the </span><span>4134</span><span> tweets with the keyword "Israel" (1</span><span>6</span><span>%) were classified as antisemitic. 97 out of the 283 tweets using the antisemitic slur "K---s" (34%) are antisemitic.</span> <span>Interestingly, many tweets featuring the slur "K---s" actually </span><span>call out</span><span> its u</span><span>s</span><span>e.</span><span> In contrast, </span><span>the majority of</span><span> tweets </span><span>using</span><span> the derogatory term "</span><span>ZioNazi</span><span>*" are antisemitic, with 467 out of 529 (88%) being classified as such. </span></span><span> </span></p> </div> <p> </p> <p><strong>File Description: </strong></p> <div> <div> <p><span><span>The dataset is provided in a csv file format, with each row </span><span>representing</span><span> a single message, including replies, quotes, and retweets. The file </span><span>contains</span><span> the following columns: </span></span><span> </span></p> </div> <div> <p><span><span> </span></span><span><span> </span><br></span><span><span>‘ID’:</span> <span>Represents</span><span> the tweet ID. </span></span><span> </span></p> </div> <div> <p><span><span>‘Username’: </span><span>Represents</span><span> the username </span><span>that posted </span><span>the tweet</span><span>. </span></span><span> </span></p> </div> <div> <p><span><span>‘Text’: </span><span>Represents</span><span> the full text of the tweet (not pre-processed).</span></span><span> </span></p> </div> <div> <p><span><span>‘</span><span>CreateDate</span><span>’: </span><span>Represents</span><span> the date </span><span>on which </span><span>the tweet was created</span><span>. </span></span><span> </span></p> </div> <div> <p><span><span>‘Biased’: </span><span>Represents</span><span> the label given by our annotations as to whether the tweet </span><span>is antisemitic or no</span><span>t</span><span>.</span></span><span> </span></p> </div> <div> <p><span><span>‘Keyword’: </span><span>Represents</span><span> the keyword that was used in the query. The keyword can be in the text, including </span><span>hashtags, </span><span>mentioned </span><span>users</span><span>, or the username</span><span> itself.</span><span> </span></span><span> </span></p> </div> </div> <p> </p> <p>Licences </p> <p>Data is published under the terms of the "Creative Commons Attribution 4.0 International" licence (https://creativecommons.org/licenses/by/4.0) </p> <p> </p> <p>Acknowledgements </p> <p>We are grateful for the support of Indiana University’s Observatory on Social Media (OSoMe) (Davis et al. 2016) and the contributions and annotations of all team members in our Social Media & Hate Research Lab at Indiana University’s Institute for the Study of Contemporary Antisemitism, especially Grace Bland, Elisha S. Breton, Kathryn Cooper, Robin Forstenhäusler, Sophie von Máriássy, Mabel Poindexter, Jenna Solomon, Clara Schilling, and Victor Tschiskale. </p> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.</p>
Hate Speech and Bias against Asians, Blacks, Jews, Latines, and Muslims: A Dataset for Machine Learning and Text Analytics
<h1>Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University Dataset on bias against Asians, Blacks, Jews, Latines, and Muslims </h1> <div> <h2> </h2> <h2>Description </h2> </div> <div> <p>The dataset is a product of a research project at Indiana University on biased messages on Twitter against ethnic and religious minorities. We scraped all live messages with the keywords "Asians, Blacks, Jews, Latinos, and Muslims" from the Twitter archive in 2020, 2021, and 2022.</p> <p>Random samples of 600 tweets were created for each keyword and year, including retweets. The samples were annotated in subsamples of 100 tweets by undergraduate students in Professor Gunther Jikeli's class 'Researching White Supremacism and Antisemitism on Social Media' in the fall of 2022 and 2023. A total of 120 students participated in 2022. They annotated datasets from 2020 and 2021. 134 students participated in 2023. They annotated datasets from the years 2021 and 2022. The annotation was done using the <a href="https://annotationportal.com/" target="_blank" rel="noreferrer noopener">Annotation Portal</a> (Jikeli, Soemer and Karali, 2024). The updated version of our portal, <a href="https://portal2.annotationportal.com/" target="_blank" rel="noreferrer noopener">AnnotHate</a>, is now publicly available. Each subsample was annotated by an average of 5.65 students per sample in 2022 and 8.32 students per sample in 2023, with a range of three to ten and three to thirteen students, respectively. Annotation included questions about bias and calling out bias. </p> </div> <div> <p>Annotators used a scale from 1 to 5 on the bias scale (confident not biased, probably not biased, don't know, probably biased, confident biased), using definitions of bias against each ethnic or religious group that can be found in the research reports from <a href="https://isca.indiana.edu/publication-research/social-media-project/Research-Report-BIAS-on-Twitter-against-Asians--Blacks-Jews-Latinos-Muslims-final-002.pdf" target="_blank" rel="noreferrer noopener">2022</a> and <a href="https://isca.indiana.edu/documents/BIAS%20Against%20Asian-Black-Hispanic-Jewish-and-%20Muslim-People%20on%20X-Twitter%20in%202021%20and%202022.pdf" target="_blank" rel="noreferrer noopener">2023</a>. If the annotators interpreted a message as biased according to the definition, they were instructed to choose the specific stereotype from the definition that was most applicable. Tweets that denounced bias against a minority were labeled as "calling out bias". </p> </div> <div> <p>The label was determined by a 75% majority vote. We classified “probably biased” and “confident biased” as biased, and “confident not biased,” “probably not biased,” and “don't know” as not biased. </p> </div> <div> <p>The stereotypes about the different minorities varied. About a third of all biased tweets were classified as general 'hate' towards the minority. The nature of specific stereotypes varied by group. Asians were blamed for the Covid-19 pandemic, alongside positive but harmful stereotypes about their perceived excessive privilege. Black people were associated with criminal activity and were subjected to views that portrayed them as inferior. Jews were depicted as wielding undue power and were collectively held accountable for the actions of the Israeli government. In addition, some tweets denied the Holocaust. Hispanic people/Latines faced accusations of being undocumented immigrants and "invaders," along with persistent stereotypes of them as lazy, unintelligent, or having too many children. Muslims were often collectively blamed for acts of terrorism and violence, particularly in discussions about Muslims in India. </p> </div> <div> <p>The annotation results from both cohorts (Class of 2022 and Class of 2023) will not be merged. They can be identified by the "cohort" column. While both cohorts (Class of 2022 and Class of 2023) annotated the same data from 2021,* their annotation results differ. The class of 2022 identified more tweets as biased for the keywords "Asians, Latinos, and Muslims" than the class of 2023, but nearly all of the tweets identified by the class of 2023 were also identified as biased by the class of 2022. The percentage of biased tweets with the keyword 'Blacks' remained nearly the same. </p> </div> <div> <p>*Due to a sampling error for the keyword "Jews" in 2021, the data are not identical between the two cohorts. The 2022 cohort annotated two samples for the keyword Jews, one from 2020 and the other from 2021, while the 2023 cohort annotated samples from 2021 and 2022.The 2021 sample for the keyword "Jews" that the 2022 cohort annotated was not representative. It has only 453 tweets from 2021 and 147 from the first eight months of 2022, and it includes some tweets from the query with the keyword "Israel". The 2021 sample for the keyword "Jews" that the 2023 cohort annotated was drawn proportionally for each trimester of 2021 for the keyword "Jews". </p> </div> <div> <h2> </h2> <h2>Content</h2> <h3>Cohort 2022 </h3> </div> <div> <p>This dataset contains 5880 tweets that cover a wide range of topics common in conversations about Asians, Blacks, Jews, Latines, and Muslims. 357 tweets (6.1 %) are labeled as biased and 5523 (93.9 %) are labeled as not biased. 1365 tweets (23.2 %) are labeled as calling out or denouncing bias. </p> </div> <div> <p>1180 out of 5880 tweets (20.1 %) contain the keyword "Asians," 590 were posted in 2020 and 590 in 2021. 39 tweets (3.3 %) are biased against Asian people. 370 tweets (31,4 %) call out bias against Asians. </p> </div> <div> <p>1160 out of 5880 tweets (19.7%) contain the keyword "Blacks," 578 were posted in 2020 and 582 in 2021. 101 tweets (8.7 %) are biased against Black people. 334 tweets (28.8 %) call out bias against Blacks. </p> </div> <div> <p>1189 out of 5880 tweets (20.2 %) contain the keyword "Jews," 592 were posted in 2020, 451 in 2021, and ––as mentioned above––146 tweets from 2022. 83 tweets (7 %) are biased against Jewish people. 220 tweets (18.5 %) call out bias against Jews. </p> </div> <div> <p>1169 out of 5880 tweets (19.9 %) contain the keyword "Latinos," 584 were posted in 2020 and 585 in 2021. 29 tweets (2.5 %) are biased against Latines. 181 tweets (15.5 %) call out bias against Latines. </p> </div> <div> <p>1182 out of 5880 tweets (20.1 %) contain the keyword "Muslims," 593 were posted in 2020 and 589 in 2021. 105 tweets (8.9 %) are biased against Muslims. 260 tweets (22 %) call out bias against Muslims. </p> </div> <div> <h3>Cohort 2023 </h3> </div> <div> <p>The dataset contains 5363 tweets with the keywords “Asians, Blacks, Jews, Latinos and Muslims” from 2021 and 2022. 261 tweets (4.9 %) are labeled as biased, and 5102 tweets (95.1 %) were labeled as not biased. 975 tweets (18.1 %) were labeled as calling out or denouncing bias. </p> </div> <div> <p>1068 out of 5363 tweets (19.9 %) contain the keyword "Asians," 559 were posted in 2021 and 509 in 2022. 42 tweets (3.9 %) are biased against Asian people. 280 tweets (26.2 %) call out bias against Asians. </p> </div> <div> <p>1130 out of 5363 tweets (21.1 %) contain the keyword "Blacks," 586 were posted in 2021 and 544 in 2022. 76 tweets (6.7 %) are biased against Black people. 146 tweets (12.9 %) call out bias against Blacks. </p> </div> <div> <p>971 out of 5363 tweets (18.1 %) contain the keyword "Jews," 460 were posted in 2021 and 511 in 2022. 49 tweets (5 %) are biased against Jewish people. 201 tweets (20.7 %) call out bias against Jews. </p> </div> <div> <p>1072 out of 5363 tweets (19.9 %) contain the keyword "Latinos," 583 were posted in 2021 and 489 in 2022. 32 tweets (2.9 %) are biased against Latines. 108 tweets (10.1 %) call out bias against Latines. </p> </div> <div> <p>1122 out of 5363 tweets (20.9 %) contain the keyword "Muslims," 576 were posted in 2021 and 546 in 2022. 62 tweets (5.5 %) are biased against Muslims. 240 tweets (21.3 %) call out bias against Muslims. </p> </div> <div> <h2> </h2> <h2>File Description</h2> </div> <div> <p>The dataset is provided in a csv file format, with each row representing a single message, including replies, quotes, and retweets. The file contains the following columns: </p> <p>'TweetID': Represents the tweet ID. </p> </div> <div> <p>'Username': Represents the username who published the tweet (if it is a retweet, it will be the user who retweetet the original tweet. </p> </div> <div> <p>'Text': Represents the full text of the tweet (not pre-processed). </p> </div> <div> <p>'CreateDate': Represents the date the tweet was created. </p> </div> <div> <p>'Biased': Represents the labeled by our annotators if the tweet is biased (1) or not (0). </p> </div> <div> <p>'Calling_Out': Represents the label by our annotators if the tweet is calling out bias against minority groups (1) or not (0). </p> </div> <div> <p>'Keyword': Represents the keyword that was used in the query. The keyword can be in the text, including mentioned names, or the username. </p> </div> <div> <p> ‘Cohort’: Represents the year the data was annotated (class of 2022 or class of 2023) </p> </div> <div> <h2> </h2> <h2>Acknowledgements </h2> </div> <div> <p>We are grateful for the technical collaboration with Indiana University's Observatory on Social Media (OSoMe). We thank all class participants for the annotations and contributions, including Kate Baba, Eleni Ballis, Garrett Banuelos, Savannah Benjamin, Luke Bianco, Zoe Bogan, Elisha S. Breton, Aidan Calderaro, Anaye Caldron, Olivia Cozzi, Daj Crisler, Jenna Eidson, Ella Fanning, Victoria Ford, Jess Gruettner, Ronan Hancock, Isabel Hawes, Brennan Hensler, Kyra Horton, Maxwell Idczak, Sanjana Iyer, Jacob Joffe, Katie Johnson, Allison Jones, Kassidy Keltner, Sophia Knoll, Jillian Kolesky, Emily Lowrey, Rachael Morara, Benjamin Nadolne, Rachel Neglia, Seungmin Oh, Kirsten Pecsenye, Sophia Perkovich, Joey Philpott, Katelin Ray, Kaleb Samuels, Chloe Sherman, Rachel Weber, Molly Winkeljohn, Ally Wolfgang, Rowan Wolke, Michael Wong, Jane Woods, Kaleb Woodworth, Aurora Young, Sydney Allen, Hundre Askie, Norah Bardol, Olivia Baren, Samuel Barth, Emma Bender, Noam Biron, Kendyl Bond, Graham Brumley, Kennedi Bruns, Leah Burger, Hannah Busche, Morgan Butrum-Griffith, Zoe Catlin, Angeli Cauley, Nathalya Chavez Medrano, Mia Cooper, Suhani Desai, Isabella Flick, Samantha Garcez, Isabella Grady, Macy Hutchinson, Sarah Kirkman, Ella Leitner, Elle Marquardt, Madison Moss, Ethan Nixdorf, Reya Patel, Mickey Racenstein, Kennedy Rehklau, Grace Roggeman, Jack Rossell, Madeline Rubin, Fernando Sanchez, Hayden Sawyer, Diego Scheker, Lily Schwecke, Brooke Scott, Megan Scott, Samantha Secchi, Jolie Segal, Katherine Smith, Constantine Stefanidis, Cami Stetler, Madisyn West, Alivia Yusefzadeh, Tayssir Aminou, Karen Fecht, Luciana Orrego-Hoyos, Hannah Pickett, and Sophia Tracy. </p> </div> <div> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296. </p> </div> <div> <p> </p> </div>
Dataset: Applications of Educational Data Mining and Learning Analytics on Data From Cybersecurity Training
<p>This repository contains supplementary materials for the following journal paper:</p> <p>Valdemar Švábenský, Jan Vykopal, Pavel Čeleda, Lydia Kraus.<br> <em>Applications of Educational Data Mining and Learning Analytics on Data From Cybersecurity Training.</em><br> In Springer Education and Information Technologies. 2022.<br> <a href="https://doi.org/10.1007/s10639-022-11093-6">https://doi.org/10.1007/s10639-022-11093-6</a></p> <p>Preprint available at: <a href="https://arxiv.org/abs/2307.08582">https://arxiv.org/abs/2307.08582</a></p> <ul> </ul> <p><strong>How to cite</strong></p> <p>If you use or build upon the materials, please use the BibTeX entry below to cite the original paper (not only this web link).</p> <pre><code>@article{Svabensky2022applications, author = {\v{S}v\'{a}bensk\'{y}, Valdemar and Vykopal, Jan and \v{C}eleda, Pavel and Kraus, Lydia}, title = {{Applications of Educational Data Mining and Learning Analytics on Data From Cybersecurity Training}}, journal = {Education and Information Technologies}, publisher = {Springer}, volume = {27}, year = {2022}, issn = {1360-2357}, url = {https://doi.org/10.1007/s10639-022-11093-6}, doi = {10.1007/s10639-022-11093-6}, }</code></pre> <p><strong>Attached content</strong></p> <p>The files included in the ZIP archive are:</p> <ul> <li>`All-discovered-papers.bib` -- a BibTeX export of the Mendeley database of all considered papers discovered by the automated search.</li> <li>`Candidate-papers-reviewer1.bib` -- a BibTeX export of the Mendeley database of the candidate papers suggested by the first investigator.</li> <li>`Candidate-papers-reviewer2.bib` -- a BibTeX export of the Mendeley database of the candidate papers suggested by the second investigator.</li> <li>`Selected-papers.bib` -- a BibTeX export of the Mendeley database of the 35 papers selected for the literature review.</li> <li>`Selected-papers.xlsx` -- an Excel spreadsheet with the extracted information about the selected papers.</li> <li>`Selected-papers.csv` -- a CSV equivalent of the Excel spreadsheet.</li> </ul>
Anonymized Responses to the Student Expectation of Learning Analytics Questionnaire (SELAQ) in 2022 and 2023
<p>Contains responses of 566 bachelor students to the Student Expectation of Learning Analytics Questionnaire (SELAQ) conducted at Bochum University of Applied Sciences in Germany in fall 2022 and summer 2023. The Questionnaire consists of 12 statements (in the dataset: enumerated from 1-12) which are evaluated by students regarding both their desire and their expectation (in the dataset: D and E). This results in 24 items (in the dataset: 1D, 1E, 2D, 2E, ..., 12D, 12E). Each item is evaluated by the students on a Likert-scale from 1 (strongly disagree) to 7 (strongly agree). Additionally, the students' study track (in the dataset: Study_Track) and their current semester (in the dataset: Semester) were also recorded. The questionnaires were carried out in paper form at the beginning of lectures. Statements 1, 2, 3, 5 and 6 deal with ethical and privacy expectations, while statements 4, 7, 8, 9, 10, 11 and 12 deal with service feature expectations regarding learning analytics. For a text version of the items, please refer to the paper related to this dataset (DOI 10.1145/3636555.3636923), the additional descriptions below or the attached description file.</p>
A network of countries collaborating on learning analytics research
<p>A network of countries collaborating on learning analytics research. It includes original research articles published in Scopus Database up to 16 January 2018. The file is an un-directed Graphml network of 76 countries. It can be opened in Social Network Analysis applications such as Gephi, or Igraph R package.</p>
TRANSFORMING CUSTOMER RETENTION IN FINTECH INDUSTRY THROUGH PREDICTIVE ANALYTICS AND MACHINE LEARNING
<p>In recent years, the fintech industry has experienced rapid growth, driven by technological advancements and evolving consumer expectations. Fintech companies offer innovative financial services, such as digital banking, investment platforms, and payment solutions, catering to the needs of a tech-savvy customer base. However, as competition intensifies, customer retention has emerged as a critical challenge for these companies. According to a study by Ransom (2021), acquiring a new customer can cost five times more than retaining an existing one, making it imperative for fintech organizations to focus on strategies that enhance customer loyalty. The financial technology (fintech) sector has experienced unprecedented growth in recent years, fundamentally transforming how individuals and businesses access and manage financial services. Characterized by the integration of technology with financial services, fintech encompasses a wide array of offerings, including digital banking, peer-to-peer lending, robo-advisory services, and payment processing. As of 2023, the global fintech market was valued at approximately $309 billion and is projected to reach around $1.5 trillion by 2030, according to a report by Fortune Business Insights. This remarkable growth is largely attributed to advancements in digital technology, increasing smartphone penetration, and a growing consumer preference for online financial solutions. Moreover, the COVID-19 pandemic accelerated the adoption of digital financial services, as consumers sought contactless transactions and remote banking options.</p>
Data and analytical codes for: Learning beyond-pairwise interactions enables the bottom-up prediction of microbial community structure
<p>Data and analytical codes for: Ishizawa et al. (2023) Learning beyond-pairwise interactions enables the bottom-up prediction of microbial community structure, bioRxiv, 2023.07.04.546222</p> <p> https://www.biorxiv.org/content/10.1101/2023.07.04.546222v1</p> <p> </p>
Analytics, Visualisation and Machine Learning of General Practitioner Prescribing using Open Health Data
<p>Open Prescription data used in Postgraduate project into Northern Ireland General Practice prescribing.</p>
The Impact of Generative AI on Student Learning Outcomes: A Statistical Analytical Approach - Dataset
Open the record for dataset details and reuse information.
Hard Rod Fluid Analytical Solution in .npy (For Operator learning)
<p>Hard rod fluid with diameter of 1.0 between 2 walls with distance L+2d=10.0,<br>under different chemical potential (mu) and external potential (V_ext(x)), with a sampling distance dz=0.1.</p> <p>Data are stored as numpy structured ndarrays, each npy file store one type of data. <br>The ndarrays has the shape of (n_sample,n_grid), with 4 fields:<br>the position "z", <br>the density profile "rho", <br>the local chemical potential "muloc"<br>the one-body direct correlation function "c1".<br>Where muloc(z)=mu-V_ext(z)</p> <p>In Operator Learning, I've tried the mapping of rho to c1 as well as muloc to rho.</p> <p>dataload.py contains some functions used for loading data<br>data_model_train.py contains the read-in function for model-training data.</p> <p>data_4groups.npy saves 8000 data, with 2000 in each group, generated from the following external potentials:<br>group 1: V_ext=0 <br>group 2: V_binary[z] = -epsilon*((a/(z+a/2))^3+(a/(L+a/2-z))^3)<br>group 3: V_binary[z] = -mg*z<br>group 4: V_binary[z] = -u0*(abs(z/z0))^a0<br>with mu, epsilon, a, mg, u0, z0, a0 are adjustable constants of the potential </p> <p>data_mixed23.npy saves 2000 data, generated from the following external potentials:<br> V_binary[z] = -epsilon*((a/(z+a/2))^3+(a/(L+a/2-z))^3) -mg*z</p> <p>Folders are taped in .zip files.</p> <p>Folder "~/R_data_generator" stores the generator of the data, analytical solution of 1D hard rod fluid.<br>Folder "~/Rdata" contains same data in .Rdata file, as R matrixs<br>Folder "~/R2py" contains the script to generate .npy file from .Rdata file</p>
Learning Analytics for Personalized Learning Environments: Visualizing Journal Publication Trends
<p>Data collected from Scopus database for the "Learning Analytics for Personalized Learning Environments: Visualizing Journal Publication Trends" entitled research.</p>
Student oriented subset of the Open University Learning Analytics dataset
<p>The Open University (OU) dataset is an open database containing student demographic and click-stream interaction with the virtual learning platform. The available data are structured in different CSV files. You can find more information about the original dataset at the following link: <a href="https://analyse.kmi.open.ac.uk/open_dataset">https://analyse.kmi.open.ac.uk/open_dataset</a>.</p> <p>We extracted a subset of the original dataset that focuses on student information. 25,819 records were collected referring to a specific student, course and semester. Each record is described by the following 20 attributes: <em> code_module, code_presentation, gender, highest_education, imd_band, age_band, num_of_prev_attempts, studies_credits, disability, resource, homepage, forum, glossary, outcontent, subpage, url, outcollaborate, quiz, AvgScore, count</em>.<br> <br> Two target classes were considered, namely Fail and Pass, combining the original four classes (Fail and Withdrawn and Pass and Distinction, respectively). The final_result attribute contains the target values.<br> <br> All features have been converted to numbers for automatic processing.<br> <br> Below is the mapping used to convert categorical values to numeric:</p> <ul> <li>code_module: 'AAA'=0, 'BBB'=1, 'CCC'=2, 'DDD'=3, 'EEE'=4, 'FFF'=5, 'GGG'=6</li> <li>code_presentation: '2013B'=0, '2013J'=1, '2014B'=2, '2014J'=3</li> <li>gender: 'F'=0, 'M'=1</li> <li>highest_education: 'No_Formal_quals'=0, 'Post_Graduate_Qualification'=1, 'HE_Qualification'=2, 'Lower_Than_A_Level'=3, 'A_level_or_Equivalent'=4</li> <li>IMBD_band: 'unknown'=0, 'between_0_and_10_percent'=1, 'between_10_and_20_percent'=2, 'between_20_and_30_percent'=3, 'between_30_and_40_percent'=4, 'between_40_and_50_percent'=5, 'between_50_and_60_percent'=6, 'between_60_and_70_percent'=7, 'between_70_and_80_percent'=8, 'between_80_and_90_percent'=9, 'between_90_and_100_percent'=10</li> <li>age_band: 'between_0_and_35'=0, 'between_35_and_55'=1, 'higher_than_55'=2</li> <li>disability: 'N'=0, 'Y'=1</li> <li>student's outcome: 'Fail'=0, 'Pass'=1</li> </ul> <p>For more detailed information, please refer to:</p> <p><br> Casalino G., Castellano G., Vessio G. (2021) Exploiting Time in Adaptive Learning from Educational Data. In: Agrati L.S. et al. (eds) Bridges and Mediation in Higher Distance Education. HELMeTO 2020. Communications in Computer and Information Science, vol 1344. Springer, Cham. <a href="https://www.google.com/url?q=https%3A%2F%2Fdoi.org%2F10.1007%2F978-3-030-67435-9_1&sa=D&sntz=1&usg=AFQjCNF7fUT9S4TcSpImSr4e_DjaLn3wtg">https://doi.org/10.1007/978-3-030-67435-9_1</a></p>
Analytical prediction of scattering properties of spheroidal dust particles with machine learning
<p>This respository includes the data used in the paper "Analytical prediction of scattering properties of spheroidal dust particles with machine learning".</p> <ol> <li>"alldata.zip" represents the extinction and absorption coefficients, and phase matrix elements of spheroids dust particles for both training and test data. These data comes from the Oleg Dubovik's group: <a href="https://www.grasp-open.com/products/spheroid-package-release">https://www.grasp-open.com/products/spheroid-package-release</a>/.</li> <li>"<a href="/api/files/7c2cdcdd-dc47-4a19-adc8-59c944e76e54/tmat_Jacall_norm_intg.pickle?versionId=59294030-26a8-4683-9b26-8e8584c2b44b">tmat_Jacall_norm_intg.pickle</a>" contains the Jacobians simulated from linearized T-matrix model used for training in the paper and "<a href="/api/files/7c2cdcdd-dc47-4a19-adc8-59c944e76e54/tmat_fine_n_intg.pickle?versionId=3d20954d-4318-40bb-a900-a93627915c2e">tmat_fine_n_intg.pickle</a>" involves Jacobians used for testing.</li> </ol>
Predicting Weather Disruptions for the ICC Champions Trophy 2025 in Pakistan Using Machine Learning and Data Analytics
Open the record for dataset details and reuse information.
Scopus_Learning Analytics in HE_Bibliometric Analysis
<p>Dataset from scopus with keywords learning analytic or learning analytics in higher education or university. </p>
A New Dataset for Streaming Learning Analytics
<p>This research introduces a novel dataset developed for streaming learning analytics, derived from the Open University Learning Analytics Dataset (OULAD). The dataset incorporates essential temporal information that captures the timing of student interactions with the Virtual Learning Environment (VLE). By integrating these time-based interactions, the dataset enhances the capabilities of stream algorithms, which are particularly well-suited for real-time monitoring and analysis of student learning behaviors.</p> <p> </p> <p>The dataset consists of 34 features and 1,718,983 samples, encompassing students' demographic information, assessment scores, and interactions with the VLE for a specific time ( T ), corresponding to each student ( S ) within a given course ( C ) and module ( M ). The target classes—'Withdrawn', 'Fail', 'Pass', and 'Distinction'—were encoded as 0, 1, 2, and 3, respectively. Notably, the data exhibits a significant imbalance, with a substantial prevalence of records associated with students who passed the final examination. The class distribution is as follows: 'Pass' (1,022,760 samples), 'Distinction' (308,642 samples), 'Fail' (227,550$ samples), and 'Withdrawn' (160,031 samples).</p> <p>For further details on the data, please refer to the manuscript: Gabriella Casalino, Giovanna Castellano, Gianluca Zaza, "Does Time Matter in Analyzing Educational Data? - A New Dataset for Streaming Learning Analytics.", CEUR Proceedings </p>
Summary ouput data - Wasteaware Cities Benchmark Indicators - WABI 2023 - Global data analytics - Machine learning vs. Non-linear Regression
<p>This is the output dataset for the research publication "<em>Socio-economic development drives solid waste management performance in cities: A global analysis using machine learning</em>". It features </p> <ul> <li>Metadata info used by R codes</li> <li>Summary of results for two modelling approaches (machine learning: Conditional random-forest and non-linear regression)</li> </ul> <p>The independent variables dataset analysed here refer to specific indicators of the WABI methodology (<a href="https://www.sciencedirect.com/science/article/pii/S0956053X14004905">https://www.sciencedirect.com/science/article/pii/S0956053X14004905</a>) that generates solid waste management and resource recovery profiles for cities. It was applied here for 40 cities around the world. The data input are available here: 10.5281/zenodo.7570174</p>
Non-invasive Point-of-care Diagnosis Using Machine Learning and Signal Analytics to Transform Early Detection of Heart Disease
ClinicalTrials.gov study NCT03864081. IPD Sharing: NO. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.