Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
10
datasets available to search
ShareScore release 0.7.1
Dataset results
10 results for “data justice”
Data for: A path forward: creating an academic culture of justice, equity, diversity and inclusion
<p>Institutions of higher education (IHE) throughout the United States have a long history of acting out various levels of commitment to diversity advancement, equity, and inclusion (DEI). Despite decades of DEI "<em>efforts</em>," the academy is fraught with legacies of racism that uphold white supremacy and prevent marginalized populations from full participation. Furthermore, politicians have not only weaponized education but passed legislation to actively ban DEI programs and censor general education curricula (<a href="https://tinyurl.com/antiDEI">https://tinyurl.com/antiDEI</a>). Ironically, systems of oppression are particularly apparent in the fields of Ecology, Evolution, and Conservation Biology (EECB)–which recognize biological diversity as essential for ecological integrity and resilience. Yet, amongst EECB faculty, people who do not identify as cis-heterosexual, non-disabled, affluent white males are poorly represented. Furthermore, IHE lack metrics to quantify DEI as a priority. Here we show that only 30.3% of US-faculty positions advertised in EECB from Jan 2019-May 2020 required a diversity statement; diversity statement requirements did not correspond with state-level diversity metrics. Though many announcements "encourage women and minorities to apply," empirical evidence demonstrates that hiring committees at most institutions did not prioritize an applicant's DEI advancement potential. We suggest a model for change and call on administrators and faculty to implement SMART (i.e., Specific, Measurable, Achievable, Realistic, and Timely) strategies for DEI advancement across IHE throughout the United States. We anticipate our quantification of diversity statement requirements relative to other application materials will motivate institutional change in both policy and practice when evaluating a candidate's potential "fit". IHE must embrace a leadership role to not only shift the academic culture to one that upholds DEI, but to educate and include people who represent the full diversity of our society. In the current context of political censure of education including book banning and backlash aimed at Critical Race Theory, which further reinforce systemic white supremacy, academic integrity and justice are more critical than ever. </p>
U.S. cities increasingly integrate justice into climate planning and create policy tools for climate justice (Diezmartínez & Short Gianotti, 2022) - Data and code
<p>This repository contains datasets and coding corresponding to the journal article titled "U.S. cities increasingly integrate justice into climate planning and create policy tools for climate justice". We include:</p> <ul> <li>DataRegressionAnalysis.csv <ul> <li>CSV file with data used for regression analysis. This file can be used directly to run R code provided in this repository.</li> </ul> </li> <li>QualitativeCodingResults.nvp <ul> <li>NVivo project with all results for the qualitative coding of urban climate action plans.</li> <li>This file also contains all climate action plans analyzed in this research.</li> </ul> </li> <li>QualitativeCodingResults_Summary.xlsx <ul> <li>Excel file with a results summary for the qualitative coding of urban climate action plans.</li> </ul> </li> <li>RegressionAnalysis.Rmd <ul> <li>Rmd file with R code used for regression analysis. </li> </ul> </li> <li>RegressionAnalysis_KnitOutput.html <ul> <li>Knit output from R code with regression analysis results, html format.</li> </ul> </li> <li>RegressionAnalysis_KnitOutput.pdf <ul> <li>Knit output from R code with regression analysis results, PDF format.</li> </ul> </li> </ul>
Data for: A path forward: creating an academic culture of justice, equity, diversity and inclusion
Open the record for dataset details and reuse information.
Data From: Emissions redistribution and environmental justice implications of California's Clean Vehicle Rebate Project
Open the record for dataset details and reuse information.
Data for manuscript: "Themes in Academic Literature: Prejudice and Social Justice"
<p>This data set contains frequency counts of target words in 175 million academic abstracts published in all fields of knowledge. We quantify the prevalence of words denoting prejudice against ethnicity, gender, sexual orientation, gender identity, minority religious sentiment, age, body weight and disability in SSORC abstracts over the period 1970-2020. We then examine the relationship between the prevalence of such terms in the academic literature and their concomitant prevalence in news media content. We also analyze the temporal dynamics of an additional set of terms associated with social justice discourse in both the scholarly literature and in news media content. A few additional words not denoting prejudice are also available since they are used in the manuscript for illustration purposes.</p> <p>The list of academic abstracts analyzed in this work was taken from the Semantic Scholar Open Research Corpus (SSORC). The corpus contains, as of 2020, over 175 million academic abstracts, and associated metadata, published in all fields of knowledge. The raw data is provided by Semantic Scholar in accessible JSON format.</p> <p>Textual content included in our analysis is circumscribed to the scholarly articles’ titles and abstracts and does not include other article elements such as main body of text or references section. Thus, we use frequency counts derived from academic articles’ titles and abstracts as a proxy for word prevalence in those articles. This proxy was used because the SSORC corpus does not provide the entire text body of the indexed articles. Targeted textual content was located in JSON data and sorted by year to facilitate chronological analysis. Tokens were lowercased prior to estimating frequency counts.</p> <p>Yearly relative frequencies of a target word or n-gram in the SSORC corpus were estimated by dividing the number of occurrences of the target word/n-gram in all scholarly articles within a given year by the total number of all words in all articles of that year. This method of estimating word frequencies accounts for variable volume of total scientific output over time. This approach has been shown before to accurately capture the temporal dynamics of historical events and social trends in news media corpora.</p> <p>It is possible that a small percentage of scholarly articles in the SSORC corpus contain incorrect or missing data. For earlier years in the SSORC corpus, abstract information is sometimes missing and only article’s title information is available. As a result, the total and target word count metrics for a small subset of academic abstracts might not be precise. In a data analysis of 175 million scientific abstracts, manually checking the accuracy of frequency counts for every single academic abstract is unfeasible and hundred percent accuracy at capturing abstracts’ content might be elusive due to a small number of erroneous outlier cases in the raw data. Overall, however, we are confident that our frequency metrics are representative of word prevalence in academic content as illustrated by Figure 2 in the main manuscript, which shows the chronological prevalence in the SSORC corpus of several terms associated with different disciplines of scientific/academic knowledge.</p> <p>Factor analysis of frequency counts time series was carried out only after Bartlett’s test of sphericity and Kaiser-Meyer-Olkin (KMO) test confirmed the suitability of the data for factor analysis. A single factor derived from the frequency counts time series of prejudice-denoting terms was extracted from each corpus (academic abstracts and news media content). The same procedure was applied for the terms denoting social justice discourse. A factor loading cutoff of 0.5 was used to ascribe terms to a factor. Chronbach alphas to determine if the resulting factors appeared coherent were extremely high (>0.95).</p> <p>The textual content of news and opinion articles from the outlets listed in Figure 5 of the main manuscript is available in the outlet's online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We used derived word frequency counts from original sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific xpath expressions. Tokens were lowercased prior to estimating frequency counts. To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1 million words of article content from an outlet. </p> <p>Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript and raw data metrics</p> <p>-scholarlyArticlesContainingTargetWords.rar contains the IDs of each analyzed abstract in the SSORC corpus and the counts of target words and total words for each scholarly article</p> <p>-targetWordsInMediaArticlesCounts.rar contains counts of target words in news outlets articles as well as total counts of words in articles</p> <p>In a small percentage of news articles, outlet specific XPath expressions can fail to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles might not be precise. </p> <p>In a data analysis of millions of news articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Rozado, Al-Gharbi, and Halberstadt, “Prevalence of Prejudice-Denoting Words in News Media Discourse" for supporting evidence).</p> <p>31/08/2022 Update: There is a new way to download the Semantic Scholar Open Research Corpus (see https://github.com/allenai/s2orc). This updated version states that the corpus contains 136M+ paper nodes. However, when I downloaded a previous version of the corpus in 2021 from http://s2-public-api-prod.us-west-2.elasticbeanstalk.com/corpus/download/ I counted 175M unique identifiers. The URL of the previous version of the corpus is no longer active, but it has been cached by the Internet Archive at https://web.archive.org/web/20201030131959/http://s2-public-api-prod.us-west-2.elasticbeanstalk.com/corpus/download/ I haven't had the time to look at the specific reason for the mismatch but perhaps the newer version of the corpus has cleaned a lot of noisy entries in the previous version which often contained entries with missing abstracts. Filtering out entries in low prevalence languages other than English might be another reason. In any case, Figure 2 of the main manuscript of this work (at https://www.nas.org/academic-questions/35/2/themes-in-academic-literature-prejudice-and-social-justice) should provide support for the validity of the frequency counts.</p> <p> </p>
Data from: A framework for sharing power in research teams and promoting justice in scientific publication
Open the record for dataset details and reuse information.
Water, Dust, and Environmental Justice: The Case of Agricultural Water Diversions - Data
<p>Replication files and code for paper "Water, Dust, and Environmental Justice: The Case of Agricultural Water Diversions". </p>
Data and code of co-created rooftop harvested rainwater study in AZ environmental justice communities
<p>See README</p>
Data from Political Actors' Twitter Accounts Related to Transitional Justice
Open the record for dataset details and reuse information.
Data from: extended use and end of life in the Global South: a transportation justice perspective of US-Mexico second-hand vehicle trade
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.