Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13
datasets available to search
ShareScore release 0.9.0
Dataset results
13 results for “author profiling”
Pinterest dataset for age and gender identification in author profiling
<p>This dataset was used for the experiments presented in the article "Reconstructive Classification for Age and Gender Identification in Social Networks" - IEEE Transactions on Computational Social Systems</p> <p>The dataset contains text data from 548,761 pins corresponding to 264 users of Pinterest.</p> <p>There are 7 files.</p> <p>The first 5 files correspond to the extracted textual features from the pins that are aggregated per user: ats, emojis/emoticons, hashtags, links, and words.</p> <p>There are 264 lines in each file (one per user), as the concatenation of the extracted features from all the pins corresponding to each user.</p> <p>The last 2 files are the labels for the age and gender of the users. There are also 264 lines (one per user).</p> <p>For age, there are 4 possible labels: 18-24, 25-34, 35-46, and 50+</p> <p>For gender, there are 2 possible labels: F and M</p>
Total Load Profiles by Balancing Authority in the Western United States for GODEEEP
<p>This dataset contains time-series of total load profiles across Balancing Authorities (BAs) in the western United States (U.S.) electricity grid interconnection for the years 2025, 2030, 2035, 2040, 2045, and 2050. The data is provided for two different socioeconomic pathways and one climate scenario. The socioeconomic pathways -- Business-As-Usual (BAU_Climate) and Net-Zero without CCS (NetZeroNoCCS_Climate) -- are described by <a href="https://doi.org/10.5281/zenodo.7838871">https://doi.org/10.5281/zenodo.7838871</a>. The climate scenario -- Representative Concentration Pathway 8.5 hotter (rcp85hotter) -- is described by <a href="https://doi.org/10.57931/1885756">https://doi.org/10.57931/1885756</a>. Transportation loads in this dataset are derived from <a href="https://doi.org/10.5281/zenodo.8065137">https://doi.org/10.5281/zenodo.8065137</a>. Non-transportation loads are derived using the Total Electricity Loads (TELL) model which is available <a href="https://github.com/IMMM-SFA/tell">here</a>. The data was post-processed using a Jupyter notebook available <a href="https://github.com/GODEEEP/load_analysis/blob/main/notebooks/process_load_data.ipynb">here</a>.</p> <p>A brief summary of the files and directories in this data package is provided below. Each file in the "<em>total_loads</em>" subdirectory is a comma-separated-value (CSV) format with the following columns:</p> <ul> <li><strong>BA</strong> - Acronym of the balancing authority (BA) for this data point</li> <li><strong>Time_UTC</strong> - Timestamp of the hourly data in UTC format</li> <li><strong>Non-Transportation_Load_MWh</strong> - Total hourly load in the BA from non-transportation sources in megawatt hours</li> <li><strong>Transportation_Load_MWh</strong> - Total hourly load in the BA from transportation sources in megawatt hours</li> <li><strong>Total_Load_MWh</strong> - Sum of the load from non-transportation and transportation sources in megawatt hours</li> </ul> <p>Data in the "<em>gridview_ready_total_loads</em>" subdirectory contains the hourly total loads by BA in a format that is ready for ingestion into the GridView production cost model.</p> <p>This research was supported by the Grid Operations, Decarbonization, Environmental and Energy Equity Platform (GODEEEP) Investment, under the Laboratory Directed Research and Development (LDRD) Program at Pacific Northwest National Laboratory (PNNL).</p> <p>PNNL is a multi-program national laboratory operated for the U.S. Department of Energy (DOE) by Battelle Memorial Institute under Contract No. DE-AC05-76RL01830.</p>
Transportation Electrification Load Profiles by Balancing Authority and State-Level Electrification Rates in the Western United States for GODEEEP
<p>Time-series hourly electric charging load profiles for the transportation sector across Balancing Authorities (BAs) in the Western Electricity Coordinating Council (WECC) interconnect, annual fleet sizes by state and vehicle type, annual transportation sector energy usage by state and fuel, and annual transportation fuel usage by state. The data is provided for three different socioeconomic pathways and two different climate pathways, resulting in four total scenarios. The socioeconomic pathways--Net-Zero (<code>nz_climate</code>), Net-Zero allowing for Carbon Capture Sequestration (CCS) technology (<code>nz_ccs_climate</code>), and Net-Zero allowing for CCS with Inflation Reduction Act (IRA) policies (<code>nz_ira_ccs_climate</code>)--are described by <a href="https://doi.org/10.5281/zenodo.10642507">https://doi.org/10.5281/zenodo.10642507</a>. The climate pathways--Representative Concentration Pathway (RCP) 4.5 cooler (<code>rcp45cooler</code>) and RCP 8.5 hotter (<code>rcp85hotter</code>)--are described by <a href="https://doi.org/10.57931/1885756">https://doi.org/10.57931/1885756</a>. The climate influence is only considered for Light Duty Vehicles (LDVs).</p> <p>For additional details please consult the paper Acharya et al 2024, Impact of the Inflation Reduction Act and Carbon Capture on Transportation Electrification for a Net-Zero Western U.S. Grid, submitted, and the code repository <a href="https://github.com/GODEEEP/transportation_electrification">https://github.com/GODEEEP/transportation_electrification</a>.</p> <p>A brief summary of the files and directories in this data package is provided below. Text within chevrons implies a multiplicity of files, one for each actual value.</p> <ul> <li>nz_climate <ul> <li>rcp45cooler <ul> <li><balancing authority>_hourly_transportation_load_<socioeconomic pathway>_<climate scenario>_<year>.csv</li> </ul> </li> <li>rcp85hotter <ul> <li><balancing authority>_hourly_transportation_load_<socioeconomic pathway>_<climate scenario>_<year>.csv</li> </ul> </li> </ul> </li> <li>nz_ccs_climate <ul> <li>rcp45cooler <ul> <li><balancing authority>_hourly_transportation_load_<socioeconomic pathway>_<climate scenario>_<year>.csv</li> </ul> </li> <li>rcp85hotter <ul> <li><balancing authority>_hourly_transportation_load_<socioeconomic pathway>_<climate scenario>_<year>.csv</li> </ul> </li> </ul> </li> <li>nz_ira_ccs_climate <ul> <li>rcp45cooler <ul> <li><balancing authority>_hourly_transportation_load_<socioeconomic pathway>_<climate scenario>_<year>.csv</li> </ul> </li> <li>rcp85hotter <ul> <li><balancing authority>_hourly_transportation_load_<socioeconomic pathway>_<climate scenario>_<year>.csv</li> </ul> </li> </ul> </li> <li>WECC_hourly_transportation_load_<socioeconomic pathway>_<climate scenario>_<year>.csv</li> <li>EV_electric_and_total_energy.csv</li> <li>LDV_fleet_size_all_fuel_types_state_wise.csv</li> <li>MDV_fleet_size_all_fuel_types_state_wise.csv</li> <li>HDV_fleet_size_all_fuel_types_state_wise.csv</li> </ul> <p> </p> <p><strong>Hourly transportation load:</strong></p> <ul> <li><code>time</code> - ISO 8601 timestamp representing the end of the hourly timestep; values are reported as the summation over the preceding hour</li> <li><code>balancing_authority</code> - Acronym of the balancing authority for this data point</li> <li><code>LDV_load_MWh</code> - Energy consumed by the charging of Light Duty Vehicles (LDVs) during the previous hour in Megawatt hours</li> <li><code>MDV_load_MWh</code> - Energy consumed by the charging of Medium Duty Vehicles (MDVs) during the previous hour in Megawatt hours</li> <li><code>HDV_load_MWh</code> - Energy consumed by the charging of Heavy Duty Vehicles (HDVs) during the previous hour in Megawatt hours</li> <li><code>passenger_rail_load_MWh</code> - Energy consumed by the charging of passenger rail vehicles during the previous hour in Megawatt hours</li> <li><code>freight_rail_load_MWh</code> - Energy consumed by the charging of freight rail vehicles during the previous hour in Megawatt hours</li> <li><code>aviation_load_MWh</code> - Energy consumed by the charging of aviation vehicles during the previous hour in Megawatt hours</li> <li><code>ship_load_MWh</code> - Energy consumed by the charging of ships during the previous hour in Megawatt hours</li> <li><code>transportation_load_MWh</code> - Total energy consumed by the charging of vehicles during the previous hour in Megawatt hours (summation of the other columns)</li> </ul> <p>The WECC files provide summations of all BAs for each scenario, with the same columns as above excepting <code>balancing_authority</code></p> <p><strong>State-wise fleet sizes by vehicle type:</strong></p> <ul> <li><code>scenario</code> - the socioeconomic pathway, one of <code>nz_climate</code>, <code>nz_ccs_climate</code>, or <code>nz_ira_ccs_climate</code></li> <li><code>state</code> - two letter abbreviation of the state within the Western U.S. Interconnection</li> <li><code>year</code> - 5 year increments from 2020 to 2050</li> <li><code>technology</code> - fuel type such as BEV (battery electric vehicle), FCEV (fuel cell electric vehicle), hybrid liquids and liquids (refined liquids)</li> <li><code>veh_type</code> - one of LDV, MDV, or HDV (Light, Medium, or Heavy Duty Vehicle)</li> <li><code>fleet_size</code> - the number of vehicles</li> </ul> <p>To calculate an electrification rate in terms of fleet size for a given scenario, state, year, and veh<em>type, we divide the fleet</em>size for BEV technology by the summation of fleet_size for all technologies.</p> <p><strong>State-wise electric and total energy for LDVs, MDVs, and HDVs:</strong></p> <ul> <li><code>state</code> - two letter abbreviation of the state within the Western U.S. Interconnection</li> <li><code>year</code> - 5 year increments from 2020 to 2050</li> <li><code>scenario</code> - the socioeconomic pathway, one of <code>nz_climate</code>, <code>nz_ccs_climate</code>, or <code>nz_ira_ccs_climate</code></li> <li><code>hdv_total</code> - energy in ExaJoules consumed by all HDVs irrespective of fuel type</li> <li><code>ldv_total</code> - energy in ExaJoules consumed by all LDVs irrespective of fuel type</li> <li><code>mdv_total</code> - energy in ExaJoules consumed by all MDVs irrespective of fuel type</li> <li><code>hdv_electric</code> - electric energy in ExaJoules consumed by HDVs</li> <li><code>ldv_electric</code> - electric energy in ExaJoules consumed by LDVs</li> <li><code>mdv_electric</code> - electric energy in ExaJoules consumed by MDVs</li> </ul> <p>To calculate the electrification rate in terms of EV energy for a given scenario, state, year, and veh_type, we divide electric energy by the total energy.</p> <p><strong>State-wise transportation fuel mix:</strong></p> <ul> <li><code>state</code> - two letter abbreviation of the state within the Western U.S. Interconnection</li> <li><code>year</code> - 5 year increments from 2020 to 2050</li> <li><code>scenario</code> - the socioeconomic pathway, one of <code>nz_climate</code>, <code>nz_ccs_climate</code>, or <code>nz_ira_ccs_climate</code></li> <li><code>hydrogen</code> - hydrogen energy in ExaJoules consumed by the transportation sector</li> <li><code>electricity</code> - electric energy in ExaJoules consumed by the transportation sector</li> <li><code>refined liquids</code> - refined liquid energy in ExaJoules consumed by the transportation sector</li> </ul> <p><br><br></p> <p><strong>Changelog:</strong></p> <ul> <li>v2.0.0 - new set of scenarios; fuel mix data added</li> </ul> <p> </p> <p>This research was supported by the Grid Operations, Decarbonization, Environmental and Energy Equity Platform (GODEEEP) Investment, under the Laboratory Directed Research and Development (LDRD) Program at Pacific Northwest National Laboratory (PNNL).</p> <p>PNNL is a multi-program national laboratory operated for the U.S. Department of Energy (DOE) by Battelle Memorial Institute under Contract No. DE-AC05-76RL01830.</p>
PAN19 Author Profiling: Bots and Gender Profiling
<p>Social media bots pose as humans to influence users with commercial, political or ideological purposes. For example, bots could artificially inflate the popularity of a product by promoting it and/or writing positive ratings, as well as undermine the reputation of competitive products through negative valuations. The threat is even greater when the purpose is political or ideological (see Brexit referendum or US Presidential elections). Fearing the effect of this influence, the German political parties have rejected the use of bots in their electoral campaign for the general elections. Furthermore, bots are commonly related to fake news spreading. Therefore, to approach the identification of bots from an author profiling perspective is of high importance from the point of view of marketing, forensics and security.</p> <p>After having addressed several aspects of <strong>author profiling</strong> in social media from 2013 to 2018 (age and gender, also together with personality, gender and language variety, and gender from a multimodality perspective), this year we aim at investigating whether the author of a Twitter feed is a <strong>bot</strong> or a <strong>human</strong>. Furthermore, in case of human, to profile the <strong>gender</strong> of the author.</p> <p>The uncompressed dataset consists in a folder per language (en, es). Each folder contains:</p> <ul> <li>A XML file per author (Twitter user) with 100 tweets. The name of the XML file correspond to the unique author id.</li> <li>A truth.txt file with the list of authors and the ground truth.</li> </ul>
PAN14 Author Profiling
<p>We provide you with a training data set that consists of blog posts, Twitter tweets and social media texts written in both English and Spanish as well as hotel reviews written in English. With regard to age, we will consider the following classes: 18-24, 25-34, 35-49, 50-64, 65-xx.</p>
PAN13 Author Profiling
<p>We provide you with a training data set that consists of documents written in both English and Spanish. With regard to age, we will consider posts of three classes: 10s (13-17), 20s (23-27), and 30s (33-47). Moreover, documents from authors who pretend to be minors will be included (e.g., documents composed of chat lines of sexual predators will be also considered). <a href="https://www.uni-weimar.de/medien/webis/events/pan-13/pan13-papers-final/pan13-author-profiling/rangel13-overview.pdf#page=3">Learn more »</a></p>
PAN16 Author Profiling
<p>We provide a training data set that consists of Twitter tweets in English, Spanish and Dutch.</p> <p>The English and Spanish datasets are labeled with age and gender, whereas the Dutch one only with gender. With regard to age, we will consider the following classes: 18-24, 25-34, 35-49, 50-64, 65-xx.</p> <p><strong>Remark.</strong> Due to Twitter's privacy policy we cannot provide tweets directly, but only URLs referring to them. You will have to download them yourself. For your convenience, we provide a download software for this. We expect participants to extract gender and age information only from the textual part of a tweet and to discard any other meta information that may be provided by Twitter's API. When we evaluate your software at our site, we do not expect it downloads tweets. We will do this beforehand.</p> <p>More information about the task: <a href="https://pan.webis.de/clef16/pan16-web/author-profiling.html">Link</a></p>
PAN18 Author Profiling
<p>We provide you with a training data set that consists of Twitter users labeled with gender. For each author, a total of 100 tweets and 10 images are provided. Authors are grouped by the language of their tweets: English, Arabic and Spanish.</p> <p>More information about the task: <a href="https://pan.webis.de/clef18/pan18-web/author-profiling.html">Link</a></p>
Post Authorization Safety Profile of Booster IndoVac COVID-19 Vaccination
ClinicalTrials.gov study NCT06690502. IPD Sharing: NO. Countries: 1. Publications: 0.
Multi-dimensional author profiling by business roles
<p>This dataset contains the data used in the paper "Multidimensional Author Profiling for Social Business Intelligence", more specifically, the gold standard (GS) and silver standard (SS) created for training and validating text classifiers for business profiling of social network users.</p> <p>The GS dataset is a CSV file with the following columns:</p> <div> <div> <ul> <li>screen-name, user-id, verified-user (boolean), <strong>multi-level-label</strong>, <strong>manual-verification</strong>, textual-description, followers (int), friends (int), source (not used)</li> </ul> </div> <div>The attribute "multi-level label" contains label represeting the user business profile, regarding the three perspectives: <em>role</em>, <em>colective-vs-individual</em>, and <em>on-domain</em> ones. The attribute "manual-verification" is a second pass from experts to validate the assigned label.</div> <div> </div> <div>The SS dataset is a "|"-separated text file with the following columns:</div> <div> <ul> <li> <div>screen-name|user-id|verified-user|<strong>multi-level-label</strong>|textual-description</div> <div> </div> </li> </ul> </div> <div>The SS dataset is generated with an unsupervised method through an initial seed of bigrams. Therefore, the dataset can contain wrong and incomplete labels, hence the name silver standard (SS).</div> <div> </div> <div>As data is captured from Twitter, we can only relase it under restricted conditions.</div> </div>
PAN17 Author Profiling
<p>We provide you with a training data set that consists of Twitter tweets in English, Spanish, Portuguese and Arabic, labeled with gender and language variety.</p> <p>More information about the task: <a href="https://pan.webis.de/clef17/pan17-web/author-profiling.html">Link</a></p>
PAN15 Author Profiling
<p>We provide a data set that consists of Twitter tweets in English, Spanish, Italian and Dutch. With regard to age, we will consider the following classes: 18-24, 25-34, 35-49, 50-xx. With regard to personality traits, for each trait we will provide scores (between -0.5 and 0.5).</p> <p>More information about the task: <a href="https://pan.webis.de/clef15/pan15-web/author-profiling.html">Link</a></p>
PAN 22 Author Profiling: Profiling Irony and Stereotype Spreaders on Twitter (IROSTEREO)
<p><strong>TASK</strong></p> <p>With irony, language is employed in a figurative and subtle way to mean the opposite to what is literally stated. In case of sarcasm, a more aggressive type of irony, the intent is to mock or scorn a victim without excluding the possibility to hurt. Stereotypes are often used, especially in discussions about controversial issues such as immigration or sexism and misogyny. At PAN’22, we will focus on profiling ironic authors in Twitter. Special emphasis will be given to those authors that employ irony to spread stereotypes, for instance, towards women or the LGTB community. The goal will be to classify authors as ironic or not depending on their number of tweets with ironic content. Among those authors we will consider a subset that employs irony to convey stereotypes in order to investigate if state-of-the-art models are able to distinguish also these cases. Therefore, given authors of Twitter together with their tweets, the goal will be to profile those authors that can be considered as ironic.</p> <p><strong>DATA</strong></p> <p><strong>Input</strong></p> <p>The uncompressed dataset consists in a folder which contains:</p> <ul> <li>A XML file per author (Twitter user) with 200 tweets. The name of the XML file correspond to the unique author id.</li> <li>A truth.txt file with the list of authors and the ground truth.</li> </ul> <p>The format of the XML files is:</p> <pre><code class="language-xml"> <author lang="en"> <documents> <document>Tweet 1 textual contents</document> <document>Tweet 2 textual contents</document> ... </documents> </author></code></pre> <p>The format of the truth.txt file is as follows. The first column corresponds to the author id. The second column contains the truth label.</p> <pre><code> 2d0d4d7064787300c111033e1d2270cc:::I b9eccce7b46cc0b951f6983cc06ebb8:::NI f41251b3d64d13ae244dc49d8886cf07:::I 47c980972060055d7f5495a5ba3428dc:::NI d8ed8de45b73bbcf426cdc9209e4bfbc:::I 2746a9bf36400367b63c925886bc0683:::NI ...</code></pre> <p><strong>Evaluation</strong></p> <p>The performance of your system will be ranked by accuracy.</p> <p> </p> <p>More info on the task: <a href="https://pan.webis.de/clef22/pan22-web/author-profiling.html">https://pan.webis.de/clef22/pan22-web/author-profiling.html </a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.