Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11,174
datasets available to search
ShareScore release 0.7.1
Dataset results
11,174 results for “identifiers”
Chromatin activity identifies differential gene regulation across human ancestries
<p>This repository contains data related to:</p> <p>Chromatin activity identifies differential gene regulation across human ancestries</p> <p>Kade P. Pettie, Maxwell Mumbach, Amanda J. Lea, Julien Ayroles, Howard Y. Chang, Maya Kasowski, Hunter B. Fraser</p> <p> </p>
LauNuts: A Knowledge Graph to identify and compare geographic regions in the European Union
<p><strong>LauNuts</strong> is a RDF Knowledge Graph consisting of:</p> <ul> <li>Local Administrative Units (LAU) and</li> <li>Nomenclature of Territorial Units for Statistics (NUTS)</li> </ul> <p><a href="https://w3id.org/launuts">https://w3id.org/launuts</a></p>
Data for: Bivariate Genome-Wide Association Scan Identifies 6 Novel Loci Associated With Lipid Levels and Coronary Artery Disease.
<p>Summary of Bivariate GWAS scan results reported in:<br> <a href="https://pubmed.ncbi.nlm.nih.gov/30525989/">Bivariate Genome-Wide Association Scan Identifies 6 Novel Loci Associated With Lipid Levels and Coronary Artery Disease. </a>Siewert KM, Voight BF. Circ Genom Precis Med. 2018 Dec;11(12):e002239. doi: 10.1161/CIRCGEN.118.002239.</p> <p>PMID: 30525989 </p>
Data For: Identifying rare variants inconsistent with identity-by-descent in population-scale whole-genome sequencing data
<p>Simulation output and Genome-wide scan for nIBD variants in UK10K data as reported in:</p> <p>Identifying rare variants inconsistent with identity-by-descent in population-scale whole-genome sequencing data</p> <p>Johnson KE, Adams CJ, Voight BF. Methods Ecol Evol 2022 Nov;13(11): 2429–2442.</p> <p>Code available at: https://github.com/kelsj/EVICORD</p>
The First Transcriptomic Atlas of the Adult Lacrimal Gland Reveals Epithelial Complexity and Identifies Novel Progenitor Cells in Mice
<p>This project contains the R objects and code to reproduce the analyses and figures presented in the research article:</p> <p>'The First Transcriptomic Atlas of the Adult Lacrimal Gland Reveals Epithelial Complexity and Identifies Novel Progenitor Cells in Mice.' <em>Cells</em> <strong>2023</strong>, <em>12</em>, 1435. https://doi.org/10.3390/cells12101435</p> <p>Raw data (FASTQ files and CellRanger output files used for the preprocessing of individual datasets) can be found on Gene Expression Omnibus database (www.ncbi.nlm.nih.gov/geo/) under accession # GSE232146.</p>
Data release for paper "Waveform systematics in identifying gravitationally lensed gravitational waves: Posterior overlap method"
<p>This is the data release for the paper "Waveform systematics in identifying gravitationally lensed gravitational waves: Posterior overlap method", which is available on https://arxiv.org/abs/2306.12908.</p> <p>These results are derived from the gravitational-wave parameter-estimation results by the LIGO-Virgo-KAGRA Collaboration, released with the GWTC-1, GWTC-2, GWTC-2.1, and GWTC-3 catalogs under the following links:</p> <ul> <li> https://dcc.ligo.org/P1800370-v5/public</li> <li> https://dcc.ligo.org/P2000223-v7/public</li> <li> https://doi.org/10.5281/zenodo.6513631</li> <li> https://doi.org/10.5281/zenodo.5546663</li> </ul> <p>For the lensed-unlensed hypothesis test posterior overlap Bayes factors, we provide the following files for event pairs from within each observing run:</p> <ul> <li> blu_all_pairs_O1.txt</li> <li> blu_all_pairs_O2.txt</li> <li> blu_all_pairs_O3.txt</li> </ul> <p>In each file, the column "event_pair" contains the names of the two events from the pair sorted chronologically, the column "data_releases" contains the names of the data releases from which the posterior samples of each event were taken, the column "waveform" contains the name of the waveform model used in the parameter estimation for both sets of posteriors, and the column "log10blu" contains the log10 of the Bayes factors.</p> <p>The differences between runs for the same event pair, only including O1-O1, O2-O2, O3-O3 pairs, where at least one run gave log10blu>0, are also given in the file "blu_differences_pairs_with_log10blu_pos.txt". The column "event_pair" contains the event pairs, the columns "waveform_{1,2}" contain the names of the waveform models used in the parameter estimation for both sets of posteriors, the columns "data_releases_{1,2}" contain the the data releases from which the posterior samples of each event were taken, the columns "log10blu_{1,2}" contain the log10 Bayes factors, and the column "difference" contains the difference between "log10blu_1" and "log10blu_2".</p> <p>We also provide the following files corresponding to the appendix of the paper, analyzing overlaps between posterior samples for individual events:</p> <ul> <li> overlap_different_runs.txt</li> <li> overlap_same_run.txt</li> <li> rescaled_difference_single_event.txt</li> </ul> <p>The file "overlap_different_runs.txt" contains Bayes factors for a single event, but comparing the posteriors from different runs. The file "overlap_same_run.txt" contains Bayes factors for the overlap of a single run on a single event with itself. The file "rescaled_difference_single_event.txt" contains the difference between the results contained in the file overlap_different_runs.txt and the results in overlap_same_run.txt, taking the ones that produce the biggest difference, as per equation (A.1) in the paper.</p> <p>In these files, the column "event_name" is the name of the event, the column "data_release" or "data_releases" contains the name(s) of the data release(s) from which the posterior samples of each run were taken, the column "waveform" or "waveform_pair" contains the name(s) of the waveform model(s) used, and the column "log10blu" is the log10 Bayes factor obtained. In the file "rescaled_difference_single_event.txt", the columns "max_run_waveform" and "max_run_data_release" identify an entry from the "overlap_same_run.txt" file from which we use the "log10blu" to compute the value listed in the "difference" column using equation (A.1).<br> </p>
Modeling islet enhancers using deep learning identifies candidate causal variants at loci associated with T2D and glycemic traits
<p>Genetic association studies have identified hundreds of independent genetic signals associated with type 2 diabetes (T2D) and related traits. Despite these successes, the identification of specific causal variants underlying a genetic association signal remains challenging. In this study, we describe a deep learning method to analyze the impact of sequence variants on enhancers. Focusing on pancreatic islets, a relevant T2D tissue, we show that our model learns islet-specific transcription factor (TF) regulatory patterns and can be used to prioritize candidate causal variants. At 101 genetic signals associated with T2D and related glycemic traits where multiple variants occur in linkage disequilibrium, our method nominates a single causal variant for each association signal, including three variants previously shown to alter reporter activity in islet-relevant cell types. For another signal associated with blood glucose levels, we biochemically test all candidate causal variants from statistical fine-mapping using a pancreatic islet beta cell line and show biochemical evidence of allelic effects on TF binding for the model-prioritized variant. To aid in future research, we publicly distribute our model and islet enhancer perturbation scores across ~67 million variants. We anticipate that deep learning methods like the one presented in this study will enhance the prioritization of candidate causal variants for functional studies.</p>
Chironomid taxa relative abundance information and lake identifiers for: Changes in midge assemblages reflect climate and trophic gradients across north temperate and boreal lakes since the pre-industrial period
<p>File 1: Relative abundances for chironomid taxa used in the manuscript: Changes in midge assemblages reflect climate and trophic gradients across north temperate and boreal lakes since the pre-industrial period. Lake_ID corresponds to the lake IDs attributed to each lake sampled as part of the LakePulse Network</p> <p>File 2: Lake_ID, lake name, latitude, longitude, sampling date, province, and ecozone for the 69 lakes examined in the manuscript: Changes in midge assemblages reflect climate and trophic gradients across north temperate and boreal lakes since the pre-industrial period. </p>
A Complement Atlas identifies interleukin 6 dependent alternative pathway dysregulation as a key druggable feature of COVID-19.
<p>Improvements in COVID-19 treatments, especially for the critically ill, require deeper understanding of the mechanisms driving disease pathology. The complement system is a crucial component of innate host defense, but can also contribute to tissue injury. Although all complement pathways have been implicated in COVID-19 pathogenesis, the upstream drivers and downstream effects on tissue injury remain poorly defined. We demonstrate that complement activation is primarily mediated by the alternative pathway, and we provide a comprehensive atlas of the complement alterations around the time of respiratory deterioration. Proteomic and single-cell sequencing mapping across cell types and tissues reveals a division of labor between lung epithelial, stromal, and myeloid cells in complement production, in addition to liver-derived factors. We identify IL-6 and STAT1/3 signaling as an upstream driver of complement responses, linking complement dysregulation to approved COVID-19 therapies. Furthermore, an exploratory proteomic study indicates that inhibition of complement C5 decreases epithelial damage and markers of disease severity. Collectively, these results support complement dysregulation as a key druggable feature of COVID-19.</p>
Dataset for "A new method for identifying weather-induced power system stress using shadow prices"
<p>These are data accompanying "A new method for identifying weather-induced power system stress using shadow prices". They consist of</p> <ul> <li>solved network files (generated with <a href="https://github.com/PyPSA/pypsa-eur/">PyPSA-Eur</a>, here v0.6.1), used for the analysis,</li> <li>necessary data to reproduce the figures in the paper and supplementary material.</li> </ul> <p>The optimised network files are of the form `workflow_data/results/stressful-weather/optimum/{weather_year}_181_90m_c1.25_Co2L0.0-1H.nc` (for weather_years in {1980,...,2019}). Unsolved ones can be found in `workflow_data/networks/...`.</p> <p>The filenames in `plot_data/` indicate which figure the data are associated to (e.g. `plot_data/fig_1_hourly_costs.csv` contains the hourly electricity costs during the winter of all networks and is necessary for Figure 1). We also added weather data for all system-defining events (mean surface level pressure, 10m wind speed anomaly, 2m temperature anomaly) in .nc files.</p> <p>Find more information about how to use these data and how they were generated in the README of the GitHub repository: <a href="https://github.com/koen-vg/stressful-weather/tree/v0">https://github.com/koen-vg/stressful-weather/tree/v0</a>.</p>
CRISPR/dCas9-mediated DNA demethylation screen identifies driver epigenetic determinants of colorectal cancer (Processed data)
<p><strong>Background:</strong> Promoter hypermethylation of tumour suppressor genes is frequently observed during the malignant transformation of colorectal cancer (CRC). However, whether this epigenetic mechanism is an actual driver of cancer or is a mere consequence of the carcinogenic process remains to be elucidated.</p> <p><strong>Results: </strong>In this work we performed an integrative multi -omic approach to identify gene candidates with strong correlations between DNA methylation and gene expression in human CRC samples and a set of 8 colon cancer cell lines. As a proof of concept, we combined recent CRISPR-Cas9 epigenome editing tools (dCas9-TET1, dCas9-TET-IM) with a custom arrayed gRNA library to modulate the DNA methylation status of 56 promoters previously linked with strong epigenetic repression in CRC, and we monitored the potential functional consequences of such DNA methylation loss by means of a high-content cell proliferation screen. Overall, the epigenetic modulation of most of these DNA methylated regions had a mild impact in the reactivation of gene expression and in the viability of cancer cells. Interestingly, we found that epigenetic reactivation of RSPO2 in the tumour context was associated with a significant impairment in cell proliferation in p53-/- cancer cell lines and further validation with human samples demonstrated that the epigenetic silencing of RSPO2 is a mid-late event in the adenoma to carcinoma sequence.</p> <p><strong>Conclusions: </strong>These results highlight the potential role of DNA methylation as a driver mechanism of CRC and open up the venue for the identification of novel therapeutic windows based on the epigenetic reactivation of certain tumour suppressor genes.</p>
Data from: Identifying priority areas for spatial management of mixed fisheries using ensemble of multi-species distribution models. Panzeri D. et al., 2023, Fish and Fisheries
<p>Panzeri D.<sup>1</sup>, Russo T., Arneri E., Carlucci R., Cossarini G., Isajlović I., Krstulović Šifner S., Manfredi C., Masnadi F., Reale M., Scarcella G., Solidoro C., Spedicato M.T., Vrgoč N., W. Zupa, Libralato S<sup>2</sup>.</p> <p><sup>1 </sup>dpanzeri@ogs.it<br> <sup>2 </sup>slibralato@ogs.it</p> <p>Spatial fisheries management is widely used to reduce overfishing, rebuild stocks, and protect biodiversity. However, the effectiveness and optimization of spatial measures depend on accurately identifying ecologically meaningful areas, which can be difficult in mixed fisheries. To apply a method generally to a range of target species, we developed an ensemble of species distribution models (e-SDM) that combines general additive models, generalized linear mixed models, random forest, and gradient-boosting machine methods in a training and testing protocol. The e-SDM was used to integrate density indices from two scientific bottom trawl surveys with the geopositional data, relevant oceanographic variables from the three-dimensional physical-biogeochemical operational model, and fishing effort from the vessel monitoring system. The determined best distributions for juveniles and adults are used to determine hot spots of aggregation based on single or multiple target species. We applied e-SDM to juvenile and adult stages of 10 marine demersal species representing 60% of the total demersal landings in the central areas of the Mediterranean Sea. Using the e-SDM results, hot spots of aggregation and grounds potentially more selective were identified for each species and for the target species group of otter trawl and beam trawl fisheries. The results confirm the ecological appropriateness of existing fishery restriction areas and support the identification of locations for new spatial management measures.</p> <p>Data (csv) for Panzeri et al. 2023</p> <p>1. <a href="https://zenodo.org/api/files/0b1b7af4-6a3b-481d-8d5f-57cf02d20eaa/Ensemble_density_F%26F_D.Panzeri_et_al_2023.csv">Ensemble_density_F&F_D.Panzeri_et_al_2023.csv: CSV file with density values (column pred) in terms of number of individuals (log N/km2) for each species (column sp) and life stage (column age) for each grid cell (X = longitude and Y = latitude).</a> </p> <p>2. <a href="https://zenodo.org/api/files/0b1b7af4-6a3b-481d-8d5f-57cf02d20eaa/Ensemble_density_F%26F_D.Panzeri_et_al_2023.csv">Getis_hotspot_F&F_D.Panzeri_et_al_2023.csv: CSV file with Getis ord Gi* values (column Gi) derived from the previous file 1, developed for each species and life stage for each grid cell (X = longitude and Y = latitude).</a></p> <p>3. <a href="https://zenodo.org/api/files/0b1b7af4-6a3b-481d-8d5f-57cf02d20eaa/Ensemble_density_F%26F_D.Panzeri_et_al_2023.csv">Multispecies_HotSpot_F&F_D.Panzeri_et_al_2023.csv: Frequency map expressed as the number of species for each grid cell (column freq) that has the hotspot (previous file 2) above the third quartile.</a></p> <p> </p> <p> </p>
Multivariate analysis of FcR-mediated NK cell functions identifies unique clustering among humans and rhesus macaques - dataset
<p>Dataset from Tuyishime M, Spreng RL, et al. Multivariate analysis of FcR-mediated NK cell functions identifies unique clustering among humans and rhesus macaques. Frontiers in Immunology 2023 doi: 10.3389/fimmu.2023.1260377</p>
University of Kansas Field Station: Forest demography, 1980 – 2015. On ten study plots established on three management units all live trees with a dbh > 7.5 cm (3 in) were identified to species, measured, and tagged. Trees were initially measured in 1980/1981 and re-measured in three successive time periods: 1993/95; 2002/03; and 2014/15. Trees will be measured again in 2025/26.
In 1980 researchers at the University of Kansas initiated a long-term experiment monitoring the composition of oak-hickory forest communities at the University’s field station near Lawrence, Kansas. The purpose of the study was to determine how forest species composition varied temporally across distinct habitats that varied in topography, elevation, sun exposure, management history and successional stage. Ten permanent sites were sampled approximately each decade with data collection periods of 1980/81, 1993/95, 2002/03, and 2014/15. Trees with a minimum diameter at breast height (dbh) of ≥ 7.5 cm were tagged, identified to species and measured. Trees will be measured again in 2025/26.
Identified invertebrate bycatch from beetle pitfall traps at SJER and SOAP, 2017 - 2018 (repackaging of occurrences published by the NEON Biorepository Data Portal)
California permit requirements necessitated a more thorough identification of beetle pitfall samples than is typical of this protocol. These invertebrate bycatch samples therefore have occurrence associations that indicate their contents in both the NEON Biorepository and main NEON data portals. See NEON prototype dataset 9bc959c-148b-aaad-aa35-2d0805327428 available here.
Identifying hotspots of soil legacy phosphorus for soil P remediation on a cattle ranch in the Headwaters of the Everglades, South Central Florida, USA, 2020.
Phosphorus (P) cycling has been altered by human activities across various scales. 'Soil legacy P,' driven by agricultural changes such as excessive P fertilization and manure input, has led to P accumulation in soils. These legacy P reserves are long-term non-point sources, causing downstream eutrophication. Despite considerable scientific and policy interest, the fine-scale spatial heterogeneity, underlying drivers, and scales of variance of legacy P remain poorly understood. This dataset comprises of 1,438 surface soils sampled in 2020 across two typical subtropical grasslands managed for livestock production in South Central Florida, USA. The types of grasslands sampled were Intensively-managed or Improved pastures (IM), and Semi-native (SN) pastures. Chemical analysis was performed on the soil samples to determine three soil legacy P measurements (total P, Mehlich-1 and Mehlich-3 extractable P representing labile P pools) across the landscape. Other variables analyzed includedsoil organic matter, pH, available Fe and Al. Additionally, aboveground biomass samples were collected at a subset of soil sites, and analyzed for P content. The key questions regarding soil legacy P related to: its spatial variability and hotspots, variance distribution, relationship to land management and soil characteristics, and correlation with aboveground plant tissue P concentration. Subsequent analysis and spatial autoregressive modeling from this dataset revealed extreme variability of soil P at small scales, with diminishing variance as spatial scale increased, and increased variance in IM vs SN pastures. These findings enhance our understanding of the underlying drivers, spatial patterns, and variances of soil legacy P. Research suggests that broad pasture- or farm-level best management practices may be limited and less efficient, particularly for high-intensity pastures. Instead, management strategies to reduce soil legacy P could be implemented at fine scales, targeting P hots
LAGOS-US LOCUS v1.0: Data module of location, identifiers, and physical characteristics of lakes and their watersheds in the conterminous U.S.
This data package, LAGOS-US LOCUS v1.0, is one of the core data modules of the LAGOS-US platform that provides an extensible research-ready platform to study the 479,950 lakes and reservoirs larger than or equal to 1 ha in the conterminous US (48 states plus the District of Columbia). This data module contains information on the location, identifiers, and physical characteristics of lakes and their watersheds. The characteristics in this module include: variables that can be obtained from GIS data such as location and geometry; variables that can be derived using GIS processing such as lake watersheds and their geometry, lake glaciation history, and lake connectivity; and commonly used identifiers from GIS and other data products useful for linking with LAGOS-US. LOCUS is based on a snapshot of the high-resolution National Hydrography Dataset product available at the initiation of the project that provided the basis for locating, identifying, and characterizing the geometry of all lakes in LAGOS-US. The database design that supports the LAGOS-US research platform was created based on several important design features. Lakes are the fundamental unit of consideration, all lakes in the spatial extent must be represented (above a minimum size) and most information is connected to individual lakes. The design is modular, interoperable (the modules can be used with each other), and extensible (future database modules can be developed and used in the LAGOS-US research platform by others). Users are encouraged to use the other 2 core data modules that are part of the LAGOS-US platform: GEO (which includes geospatial ecological context at multiple spatial and temporal scales for lakes and their watersheds) and LIMNO (in situ lake surface-water physical, chemical, and biological measurements through time) that are each found in their own data packages.
Identifying the drivers and responses of abrupt changes across spatial and temporal scales in ecology: a review
Recently, the theoretical basis for understanding abrupt changes in ecosystems relative to regime shifts has emerged (Ratajczak et al. 2018). Abrupt changes are defined as, “substantial changes in the mean or variability of a system that occur in a short period of time relative to typical rates of change” (Ratajczak et al. 2018). Despite a driver-response framework to guide the environmental conditions under which abrupt changes are likely to occur coupled with many examples of unexpected changes from long-term ecological research, our theoretical basis of understanding of abrupt changes doesn’t include long-term scales, variability in drivers and responses, changes in the magnitude or direction of drivers, or the interactions among multiple drivers across spatiotemporal scales (sensu Ratajczak et al. 2018). Further, a critical review of the literature is lacking and essential to further understanding how common abrupt changes are detected and reported, as well as patterns and scales of drivers and responses of abrupt change in ecosystems. To address this knowledge gap, we searched the existing ecological literature for evidence and commonalities of abrupt change across ecosystems to identify commonalities and differences of abrupt change drivers and responses across terrestrial, freshwater, and marine ecosystems. We specifically asked the following questions: (1) How common are abrupt changes reported in the ecological literature? (2) How do driver and response temporal and spatial scales of abrupt changes compare and vary across terrestrial, freshwater, and marine ecosystem types? (3) Is there relative congruence between the temporal and spatial scale of drivers and responses? (3) What are common types of drivers and responses to abrupt changes, and how do they vary across ecosystem types? (4) What terms are most associated with drivers and responses of abrupt changes among ecosystem types?
ISPON: A New Dataset for Identifying Sources in Political Online News
<p>This dataset contains a set of annotations for informational news sources (such as eyewitnesses, public officials, academic experts, reports, or other documentation) that provide support for claims made within online political news articles. Our dataset contains fine-grained annotations on the sources cited within each article, including in-text notations highlighting the words or phrases signaling a source. The dataset comprises annotations for nearly 2,500 articles covering 47 outlets. In addition, the dataset includes a larger set of >150,000 URLs from 92 outlets.</p>
Identifying Machine-Paraphrased Plagiarism
<p>README.txt</p> <p>Title: <em>Identifying Machine-Paraphrased Plagiarism</em><br> Authors: Jan Philip Wahle, Terry Ruas, Tomas Foltynek, Norman Meuschke, and Bela Gipp<br> contact email: wahle@gipplab.org; ruas@gipplab.org;<br> Venue: iConference<br> Year: 2022<br> ================================================================<br> <strong>Dataset Description:</strong></p> <p><em><strong>Training:</strong></em><br> 200,767 paragraphs (98,282 original, 102,485paraphrased) extracted from 8,024 Wikipedia (English) articles (4,012 original, 4,012 paraphrased using the SpinBot API).</p> <p><em><strong>Testing:</strong></em><br> SpinBot: <br> arXiv - Original - 20,966; Spun - 20,867<br> Theses - Original - 5,226; Spun - 3,463<br> Wikipedia - Original - 39,241; Spun - 40,729<br> <br> SpinnerChief-4W: <br> arXiv - Original - 20,966; Spun - 21,671<br> Theses - Original - 2,379; Spun - 2,941<br> Wikipedia - Original - 39,241; Spun - 39,618<br> <br> SpinnerChief-2W: <br> arXiv - Original - 20,966; Spun - 21,719<br> Theses - Original - 2,379; Spun - 2,941<br> Wikipedia - Original - 39,241; Spun - 39,697</p> <p>================================================================<br> Dataset Structure:</p> <p><strong>[human_evaluation]</strong> folder: human evaluation to identify human-generated text and machine-paraphrased text. It contains the files (original and spun) as for the answer-key for the survey performed with human subjects (all data is anonymous for privacy reasons).</p> <p>NNNNN.txt - whole document from which an extract was taken for human evaluation<br> key.txt.zip - information about each case (ORIG/SPUN)<br> results.xlsx - raw results downloaded from the survey tool (the extracts which humans judged are in the first line)<br> results-corrected.xlsx - at the very beginning, there was a mistake in one question (wrong extract). These results were excluded.</p> <p><br> <strong>[automated_evaluation]: </strong>contains all files used for the automated evaluation considering [spinbot] (https://spinbot.com/API) and [spinnerchief] (http://developer.spinnerchief.com/API_Document.aspx).</p> <ul> <li>Each paraphrase tool folder contains:</li> <li><strong>[corpus] </strong>and<strong> [vectors]</strong> sub-folders.</li> <li>For [spinnerchief], two variations are included, with 4-word-chaging ratio (default) and 2-word-chaging ratio. </li> </ul> <p><strong>[vectors] sub-folder</strong> contains the average of all word vectors for each paragraph. Each line has the number of dimensions of the word embeddings technique used (see paper for more details) followed by its respective class (i.e., label mg or og). Each file belongs to one class, either "mg" or "og". The values are comma-separated (.csv). The extension is .arff can be read as a normal .txt file.</p> <ul> <li>The word embedding technique used is described in the file name with the following structure: <technique>-<type>-mean-<data>.arff . Where</li> </ul> <p><em><technique></em> - d2v - doc2vec<br> google - word2vec<br> fasttextnw - fastText without subwording<br> fasttextsw - fastText with subwording<br> glove - Glove<br> <br> Details for each technique used can be found in the paper.<br> <br> <em><type> - </em> arxivp - arXiv paragraph split<br> thesisp - Theses paragraph split<br> wikip - Wikipedia paragraph split (wikipedia_paragraph_vector_train are the vectors used for training. It follows the same wikip structure) </p> <p>Details for each technique used can be found in the paper referenced at the start of this README file.</p> <p><strong>[corpus] sub-folder:</strong> contains de raw text (No pre-processing) used for train and test at a paragraph level.</p> <ul> <li>The Spun paragraphs used for <strong>training</strong> are only generated using the <strong>SpinBot tool</strong>. For test both SpinBot and SpinnerChief are used. </li> <li>The paragraph split is generated by selecting paragraphs from the original documents with 3 or more sentences. Each folder is divided in mg (i.e., machine-generated through SpinBot and SpinnerChief) and og (i.e., original-generated file). the document split is not avaiable since our experiments only use the paragraph level.</li> <li>Machine Learning models: SVM, Naive Bayes, and Logistic Regression. The grid search for hyperparameter adjustments for the machine learning classifiers is described in the paper.</li> </ul> <p>@incollection{WahleRFM22,<br> title = {Identifying {{Machine-Paraphrased Plagiarism}}},<br> booktitle = {Information for a {{Better World}}: {{Shaping}} the {{Global Future}}},<br> author = {Wahle, Jan Philip and Ruas, Terry and Folt{\’y}nek, Tom{\’a}{\v s} and Meuschke, Norman and Gipp, Bela},<br> editor = {Smits, Malte},<br> year = {2022},<br> volume = {13192},<br> pages = {393--413},<br> publisher = {{Springer International Publishing}},<br> address = {{Cham}},<br> doi = {10.1007/978-3-030-96957-8_34},<br> isbn = {978-3-030-96956-1 978-3-030-96957-8},<br> }</p> <p> </p> <p>For our previous publication using only SpinBot and Wikipedia articles for document and paragraph split, please see the following publication. The dataset used is hosted in <a href="https://deepblue.lib.umich.edu/data/concern/data_sets/2801pg45f?locale=en">DeepBlue</a></p> <p><br> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.