Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Data set of the article: Language Bias in the Google Scholar Ranking Algorithm
<p>Data of investigation published in the article Cristòfol Rovira; Lluís Codina; Carlos Lopezosa Language Bias in the Google Scholar Ranking Algorithm. Future Internet, 2021, 13.</p> <p><strong>Abstract: </strong>The visibility of academic articles or conference papers depends on their being easily found in academic search engines, above all in Google Scholar. To enhance this visibility, search engine optimization (SEO) has been applied in recent years to academic search engines in order to optimize documents and, thereby, ensure they are better ranked in search pages (i.e., academic search engine optimization or ASEO). To achieve this degree of optimization, we first need to further our understanding of Google Scholar’s relevance ranking algorithm, so that, based on this knowledge, we can highlight or improve those characteristics that academic documents already present and which are taken into account by the algorithm. This study seeks to advance our knowledge in this line of research by determining whether the language in which a document is published is a positioning factor in the Google Scholar relevance ranking algorithm. Here, we employ a reverse engineering research methodology based on a statistical analysis that uses Spearman’s correlation coefficient. The results obtained point to a bias in multilingual searches conducted in Google Scholar with documents published in languages other than in English being systematically relegated to positions that make them virtually invisible. This finding has important repercussions, both for conducting searches and for optimizing positioning in Google Scholar, being especially critical for articles on subjects that are expressed in the same way in English and other languages, the case, for example, of trademarks, chemical compounds, industrial products, acronyms, drugs, diseases, etc.</p>
A set of generated Instagram Data Download Packages (DDPs) to investigate their structure and content
<p><strong>Instagram data-download example dataset</strong></p> <p>In this repository you can find a data-set consisting of 11 personal Instagram archives, or Data-Download Packages (DDPs).</p> <p> </p> <p><strong>How the data was generated</strong></p> <p>These Instagram accounts were all new and generated by a group of researchers who were interested to figure out in detail<br> the structure and variety in structure of these Instagram DDPs. The participants user the Instagram account extensively for approximately a week. The participants also intensively communicated with each other so that the data can be used as an example of a network. </p> <p>The data was primarily generated to evaluate the performance of de-identification software. Therefore, the text in the DDPs particularly contain many randomly chosen (Dutch) first names, phone numbers, e-mail addresses and URLS. In addition, the images in the DDPs contain many faces and text as well. The DDPs contain faces and text (usernames) of third parties. However, only content of so-called `professional accounts' are shared, such as accounts of famous individuals or institutions who self-consciously and actively seek publicity, and these sources are easily publicly available. Furthermore, the DDPs do not contain sensitive personal data of these individuals. </p> <p><br> <strong>Obtaining your Instagram DDP</strong></p> <p>After using the Instagram accounts intensively for approximately a week, the participants requested their personal Instagram DDPs by using the following steps. You can follow these steps yourself if you are interested in your personal Instagram DDP. </p> <p>1. Go to www.instagram.com and log in<br> 2. Click on your profile picture, go to *Settings* and *Privacy and Security*<br> 3. Scroll to *Data download* and click *Request download*<br> 4. Enter your email adress and click *Next*<br> 5. Enter your password and click *Request download*</p> <p>Instagram then delivered the data in a compressed zip folder with the format **username_YYYYMMDD.zip** (i.e., Instagram handle and date of download) to the participant, and the participants shared these DDPs with us.</p> <p> </p> <p><strong>Data cleaning</strong></p> <p>To comply with the Instagram user agreement, participants shared their full name, phone number and e-mail address. In addition, Instagram logged the i.p. addresses the participant used during their active period on Instagram. After colleting the DDPs, we manually replaced such information with random replacements such that the DDps shared here do not contain any personal data of the participants.</p> <p> </p> <p><strong>How this data-set can be used</strong></p> <p>This data-set was generated with the intention to evaluate the performance of the de-identification software. We invite other researchers to use this data-set for example to investigate what type of data can be found in Instagram DDPs or to investigate the structure of Instagram DDPs. The packages can also be used for example data-analyses, although no substantive research questions can be answered using this data as the data does not reflect how research subjects behave `in the wild'. </p> <p><br> <strong>Authors</strong></p> <p>The data collection is executed by Laura Boeschoten, Ruben van den Goorbergh and Daniel Oberski of Utrecht University. For questions, please contact l.boeschoten@uu.nl. </p> <p> </p> <p><strong>Acknowledgments</strong></p> <p>The researchers would like to thank everyone who participated in this data-generation project.</p>
Data Set Knowledge Graph (DSKG)
<p>We present the <strong>Data Set Knowledge Graph (<a href="http://dskg.org">DSKG.org</a>)</strong>, an <strong>RDF</strong> <strong>dataset about datasets </strong>that are <strong>linked to publications</strong> (modeled in the Microsoft Academic Knowledge Graph, MAKG) that mention the datasets. The metadata of the datasets is based on datasets that are registered in <strong>OpenAIRE</strong> and <strong>Wikidata</strong>.</p> <p><strong>What exactly do we provide?</strong></p> <ol> <li>Periodically updated <strong><a href="http://dskg.org">RDF dump files</a></strong> of the Data Set Knowledge Graph.</li> <li><strong><a href="http://dskg.org">URI resolution</a></strong> of the Data Set Knowledge Graph within the Linked Open Data.</li> <li>A publicly accessible <strong><a href="http://dskg.org">SPARQL endpoint</a></strong> containing the latest Dataset Knowledge Graph data.</li> </ol> <p><strong>How big is the Dataset Knowledge Graph?</strong></p> <p>The <a href="http://dskg.org">Dataset Knowledge Graph</a> models, among others,</p> <ul> <li>2,208 datasets from all scientific disciplines</li> <li>813,551 links to 634,803 unique papers</li> <li>1,169 authors of datasets</li> <li>208 ORCID IDs.</li> </ul> <p><strong>Potential use cases:</strong></p> <ul> <li>Use the DSKG for the development of semantic search engines (e.g. use the metadata of the linked publications of the datasets for advanced search capabilities)</li> <li>Easier data integration by using the RDF standard vocabulary DCAT and by linking resources to other data sources (e.g., combining the DSKG with other dataset collections in RDF).</li> <li>Data analysis to measure and award the provisioning of datasets (e.g., determine the scientific influence of datasets and authors).</li> </ul>
Regional Revised River Runoff Reanalysis (R5): historical and projected river runoff data set for the northwest of the European part of Russia
<p>This data set presents a uniform spatio-temporal assessment of projected river runoff for the northwest of the European part of Russia, which is based on two hydrological models (GR4J-REG and LSTM-REG), four General Circulation models (GFDL-ESM2M, HadGEM2-ES, IPSL-CM5A, and MIROC5), and three Representative Concentration Pathways (RCP2.6, RCP6.0, and RCP8.5). Each of the 24 gridded runoff data sets has daily temporal and 0.5° spatial resolution. They cover the geographical domain of 25–57° East and 55–70° North, and the temporal period from 2006 (2007 for LSTM-REG) to 2099.</p>
Data set on the main text of "A bright and fast source of coherent single photons"
<p>The data set that is presented in the main text is uploaded to the repository. Please note that all the data is scaled according to the axis on the paper, that means if the axis has a multiplication by 1e3 then the data is divided by 1e3.</p> <p>Each file is named after the corresponding subfigure.</p> <p>The preprint version of the article can be found in: https://arxiv.org/abs/2007.12654</p>
Fig. 27. Maximum likelihood tree from the concatenated data set with COI, 28S and 18S in Revision of the Merodon bombiformis group (Diptera: Syrphidae) - rare and endemic African hoverflies
Fig. 27. Maximum likelihood tree from the concatenated data set with COI, 28S and 18S rRNA gene sequences.
Data from: Enriching the ant tree of life: enhanced UCE bait set for genome-scale phylogenetics of ants and other Hymenoptera
1. Targeted enrichment of conserved genomic regions (e.g., ultraconserved elements or UCEs) has emerged as a promising tool for inferring evolutionary history in many organismal groups. Because the UCE approach is still relatively new, much remains to be learned about how best to identify UCE loci and design baits to enrich them. 2. We test an updated UCE identification and bait design workflow for the insect order Hymenoptera, with a particular focus on ants. The new strategy augments a previous bait design for Hymenoptera by (a) changing the parameters by which conserved genomic regions are identified and retained, and (b) increasing the number of genomes used for locus identification and bait design. We perform in vitro validation of the approach in ants by synthesizing an ant-specific bait set that targets UCE loci and a set of "legacy" phylogenetic markers. Using this bait set, we generate new data for 84 taxa (16/17 ant subfamilies) and extract loci from an additional 17 genome-enabled taxa. We then use these data to examine UCE capture success and phylogenetic performance across ants. We also test the workability of extracting legacy markers from enriched samples and combining the data with published data sets. 3. The updated bait design (hym-v2) contained a total of 2,590-targeted UCE loci for Hymenoptera, significantly increasing the number of loci relative to the original bait set (hym-v1; 1,510 loci). Across 38 genome-enabled Hymenoptera and 84 enriched samples, experiments demonstrated a high and unbiased capture success rate, with the mean locus enrichment rate being 2,214 loci per sample. Phylogenomic analyses of ants produced a robust tree that included strong support for previously uncertain relationships. Complementing the UCE results, we successfully enriched legacy markers, combined the data with published Sanger data sets, and generated a comprehensive ant phylogeny containing 1,060 terminals. 4. Overall, the new UCE bait design strategy resulted in an enhanced bait set for genome-scale phylogenetics in ants and likely all of Hymenoptera. Our in vitro tests demonstrate the utility of the updated design workflow, providing evidence that this approach could be applied to any organismal group with available genomic information.
Data set for "Quantitative and Qualitative bibliometric scope toward the Synthesis of Rose Oxide as a Natural Product in perfumery"
<p>This is the bibliometric data for "Quantitative and Qualitative bibliometric scope toward the Synthesis of Rose Oxide as a Natural Product in perfumery" study which were derived from SCOPUS database, on 23<sup>rd</sup> September 2019, based on title search.</p>
An elaborate data set on human gait and the effect of mechanical perturbations
<p>This data set includes measurements during trials of human walking on a treadmill with a movable base intended to provide data rich in content for the identification of the human's control system. The primary experiments were designed to longitudinally perturb the subject at the ground contact by randomly accelerating the belt. The marker locations (treadmill and human), treadmill accelerations, treadmill belt speeds, and the forces and moments from the dual force plates were measured during the trials.</p> <p>PeerJ Article: https://peerj.com/articles/918</p> <p>PeerJ Preprint: https://peerj.com/preprints/700/</p> <p>Paper source repository: https://github.com/csu-hmc/perturbed-data-paper</p>
Data sets for orthologous target pair analysis
<p>The set of all 803 originally identified orthologous target pairs (OTPs) and the subset of 222 OTPs with at least 10 shared compounds are provided herein. For each OTP both organisms, the target, the number of shared compounds,the OTP category, and the number of reference articles is reported. In addtion, the list of all 1149 candidate compounds and their human target assignments is provided. </p>
Dynamic testing of a four-storey building with reinforced concrete and unreinforced masonry wall: Data set
<p>This paper presents a publically available data set recorded during the shake-table test of a structure with reinforced concrete (RC) and unreinforced masonry (URM) walls. The shake-table test, performed at the TREES laboratory of EUCENTRE (Pavia, Italy), was part of a larger research initiative at EPFL (Lausanne, Switzerland) that addresses the seismic behaviour of mixed RC-URM structures. The half-scale test unit was subjected to several shakings of different intensity levels. The paper presents the geometry of the test unit, the properties of the construction materials, the instrumentation, and outlines the organization of the recorded data. Two sets of data are available: the unprocessed data and a second set of processed data where conventional and optical measurements are synchronised. This second set contains also some derived data, which allows to quickly plot key quantities such as base shear and top displacement. The aim of the paper is to provide all information required by the reader for analysing the test data and using it for validation purposes of numerical and mechanical models. The performance of the test unit is described in a companion paper.</p>
Pregnancy advertisement Japanese macaques, Data Set
<p>Data set used for the analyses of female Japanese macaques (<em>Macaca fuscata</em>) sexual signals of pregnancy (variations in behaviors, estrus calls and face color).</p> <p>Here are some of the variables tested: ecall=estrus calls, contactm=contact made, contactb=contact borken, apf=female approaches, apm=male approaches, rd=R/G ratio (redness), lum=luminance, pregmonth=period of interest with pcp:pre-conceptive, m1:1<sup>st</sup> month of pregnancy, m2: 2<sup>nd</sup> month of pregnancy.</p>
Data sets for SAR progression analysis
<p>Four compound data sets assembled from ChEMBL are provided that have been subjected to SAR progression analysis, to be published in Journal of Medicinal Chemistry. </p>
Supporting data: Grain-dependent responses of mammalian diversity to land-use and the implications for conservation set-aside
<p>Camera trap and live trap datasets underlying the analyses in an <em>Ecological Applications </em>paper (http://onlinelibrary.wiley.com/doi/10.1890/15-1363/abstract), provided in .csv format. Each row consists of a single trap night at a given location, with species in different columns. Old-growth forest, logged forest and oil palm plantation locations have the prefixes "Old", "Log" and "Palm", respectively. Values in each cell are the number of independent captures, as defined in the paper. </p>
Orthology guided transcriptome assembly of Italian ryegrass and meadow fescue for single nucleotide polymorphisms discovery (data set)
<p>Transcriptome sequencing was performed on ten samples (corresponding to six genotypes) of <em>Festuca pratensis</em> and ten samples (corresponding to six genotypes) of <em>Lolium multiflorum</em> and fourteen samples of<em> Lolium perenne</em> (corresponding to fourteen genotypes). Using the OGA approach, 18,952 non-redundant <em>F. pratensis</em> transcripts were assembled by combining the contigs of all six genotypes based on orthology with the <em>Brachypodium distachyon </em>proteome. Similarly, <em>19,036</em> non-redundant<em> L. multiflorum</em> transcripts were assembled and annotated. In total, 17,455 orthologous transcripts were shared between the transcriptomes of the two species. Out of these, 16,613 orthologous transcripts overlap with the previously published<em> L. perenne</em> transcriptome containing 19,279 non-redundant transcripts(fasta files). We identified SNPs, the following criteria were used to classify it as one of following three classes (1) intraspecific SNPs (INTRA), (2) interspecific SNPs in two-way comparison (INTER-2W) and (3) interspecific SNPs in three-way comparison (INTER-3W) (GFF files).</p>
Orthology guided transcriptome assembly of Italian ryegrass and meadow fescue (update data set)
<p>Transcriptome sequencing was performed on ten samples (corresponding to six genotypes) of <em>Festuca pratensis</em> and ten samples (corresponding to six genotypes) of <em>Lolium multiflorum</em> and fourteen samples of<em> Lolium perenne</em> (corresponding to fourteen genotypes). Using the OGA approach, 18,952 non-redundant <em>F. pratensis</em> transcripts were assembled by combining the contigs of all six genotypes based on orthology with the <em>Brachypodium distachyon </em>proteome. Similarly, <em>19,036</em> non-redundant<em> L. multiflorum</em> transcripts were assembled and annotated. In total, 17,455 orthologous transcripts were shared between the transcriptomes of the two species. Out of these, 16,613 orthologous transcripts overlap with the previously published<em> L. perenne</em> transcriptome containing 19,279 non-redundant transcripts(fasta files). We identified SNPs, the following criteria were used to classify it as one of following three classes (1) intraspecific SNPs (INTRA), (2) interspecific SNPs in two-way comparison (INTER-2W) and (3) interspecific SNPs in three-way comparison (INTER-3W) (GFF files).</p>
solar_home_system_data_log: Initial release of data sets and script
<p>In this first release, this repository includes three sets of data (date/time, temperature, current, and voltage) of over 6 months of electricity consumption of three households in an off-grid area in the state of Jharkhand, India. The goal of this data set is to be made open so that the community working in off-grid electricity access can get a sense of electricity consumption patterns and apply various analytical and visualization techniques.</p>
Secretin-like class B G protein-coupled receptor (GPCR) mutation data set
<p>Curated set of 2463 quantitative mutation data points covering 13 secretin-like class B G protein-coupled receptors.</p> <p>Annotated according to GPCRdb standards (http://gpcrdb.org/), complemented by GPCRdb Ballesteros-Weinstein numbers assigned using GPCRdb KNIME nodes (https://github.com/3D-e-Chem/knime-gpcrdb).</p> <p>The data set is build from data sets published in:</p> <p>- Siu et al. Nature 2013, 499: 444-449. doi:10.1038/nature12393</p> <p>- Hollenstein, de Graaf et al. Tr Pharmacol Sci 2014, 35: 12-22. doi: 10.1016/j.tips.2013.11.001</p> <p>- Yang, de Graaf et al. J Biol Chem 2016, 291: 12991-3004. doi: 10.1074/jbc.M116.721977</p>
Data set published in the IEEE TCAD article "Custom Multi-Cache Architectures for Heap-Manipulating Programs"
<p>This data set contains the results presented in the paper "Custom Multi-Cache Architectures for Heap-Manipulating Programs", published in the IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) in 2016.</p> <p>The data set consists of two parts, a Microsoft Excel file ('FPGA_implementation_results.xlsx') and a Matlab script ('plot_cache_performance.m', in combination with measurement results in an ascii file).</p> <p>The Excel file contains<br /> - the FPGA resource utilisation,<br /> - execution time measurements,<br /> - hit rate measurement of the multi-cache system,<br /> - and power measurements</p> <p>of different FPGA designs with different on-chip cache configurations. The resource utilisation is split into FPGA slices, LUTs, FlipFlops, DSP slices and block RAMs. Results in this file can be found in Table I-IV in the paper. Please refer to the paper for more information or email f.winterstein12@imperial.ac.uk.</p> <p>The Matlab script loads a data file ('cache_performance_N16384_L1') containing the hit rate measurements for different cache sizes of two direct-mapped cache with 64bit line width. The script produces a 3D 'skyscraper' plot, i.e. a grid of coloured bars. Each bar corresponds to the hit rate measured at the particular cache size configuration. The plot is saved in the file 'surf.pdf'. The script was used to produce Figure 4 of the paper. Please refer to the paper for more information or email f.winterstein12@imperial.ac.uk.</p> <p>In addition to this description, we include an author copy of the paper. Note that this is not the official version of the paper. Please cite the original IEEE TCAD article if you use the data.</p>
LIBER Academic Librarians RDS Data Set
<p>Included are the data set used in the LIBER/Second Academic Libraries Follow-Up Assessment, the survey instrument, and an about file are available.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.