Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
192
datasets available to search
ShareScore release 0.9.0
Dataset results
192 results for “website”
Repository of fact-checking websites and resources to combat climate mis/disinformation
<p>This dataset is the result of collaborative work for Deliverable 1.3 (WP1; T1.3) of the AGORA project. It compiles a list of fact-checking websites and resources dedicated to debunking climate change mis/disinformation. The identification of these resources was achieved by leveraging the expertise of our consortium and, therefore, most of the resources listed are in English, German, Italian and Spanish.</p>
BIObec's website analytics
<p>This document shows all the analytics from the BIObec's webpage, from the beginning of the project to its month 32 (January 2024).</p>
Survey of tourists and website visits in Öregrund 2023 from INCULTUM Sweden Pilot
<p>The data were collected using QR codes placed in different locations in Öregrund in summer 2023, sending the visitors to a dedicated website, and then to a survey. The dataset contains answers to the survey and some statistics about the website visits.</p>
CENTRINNO Dissemination - website and social media material
<p>The data bundle contains all supporting visual material for the website and all social media<br>publications deployment, including:<br>- Visual assets, photos, videos, social media posts, newsletters<br>- Pilot contact persons data</p>
Muestra de universidades norteamericanas (Estado de Massachusetts) y españolas con servicio de Diversidad. Universidad, titularidad, identificación del servicio y enlace website
<p>Muestra de estudio comparativo de website de diversidad de universidades españolas y estadounidenses</p>
Failure Handling Website States for Potato GUSC IoT System
<p><span>Screenshots of the Statuses section of </span><span>the website in ERS-FH after sensor node software and sensor </span><span>failures occur. The screenshots of the website after the sensor </span><span>node hardware and lost data failures are shown in Figure 7 of the paper.</span></p>
Withdrawn: DZD Core Data Set - first Version published at DZD Website for internal use (obsoleted by DOI 10.21961/mdm:45923)
<p>The German Center for Diabetes Research (DZD) conducts large clinical multicenter studies in the field of diabetes and metabolic research. In this vein, a core data set (CDS) which contains a list of clinical parameters relevant for joint studies in diabetes research was established in 2021 and published for internal use at the DZD website (https://www.dzd-ev.de/en/). In 2022 a FAIRified version of the data set was published at MDM portal (https://medical-data-models.org/). This entry shows the very first version of the core data set, published as an excel file before the FAIRification. It is intended as a supplement for an article about the FAIRification process: "The Journey to a FAIR CORE DATA SET for Diabetes Research in Germany"</p>
Charter School Websites: A Physical Education and Physical Activity Content Analysis
<p>Excel file for data set associated with analysis of 520 California elementary charter school websites' mentioning of physical education and physical activity opportunities.</p>
Crunchbase in RDF: A Large Data Set About Jobs, Websites, Organizations, News, People, Products, and Acquisitions
<p><strong>CrunchBase</strong> in an online platform providing information about startups and technology companies, including related entities such as the products they sell, key people they employ, and investments they made and received.</p> <p>We provide here an <strong>RDF data set of Crunchbase</strong> as of October 2015. The data set contains information about</p> <ul> <li>1,946,435 jobs</li> <li>1,348,449 websites</li> <li>567,937 organizations</li> <li>519,763 news</li> <li>430,093 people</li> <li>60,076 products, and</li> <li>33,127 acquisitions.</li> </ul> <p>The data set has been used, among other things, for data integration with financial data sources to evaluate the performance of particular companies and for monitoring news to find statements that are not in Crunchbase as an RDF knowledge graph yet.</p> <p>Note that the provided data set was created in October 2015 when all Crunchbase data was <strong>licensed under Creative Commons Attribution-NonCommercial License 4.0 (CC-BY-NC) and partly under Creative Commons Attribution License 4.0 (CC-BY)</strong>. Also the provied<strong> data set is licensed under these licenses.</strong> Concerning licensing of current Crunchbase data, we can refer to <a href="https://about.crunchbase.com/terms-of-service/">https://about.crunchbase.com/terms-of-service/</a>.</p> <p>For <strong>more information</strong> about the data set, see our paper <a href="http://dbis.informatik.uni-freiburg.de/content/team/faerber/papers/CrunchBaseWrapper_SWJ2017.pdf">A Linked Data Wrapper for CrunchBase.</a></p> <p>When you use the data set, please <strong>cite</strong> us as follows:</p> <blockquote> <p>Michael Färber, Carsten Menne, Andreas Harth. “A Linked Data Wrapper for CrunchBase”. In: Semantic Web Journal 9(4). IOS Press, 2018, pp. 505–5015. (<a href="https://dblp.org/rec/bibtex/journals/semweb/FarberMH18">BibTeX entry at DBLP</a>)</p> </blockquote>
Academic Excellence, Website Quality, SEO Performance: Is there a Correlation? - Dataset of measurements, test results and calculated ratings.
<p>This Dataset, in two files of xlsx format, contains the data of all measurements, test results and calculated ratings as they are described in the methodology of the research article "Academic Excellence, Website Quality, SEO Performance: Is there a Correlation".</p>
Web tracking data for 500 websites popular among Finnish web users
<p>This dataset includes observations of trackers present on the top 500 pages popular among Finnish web users as per Alexa. The data collection was conducted using TrackerTracker in five separate requests for five subsets of 100 sites each between 19.8.2017 and 20.8.2017. The tool used a tracker database from March 24, 2017. More methodology details are described in the associated journal article <a href="https://doi.org/10.23978/inf.87841">https://doi.org/10.23978/inf.87841</a></p>
.bed / .bim / .fam files, for 1kg, converted from the raw data on the PLINK website
<div> <div># Download the hg38 genome reference files from the PLINK website</div> <div>RUN wget -L https://www.dropbox.com/s/j72j6uciq5zuzii/all_hg38.pgen.zst</div> <div>RUN wget -L https://www.dropbox.com/scl/fi/fn0bcm5oseyuawxfvkcpb/all_hg38_rs.pvar.zst?rlkey=przncwb78rhz4g4ukovocdxaz -O all_hg38.pvar.zst</div> <div>RUN wget -L https://www.dropbox.com/scl/fi/u5udzzaibgyvxzfnjcvjc/hg38_corrected.psam?rlkey=oecjnk4vmbhc8b1p202l0ih4x -O all_hg38.psam</div> <br> <div># Download the hg38 related samples file from the PLINK website</div> <div>RUN wget -L https://www.dropbox.com/s/4zhmxpk5oclfplp/deg2_hg38.king.cutoff.out.id</div> <br> <div># Decompress the genome reference files</div> <div>RUN /plink-ng-master/2.0/bin/plink2 --zst-decompress all_hg38.pgen.zst all_hg38.pgen</div> <div>RUN rm all_hg38.pgen.zst</div> <br> <div>RUN /plink-ng-master/2.0/bin/plink2 --pfile all_hg38 vzs --allow-extra-chr --chr 1-22 --max-alleles 2 --remove deg2_hg38.king.cutoff.out.id --memory 6000 --make-bed --out 1kg_hg38</div> <br> <div># Replace rsIDs with chr:pos:ref:alt</div> <div>RUN awk 'BEGIN{OFS="\t"} {print $1,$1":"$4":"$6":"$5,$4,$6,$5}' 1kg_hg38.bim > 1kg_hg38_clean.bim</div> <div>RUN mv 1kg_hg38_clean.bim 1kg_hg38.bim</div> <div> </div> <div># Apply PLINK filtering (mAF > 0.1%, HWE p-value <1e-12, keep SNPs only)</div> <div>/plink-ng-master/2.0/bin/plink2 --bfile 1kg_hg38 --maf 0.001 --hwe 1e-12 --snps-only --make-bed --out 1kg_hg38_filtered --memory 6000</div> <div> </div> <div># Compress output files</div> <div>gzip 1kg_hg38_filtered.* -v --force</div> <div> </div> </div>
Replication package for "An Empirical Study of Q&A Websites for Game Developers"
<p><strong>Replication package for the paper "An Empirical Study of Q&A Websites for Game Developers"</strong></p> <p>This repository contains the datasets and scripts used to replicate the results from the paper "An Empirical Study of Q&A Websites for Game Developers".</p> <p>This is an exact copy of the repository on GitHub: https://github.com/asgaardlab/done-21-arthur-gamedev_qa_websites-code</p> <p><strong>Replication data</strong></p> <p>The datasets used to replicate the results for the paper can be found in the data directory (data/). These are the datasets we obtained after running all of the notebooks in this repository.</p> <p>Two of the studied websites are owned by companies (Epic and Unity) and we are not legally allowed to share the textual contents of the questions and answers as they are considered intellectual property. Therefore, instead of sharing the content of those posts, we included the URLs to all of the pages where the information used in the paper can be found, so that they can be crawled by future researchers.</p> <p>This is not an issue for Stack Overflow and the Game Development Stack Exchange, since that data is provided by Stack Exchange in the Stack Exchange Data Dump (https://archive.org/details/stackexchange).</p> <p><strong>Survey data:</strong> Unfortunately, our University's ethics board only allows us to share the survey responses in aggregated format, which is done in the paper. In this repository, we added the list of communities in which we shared the survey (data/surveyed_communities.csv).</p> <p><strong>Using this repository</strong></p> <p>If you are using the datasets provided in this repository, you just need to run the analysis notebook (code/analysis/paper_results.ipynb) to obtain the results as shown in the paper.</p> <p>Otherwise, if you want to run the whole pipeline from scratch, follow these steps:</p> <p>1. Download the data from Unity Answers and the UE4 AnswerHub from their websites (you can use the URLs provided in our datasets). Parse the HTML pages and extract the required information.</p> <p>2. Download the data from Stack Overflow and the Game Development Stack Exchange from the Stack Exchange Data Dump (https://archive.org/details/stackexchange). Run the notebooks to process the XML files from the Stack Exchange data dump (code/process_xml). For Stack Overflow, run the select_gamedev_posts.ipynb (code/process_xml/stackoverflow/select_gamedev_posts.ipynb) first.</p> <p>3. Run the text processing notebook (code/text_processing.ipynb).</p> <p>4. Run the topic modelling notebook (code/topic_modelling.ipynb).</p> <p>5. Run the topic comparisons notebook (code/topic_comparisons.ipynb).</p> <p>6. Finally, run the analysis notebook (code/analysis/paper_results.ipynb) to get the results as shown on the paper.</p>
Demonstaration of the use of Zenodo website to upload data and link it to the PersonalizeAF project
<p>This video demonstarates the use of Zenodo website to upload data and link it to the PersonalizeAF project</p>
1998 World Cup Website Access Logs
<p><strong>Description:</strong></p> <p>The access logs, as well as the accompanying description, are directly taken from [1] and include traffic of the 1998 World Cup website on three days as follows. The log files have the following naming format "wc_dayX_Y.gz"</p> <p>where:</p> <ul> <li>X is an integer that represents the day the access log was collected</li> <li>Y is an integer that represents the subinterval for a particular day</li> </ul> <p>This collection includes <em>three</em> log files containing the access traffic on three different days as listed below:</p> <pre><a>wc_day25_1.gz</a> May 20, 1998 -> TR1 <a>wc_day9_1.gz</a> May 4, 1998 -> TR2 <a>wc_day28_1.gz</a> May 23, 1998 -> TR3</pre> <p><strong>Format</strong></p> <p>The access logs from the 1998 World Cup Web site were originally in the Common Log Format. In order to reduce both the size of the logs and the analysis time the access logs were converted to a binary format (big endian = network order). Each entry in the binary log is a fixed size and represents a single request to the site. The format of a request in the binary log looks like:</p> <pre>struct request { uint32_t timestamp; uint32_t clientID; uint32_t objectID; uint32_t size; uint8_t method; uint8_t status; uint8_t type; uint8_t server; };</pre> <p>The fields of the request structure contain the following information:</p> <p><em>timestamp </em>- the time of the request, stored as the number of seconds since the Epoch. The timestamp has been converted to GMT to allow for portability. During the World Cup the local time was 2 hours ahead of GMT (+0200). In order to determine the local time, each timestamp must be adjusted by this amount.</p> <p><em>clientID </em>- a unique integer identifier for the client that issued the request (this may be a proxy); due to privacy concerns these mappings cannot be released; note that each clientID maps to exactly one IP address, and the mappings are preserved across the entire data set - that is if IP address 0.0.0.0 mapped to clientID X on day Y then any request in any of the data sets containing clientID X also came from IP address 0.0.0.0</p> <p><em>objectID</em> - a unique integer identifier for the requested URL; these mappings are also 1-to-1 and are preserved across the entire data set</p> <p><em>size</em> - the number of bytes in the response</p> <p><em>method</em> - the method contained in the client's request (e.g., GET).</p> <p><em>status</em> - this field contains two pieces of information; the 2 highest order bits contain the HTTP version indicated in the client's request (e.g., HTTP/1.0); the remaining 6 bits indicate the response status code (e.g., 200 OK).</p> <p><em>type</em> - the type of file requested (e.g., HTML, IMAGE, etc), generally based on the file extension (.html), or the presence of a parameter list (e.g., '?' indicates a DYNAMIC request). If the url ends with '/', it is considered a DIRECTORY.</p> <p><em>server</em> - indicates which server handled the request. The upper 3 bits indicate which region the server was at (e.g., SANTA CLARA, PLANO, HERNDON, PARIS); the remaining bits indicate which server at the site handled the request. All 8 bits can also be used to determine a unique server.</p> <p><strong>Reference</strong></p> <p>[1]<strong> </strong>M. Arlitt and T. Jin, "1998 World Cup Web Site Access Logs", August 1998. </p>
All data for proChIPdb database website
<p>All data for proChIPdb (previously ChIP-pro) database available at <a href="https://prochipdb.org/">prochipdb.org</a> as of 10/01/2021. The GitHub repository <a href="https://github.com/SBRG/ChIPdb.git">here</a> is also publicly available and contains the most up-to-date data.</p>
Hong Kong Jellyfish Project 2021 Website and iNaturalist observations
<p>Jellyfish are important organisms within marine ecosystems, although the extent of their occurrence and diversity is likely underestimated, particularly for biodiversity-rich locations such as the Indo-Pacific. The potential for citizen science to monitor phenomena associated with these Cnidarians over large spatial scales has been recognized in an increasingly broad array of locations, including the Mediterranean, South Africa, and the UK. Here, we were interested if such an approach could be used to understand more about their presence, seasonal occurrence and distribution in Hong Kong. To address these areas, the Hong Kong Jellyfish Project was launched in early 2021 with citizen scientists invited to submit photographs and simple information (date, time, location, number of jellyfish) to a project website or a collection project on the online citizen science biodiversity platform iNaturalist . For some of these observations, jellyfish were sampled and DNA analysis conducted to confirm morphologically-based identification. During 2021, over 380 observations of jellyfish were submitted through the website and iNaturalist project, with 19 species recorded as present. The species most frequently recorded were Cyanea nozakii and Rhopilema hispidum , and two new species records of Thysanostoma loriferum and Netrostoma setouchianum were documented for Hong Kong . There was a seasonal trend in observations, with most jellyfish seen in March-May. Finally, there was a broad geographical distribution of observations throughout Hong Kong’s coastal waters, with more observations made in the south/east. Together, these observations gathered by citizen scientists indicate a broad distribution of jellyfish – and the occurrence of previously undocumented species – in Hong Kong’s waters. </p>
Maryland private school websites' physical activity promotion
<p>Excel file of 387 Maryland (USA) private schools' websites' mentioning of PE (the term physical education, PE teacher, PE curriculum, PE dosage) and physical activity (intramural physical activity, interscholastic sports, physical activity images)</p>
A dataset for Customer Churn Prediction for Video Websites Incorporating Behavioral Sequence Features
<p>In order to study the issue of network customer churn, the iQiyi customer dataset was collected. Behavioral sequence features were extracted from it to build a deep learning model and experiments were conducted.Here, we provide the corresponding raw dataset, including all the data we used.</p>
Collection of datasets of the MCR-ALS website
<p>Below is the list of data sets from the mcrals.info website. Each dataset is identified by their original topic (see more information of the MCR-ALS website).</p> <p> </p> <p>* HPLC with diode array detection experiment data set</p> <p>1) Simulated HPLC-DAD non trilinear data sets: ntdata.zip</p> <p>2) Real HPLC-DAD data set (A): adataset.zip</p> <p>3) Real HPLC-DAD data set (B): bdataset.zip</p> <p><br> * Flow Injection Analysis data set</p> <p>Data set: fia_data.zip</p> <p><br> * UV and CD data of a DNA system</p> <p>Data set: dna_data.zip</p> <p><br> * 1H-NMR data of a platination reaction</p> <p>Data set: nmr_data.zip</p> <p><br> * MCR-ALS GUI (als2004) example data set</p> <p>Data set: data_gui2004.zip</p> <p><br> * Examples of the tutorial at Analytical Methods Special Issue</p> <p>Data set: AMtutorial.zip</p> <p><br> * Examples of the MCR-ALS GUI 2.0 manuscript</p> <p>Data set: MCR2datasets.zip</p> <p><br> * Examples of HSI images</p> <p>Data set: examplesHSI.zip</p> <p><br> * Examples of the ROIMCR manuscript</p> <p>Data set: data roimcr gui paper.zip</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.