Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,359

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,359 results for “Online”

Learn how ShareScore rates datasets ↗
zenodo44/100

Hanslick-Online/hsl-vms-data: Hanslick "Vom Musikalisch-Schönen" - TEI/XML v1.0.0

<p>Die TEI/XML-Annotation beinhaltet den ausgehend von Druckvorlagen transkribierten Text der 10 Auflagen von Eduard Hanslicks &Auml;sthetik-Traktat &quot;Vom Musikalisch-Sch&ouml;nen. Ein Beitrag zur Revision der &Auml;sthetik der Tonkunst&quot; mit zus&auml;tzlich inhaltlicher Erschlie&szlig;ung von Personen, Orten und Werken und der M&ouml;glichkeit von textlichen Vergleichen. Der Index bietet Namen von Personen und Orten gem&auml;&szlig; der Schreibweise von <a href="https://lobid.org/gnd">GND</a> und <a href="https://www.geonames.org/">Geonames</a>. Ebenfalls bereitgestellt werden Faksimiles der gedruckten Publikation. Die Traktat-Auflagen werden originalgetreu mit beibehaltener Rechtschreibung wiedergegeben. Lediglich eindeutige Druckfehler wurden stillschweigend richtiggestellt.</p> <p>Die Abs&auml;tze wurden (mit Ausnahme der Gedichte) beibehalten und &Auml;nderungen zwischen einzelnen Auflagen mit folgenden Symbolen angezeigt: komprimiert &rarr; &amp;; segmentiert &rarr; a und b; hinzugef&uuml;gt &rarr; 3. Dezimalstelle; gestrichen &rarr; Zahlenl&uuml;cke. Die Seitenzahlen und Seitenumbr&uuml;che der Textvorlagen k&ouml;nnen jeweils im Editor-Men&uuml; eingeblendet oder ausgeblendet werden.</p> <p>Versumbr&uuml;che werden mittels Schr&auml;gstrich dargestellt. Die Fu&szlig;noten wurden in Endnoten mit Sprungmarke umgewandelt und zum Zweck einer exakteren Zuordnung statt mit einem Asterisk *, mit fortlaufender Nummerierung ausgezeichnet. Die &Uuml;berschriften der Einzelkapitel sind bis zur 7. Auflage nur im Inhaltsverzeichnis zu finden und wurden bei der Transkription von Auflage 1&ndash;7 zur leichteren Navigation eingef&uuml;gt. Die verschiedenen Schriftarten der Original-Auflagen (Fraktur f&uuml;r Deutsch, lateinischer Schrifttypus f&uuml;r andere Sprachen) wurden nicht beibehalten, da es sich hier um Konventionen des neunzehnten Jahrhunderts ohne semantische Aussagekraft handelt: Hervorhebungen (Sperrdruck in Original-Auflagen) werden kursiv wiedergegeben, Fettdrucke beibehalten.</p> <p><strong>Full Changelog</strong>: <a href="https://github.com/Hanslick-Online/hsl-vms-data/commits/v1.0.0">https://github.com/Hanslick-Online/hsl-vms-data/commits/v1.0.0</a></p>

openother-openApr 2023View details →
zenodo44/100

Anonymisation for data sharing in practice [Online Workshop. Recording]

<p>The goal of this event was to show trainers the tools they need to teach the fundamentals of data anonymisation and disclosure control in training sessions while also giving them hands-on experience with current open source technologies (sdcMicro). Some of the concepts and techniques presented, included k-anonymity, top/bottom coding and aggregation with practical examples and recommendations on incorporating anonymisation into research designs.</p> <p>&nbsp;</p> <p>The video is available on<a href="https://www.youtube.com/watch?v=JeJ6OOxXZwo&amp;t=328s"> the&nbsp;CESSDA Training&nbsp;YouTube channel</a>.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

How to Ensure Researchers Share Their FAIR Data: Practical Tips and Tools [Online Workshop, Recording]

<p>The online hands-on workshop was aimed at trainers and support staff covering critical elements of data sharing and available tools and resources for supporting Open Science including:<br> &bull; Open Science resources and Data Management Planning<br> &bull; Consent and Ethical considerations<br> &bull; Legislation and Licence frameworks<br> The objectives of the workshop were i) to raise awareness of key tools and resources available for Open Science training ii) to enable a platform to exchange ideas regarding key training topics and iii)n to provide training materials and worksheets for future reuse.<br> The workshop consisted of presentations, demos, a roundtable discussion on ethical considerations, a showcase of licence frameworks at different European archives and an exercise with all participants fostering an exchange of experiences focused on learnt lessons.</p> <p>The video is available on<a href="https://www.youtube.com/watch?v=uztTCRFRZHg"> the&nbsp;CESSDA Training&nbsp;YouTube channel</a>.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Journal and Data Archive Collaboration Forum [online event recording]

<p>The availability of research data underlying articles published in journals is becoming a common practice in scientific communication. The European Commission and other funders of scientific research have set high expectations for scientists towards openness and availability of scientific work and results. Scientific publishers, through journals and scholarly publications are the main point of realising open science in practice.<br> <br> This event was part of the continuous Journals Outreach initiative (<a href="https://www.cessda.eu/Training/Journals-outreach">https://www.cessda.eu/Training/Journals-outreach</a>), bringing together CESSDA service providers (SPs) with Social Science &amp; Humanities Journals. <strong>Its target audiences were publishers, editors, researchers, and CESSDA Service providers.&nbsp;</strong>The event was also an opportunity for publishers/journals to highlight new initiatives in research data services linked to scientific publications.<br> <br> The video is available on<a href="https://www.youtube.com/watch?v=zCKoyzLifkg"> the&nbsp;CESSDA Training&nbsp;YouTube channel</a>.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Italian TikTok users online behaviour patterns and social attitudes (survey)

<p>Survey of 500 young TikTok users (18-35) in Italy covering online behaviour patterns and social attitudes</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

TDMentions: A Dataset of Technical Debt Mentions in Online Posts

<p># TDMentions: A Dataset of Technical Debt Mentions in Online Posts (version 1.0)</p> <p>TDMentions is a dataset that contains mentions of technical debt from Reddit, Hacker News, and Stack Exchange. It also contains a list of blog posts on Medium that were tagged as technical debt. The dataset currently contains approximately 35,000 items.&nbsp;</p> <p>## Data collection and processing</p> <p>The dataset is mainly collected from existing datasets. We used data from:</p> <p>- the archive of Reddit posts by Jason Baumgartner (available at [https://pushshift.io](https://pushshift.io),&nbsp;<br> - the archive of Hacker News available at Google&#39;s BigQuery (available at [https://console.cloud.google.com/marketplace/details/y-combinator/hacker-news](https://console.cloud.google.com/marketplace/details/y-combinator/hacker-news)), and the Stack Exchange data dump (available at [https://archive.org/details/stackexchange](https://archive.org/details/stackexchange)).&nbsp;<br> - the [GHTorrent](http://ghtorrent.org) project&nbsp;<br> - the [GH Archive](https://www.gharchive.org)</p> <p>The data set currently contains data from the start of each source/service until 2018-12-31. For GitHub, we currently only include data from 2015-01-01.</p> <p>We use the regular expression `tech(nical)?[\s\-_]*?debt` to find mentions in all sources except for Medium. We decided to limit our matches to variations of technical debt and tech debt. Other shorter forms, such as TD, can result in too many false positives. For Medium, we used the tag `technical-debt`.&nbsp;</p> <p>## Data Format</p> <p>The dataset is stored as a compressed (bzip2) JSON file with one JSON object per line. Each mention is represented as a JSON object with the following keys.</p> <p>- `id`: the id used in the original source. We use the URL path to identify Medium posts.<br> - `body`: the text that contains the mention. This is either the comment or the title of the post. For Medium posts this is the title and subtitle (which might not mention technical debt, since posts are identified by the tag).<br> - `created_utc`: the time the item was posted in seconds since epoch in UTC.&nbsp;<br> - `author`: the author of the item. We use the username or userid from the source.<br> - `source`: where the item was posted. Valid sources are:<br> &nbsp;&nbsp; &nbsp;- HackerNews Comment<br> &nbsp;&nbsp; &nbsp;- HackerNews Job<br> &nbsp;&nbsp; &nbsp;- HackerNews Submission<br> &nbsp;&nbsp; &nbsp;- Reddit Comment<br> &nbsp;&nbsp; &nbsp;- Reddit Submission<br> &nbsp;&nbsp; &nbsp;- StackExchange Answer<br> &nbsp;&nbsp; &nbsp;- StackExchange Comment<br> &nbsp;&nbsp; &nbsp;- StackExchange Question<br> &nbsp;&nbsp; &nbsp;- Medium Post<br> - `meta`: Additional information about the item specific to the source. This includes, e.g., the subreddit a Reddit submission or comment was posted to, the score, etc. We try to use the same names, e.g., `score` and `num_comments` for keys that have the same meaning/information across multiple sources.</p> <p>This is a sample item from Reddit:</p> <p>```JSON<br> {<br> &nbsp; &quot;id&quot;: &quot;ab8auf&quot;,<br> &nbsp; &quot;body&quot;: &quot;Technical Debt Explained (x-post r/Eve)&quot;,<br> &nbsp; &quot;created_utc&quot;: 1546271789,<br> &nbsp; &quot;author&quot;: &quot;totally_100_human&quot;,<br> &nbsp; &quot;source&quot;: &quot;Reddit Submission&quot;,<br> &nbsp; &quot;meta&quot;: {<br> &nbsp; &nbsp; &quot;title&quot;: &quot;Technical Debt Explained (x-post r/Eve)&quot;,<br> &nbsp; &nbsp; &quot;score&quot;: 1,<br> &nbsp; &nbsp; &quot;num_comments&quot;: 0,<br> &nbsp; &nbsp; &quot;url&quot;: &quot;http://jestertrek.com/eve/technical-debt-2.png&quot;,<br> &nbsp; &nbsp; &quot;subreddit&quot;: &quot;RCBRedditBot&quot;<br> &nbsp; }<br> }<br> ```</p> <p>## Sample Analyses</p> <p>We decided to use JSON to store the data, since it is easy to work with from multiple programming languages. In the following examples, we use [`jq`](https://stedolan.github.io/jq/) to process the JSON.</p> <p>### How many items are there for each source?</p> <p>```<br> lbzip2 -cd postscomments.json.bz2 | jq &#39;.source&#39; | sort | uniq -c<br> ```</p> <p>### How many submissions that mentioned technical debt were posted each month?</p> <p>```<br> lbzip2 -cd postscomments.json.bz2 | jq &#39;select(.source == &quot;Reddit Submission&quot;) | .created_utc | strftime(&quot;%Y-%m&quot;)&#39; | sort | uniq -c<br> ```</p> <p>### What are the titles of items that link (`meta.url`) to PDF documents?</p> <p>```<br> lbzip2 -cd postscomments.json.bz2 | jq &#39;. as $r | select(.meta.url?) | .meta.url | select(endswith(&quot;.pdf&quot;)) | $r.body&#39;<br> ```</p> <p>### Please, I want CSV!</p> <p>```<br> lbzip2 -cd postscomments.json.bz2 | jq -r &#39;[.id, .body, .author] | @csv&#39;<br> ```</p> <p>Note that you need to specify the keys you want to include for the CSV, so it is easier to either ignore the meta information or process each source.</p> <p>Please see [https://github.com/sse-lnu/tdmentions](https://github.com/sse-lnu/tdmentions) for more analyses</p> <p># Limitations and Future updates</p> <p>The current version of the dataset lacks GitHub data and Medium comments. GitHub data will be added in the next update. Medium comments (responses) will be added in a future update if we find a good way to represent these.</p>

opencc-by-4.0May 2019View details →
zenodo40/100

Online supplementary data linked to the publication "Aubenas-les-Alpes (S-E France). Part III – Last and final part of the mammalian assemblage with some comments on the palaeoenvironment and palaeobiogeography" doi:10.1016/j.annpal.2019.03.001

<p>Online supplementary appendix including the list of Oligocene localities and associated faunal lists compared to Aubenas-les-Alpes, and the size estimation of the non-predatory species for the construction of Fig.10.</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

ISPON: A New Dataset for Identifying Sources in Political Online News

<p>This dataset contains a set of annotations for informational news sources (such as eyewitnesses, public officials, academic experts, reports, or other documentation) that provide support for claims made within online political news articles. Our dataset contains fine-grained annotations on the sources cited within each article, including in-text notations highlighting the words or phrases signaling a source. The dataset comprises annotations for nearly 2,500 articles covering 47 outlets. In addition, the dataset includes a larger set of &gt;150,000 URLs from 92 outlets.</p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

Data Review Handphone Toko Online XYZ Market Place XYZ

<p>Data set review Handphone di toko online di market place yang dipergunakan untuk natural language processing.&nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo40/100

Conclusiones del debate sobre evaluación online. Jornadas Conversación casUSAL

<p>Conclusiones del debate sobre evaluaci&oacute;n <em>online</em> celebrado el 2 de julio de 2020 en las Jornadas Conversaci&oacute;n casUSAL&nbsp;(<a href="https://facultadcero.org/encuentroUSAL/">https://facultadcero.org/encuentroUSAL/</a>)</p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

SAN Online Databases

<p>Collection of online resources of interest to SAN topics. The database is available at SAN website. This is an evolving database. Latest version=V1</p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

Online Supplemental Materials for: "Total Error and Variability Measures for the Quarterly Workforce Indicators and LEHD Origin Destination Employment Statistics in OnTheMap"

<p>This archive contains supplementary materials for the published manuscript.</p> <p>We report results from the first comprehensive total quality evaluation of five major indicators in the U.S. Census Bureau&#39;s Longitudinal Employer-Household Dynamics (LEHD) Program Quarterly Workforce Indicators (QWI): total flow-employment, beginning-of-quarter employment, full-quarter employment, average monthly earnings of full-quarter employees, and total quarterly payroll. Beginning-of-quarter employment is also the main tabulation variable in the LEHD Origin-Destination Employment Statistics (LODES) workplace reports as displayed in OnTheMap (OTM), including OnTheMap for Emergency Management. We account for errors due to coverage; record-level non-response; edit and imputation of item missing data; and statistical disclosure limitation. The analysis reveals that the five publication variables under study are estimated very accurately for tabulations involving at least 10 jobs. Tabulations involving three to nine jobs are a transition zone, where cells may be fit for use with caution. Tabulations involving one or two jobs, which are generally suppressed on fitness-for-use criteria in the QWI and synthesized in LODES, have substantial total variability but can still be used to estimate statistics for untabulated&nbsp; aggregates as long as the job count in the aggregate is more than 10.</p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

Supplementary Online Material to the paper: Modelling and empirical validation of carbon stock accumulation during the forest transition in France 1850-2015

<p><strong>Supplementary Online Material to the paper:</strong></p> <p><strong>Modelling and empirical validation of carbon stock accumulation during the forest transition in France 1850-2015</strong></p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

Popularity Dataset for Online Stats Training

<p>This is a dataset&nbsp;used for the online stats training website (<a href="https://www.rensvandeschoot.com/tutorials/">https://www.rensvandeschoot.com/tutorials/</a>) and is based on the data used by&nbsp;&nbsp;<a href="https://doi.org/10.1016/j.adolescence.2009.12.004">van de Schoot, van der Velden, Boom, and Brugman (2010)</a>.</p> <p>The dataset is based on a study that investigates an association between popularity status and antisocial behavior from at-risk adolescents (n = 1491), where gender and ethnic background are moderators under the association. The study distinguished subgroups within the popular status group in terms of overt and covert antisocial behavior.For more information on the sample, instruments, methodology, and research context, we refer the interested readers to <a href="https://doi.org/10.1016/j.adolescence.2009.12.004">van de Schoot, van der Velden, Boom, and Brugman (2010)</a>.</p> <p>&nbsp;</p> <p>Variable name&nbsp;&nbsp; Description</p> <p>Respnr =&nbsp; Respondents&rsquo; number</p> <p>Dutch =&nbsp; Respondents&rsquo; ethnic background (0 = Dutch origin, 1 = non-Dutch origin)</p> <p>gender&nbsp; = Respondents&rsquo; gender (0 = boys, 1 = girls)</p> <p>sd =&nbsp;&nbsp;Adolescents&rsquo; socially desirable answering patterns</p> <p>covert =&nbsp;Covert antisocial behavior</p> <p>overt =&nbsp; Overt antisocial behavior</p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

Emotion and Diversity in Online Publics - Supporting Materials

<p>Supporting materials for the article &quot;Counterpublic, Deliberative Sphere, Incubator, Battleground: Emotion and Diversity in Online Publics&quot;.</p> <p>&quot;</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Conclusions of the online assessment debate held on July 2, 2020, at the University of Salamanca (Spain)

<p>Conclusions of the online assessment debate held on July 2, 2020, at the University of Salamanca (Spain)</p>

opencc-by-4.0Oct 2020View details →
zenodo40/100

Dataset and trained models belonging to the article 'Distant reading patterns of iconicity in 940.000 online circulations of 26 iconic photographs'

<p>Quantifying Iconicity - Zenodo</p> <p><br> ## The Dataset<br> This dataset contains the material collected for the article &quot;Distant reading 940,000 online circulations of 26 iconic photographs&quot; (to be) published in New Media &amp; Society (DOI: 10.1177/14614448211049459). We identified 26 iconic photographs based on earlier work (Van der Hoeven, 2019). The Google Cloud Vision (GCV) API was subsequently used to identify webpages that host a reproduction of the iconic image. The GCV API uses computer vision methods and the Google index to retrieve these reproductions. The code for calling the API and parsing the data can be found on GitHub: https://github.com/rubenros1795/ReACT_GCV.</p> <p>The core dataset consists of .tsv-files with the URLs that refer to the webpages. Other metadata provided by the GCV API is also found in the file and manually generated metadata. This includes:<br> - the URL that refers specifically to the image. This can be an URL that refers to a full match or a partial match<br> - the title of the page<br> - the iteration number. Because the GCV API puts a limit on its output, we had to reupload the identified images to the API to extend our search. We continued these iterations until no more new unique URLs were found<br> - the language found by the ``langid`` Python module [link](https://github.com/saffsd/langid.py), along with the normalized score.<br> - the labels associated with the image by Google<br> - the scrape date</p> <p>Alongside the .tsv-files, there are several other elements in the following folder structure:</p> <p>```<br> ├── data<br> │&nbsp;&nbsp; ├── embeddings<br> │&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── doc2vec<br> │&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── input-text<br> │&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── metadata<br> │&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── umap<br> │&nbsp;&nbsp; └── evaluation<br> │&nbsp;&nbsp; └── results<br> │&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── diachronic-plots<br> │&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; └── top-words<br> │&nbsp;&nbsp; └── tsv<br> ```</p> <p>1. The ```/embeddings``` folder contains the doc2vec models, the training input for the models, the metadata (id, URL, date) and the UMAP embeddings used in the GMM clustering. Please note that the date parser was not able to find dates for all webpages and for this reason not all training texts have associated metadata.<br> 2. The ```/evaluation``` folder contains the AIC and BIC scores for GMM clustering with different numbers of clusters.<br> 3. The ```/results``` folder contains the top words associated with the clusters and the diachronic cluster prominence plots.</p> <p>## Data Cleaning and Curation<br> Our pipeline contained several interventions to prevent noise in the data. First, in between the iterations we manually checked the scraped photos for relevance. We did so because reuploading an iconic image that is paired with another, irrelevant, one results in reproductions of the irrelevant one in the next iteration. Because we did not catch all noise, we used Scale Invariant Feature Transform (SIFT), a basic computer vision algorithm, to remove images that did not meet a threshold of ten keypoints. By doing so we removed completely unrelated photographs, but left room for variations of the original (such as painted versions of Che Guevara, or cropped versions of the Napalm Girl image). Another issue was the parsing of webpage texts. After experimenting with different webpage parsers that aim to extract &#39;relevant&#39; text it proved too difficult to use one solution for all our webpages. Therefore we simply parsed all the text contained in commonly used html-tags, such as ```&lt;p&gt;```, ```&lt;h1&gt;``` etc.</p>

openNov 2020View details →
zenodo40/100

Productos a la venta online de la marca española "Mustang"

<p>Dataset que contiene informaci&oacute;n referente a todos los productos (ropa y accesorios) que est&aacute;n a la venta en la p&aacute;gina web de la marca espa&ntilde;ola Mustang. Estos atributos se han obtenido mediante el uso de WebScraping en una pr&aacute;ctica de la UOC</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

I-BiDaaS - CAIXA - Online Banking - Tokenised Dataset

<p>This dataset contains the information of mobile-to-mobile bank transfers ordered through online banking (web or application) by CaixaBank&rsquo;s customers. It was generated for the assessing that the controls applied to user authentication are applied adequately (e.g. second factor authentication) in accordance with PSD2 regulation and depending on the context of the bank transfer.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

GECCO Industrial Challenge 2019 Dataset: A water quality dataset for the 'Internet of Things: Online Event Detection for Drinking Water Quality Control' competition at the Genetic and Evolutionary Computation Conference 2019, Prague, Czech Republic.

<p>Dataset &nbsp;of the &#39;Internet of Things: Online Event Detection for Drinking Water Quality Control&#39; competition hosted at&nbsp;The Genetic and Evolutionary Computation Conference (GECCO)&nbsp;July 13th-17th 2019, Prague, Czech Republic</p> <p>&nbsp;</p> <p>The task of the&nbsp;competition was&nbsp;to develop an anomaly detection algorithm for a water- and environmental data set.</p> <p>&nbsp;</p> <p>Included in zenodo:&nbsp;</p> <p>1. Original train dataset of water quality data provided to participants (identical to&nbsp;gecco2019_train_water_quality.csv)</p> <p>2.&nbsp;Call for Participation</p> <p>3. Rules and Description of the Challenge</p> <p>4. Resource Package provided to&nbsp;participants</p> <p>5. The complete dataset, consisting of train, test and validation merged together&nbsp;(gecco2019_all_water_quality.csv)</p> <p>6.&nbsp;The&nbsp;test&nbsp;dataset, which was used for creating the leaderboard on the server&nbsp; (gecco2019_test_water_quality.csv)</p> <p>7.&nbsp;The train dataset, which participants had available for training their models&nbsp; (gecco2019_train_water_quality.csv)</p> <p>8.&nbsp;The&nbsp;&nbsp;validation dataset, which was used for the end results for the challenge (gecco2019_valid_water_quality.csv)</p> <p>&nbsp;</p> <p>The challenge required the participants to submit a program for event detection. A training dataset was available to the participants (gecco2019_train_water_quality.csv). During the challenge the participants were able to upload a version of their program to out online platform, where this version was scored against the testing dataset (gecco2019_test_water_quality.csv), thus an intermediate leaderboard was available. To avoid overfitting against this dataset, at the end of the challenge, the end result was created from scoring with the validation dataset (gecco2019_valid_water_quality.csv).&nbsp;</p> <p>Train, Test, Validation dataset are from the same measuring station and are in chronological order. So the timestamps from the test dataset begin directly after the train timestamps, while the validation timestamps begin directly after the test timestamps.&nbsp;</p> <p>&nbsp;</p> <p>The competition was organized by:</p> <p>F. Rehbach, S. Moritz,&nbsp;T. Bartz-Beielstein (TH K&ouml;ln)</p> <p>&nbsp;</p> <p>The dataset was provided by:</p> <p>Th&uuml;ringer Fernwasserversorgung and&nbsp;IMProvT research project</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>Internet of Things: Online Event Detection for Drinking Water Quality Control</p> <p>&nbsp;</p> <p>Description:</p> <p>For the 8th time in GECCO history, the SPOTSeven Lab is hosting an industrial challenge in cooperation with various industry partners. This years challenge, based on the 2018 challenge, is held in cooperation with &quot;Th&uuml;ringer Fernwasserversorgung&quot; which provides their real-world data set. The task of this years competition is to develop an anomaly detection algorithm for the water- and environmental data set. Early identification of anomalies in water quality data is a challenging task. It is important to identify true undesirable variations in the water quality. At the same time, false alarm rates have to be very low.</p> <p><br> Competition Opens: End of January/Start of February 2019<br> Final Submission: 30 June 2019</p> <p>Official webpage:</p> <p><a href="https://www.th-koeln.de/informatik-und-ingenieurwissenschaften/gecco-challenge-2019_63244.php">https://www.th-koeln.de/informatik-und-ingenieurwissenschaften/gecco-challenge-2019_63244.php</a></p> <p>&nbsp;</p>

opencc-by-4.0Jan 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record