Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,085

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,085 results for “Documentation”

Learn how ShareScore rates datasets ↗
zenodo40/100

Rapid Documentation Of Avifaunal Diversity of Mahananda Wildlife Sanctuary, Darjeeling, West Bengal, India

<p>The dataset &ldquo;Rapid Documentation Of Avifaunal Diversity of Mahananda Wildlife Sanctuary, Darjeeling, West Bengal, India&rdquo; is published by Nature Mates Nature Club</p> <p>The Mahananda Wildlife Sanctuary is situated in the Darjeeling district of West Bengal, India, on the slopes of the Himalayas, bounded by the Teesta and Mahananda rivers.</p> <p>The sanctuary encompasses an expansive area of 159 square kilometres within a reserve forest and was initially established as a game sanctuary in the year 1955. In 1959, the sanctuary was designated with the purpose of safeguarding the Indian Gaur and royal Bengal tiger, both of which were confronted with the imminent risk of extinction.</p> <p>The entire land area is partitioned into 33 distinct forest blocks, which are further categorised into four ranges: East, West, North, and South. The forest blocks encompass the following areas: Punding, Bandar jhola, Jogi jhora, Kuni, Choklong, Upper Champasari, Gulma valley, Silihhita, West Sevoke, East Sevoke, North Sevoke, Jhenaikuri, Lower Ghoramara, Upper Ghoramara, Gola, Ruyem, Andera, Chawa, Samaardanga, Lower Champasari, Singimari, Gulma, Mahanadi, Sukna (Part 1), Rongdong, Kaklong, Mohorganj, Panchenai, Hatisar, Kyananuka, Adalpur, Chumta, and Laltong.&nbsp;</p> <p>The geographical scope of our investigation includes the Rong Tong block within the sanctuary.</p> <p>The dataset presented encompasses the avian species documented within the Rong Tong region during a comprehensive biodiversity survey conducted on the 9th and 10th of June in the year 2023.</p> <p>At the taxonomic level, each species has been identified and categorized at the species or genus level. The bird community encompasses a wide array of 72 species, each of which is methodically categorized into 32 families and 11 orders.</p> <p>Resource Contacts</p> <p>Name: <strong>Tarak Samanta</strong></p> <p>Position: Research Affiliate</p> <p>Organization: Nature Mates-Nature Club</p> <p>Address: 6/7 Bijoygarh Kolkata-700032</p> <p>Email: taraksamanta995@gmail.com</p> <p>Orcid :<a href="https://orcid.org/0000-0001-6809-0549"> https://orcid.org/0000-0001-6809-0549</a></p> <p>Home page:<a href="http://www.naturematesindia.org/"> http://www.naturematesindia.org/</a></p> <p>&nbsp;</p> <p>Name: <strong>Asim Giri</strong></p> <p>Position: Field Assistant</p> <p>Organization: Padmaja Naidu Himalayan Zoological Park</p> <p>Address: Darjeeling, West Bengal, IN</p> <p>Email: <a href="mailto:giriasim2013@gmail.com">giriasim2013@gmail.com</a></p> <p>Orcid :<a href="https://orcid.org/0000-0002-4450-3609"> https://orcid.org/0000-0002-4450-3609</a></p> <p>&nbsp;</p> <p>Name:<strong> Nabarun Mondal</strong></p> <p>Position: Laboratory Technician&nbsp;</p> <p>Organization: Kothari Medical Centre</p> <p>Email: <a href="mailto:nabarunmondal865@gmail.com">nabarunmondal865@gmail.com</a></p> <p>&nbsp;</p> <p>Name: <strong>Sayandeep Mondal</strong></p> <p>Position: System Engineer&nbsp;&nbsp;</p> <p>Organization: Siemens Mobility</p> <p>Email: <a href="mailto:sayandeepmaity10@gmail.com">sayandeepmaity10@gmail.com</a></p> <p>&nbsp;</p> <p>Name: <strong>Arjan Basu Roy</strong></p> <p>Position: Secretary</p> <p>Organization: Nature Mates-Nature Club</p> <p>Address: 6/7 Bijoygarh Kolkata-700032</p> <p>Email: basuroyarjan@gmail.com</p> <p>Orcid :<a href="https://orcid.org/0000-0001-9872-3562"> https://orcid.org/0000-0001-9872-3562</a></p> <p>Home page:<a href="http://www.naturematesindia.org/"> http://www.naturematesindia.org/</a></p> <p>&nbsp;</p> <p>Name: <strong>Rishin Basu Roy</strong></p> <p>Position: Researcher</p> <p>Organization: Nature Mates-Nature Club</p> <p>Address: 6/7 Bijoygarh Kolkata-700032</p> <p>Email: <a href="mailto:wildrishin@gmail.com">wildrishin@gmail.com</a></p> <p>Orcid : <a href="https://orcid.org/0000-0002-9638-818X">https://orcid.org/0000-0002-9638-818X</a></p> <p>Home page: <a href="http://www.naturematesindia.org/">http://www.naturematesindia.org/</a></p> <p>&nbsp;</p> <p>Name: <strong>Lina Chatterjee</strong></p> <p>Position: Research Affiliate</p> <p>Organization: Nature Mates-Nature Club</p> <p>Address: 6/7 Bijoygarh Kolkata-700032</p> <p>Email: lina.linachatterjee@gmail.com</p> <p>Orcid :<a href="https://orcid.org/0000-0002-5626-5046"> https://orcid.org/0000-0002-5626-5046</a></p> <p>Home page:<a href="http://www.naturematesindia.org/"> http://www.naturematesindia.org/</a></p> <p><br> Name:<strong> Nivedita Sengupta</strong></p> <p>Position: Intern</p> <p>Organization: Nature Mates-Nature Club</p> <p>Address: 6/7 Bijoygarh Kolkata-700032</p> <p>Email: niveditasngpta.ns@gmail.com</p> <p>Orcid :<a href="https://orcid.org/0000-0003-1085-7385"> https://orcid.org/0000-0003-1085-7385</a></p> <p>Home page:<a href="http://www.naturematesindia.org/"> http://www.naturematesindia.org/</a></p> <p><br> Name: <strong>Vijay Barve</strong></p> <p>Position: Research Advisor</p> <p>Organization: Nature Mates-Nature Club</p> <p>Address: 6/7 Bijoygarh Kolkata-700032</p> <p>Email: vijay.barve@gmail.com</p> <p>Orcid :<a href="https://orcid.org/0000-0002-4852-2567"> https://orcid.org/0000-0002-4852-2567</a></p> <p>Home page:<a href="http://www.naturematesindia.org/"> http://www.naturematesindia.org/</a></p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Forschungsprojekt: Digitalisierung, Klassifikationen und Gesundheits-Apps. Dataset B - Document analysis for health app users and developers

<p><strong>Allgemeine Hinweise Data Set B</strong></p> <ol> <li>Titel des Forschungsprojekts</li> </ol> <p>Digitale Gesundheitsklassifikationen in Apps - Praktiken und Probleme ihrer Entwicklung und situativen Anwendung. Projektleitung: Prof. Dr. Rainer Diaz-Bone. Bearbeitung: Valeska Cappel Dipl. Soz., Miriam Kutt (Hilfsassistenz). Laufzeit: 2019-2023. Finanzierung: Schweizer Nationalfonds.</p> <p>2. Prim&auml;rforscherInnen:</p> <ul> <li>Rainer Diaz-Bone</li> <li>Valeska Cappel</li> <li>Miriam Kutt</li> </ul> <p>3. Publikationsjahr:</p> <ul> <li>2023</li> </ul> <p>4. Hinweise zur Verf&uuml;gbarkeit</p> <p>Die Daten werden &uuml;ber LORY (Lucerne Open Repository) dauerhaft zug&auml;nglich gemacht.</p> <ul> <li>Alle Dateien die sich auf die App-Entwicklung beziehen beginnen mit E</li> <li>Alle Dateien die sich auf die App-Nutzung beziehen beginnen mit N</li> </ul> <p>5. Fachgebiet</p> <ul> <li>Soziologie</li> </ul> <p>6. Kategorie und Schlagw&ouml;rter</p> <ul> <li>Gesundheitswesen</li> <li>Selbstvermessung</li> <li>Gesundheits-Apps</li> <li>Klassifikationen</li> <li>Pragmatismus</li> <li>Economics of convention</li> <li>Soziologie der Konventionen</li> <li>Digitalisierung</li> </ul> <p>7. Abstract, wozu die Daten erhoben wurden</p> <p>Die Daten wurden f&uuml;r eine Dokumentenanalyse erhoben. Ausgew&auml;hlt wurden (1) Daten von forschungsrelevanten Akteuren oder Institutionen, die nicht f&uuml;r ein Interview gewonnen werden konnten und (2) Medienberichte, Anleitungen, Beschreibungen von pr&auml;ventiven Gesundheits-Apps, Wissenschaftliche Berichte (White Papers, Ausschreibungen) sowie Bewertungen von diesen Gesundheits-Apps aus dem &bdquo;App-Store&ldquo; und &bdquo;Google-Play-Store&ldquo; (Plattformen zum Download von Apps). Ziel war es mit diesen Dokumenten weitere Analysen durchzuf&uuml;hren, die diskursanlytische Aussagen &uuml;ber die Entstehung und Nutzung von pr&auml;ventiven Gesundheits-Apps sowie Entwicklungen im Feld der digitalen Gesundheit zulassen. Gleichzeitig wurden auch Dokumente, wie Nutzer-Bewertungen erhoben, um die Interviews zu erg&auml;nzen aus einer pragmatischen Perspektive Aushandlungs- und Probleml&ouml;sungsprozesse im Umgang mit Gesundheits-Apps und der damit verbundenen Technologie zu untersuchen.</p> <p>8. Untersuchungsgebiet</p> <p>Das Untersuchungsgebiet liegt im Bereich der digitalen Gesundheit und beschr&auml;nkt sich im Speziellen auf die Prozesse der Entwicklung von Gesundheits-Apps, sowie die Nutzung der Gesundheits-Apps. Untersuchungsgebiet waren &ouml;ffentlich zug&auml;ngliche Medien sowie die Gesundheits-App selbst. Dabei wurden Unternehmen, Zeitschriften, Blogs und Apps, die sich konkret mit der App-Entwicklung, der App-Nutzung oder der Berichterstattung &uuml;ber pr&auml;ventive Gesundheits-Apps besch&auml;ftigen ausgew&auml;hlt zur Dokumentenerhebung ausgew&auml;hlt. Bei der Auswahl wurden der Fokus darauf gelegt Inhalte auszuw&auml;hlen, die sich nicht auf Gesundheits-App als Medizinprodukte konzentrieren, sondern auf pr&auml;ventive Gesundheits-Apps, die kein spezifisches Krankheitsbild adressieren, sondern einen allgemeinen positiven Gesundheitszustand herstellen oder erhalten sollen.&nbsp;</p> <p>9. Gesamtheit auf die generalisiert werden k&ouml;nnte (&bdquo;Grundgesamtheit&ldquo;)</p> <p>Insgesamt wurden ca. 300 Dokumente erhoben.</p> <p>10. Auswahlverfahren und Stichproben</p> <p>Die Daten im Projekt wurden anhand qualitativer Methoden gewonnen. Die F&auml;lle und Daten wurden &uuml;ber die Methode der &bdquo;Theoretical-Sampling-Technik&ldquo; ausgew&auml;hlt. Die genauen Begr&uuml;ndungen zur Auswahl der F&auml;lle wurden im Verlauf des Projektes theoretisch erarbeitet. Dabei wurde sich dem Forschungsfeld der pr&auml;ventiven Gesundheits-Apps mit heuristischen Vermutungen angen&auml;hert, die anhand der Konzepte der Theorie der Konventionen und einer machttheoretischen Perspektive Foucaults entwickelt wurden. Gesundheit und Gesundheitshandlungen wurden dabei aus einer pragmatischen Perspektive als ein Ergebnis von Koordinationsbem&uuml;hungen zwischen Akteuren, Gegenst&auml;nden, Technologien und Machtstrukturen verstanden. Die Auswahl der Dokumente st&uuml;tzte sich besonders auf die Memos und Inhalte der vorher gef&uuml;hrten Interviews und den daraus entwickelten heuristische Fragestellungen w&auml;hrend des Forschungsprozesses.</p> <p>11. Erhebungszeitraum</p> <p>Die Dokumente wurden in dem Zeitraum 2019-2023 erhoben. Das Datum in den Dateinamen bezieht sich immer auf den Erhebungszeitpunkt</p> <p>Sprache</p> <ul> <li>Deutsch und Englisch</li> </ul> <p>12. Gr&ouml;&szlig;e des Datensatzes</p> <ul> <li>ca. 45 MB</li> </ul> <p>13. Verwendete Dateiformate und notwendige Software</p> <ul> <li>Dateiformat: PDF (Portable Document Format) und RTF (Rich Text Format)</li> </ul> <p>Software:</p> <ul> <li>RTF Standard-Textprogrammen auf unterschiedlichen Betriebssystemen (Bspw. Word, Wordpad, LibreOffice, OpenOffice)</li> <li>PDF: PDF-Programme/Reader oder auch ATLAS</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

TexBiG Dataset for Analysing Complex Document Layouts in the Digital Humanities

<p>This is the dataset for the paper&nbsp;&quot;A Dataset for Analysing Complex Document Layouts in the Digital Humanities and its Evaluation with Krippendorff &rsquo;s Alpha&quot; in its second version, containing an update of the test images (without annotations) from the paper &quot;Drawing the Same Bounding Box Twice? Coping Noisy Annotations in Object Detection with Repeated Labels&quot;. Organization of the dataset is also updated to make it easier to use.</p> <p>TexBiG (from the German Text-Bild-Gef&uuml;ge, meaning Text-Image-Structure) is a document layout analysis dataset for historical documents in the late 19th and early 20th century. The dataset provides instance segmentation (bounding boxes and polygons/masks) annotations for 19 different classes with more then 52.000 instances. The added test images can be used to make submission on the leaderboard on <a href="https://eval.ai/web/challenges/challenge-page/2078/overview">EvalAI</a>.&nbsp;</p> <p>The <a href="https://zenodo.org/record/6885144/files/Annotations_Guideline.pdf?download=1">annotation guideline</a> can be found in the first of the dataset.</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

MEDDOPLACE Corpus: Gold Standard annotations for Medical Documents Place-related Content Extraction

<p><strong>MEDDOPLACE</strong>&nbsp;stands for MEDical DOcument PLAce-related Content Extraction. It is a shared task and set of resources focused on the detection, normalization (entity linking/toponym resolution) and classification of different kinds of places, as well as related types of information such as clinical departments, nationalities or patient movements, in medical documents in Spanish.</p> <p>This repository includes the corpus' <strong>train and test sets</strong> in multiple formats, as well as the <strong>SNOMED gazetteer</strong>, <strong>cross-mapping</strong> between SNOMED and MeSH and the <strong>multilingual silver standard in 8 languages&nbsp;</strong>(Catalan, English, French, Italian, Dutch, Portuguese, Romanian and Swedish). For more information, please check the attached README file.</p> <p>MEDDOPLACE was developed by the Barcelona Supercomputing Center's NLP for Biomedical Information Analysis and used as part of IberLEF 2023. For more information on the corpus, annotation scheme and task in general, please visit: <a href="https://temu.bsc.es/meddoplace">https://temu.bsc.es/meddoplace</a>.</p> <p>&nbsp;</p> <p><strong>Please cite if you use this resource:</strong></p> <p>Salvador Lima-L&oacute;pez, Eul&agrave;lia Farr&eacute;-Maduell, Antonio Miranda-Escalada, Vicent Briv&aacute;-Iglesias and Martin Krallinger. NLP applied to occupational health: MEDDOPROF shared task at IberLEF 2021 on automatic recognition, classification and normalization of professions and occupations from medical texts. In Procesamiento del Lenguaje Natural, 67. 2021.</p> <pre><code>@article{meddoplace, title={MEDDOPLACE Shared Task overview: recognition, normalization and classification of locations and patient movement in clinical texts}, author={Lima-L&oacute;pez, Salvador and Farr&eacute;-Maduell, Eul&agrave;lia and Briv&aacute;-Iglesias, Vicent and Gasco-Sanchez, Luis and Krallinger, Martin}, journal = {Procesamiento del Lenguaje Natural}, volume = {71}, year={2023}, issn = {1135-5948},<br>DOI = {10.26342/2023-71-23}, url = {http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561/3961}, pages = {301--311} }</code></pre> <p><strong>Related Links:</strong></p> <p>- MEDDOPLACE website: <a href="https://temu.bsc.es/meddoplace">https://temu.bsc.es/meddoplace</a></p> <p>- MEDDOPLACE overview paper: <a href="http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561">http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561</a></p> <p>- Annotation Guidelines (Spanish): <a href="https://doi.org/10.5281/zenodo.7775234">https://doi.org/10.5281/zenodo.7775234</a></p> <p>- Annotation Guidelines (English): <a href="https://doi.org/10.5281/zenodo.7928145">https://doi.org/10.5281/zenodo.7928145</a></p> <p><strong>License</strong></p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <p><strong>Contact</strong></p> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-L&oacute;pez (&lt;salvador [dot] limalopez [at] gmail [dot] com&gt;)<br>- Martin Krallinger (&lt;krallinger [dot] martin [at] gmail [dot] com&gt;)</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

"It was recorded on Sunday, morning of the 28th of September as some of the slower runners of the Berlin Marathon made it past Torstrasse near my flat. Iwas out to buy some bread for breakfast, but Iusually bring a camera and my Edirol R-1 recorder whenever Igo out. Since Iwas freshly returned to Berlin Iguess Iwas sensitive to the more antiquated sounds which still survive there, like that of the organ grinder. Iam generally interested in how human beings are replacing the presence of Nature with an artificial environment made entirely by human hands (and thus far more understandable, it is hoped). In this new Human Nature, the sounds of Nature are also Human made. Iwrite about these things, but Ialso use the sounds in my videos and my interactive and generative media work, so generally Iam wandering around building up my archive of media documents for use as material in future works." [Baruch/ gottlieb]17 in Collecting Sounds. Online Sharing of Field Recordings as Cultural Practice

"It was recorded on Sunday, morning of the 28th of September as some of the slower runners of the Berlin Marathon made it past Torstrasse near my flat. Iwas out to buy some bread for breakfast, but Iusually bring a camera and my Edirol R-1 recorder whenever Igo out. Since Iwas freshly returned to Berlin Iguess Iwas sensitive to the more antiquated sounds which still survive there, like that of the organ grinder. Iam generally interested in how human beings are replacing the presence of Nature with an artificial environment made entirely by human hands (and thus far more understandable, it is hoped). In this new Human Nature, the sounds of Nature are also Human made. Iwrite about these things, but Ialso use the sounds in my videos and my interactive and generative media work, so generally Iam wandering around building up my archive of media documents for use as material in future works." [Baruch/ gottlieb]17

opencc-by-4.0Dec 2019View details →
dryad40/100

Documenting twenty years of the contracted labor-intensive forestry workforce on National Forest System lands in the United States

Open the record for dataset details and reuse information.

publicJun 2023View details →
dryad40/100

Improving distribution models of sparsely-documented disease vectors by incorporating information on related species via joint modeling

Open the record for dataset details and reuse information.

publicApr 2024View details →
edi40/100

Photographic documentation of forestry treatments of 77 forest plots at Sagehen Creek Field Station, 2016-2019

Data package contains sets of photographs taken at 77 forest monitoring plots within the Sagehen Experimental Forest. These plots are a subset of 500+ forest monitoring plots established in 2004 and 2005 for the purpose of testing strategically-placed land area treatments (SPLATS) that impede forest fire progression (Vaillant 2008, UC Berkeley Doctoral Dissertation). Sites were photographed at various intervals before and after prescribed forestry treatments. Two major types of treatment occurred and additional activity is ongoing. In Summer/Fall 2016, hand thinning or mastication was performed on a subset of plots, and in Summer/Fall 2018, logging was performed on another subset of plots. Some plots received one treatment, and others received none. Information about photo dates and treatment status can be found in the FMP details csv file. This monitoring is ongoing through the Sagehen Forest Project.

openCC0Sep 2019View details →
edi40/100

SGS-LTER GIS layer of Level 2 Soil Survey and Related Document on Central Plains Experimental Range, Nunn, Colorado, USA 2012

This data package was produced by researchers working on the Shortgrass Steppe Long Term Ecological Research (SGS-LTER) Project, administered at Colorado State University. Long-term datasets and background information (proposals, reports, photographs, etc.) on the SGS-LTER project are contained in a comprehensive project collection within the Digital Collections of Colorado (http://digitool.library.colostate.edu/R/?func=collections&collection_id=3429). The data table and associated metadata document, which is generated in Ecological Metadata Language, may be available through other repositories serving the ecological research community and represent components of the larger SGS-LTER project collection. No Abstract Available

openOpenJan 2020View details →
zenodo36/100

Dataset and documented R code for "Nouns and verbs in the speech signal"

<p>The files available constitute supplementary material to the following article:</p> <p>Lohmann, Arne. Nouns and verbs in the speech signal: Are there phonetic correlates of grammatical category? <em>Linguistics</em> - <em>An Interdisciplinary Journal of the Language Sciences</em>.</p> <p>The article is to be published online in 2020, and in 2021 in the print version of the journal.</p>

opencc-by-4.0Jan 2020View details →
zenodo36/100

Salep orchid patent documents

<p>Spreadsheet of patent documents referring to salep, used in publication: &#39;Patent analysis as a novel method for exploring commercial interest in wild harvested species&#39; Biological Conservation&nbsp;</p>

opencc-by-4.0Jan 2020View details →
zenodo36/100

EUROMALT Briefing note No. 2 Ochratoxin A - Additional document from EUROMALT supporting the comments submitted during the public consultation on the Risk assessment of ochratoxin A in food

<p>This document has been submitted by EUROMALT as additional document supporting the comments submitted during the public consultation organised by the European Food Safety Authority in relation to the draft scientific opinion on the Risk assessment of ochratoxin A in food. The text of the comments is published in the Technical report of the public consultation - see related identifiers section, http://doi.org/10.2903/sp.efsa.2020.1845</p> <p>&nbsp;</p> <p>EUROMALT has agreed by email submitted to EFSA to disclose this confidential document and to have it published on Zenodo.</p> <p>.</p>

opencc-by-4.0May 2020View details →
zenodo36/100

Supplementary Material for the paper: Automatic Document Screening of Medical Literature Using Word and Text Embeddings in an Active Learning Setting

<p>This is the dataset used in the paper:&nbsp;Automatic Document Screening of Medical Literature Using Word and Text Embeddings in an Active Learning Setting.&nbsp;</p> <p>It is composed of:&nbsp;</p> <p>- Pre-trained models using active learning for document screening on HealthCLEF and Epistemonikos datasets.&nbsp;</p> <p>- Epistemonikos and HealthCLEF datasets containing medical questions and relevant/non relevant articles.&nbsp;</p> <p>- Embeddings and Document Representations used for experiments on both datasets.&nbsp;</p> <p>Scripts to run experiments can be found at:&nbsp;<a href="https://github.com/afcarvallo/active_learning_document_screening">https://github.com/afcarvallo/active_learning_document_screening</a></p> <p>&nbsp;</p> <p><strong>Paper abstract:</strong></p> <p>Document screening is a fundamental task within Evidence-based Medicine (EBM), a practice that provides scientific evidence to support medical decisions. Several approaches have tried to reduce physicians&#39; workload of screening and labeling vast amounts of documents to answer clinical questions. Previous works tried to semi-automate document screening, reporting promising results, but their evaluation was conducted on small datasets, which hinders generalization. Moreover, recent works in natural language processing have introduced neural language models, but none have compared their performance in EBM. In this paper, we evaluate the impact of several document representations such as TF-IDF along with neural language models (BioBERT, BERT, Word2vec, and GloVe) on an active learning-based setting for document screening in EBM. Our goal is to reduce the number of documents that physicians need to label to answer clinical questions. We evaluate these methods using both a small challenging dataset (HealthCLEF 2017) as well as a larger one but easier to rank (Epistemonikos). Our results indicate that word as well as textual neural embeddings always outperform the traditional TF-IDF representation. When comparing among neural and textual embeddings, in the HealthCLEF dataset the models BERT and BioBERT yielded the best results. On the larger dataset, Epistemonikos, Word2Vec and BERT were the most competitive, showing that BERT was the most consistent model across different corpuses. In term of active learning, an uncertainty sampling strategy combined with logistic regression achieved the best performance overall, above other methods under evaluation, and in fewer iterations.</p>

opencc-by-4.0Mar 2020View details →
zenodo36/100

What Is Metadata and How Do I Document My Data?

<p>This video has been realized in occasion of the CESSDA Training Days. The CESSDA Training Days were a two-day training event showcasing diverse training resources on both CESSDA tools and services. They took place on November 27 and 28, 2019, and were hosted by the GESIS &ndash; Leibniz-Institute for the Social Sciences in Cologne, Germany.</p> <p>In the video, Alexander Jedinger discusses metadata and data documentation.&nbsp;</p> <p>The video is also available for <a href="https://youtu.be/cjGz-I0GgKk">viewing on Youtube</a>.</p>

opencc-by-4.0Jun 2020View details →
zenodo36/100

An Empirical Validation of Cognitive Complexity as a Measure of Source Code Understandability - Data, Code and Documentation

<p>Release version of the data, code and documentation used in and generated by our data analysis and literature search to ensure reproducibility, repeatability, and transparency, to be published alongside our paper &quot;An Empirical Validation of Cognitive Complexity as a Measure of Source Code Understandability&quot;.</p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

Example FAIRtracks JSON document - augmented

<p><strong>Background</strong></p> <p>Many types of data from genomic analyses can be represented as genomic tracks, i.e. features linked to the genomic coordinates of a reference genome. Examples of such data are epigenetic DNA methylation data, ChIP-seq peaks, germline or somatic DNA variants, or RNA-seq expression levels. Researchers often face difficulties in locating, accessing and combining relevant tracks from external sources, as well as locating the raw data, reducing the value of the generated information.&nbsp;</p> <p><strong>FAIRtracks software ecosystem</strong></p> <p>We have, as an output of the ELIXIR Implementation Study &quot;FAIRification of Genomic Tracks&quot;, developed a basic set of recommendations for genomic track metadata together with an implementation&nbsp;called FAIRtracks in the form of&nbsp;a JSON Schema. We propose&nbsp;FAIRtracks as a draft standard for genomic track metadata in order to&nbsp;advance the application of FAIR data principles (Findable, Accessible, Interoperable, and Reusable). We have demonstrated practical usage of this approach by designing a software ecosystem around the FAIRtracks draft standard, integrating&nbsp;globally identifiable metadata from various track hubs in the Track Hub Registry and other relevant repositories into a novel track search service, called TrackFind. The software ecosystem also&nbsp;includes the FAIRtracks augmentation service, which&nbsp;assists&nbsp;metadata producers by automatically augmenting minimal machine-readable metadata with their human-readable counterparts, as well as the FAIRtracks validation service, which extends basic JSON Schema validation to include&nbsp;FAIR-related features (global identifiers, ontology terms, and object references).&nbsp;Finally, we have implemented track metadata search and import functionality into relevant analytical tools: EPICO and the GSuite HyperBrowser. For an overview of the FAIRtracks software ecosystem, please visit:&nbsp;<a href="http://fairtracks.github.io/">http://fairtracks.github.io/</a></p> <p><strong>Example FAIRtracks JSON document - augmented</strong></p> <p>The &quot;Example FAIRtracks JSON document - augmented&quot; is generated as part of the build process of the FAIRtracks draft standard JSON Schema (source code: <a href="https://github.com/fairtracks/fairtracks_standard/">https://github.com/fairtracks/fairtracks_standard/</a>). The example FAIRtracks document contains a small selection of tracks and objects from the ENCODE project metadata (<a href="https://www.encodeproject.org/">https://www.encodeproject.org/</a>), adapted to align with&nbsp;the FAIRtracks draft standard. In addition to being available in the above-mentioned GitHub repository, the &quot;Example FAIRtracks JSON document - augmented&quot; is also published here on Zenodo in order for the document to be globally uniquely identifiable by a&nbsp;Digital Object Identifier (DOI).</p>

opencc-by-4.0Aug 2020View details →
zenodo36/100

Documented socio-economic utitity of Asia-Pacific mangroves and mangrove-associate species

<p>We here present a table of listing non-timber uses&nbsp;of 203 different Asia-Pacific mangroves and mangrove-associate species. The table was assembled based on a selection of key published sources. The 409 uses compiled include food uses (staple/starch, vegetables, fruits, nuts and seeds, sweets and sugars, beverages, spices and as nectar source for beekeeping), domestic uses&nbsp;such as crafts, materials (stains, resins, fragrance, oil, roofing material and&nbsp;fiber) and as ornamental plants, animal husbandry uses (livestock feed or fodder)&nbsp;and uses for medicinal or personal care purposes.&nbsp;</p>

opencc-by-4.0Dec 2019View details →
zenodo36/100

IPBES Data Management Tutorials - Session 3.3: Data management report details: Data and metadata documentation and curation

<p>The&nbsp;<em>IPBES data management tutorials</em>&nbsp;are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The&nbsp;<em>IPBES data management reports </em>chapter&nbsp;provides an overview and discussion of specific elements of IPBES data management reports.</p> <p>This session&nbsp;<em>Data management report details: Data and metadata documentation and curation&nbsp;</em>reviews what should be included in metadata and why it should be tracked in a data management report.</p>

opencc-by-4.0Nov 2020View details →
zenodo36/100

A 4.5 year-long record of Svalbard water vapor isotopic composition documents winter air mass origin

<p>Isotopic data from the article in JGR A: A 4.5 year-long record of Svalbard water vapor isotopic composition documents winter air mass origin</p> <p><strong>CLDS_2020_ground_iso_vapor_1h.dat</strong></p> <p><strong>CLDS_2020_zeppelin_iso_vapor_1h.dat </strong></p> <p>first column: date in matlab format</p> <p>second column: humidity in ppmv</p> <p>thrid column: d18O in per mil</p> <p>forth column: dD in per mil</p> <p>fifth column: not to take into account</p> <p><strong>CLDS_2020_iso_precip.dat </strong></p> <p>first column: date in matlab format</p> <p>second column: temperature at noon in degre C</p> <p>thrid column: d18O in per mil</p> <p>forth column: dD in per mil</p> <p>fifth column:type of precip 1: water &amp; 2 : snow &amp; 3 : other (melt, etc.)</p>

opencc-by-4.0Feb 2020View details →
zenodo36/100

Données de l'enquête Couperin sur l'accompagnement à la gestion des données de recherche par les services de documentation

<p>Ce jeu de donn&eacute;es pr&eacute;sente les r&eacute;sultats obtenus lors de l&#39;enqu&ecirc;te sur les services et l&#39;accompagnement propos&eacute; par les services de documentation et d&#39;information scientifique et technique des &eacute;tablissements de l&#39;enseignement sup&eacute;rieur et de la recherche fran&ccedil;ais. Cette enqu&ecirc;te a &eacute;t&eacute; men&eacute;e par le groupe Donn&eacute;es du GTSO-Couperin en septembre et octobre 2020 aupr&egrave;s des &eacute;tablissements membres de Couperin.</p>

opencc-by-4.0Dec 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record