Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

42

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

42 results for “directory”

Learn how ShareScore rates datasets ↗
zenodo44/100

Crow's Directories Data (1935)

<p>This set of two files contains the data extracted from Carl Crow&#39;s <em>Newspaper Directory of China</em> (1935&nbsp;edition). The first file contains the raw extraction from Freizo (Crow_1935_raw). The second is a clean version of the same dataset&nbsp;after correcting, cleaning,&nbsp;and standardizing the data initially extracted (Crow_zenodo_1935). Each line corresponds to a unique title. For each title, we provide the following information: name (full name, English, Wade-Giles, Chinese, pinyin transliteration), place of publication (city, province), year of establishment, circulation, publisher&#39;s name and profile, number&nbsp;and size of pages, number and size of columns.&nbsp;</p> <p>In the cleaned version, the first tab contains the data (list of the periodicals). Additional tabs describe the variables and the classification used for analytical&nbsp;purposes (period, format, etc).&nbsp;</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Prevalence of Creative Commons licenses in the Directory of Open Access Journals by discipline, author fees, number of journals per country and publisher

<p>An analysis on the prevalence of Creative Commons licenses in the Directory of Open Access Journals by discipline, author fees, country and publisher according to the number of journals.</p>

opencc-by-sa-4.0Aug 2014View details →
zenodo44/100

Numbers and shares of Open Access Journals from all disciplines and from the discipline Sociology using Creative Commons Licenses as listed by the Directory of Open Access Journals (2014-06-08)

<p>These files contains the data on frequencies and shares of Open Access journals using Creative Commons Licenses (comparing journals from all disciplines and sociologial journals). Date of data collection: 2014-06-08.</p>

opencc-by-sa-4.0Jun 2014View details →
zenodo44/100

EPDnew: FTP directory, Nov 2023

<p>Snapshot of the EPDnew FTP directory of November 2023. The acronym EPDnew refers to the newer versions of the Eukaryotic Promoter Database published in the format introduced in 2013. The directory contains all database versions for 15 model organisms released since then.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Github Directory Listings Dataset 2021

<p>Directory listings for the HEAD revisions of all publically-accessible Github repositories in the <a href="https://ghtorrent.org/">GHTorrent</a> database dump from 2021-03-06. See the Readme.md file for additional details.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

A Dataset of French Trade Directories from the 19th Century (FTD)

<p>This dataset is composed of pages and entries extracted from French directories published between 1798 and 1861.</p> <p>The purpose of this dataset is to evaluate the performance of Optical Character Recognition (OCR) and Named Entity Recognition (NER) on 19th century French documents.</p> <p><br> This dataset is divided into two parts:</p> <ol> <li>A <strong>labeled dataset</strong>, which contains 8765 manually corrected entries from 78 pages (18 different directories), and which is designed for supervised training.</li> <li>An <strong>unlabeled dataset</strong>, containing 1058196 raw entries from 6887 pages (13 different directories), and which is designed for self-supervised pre-training.</li> </ol> <p>For the <strong>labeled dataset</strong>, we provide:</p> <ul> <li>Original pages and cropped images</li> <li>Human-corrected positions, transcriptions and entity tagging for each entry</li> <li>OCR prediction from 3 systems (Tesseract v4, PERO OCR v2020 and Kraken)</li> <li>Projected NER reference from clean text to OCR predictions, making it suitable to evaluate the performance of NER systems on real, noisy OCR predictions</li> </ul> <p>For the <strong>unlabeled dataset</strong>, we provide:</p> <ul> <li>Automatically detected positions for each entry (lot of noise)</li> <li>OCR predictions for each entry (PERO OCR engine)</li> </ul> <p>&nbsp;</p> <p><strong>How to cite this dataset</strong><br> Please cite this dataset as:</p> <blockquote> <p>N. Abadie, S. Baciocchi, E. Carlinet, J. Chazalon, P. Cristofoli, B. Dum&eacute;nieu and J. Perret, A Dataset of French Trade Directories from the 19th Century (FTD), version 1.0.0, May 2022, online at https://doi.org/10.5281/zenodo.6394464.</p> </blockquote> <pre><code>@dataset{abadie_dataset_22, author = {Abadie, Nathalie and Bacciochi, St{\'e}phane and Carlinet, Edwin and Chazalon, Joseph and Cristofoli, Pascal and Dum{\'e}nieu, Bertrand and Perret, Julien}, title = {{A} {D}ataset of {F}rench {T}rade {D}irectories from the 19th {C}entury ({FTD})}, month = mar, year = 2022, publisher = {Zenodo}, version = {v1.0.0}, doi = {10.5281/zenodo.6394464}, url = {https://doi.org/10.5281/zenodo.6394464} }</code></pre> <p><br> You may also be interested in <strong>our paper presented at DAS 2022</strong> (15th IAPR International Workshop on Document Analysis Systems), which <strong>compares the performance of OCR and NER</strong> systems <strong>on this dataset</strong>:</p> <blockquote> <p>N. Abadie, E. Carlinet, J. Chazalon and B. Dum&eacute;nieu, A Benchmark of Named Entity Recognition Approaches in Historical Documents &mdash; Application to 19th Century French Directories, May 2022, La Rochelle, France, Springer.</p> </blockquote> <pre><code>@inproceedings{abadie_das_22, author = {Abadie, Nathalie and Carlinet, Edwin and Chazalon, Joseph and Dum{\'e}nieu, Bertrand}, title = {{A} {B}enchmark of {N}amed {E}ntity {R}ecognition {A}pproaches in {H}istorical {D}ocuments — {A}pplication to 19th {C}entury {F}rench {D}irectories}, month = may, year = 2022, publisher = {Springer}, place = {La Rochelle, France} }</code></pre> <p><br> <strong>Copyright and License</strong><br> The images were extracted from the original source <a href="https://gallica.bnf.fr">https://gallica.bnf.fr</a>, owned by the <em>Biblioth&egrave;que nationale de France</em> (French national library).<br> Original contents from the <em>Biblioth&egrave;que nationale de France</em> can be reused non-commercially, provided the mention &quot;Source gallica.bnf.fr / Biblioth&egrave;que nationale de France&quot; is kept. &nbsp;<br> <strong>Researchers do not have to pay any fee for reusing the original contents in research publications or academic works. </strong>&nbsp;<br> <em>Original copyright mentions extracted from <a href="https://gallica.bnf.fr/edit/und/conditions-dutilisation-des-contenus-de-gallica">https://gallica.bnf.fr/edit/und/conditions-dutilisation-des-contenus-de-gallica</a> on March 29, 2022.</em></p> <p>The original contents were significantly transformed before being included in this dataset.<br> All derived content is licensed under the permissive <strong>Creative Commons Attribution 4.0 International</strong> license.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Directory of the European Parliament members

<p>Over the past twenty-five years, a field of research into the careers of Members of European Parliament (MEPs) has developed. Drawing on a massive amount of accessible open data, we have assembled an updated database comprising all MEPs between 1979 and 2025.</p> <p>This dataset contains (some) socio-demographic informations about MEP&rsquo;s and their carreer paths in EP. Data were scraped from the EP website.</p> <p>Metadata are filled separately in a rich text format (.rtf) document.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

A data directory to facilitate investigations on worldwide wildlife trafficking

<p>We describe a novel, open-access data directory &nbsp;on wildlife trafficking and a corresponding visualization tool that can be used to identify data for multiple purposes, such as exploring wildlife trafficking hotspots and convergence points with other crime, discovering key drivers or deterrents of wildlife trafficking, and uncovering structural patterns. Keyword searches, expert elicitation, and peer-reviewed publications were used to search for extant sources used by industry and non-profit organizations, as well as those leveraged to publish academic research articles.&nbsp; The open-access data directory is designed to be a living document and searchable according to multiple measures. The directory can be instrumental in the data-driven analysis of unsustainable illegal wildlife trade, supply chain structure via link prediction models, the value of demand and supply reduction initiatives via multi-item knapsack problems, or trafficking behavior and transportation choices via network interdiction problems.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

A Dataset of French Trade Directories from the 19th Century for Nested NER task

<p>This dataset is composed of pages and entries extracted from French directories published between 1798 and 1861.</p> <p>The purpose of this dataset is to evaluate the performance of Nested Named Entity Recognition approaches on 19th century French documents, regarding both clean and noisy texts (due to the OCR engine).</p> <p><strong>Source dataset</strong></p> <p>This dataset has been built from this source dataset :</p> <pre><code class="language-markdown">N. Abadie, S. Baciocchi, E. Carlinet, J. Chazalon, P. Cristofoli, B. Duménieu and J. Perret, A Dataset of French Trade Directories from the 19th Century (FTD), version 1.0.0, May 2022, online at https://doi.org/10.5281/zenodo.6394464.</code></pre> <p><strong>Our experiments // Paper</strong></p> <p>Details about our experiments on&nbsp;nested NER approaches are given in our paper (<a href="https://hal.science/hal-03994759v2">the pre-print version is available here</a>).</p> <pre><code class="language-markdown">Tual, S., Abadie, N., Chazalon, J., Duménieu, B., &amp; Carlinet, E. (2023). A Benchmark of Nested NER Approaches in Historical Structured Documents. Proceedings of the 17th International Conference on Document Analysis and Recognition, San José, California, USA. 2023. Springer. https://hal.science/hal-03994759v2</code></pre> <p>Our code is available on <a href="https://github.com/soduco/paper-nestedner-icdar23-code">Git-Hub</a>.</p> <p><strong>Dataset overview</strong></p> <p>The following list describes the <strong>keys of the .JSON</strong>&nbsp;file which contain the complete materials of our experiments.</p> <p>- id&nbsp;: Entry unique ID in a given page</p> <p>- box : Bounding box of the entry in the scanned directory page</p> <p>- book&nbsp;: Source directory of the entry (*see more information bellow*)</p> <p>- page&nbsp;: Page ID in a given directory</p> <p>- valid_box&nbsp;: Is the bbox of the entry valid ? (*all bbox are valid here*)</p> <p>- text_ocr_ref`&nbsp;: OCR extracted and manually corrected text of the entry</p> <p>- nested_ner_xml_ref&nbsp;:&nbsp;<em> text_ocr_ref</em>&nbsp;with nested ner entities</p> <p>- text_ocr_pero&nbsp;: OCR extracted text of the entry with PERO-OCR engine (best engine according to Abadie et al.&nbsp;experiment)</p> <p>- has_valid_ner_xml_pero&nbsp;: Is entities mapping between nested-ner entities annotated by hand on the ref&nbsp;text and pero ocr text correct ? (in our experiments, we only use entries with True value)</p> <p>- nested_ner_xml_pero&nbsp;: Annotated noisy entries produced with PERO OCR</p> <p>- text_ocr_tess&nbsp;: OCR extracted text of the entry with Tesseract engine (*not used in our expriments*)</p> <p>- nested_ner_xml_tess&nbsp;: Is entities mapping between nested-ner entities annotated by hand on the ref&nbsp;text and tesseract text correct? (not used in our experiments)</p> <p>- has_valid_ner_xml_tess&nbsp;: Annotated noisy entries produced with Tesseract. (not used in our experiments)</p> <p>Nested entities are annotated using XML tags.&nbsp;Our hierachy of entities is a *Part Of* a two-levels hierarchy. It means that bottom entities are contained in a top level entity.</p> <p>&nbsp;</p> <p><strong>Source documents // Copyright and licence</strong></p> <p><em>This section has been copied from the <a href="https://zenodo.org/record/6394464">original dataset description</a>.</em></p> <p>The images were extracted from the original source https://gallica.bnf.fr, owned by the *Biblioth&egrave;que nationale de France* (French national library).</p> <p>Original contents from the <em>Biblioth&egrave;que nationale de France</em>&nbsp;can be reused non-commercially, provided the mention &quot;Source gallica.bnf.fr / Biblioth&egrave;que nationale de France&quot; is kept. &nbsp;</p> <p>=&gt;&nbsp;<strong>Researchers do not have to pay any fee for reusing the original contents in research publications or academic works.</strong></p> <p>Original copyright mentions extracted from <a href="https://gallica.bnf.fr/edit/und/conditions-dutilisation-des-contenus-de-gallica">https://gallica.bnf.fr/edit/und/conditions-dutilisation-des-contenus-de-gallica</a> on March 29, 2022.</p> <p>The original contents were significantly transformed before being included in this dataset.</p> <p>All derived content is licensed under the permissive *Creative Commons Attribution 4.0 International* license.</p> <p>Links to original contents are given in the window bellow :</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Total numbers and shares of Open Access Journals using Creative Commons Licenses as listed by the Directory of Open Access Journals

<p>This table provides information on the number and percentage of Open Access Journals listed by the Directory of Open Access Journals using a Creative Commons License. It identifies also the number and share of Journals using CC-Licenses that are compatible to the Open Definition.</p>

opencc-by-4.0Feb 2014View details →
zenodo40/100

Animal Communicator International Scan of Books, Websites, Animal Communicator Directory

<p>Four datasets are included focusing on animal communicators&nbsp;(AC): practitioners of intuitive interspecies communication (IIC).</p> <p><strong>Dataset #1:</strong> International English language websites data set (n = 400). To be included and coded, websites had to meet the following 3 criteria: (1) have an English language version of the website, (2) currently offer private AC consultations, and (3) be identified before website analysis was determined comprehensive enough to represent the international scope of AC (Oct. 30, 2020). CITE AS:<strong>&nbsp;</strong>Barrett, M. J., Zmud, L., Mathur, A., &amp; &nbsp;Hoessler, C. (2024). Animal communicator website scan [Data set]. Zenodo. DOI 10.5281/zenodo.15131964</p> <p>&nbsp;<strong>Dataset #2:</strong> International English language published books. Total books (n = 191): Books were identified through internet searches, including searches on Amazon and used bookstores such as Abe Books, from practicing AC websites, and from the directory: Book Authority Website for &ldquo;53 Best Animal Communication Books of All Time,&rdquo; https://bookauthority.org/books/best-animal-communication-books. Dataset includes books&rsquo; titles, descriptions, front and back covers and tables of contents where available. Where it was not clear whether the author was an animal communicator who consulted with clients, we did additional online searches to make this determination. All books had an English language printed copy of the book. We excluded books that were available in electronic copy only. We expanded our initial content inclusion criteria used in the website report beyond individuals currently offering consultations as professional animal communicators to include: (1) professional animal communicators who have since retired; (2) individuals who work intensively with animals in other capacities such as healing, but also report instances of IIC; (3) individuals who may not have worked as professional animal communicators, but write about their own lived experience of the phenomena; and (4) books written by individuals who are not ACs but have interviewed ACs. Where a book was written by two authors, or in some cases, an animal communicator with an additional author, we included both authors in our formal citation. Books were written by ACs who offered professional services (182), or authors who were not ACs (9). CITE AS:<strong> </strong>Barrett, M. J., Mathur, A., &amp;&nbsp; Ghoreishi, Z., Hoessler, C., Kuppenbender, S. (2024). Animal communicator Book Scan [Data set]. Zenodo. DOI 10.5281/zenodo.15131964</p> <p><strong>Dataset #3:</strong> Directory of practicing ACs (1990-2011; n = 80&nbsp;issues). To provide a snapshot of growth in numbers of practitioners over time, we compiled listings of ACs published in the Animal Communicator Directory from <em>Species Link: The Journal of Interspecies Telepathic Communication</em>. From 1990-2011, the publication included a directory of practicing ACs; after 2011, the directory went fully online, and data from each year is not available. CITE AS:<strong> </strong>Barrett, M.J. &amp; Hoessler, C. (2022).&nbsp; Animal Communicator Directory. [Data set]. Zenodo. DOI 10.5281/zenodo.15131964</p> <p><strong>Dataset #4:&nbsp;</strong>Reference List of 15 Analyzed Animal Communicator Books with formal or informal &ldquo;How-To&rdquo; Sections. Compiled to identify and summarize the ways in which ACs were describing essential processes for conducting a successful intuitive communication session with animals, and thus provide an overview of what happens in an IIC session. Selection criteria: We were seeking succinct summaries from ACs. The intent was not to dig into and analyze processes in detail, but rather to summarize and synthesize the essential processes and steps as the ACs were reporting them.&nbsp; As such, it was beyond the scope of this study to analyze reported example communications or analyze processes in books where the entirety of the book was describing communications with animals. Publication date range: 1998-2019. Close to 300 pages were analyzed. The actual "how-to" excerpts are not included as they are subject to copyright. CITE AS:<strong> </strong>Barrett, M.J. &amp; Kuppenbender, S. (2022).&nbsp; Reference List of 15 Analyzed Animal Communicator Books [Data set]. Zenodo. DOI 10.5281/zenodo.15131964</p> <p>For further information on data collection and analysis details contact M.J. Barrett, PhD.&nbsp;&nbsp;mj.barrett@usask.ca&nbsp;</p> <p>Funded by the Social Sciences and Humanities Research Council of Canada<em>&nbsp;</em>Insight Development Grant: <em>Deepening Connection in Pursuit of Environmental Sustainability: Assessing a Promising Lever for Shifting Assumptions of Separation </em>(Grant # 430-2019-01023).</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Books per Publisher: Data from the Directory of Open Access Books (DOAB)

<p>Quantitative information on publishers and published books from the Directory of Open Access Books (DOAB).</p>

opencc-by-4.0Jan 2018View details →
zenodo36/100

data/ directory associated with https://github.com/ericagol/TRAPPIST1_Spitzer

<p>Data directory for the repository associated with the paper &quot;Refining the transit timing and photometric analysis of TRAPPIST-1: Masses, radii, densities,&nbsp;dynamics, and ephemerides.&quot;</p>

opencc-by-4.0Sep 2020View details →
zenodo36/100

Dados coletados das revistas de revisão aberta do Directory of Open Access Journals

<p>Dados coletados para pesquisa que visa analisar os peri&oacute;dicos cient&iacute;ficos adeptos da revis&atilde;o por pares aberta e que est&atilde;o indexados no Directory of Open Access Journals. Os dados apontam as caracter&iacute;sticas descritivas e caracter&iacute;sticas de revis&atilde;o dos peri&oacute;dicos.</p>

opencc-by-4.0Nov 2020View details →
zenodo36/100

Crow's Directories Data (1931)

<p>This set of two files contains the data extracted from Carl Crow&#39;s Newspaper Directory of China (1931 edition). The first file contains the raw extraction from Freizo (Crow_1931_raw). The second is a clean version of the same dataset&nbsp;after correcting, cleaning,&nbsp;and standardizing the data initially extracted (Crow_zenodo_1931). Each line corresponds to a unique title. For each title, we provide the following information: name (full name, English, Wade-Giles, Chinese, pinyin transliteration), place of publication (city, province), year of establishment, circulation, publisher&#39;s name and profile, number&nbsp;and size of pages, number and size of columns.&nbsp;</p> <p>In the cleaned version, the first tab contains the data (list of the periodicals). Additional tabs describe the variables and the classification used for analytical&nbsp;purposes (period, format, etc).&nbsp;</p>

opencc-by-4.0Nov 2020View details →
zenodo36/100

Example 6-hour directory for FV3GFS atmospheric model

<p>This dataset serves as an example run directory for the FV3GFS atmospheric model, used in publications to show execution of the Python-wrapped FV3GFS atmospheric model.</p> <p>Files in rundir/grb provided by the National Oceanic and Atmospheric Administration (NOAA) Environmental Modeling Center (EMC) are public domain.</p> <p>All other data and configuration files are distributed under a Creative Commons Attribution-ShareAlike 4.0 International License.</p> <p>SHiELD/FV3 model data and input files produced by the Geophysical Fluid Dynamics Laboratory is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License (https://creativecommons.org/licenses/).&nbsp; Consult https://pcmdi.llnl.gov/CMIP6/TermsOfUse for terms of use governing CMIP6 output, including citation requirements and proper acknowledgment. The data producers and data providers make no warranty, either express or implied, including, but not limited to, warranties of merchantability and fitness for a particular purpose. All liabilities arising from the supply of the information (including any liability arising in negligence) are excluded to the fullest extent permitted by law.</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Harmonized directory of beneficiaries of the PDAC project - Congo-Brazzaville

<p><strong>Description:</strong> This datasets contains a list of all groups of agricultural, livestock, fisheries and agrifoodstuff initiatives that benefitted from the "Projet d'appui au D&eacute;v&eacute;loppement de l'Agriculture Commerciale" (PDAC) project from the Minist&egrave;re de l'Agriculture de l'Elevage et de la P&ecirc;che (MAEP) of the Republic of Congo, along with their locations at departmental and district levels, and initiative names. For more details see the project details at&nbsp;<a href="https://pdacmaep.cg">https://pdacmaep.cg</a></p> <p><strong>Scope of data :</strong> This dataset lists 850+ beneficiary initiatives country-wide, covering 33 activities along the 12 departments and 99 districts. List of activities is as follows :</p> <ul> <li>aliment pour b&eacute;tail</li> <li>ananas</li> <li>apiculture</li> <li>arachide</li> <li>arboriculture</li> <li>banane</li> <li>cacao</li> <li>chambre froide</li> <li>commercialisation</li> <li>&eacute;levage bovin</li> <li>&eacute;levage de cailles</li> <li>&eacute;levage ovins</li> <li>&eacute;levage pondeuse</li> <li>&eacute;levage porcin</li> <li>fumure</li> <li>gingembre</li> <li>grenadille</li> <li>haricot</li> <li>huile de palme</li> <li>igname</li> <li>ma&iuml;s</li> <li>manioc</li> <li>mara&icirc;chage</li> <li>p&ecirc;che</li> <li>pisciculture</li> <li>pois d'angol</li> <li>pomme de terre</li> <li>prestation</li> <li>production</li> <li>service v&eacute;t&eacute;rinaire</li> <li>soja</li> <li>taro</li> <li>transformation</li> </ul> <p><strong>Data Collection and Processing</strong>: This dataset was assembled from the different departmental directories of beneficiaries openly available online on April 18th, 2024 from the project website at&nbsp;<a href="https://pdacmaep.cg/index.php?page=documents">https://pdacmaep.cg/index.php?page=documents</a> and contains harmonized activity type, initiative name, department and district. Contact names and phone coordinates were removed for privacy.</p> <p><strong>Project Context</strong>: This work is part of the&nbsp;<a href="https://www.cirad.fr/dans-le-monde/cirad-dans-le-monde/projets/projet-pudt-congo">PUDT Congo</a> project.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Microbe Directory Data v1.0.0

<p>This upload contains the data from v1.0.0 of The Microbe Directory.</p> <p>&nbsp;</p> <p>The Microbe Directory is a collective research effort to profile and annotate more than 7,500 unique microbial species from the MetaPhlAn2 database that includes bacteria, archaea, viruses, fungi, and protozoa. By collecting and summarizing data on various microbes&rsquo; characteristics, the project comprises a database that can&nbsp;be used downstream of large-scale metagenomic taxonomic analyses, allowing one to interpret and explore their taxonomic classifications to have a deeper understanding of the microbial ecosystem they are studying. Such characteristics include, but are not limited to: optimal pH, optimal temperature, Gram stain, biofilm-formation, spore-formation, antimicrobial resistance, and COGEM class risk rating. The database has been manually curated by trained student-researchers from Weill Cornell Medicine and CUNY&mdash;Hunter College, and its analysis remains an ongoing effort with open-source capabilities so others can contribute. Available in SQL, JSON, and CSV (i.e. Excel) formats, the Microbe Directory can be queried for the aforementioned parameters by a microorganism&rsquo;s taxonomy. In addition to the raw database, The Microbe Directory has an online counterpart (https://microbe.directory/) that provides a user-friendly interface for storage, retrieval, and analysis into which other microbial database projects could be incorporated. The Microbe Directory was primarily designed to serve as a resource for researchers conducting metagenomic analyses, but its online web interface should also prove useful to any individual who wishes to learn more about any particular microbe.</p>

openother-openDec 2017View details →
zenodo36/100

SHEMAT-Suite output directories for WRR manuscript 2018WR023374

<p>The directories in the compressed archives contain the raw data used in the Water Resources Research publication:</p> <blockquote> <p>Comparing seven variants of the Ensemble Kalman Filter: How many<br> synthetic experiments are needed?</p> </blockquote> <p>The article can be found under</p> <p>&nbsp; &nbsp; &nbsp; <a href="https://doi.org/10.1029/2018WR023374">https://doi.org/10.1029/2018WR023374</a></p> <p><br> <strong>Extracting the data</strong><br> &nbsp;</p> <p>At first, extract the directory of the figure:</p> <pre><code class="language-bash"> tar -xzvf figure_1.tar.gz</code></pre> <p>Move the compressed archives to a directory of the form:</p> <p>1) Tracer:</p> <pre><code class="language-bash"> $HOME/shematOutputDir/wavereal_output/</code></pre> <p>2) Well:</p> <pre><code class="language-bash"> $HOME/shematOutputDir/wavewell_output/</code></pre> <p><br> Then run a command of the following form</p> <pre><code class="language-bash"> tar -xzvf 2010_01_30.tar.gz</code></pre> <p>or, to suppress output and return to the command line immediately</p> <pre><code class="language-bash"> tar -xzf 2010_01_30.tar.gz &amp;</code></pre> <p>Now the output files of 2010_01_30.tar.gz are readable for the Python<br> Scripts.</p> <p><br> <strong>Python Scripts</strong><br> &nbsp;</p> <p>The Python Scripts for generation of the figures in the manuscripts can be found in the Github repository:</p> <pre><code> pyshemkf</code></pre> <p>under</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; <a href="https://doi.org/10.5281/zenodo.1344336">https://doi.org/10.5281/zenodo.1344336</a></p> <p>or for the up-to-date version of pyshemkf:</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; <a href="https://github.com/jjokella/pyshemkf">https://github.com/jjokella/pyshemkf</a></p> <p>Documentation of and links to the scripts can be found under:</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; <a href="https://github.com/jjokella/pyshemkf#manuscript-scripts">https://github.com/jjokella/pyshemkf#manuscript-scripts</a></p> <p>&nbsp;</p>

opencc-by-4.0Aug 2018View details →
zenodo36/100

Webis Open Directory Project Corpus 2010 (Webis-ODP-10)

<p>The Webis Open Directory Project Corpus 2010 (Webis-ODP-10) is a corpus for the evaluation of cluster labeling algorithms. The corpus contains about 5,000 web pages which are grouped into 4 main categories and 12 subcategories based on the human made <a href="http://www.dmoz.org/">Open Directory Project (ODP)</a> classification. The web pages were collected in May 2010.</p> <p>In order to create the corpus, 4 different categories of the Open Directory Project were chosen: movies, databases, health and recreation. Each of these categories is again consists of up to six subcategories. In the case of movies, these subcategories correspond to six different titles of two directors. The categories were chosen to allow for different settings of experiments. Example for tasks are to distinguish the two directors or to separate web pages about recreation from those about health.</p> <p>For each category, the linked web pages were downloaded if they still existed. Furthermore, web pages which are linked from these web pages were downloaded as well and considered to be in the same category. However, only web pages of a length of 1,000 or more characters (excluding HTML tags and scripts) were kept.</p>

opencc-by-4.0Jun 2010View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record