Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,200

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3,200 results for “Index”

Learn how ShareScore rates datasets ↗
zenodo44/100

Law Indexes: Virginia

<p>As part of efforts to expand the Local Geohistory Project, which aims to educate users and disseminate information concerning the geographic history and structure of political subdivisions and local government, this repository has been created to disseminate law index data for the Commonwealth of Virginia.</p> <p>This index currently covers enrolled bills from 1776 through 1910. Prior to the 20th century, private and special laws were often the primary method used to alter municipal and county boundaries and forms of government in Virginia. The law index is released as a tab-separated values (TSV) file, <strong>output/VaLawIndex.tsv</strong>.</p> <p>Note that the Citation Page references are not to the published session laws, but to the Enrolled bills of the General Assembly collection at the <a href="https://lva.primo.exlibrisgroup.com/permalink/01LVA_INST/altrmk/alma990004934790205756">Library of Virginia</a>.</p> <p>This repository does not contain the full text of the laws, nor does it currently contain links to the full text. Because the index was created using OCR technology, it may contain uncaptured errors.</p>

opencc-zeroJan 2024View details →
zenodo44/100

Law Indexes: Connecticut

<p>As part of efforts to expand the Local Geohistory Project, which aims to educate users and disseminate information concerning the geographic history and structure of political subdivisions and local government, this repository has been created to disseminate law index data for the State of Connecticut. This index currently covers private and special laws only from 1789 through 1943.</p> <p>Private and special laws concern one or several individuals, entities, or localities, unlike public or general laws, which apply to all similarly situated individuals, entities, and localities within a jurisdiction. Throughout New England, private and special laws were often the primary method used to alter municipal and county boundaries and forms of government.</p> <p>The law index is released as a tab-separated values (TSV) file, <strong>output/ConnLawIndex.tsv</strong>. The Detail column may contain additional Prefix information that is repeated from prior entries.</p> <p>This repository does not contain the full text of the laws, nor does it currently contain links to the full text. Because the index was created using OCR technology, it may contain uncaptured errors.</p>

opencc-zeroJan 2024View details →
zenodo44/100

Hourly values of an advanced human-biometeorological index for diverse populations from 1991 to 2020

<p>The presented human thermal bioclimate dataset was created in the frame of the <a href="https://theheatalarm.wordpress.com/">HEAT-ALARM</a>&nbsp;("Development of a heat-health warning system in Greece") research project.</p> <p><strong>Initially developed for Greece</strong>,&nbsp; it consists of hourly values of population-weighted mPET (modified physiologically equivalent temperature), simulated by the RayMan Pro model for the period 1991-2020 and for 10 population subsets in 72 regional units and combinations thereof, which are based on the NUTS-3 (Nomenclature of Territorial Units for Statistics-3) classification in Greece, using the Copernicus European Regional Reanalysis (CERRA) at 5.5 km spatial resolution. The dataset also includes the main environmental drivers of mPET (e.g. temperature) at the same spatiotemporal resolution.</p> <p>In the framework of <strong>replicating</strong> the original dataset, the current version includes&nbsp;population-weighted values of mPET and its environmental drivers for six populations in five districts of <strong>Cyprus</strong> at the LAU-1 (Local Administrative Units-1) level, covering the period from 1991 to 2020.&nbsp;</p> <p>The code used to produce the presented data is available at: <a href="https://doi.org/10.5281/zenodo.10793067">https://doi.org/10.5281/zenodo.10793067</a>. It can be used to replicate the dataset not only directly in Greece, but also in any other country included in the CERRA domain after appropriate adjustments, as in the case of Cyprus above.</p> <p><em>Compared to the previous version of the dataset for Greece, this version includes vapor pressure (VP) instead of relative humidity (see README.txt for more details), as VP is more relevant for human-biometerological and health-related studies.</em><em>&nbsp;</em></p> <p><strong>References</strong></p> <p>Giannaros, C., Agathangelidis, I., Galanaki, E.&nbsp;<em>et al.</em>&nbsp;Hourly values of an advanced human-biometeorological index for diverse populations from 1991 to 2020 in Greece.&nbsp;<em>Sci Data</em>&nbsp;<strong>11</strong>, 76 (2024). <a href="https://doi.org/10.1038/s41597-024-02923-y">https://doi.org/10.1038/s41597-024-02923-y</a>&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

2-meter Universal Thermal Climate Index (UTCI) and Human Heat Health Index (H3I) hazard for Austin, Texas

<p>Universal Thermal Climate Index (UTCI) is a physiological temperature that is widely used in biometeorological studies to assess the heat stress felt by humans. UTCI considers the shortwave and longwave radiation incident on humans from the six cubical directions as well as air temperature, humidity, wind speed and clothing. As a part of NOAA National Integrated Heat Health Information System (NIHHIS) and NASA Interdisciplinary Research in Earth Science (IDS) project, we have generated the UTCI data for Austin, Texas and surrounding peri-urban area at 2-meters spatial resolution for the year 2017. Details on data generation and methodology can be found in Kamath et al., (2023) but are summarized here.&nbsp;</p> <p><strong>1. Datasets and model used</strong></p> <p>The solar and longwave environmental irradiance geometry (SOLWEIG) model was used to simulate shadows, mean radiant temperature (T<sub>MRT</sub>) and the UTCI (Lindberg et al., 2008). T<sub>MRT</sub> is the equivalent temperature due to exposure to absorbed shortwave and longwave radiation from all directions in a standing position. SOLWEIG was forced using near-surface ERA-5 data available at a spatial resolution of 0.25&deg;x 0.25&deg;. Building, vegetation heights, and digital terrain model were again derived from 3DEP LiDAR point cloud data.&nbsp; SOLWEIG was run using the urban multi-scale environment predictor (UMEP) (Lindberg et al., 2018) plug-in with QGIS.&nbsp;&nbsp;</p> <p><strong>2. Data availability</strong></p> <p>Diurnal UTCI data were calculated for typical meteorological clear sky days corresponding to Summer and Fall. The typical clear sky day was selected using the 10-year Typical meteorological Year (TMY) for Austin, Texas (30.2672&deg; N, 97.7431&deg; W) provided by National Solar Radiation Database (NSRDB). More details on TMY files can be found at: https://nsrdb.nrel.gov/data-sets/tmy</p> <p>Additionally, data is developed for heat hazard for daytime Human Heat Health Index (H3I) calculation as defined by Kamath et al., (2023). Briefly, this heat hazard is defined as the fraction of the day when the UTCI exceeds certain threshold. The threshold used to calculate heat hazard for Summer and Fall were 35&deg; C and 32&deg;C, respectively that imply strong heat stress (Jendritzky et al., 2012). Note that UTCI is on a different scale compared to air temperature, and could yield different heat stress levels.</p> <p><strong>3. Data format</strong></p> <p>The georeferenced UTCI and heat hazard data are available in the geoTIFF file format. The files can be readily visualized using GIS software such as QGIS and ArcGIS, as well as programing languages such as Python.</p> <p>&nbsp;<strong>4. Companion dataset</strong></p> <p>Based on the calculated UTCI here, the potential locations for tree planting were calculated to increase the shade to reduce heat vulnerability for Austin, Texas. [https://doi.org/10.5281/zenodo.6363494]</p> <p><strong>References</strong></p> <ol> <li>Kamath, H. G., Martilli, A., Singh, M., Brooks, T., Lanza, K., Bixler, R. P., ... &amp; Niyogi, D. (2023). Human heat health index (H3I) for holistic assessment of heat hazard and mitigation strategies beyond urban heat islands. Urban Climate, 52, 101675.</li> <li>Lindberg, F., Holmer, B., &amp; Thorsson, S. (2008). SOLWEIG 1.0&ndash;Modelling spatial variations of 3D radiant fluxes and mean radiant temperature in complex urban settings.&nbsp;<em>International journal of biometeorology</em>,&nbsp;<em>52</em>, 697-713.</li> <li>Lindberg, F., Grimmond, C. S. B., Gabey, A., Huang, B., Kent, C. W., Sun, T., ... &amp; Zhang, Z. (2018). Urban Multi-scale Environmental Predictor (UMEP): An integrated tool for city-based climate services.&nbsp;<em>Environmental modelling &amp; software</em>,&nbsp;<em>99</em>, 70-87.</li> <li>Jendritzky, G., de Dear, R., &amp; Havenith, G. (2012). UTCI&mdash;why another thermal index?.&nbsp;<em>International journal of biometeorology</em>,&nbsp;<em>56</em>, 421-428.</li> <li>Bixler, R. P., Coudert, M., Richter, S. M., Jones, J. M., Llanes Pulido, C., Akhavan, N., ... &amp; Niyogi, D. (2022). Reflexive co-production for urban resilience: Guiding framework and experiences from Austin, Texas. Frontiers in Sustainable Cities, 4, 1015630.</li> <li>Lanza, K., Jones, J., Acu&ntilde;a, F., Coudert, M., Bixler, R. P., Kamath, H., &amp; Niyogi, D. (2023). Heat vulnerability of Latino and Black residents in a low-income community and their recommended adaptation strategies: A qualitative study.&nbsp;<em>Urban Climate</em>,&nbsp;<em>51</em>, 101656.</li> </ol>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Evaluation datasets and results of the paper "Efficient Online Computation of Business Process State From Trace Prefixes via N-Gram Indexing"

<p>Event logs, process models, and results corresponding to the paper "Efficient Online Computation of Business Process State From Trace Prefixes via N-Gram Indexing".</p> <p><em><strong>Inputs</strong></em>: preprocessed event logs and discovered process models (and their characteristics) used in the evaluation.</p> <ul> <li><em><strong>Real-life</strong></em>: preprocessed event logs (<em>xes</em> and <em>csv</em>) corresponding to the real-life processes used in the evaluation. Process models (<em>pnml</em>) discovered with the Inductive Miner infrequent for thresholds of 10%, 20%, and 50%. Characteristics (<em>txt</em>) of the event logs and process models. Ongoing cases result from splitting each case in the preprocessed event logs (under folder <em>split</em>).</li> <li><em><strong>Synthetic</strong></em>: simulated&nbsp;event logs (<em>csv</em>) corresponding to the synthetic processes used in the evaluation. Designed process models (<em>bpmn</em> and&nbsp;<em>pnml</em>). Ongoing cases result from splitting each case in the preprocessed event logs (under folder <em>split</em>). Ongoing cases with injected noise as described in the publication (under folders <em>noise_1</em>, <em>noise_2</em>, and <em>noise_3</em>).</li> </ul>

opencc-by-4.0Aug 2024View details →
zenodo44/100

MESINESP2 Corpora: Annotated data for medical semantic indexing in Spanish

<p>Gold Standard annotations of the MESINESP2 corpora (training, development and test sets).&nbsp;</p> <p><strong>Please cite this paper if you use this dataset:</strong></p> <pre><code class="language-bash">@inproceedings{gasco2021overview, title={Overview of BioASQ 2021-MESINESP track. Evaluation of advance hierarchical classification techniques for scientific literature, patents and clinical trials}, author={Gasco, Luis and Nentidis, Anastasios and Krithara, Anastasia and Estrada-Zavala, Darryl and Murasaki, Renato Toshiyuki and Primo-Pe{\~n}a, Elena and Bojo Canales, Cristina and Paliouras, Georgios and Krallinger, Martin and others}, year={2021}, organization={CEUR Workshop Proceedings} }</code></pre> <p>&nbsp;</p> <p><strong>Introduction</strong></p> <p>The main aim of MESINESP2 is to promote the development of practically relevant semantic indexing tools for biomedical content in non-English language. We have generated a manually annotated corpus, where domain experts have labeled a set of scientific literature, clinical trials, and patent abstracts. All the&nbsp;documents were labeled with DeCS descriptors, which is a structured controlled vocabulary created by BIREME to index scientific publications on BvSalud,&nbsp;the largest database of scientific documents in Spanish, which hosts records from the databases LILACS, MEDLINE, IBECS, among others.&nbsp;</p> <p>MESINESP track at BioASQ9 explores the efficiency of systems for assigning DeCS to different types of biomedical documents. To that purpose, we have divided the task into three subtracks depending on the document type. Then,&nbsp;for each one we generated an annotated corpus which was provided to participating teams:</p> <ul> <li><strong>[Subtrack 1 corpus] MESINESP-L &ndash; Scientific Literature:&nbsp;</strong>It contains all Spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish.</li> <li><strong>[Subtrack 2 corpus] <strong>MESINESP-T- Clinical Trials&nbsp;</strong></strong>contains records from&nbsp;<a href="https://reec.aemps.es/reec/public/web.html">Registro Espa&ntilde;ol de Estudios Cl&iacute;nicos (REEC)</a>. REEC doesn&#39;t&nbsp;provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC&nbsp;<a href="https://github.com/luisgasco/REECapi">API</a>.&nbsp;</li> <li><strong>[Subtrack 3 corpus] MESINESP-P &ndash; Patents:&nbsp;</strong>This corpus&nbsp;includes patents in Spanish extracted from Google Patents which have the IPC code &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;.</li> </ul> <p>In addition, we also provide a set of complementary data such as: the DeCS terminology file, a silver standard with the participants&#39; predictions to the task background set and the entities of medications, diseases, symptoms and medical procedures extracted from the BSC NERs documents.</p> <p>&nbsp;</p> <p><strong>Files structure:</strong></p> <p><strong>Silver_Standard_Mesinesp2.zip </strong>contains two separate sections. On the one hand, the union of the labels of the best model of each participating team as long as this model had obtained at least an F-score of 0.2 (folder <em>join</em>). On the other hand, the predictions of the best models of each participant have been included individually and anonymized&nbsp;(folder <em>separated</em>).&nbsp;This silver standard contains a set of <em>8642 scientific articles</em>, <em>1537 text sections from Clinical Practice Guidelines</em>, a set of <em>8458 text segments from Medication Data Sheets</em>, <em>461 clinical trials from REEC and 5170 patents</em>.&nbsp;</p> <p><strong>Subtrack1-Scientific_Literature.zip</strong> contains the corpora generated for subtrack 1. Content:</p> <ul> <li>Subtrack1: <ul> <li>Train:&nbsp; <ul> <li>training_set_track1_all.json: Full training set for subtrack 1.&nbsp;</li> <li>training_set_track1_only_articles.json:&nbsp;Articles training set for subtrack 1.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack1.json:&nbsp;</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack1.json: Test set for subtrack 1.&nbsp;</li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack2-Clinical_Trials.zip</strong> contains the corpora generated for subtrack 2. Content:</p> <ul> </ul> <ul> <li>Subtrack2: <ul> <li>Train <ul> <li>training_set_subtrack2.json: Training set for subtrack 2.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack2.json:&nbsp;Manually annotated&nbsp;development set for subtrack 2.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack2.json: Test set for subtrack 2.</li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack3-Patents.zip</strong> contains the corpora generated for subtrack 3. Content:</p> <ul> </ul> <ul> <li>Subtrack3: <ul> <li>Development <ul> <li>development_set_subtrack3.json:&nbsp;Manually annotated&nbsp;development set for subtrack 3.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack3.json: Test set for subtrack 3.</li> </ul> </li> </ul> </li> </ul> <p><strong>Additional data.zip&nbsp;</strong>contains the corpora with additional data for each subtrack of MESINESP2.</p> <p><strong>DeCS2020.tsv</strong> contains a DeCS table with the following structure:</p> <ul> <li>DeCS code</li> <li>Preferred descriptor (the preferred label in the Latin Spanish DeCS 2020&nbsp;set)</li> <li>List of synonyms (the descriptors and synonyms from&nbsp; Latin Spanish DeCS 2020&nbsp;set, separated by pipes.</li> </ul> <p><strong>DeCS2020.obo&nbsp;</strong>contains the *.obo file with the hierarchical relationships between DeCS descriptors.</p> <p>*Note: The <em>obo </em>and <em>tsv </em>files with DeCS2020 descriptors contain some additional COVID19 descriptors that will be included in future versions of DeCS. These items were provided by the Pan American Health Organization (PAHO), which has kindly shared this content to improve the results of the task by taking these descriptors into account.</p> <p>&nbsp;</p> <p><strong>Data format&nbsp;description</strong></p> <p>The&nbsp;<strong>input text files</strong>&nbsp;for the MESINESP track are JSON files with the following structure:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "title": "Metas de control de la presión arterial e impacto sobre desenlaces cardiovasculares en pacientes con diabetes mellitus tipo 2: un análisis crítico de la literatura", "abstractText": "La hipertensión arterial en individuos con diabetes mellitus tipo2 incrementa el riesgo de eventos cardiovasculares. Las guías internacionales de manejo recomiendan iniciar tratamiento farmacológico con valores de presión arterial &gt;140/90mmHg Sin embargo, no existe un punto de corte óptimo a partir del cual se logre reducir los eventos cardiovasculares sin originar eventos adversos; un rango de presión arterial &gt;130/80 y &lt;140/90mmHg parece ser el adecuado. Estos valores pueden alcanzarse mediante intervenciones no farmacológicas (dieta, ejercicio) y farmacológicas (por fármacos que hayan demostrado reducir eventos cardiovasculares). La elección de uno o varios fármacos debe ser individualizada, de acuerdo con factores como etnia, edad, comorbilidades asociadas, entre otros", "journal": "Clín. investig. arterioscler. (Ed. impr.)", "year": 2019, "db": "IBECS", "decsCodes": [ "D006973", "D000959", "D002318", "D003924", "D012307" ] } ] }</code></pre> <p>MESINESP&nbsp;<strong>entity mention files</strong>&nbsp;contain automatically generated mention annotations of medications, diseases, syntoms and medical procedures with the following JSON format:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "diseases": [ {"span": "hipertensión arterial", "start": "3", "end": "24"}, {"span": "diabetes mellitus tipo2", "start": "43", "end": "66"}, {"span": "eventos cardiovasculares", "start": "91", "end": "115"}], "medications": [], "procedures": [], "symptoms": []}] } ] }</code></pre> <p>&nbsp;</p> <p><strong>Dataset description:</strong><br> These corpora contain the data for each of the subtracks of MESINESP2 shared-task:</p> <ul> <li><strong>[Subtrack 1] MESINESP-L &ndash; Scientific Literature&nbsp;</strong>: &nbsp; <ul> <li><em><strong>Training set:&nbsp;</strong></em>It contains all spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish.&nbsp;We have filtered out empty abstracts and non-Spanish abstracts.&nbsp;&nbsp;We have built the training dataset with the data crawled on 01/29/2021. This means that the data is a snapshot of that moment and that may change over time since LILACS and IBECS usually add or modify indexes after the first inclusion in the database.&nbsp;We distribute two different datasets: <ul> <li><strong>Articles training set:&nbsp;</strong>This corpus contains the set of 237574 Spanish scientific papers in VHL that have at least one DeCS code assigned to them.</li> <li><strong>Full training set</strong>: This corpus contains the whole set of 249474 Spanish documents from VHL that have at leas one DeCS code assigned to them.</li> </ul> </li> <li><strong>Development set:&nbsp;</strong>We provided a development set manually indexed by our expert annotators (not VHL ones). This dataset includes 1065 articles annotated with DeCS by three expert indexers in this controlled vocabulary. The articles were initially indexed by 7 annotators, after analyzing the Inter-Annotator Agreement among their annotations we decided to select the 3 best ones, considering their annotations the valid ones to build the test set. From those 1065 records: <ul> <li>213 articles were annotated by more than one annotator. We have selected de union between annotations.</li> <li>852 articles were annotated by only one of the three selected annotators with better performance.</li> </ul> </li> <li><strong>Test set:</strong> We provide a test set containing 491 abstracts&nbsp;from LILACS and IBECS. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 2] <strong>MESINESP-T- Clinical Trials</strong></strong>: &nbsp; <ul> <li><strong>Training set:&nbsp;</strong>The training dataset contains records from&nbsp;<a href="https://reec.aemps.es/reec/public/web.html">Registro Espa&ntilde;ol de Estudios Cl&iacute;nicos (REEC)</a>. REEC doesn&#39;t&nbsp;provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC&nbsp;<a href="https://github.com/luisgasco/REECapi">API</a>.&nbsp;Clinical trials are not indexed with DeCS terminology, we have used as training data a set of 3560 clinical trials that were automatically annotated in the first edition of MESINESP and that were published as a&nbsp;<a href="https://zenodo.org/record/3946558#.YFHyhZ1KiUk">Silver Standard outcome</a>. Because the performance of the models used by the participants was variable, we have only selected predictions from runs with a MiF higher than 0.41, which corresponds with the submission of the best team.&nbsp;</li> <li><strong>Development set: </strong>We provide a development set manually indexed by expert annotators. This dataset includes 147 clinical trials annotated with DeCS by seven expert indexers in this controlled vocabulary.</li> <li><strong>Test set:&nbsp;</strong>The test dataset contains a collection of 248 items. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 3] MESINESP-P &ndash; Patents:&nbsp;</strong> <ul> <li><strong>Development set: </strong>We provide a Development set manually indexed by expert annotators. This dataset includes 115 patents in Spanish extracted from Google Patents which have the IPC code &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;. We have selected these patents based on semantic similarity to the MESINESP-L training set to facilitate model generation and to try to improve model performance.</li> <li><strong>Test set:&nbsp;</strong>We provide a&nbsp;<strong>test set</strong>&nbsp;containing 119 records that correspond to a subset of patents published in Spanish with the IPC codes &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;.Similarly to the development set, we selected these records based on semantic similarity to the MESINESP-L training set.&nbsp;We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>Additional data:</strong> <ul> <li>&nbsp;We provide this information to the participants as additional data in the &ldquo;Additional Data&rdquo; folder. For each training, development, and test set there is an additional JSON file with the structure shown <a href="https://temu.bsc.es/mesinesp2/resources/">here</a>. Each file contains&nbsp;entities related to medications, diseases, symptoms, and medical procedures extrated with the BSC NERs.</li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Summary statistics:</strong></p> <table align="center"> <caption>MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP-L</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">Training</th> <td>237574</td> <td>1988684</td> <td>22434</td> <td>43106663</td> </tr> <tr> <th scope="row">Development</th> <td>1065</td> <td>11283</td> <td>3750</td> <td>211420</td> </tr> <tr> <th scope="row">Test</th> <td>491</td> <td>5398</td> <td>2124</td> <td>93645</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <th scope="row">Training</th> <td>3560</td> <td>52257</td> <td>3940</td> <td>4133166</td> </tr> <tr> <th scope="row">Development</th> <td>147</td> <td>2038</td> <td>771</td> <td>146791</td> </tr> <tr> <th scope="row">Test</th> <td>248</td> <td>3271</td> <td>905</td> <td>267031</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <th scope="row">Development</th> <td>109</td> <td>1092</td> <td>520</td> <td>38564</td> </tr> <tr> <th scope="row">Test</th> <td>119</td> <td>1176</td> <td>629</td> <td>9065</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> </tbody> </table> <p>&nbsp; </p><table align="center"> <caption>General MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">MESINESP-L</th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>243313</td> <td>2065199</td> <td>22641</td> <td>48006345</td> </tr> </tbody> </table> <p></p> <p><strong>Related resources:</strong></p> <ul> <li><a href="http://temu.bsc.es/mesinesp2/">MESINESP2&nbsp;Web</a></li> <li><a href="https://github.com/BioASQ/Evaluation-Measures">Evaluation library</a></li> <li><a href="http://metodologia.lilacs.bvsalud.org/download/E/LILACS-4-ManualIndexacao-es.pdf">Annotation guidelines</a></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhCNKd8zhgD0rLwbhxGqF_wX">Participating teams Youtube Videos</a></li> <li><a href="http://ceur-ws.org/Vol-2936/">Proceedings of BioASQ@CLEF2021</a></li> <li><a href="http://bioasq.org/">BioASQ Web</a></li> </ul> <p>&nbsp;</p> <p>For further information, please&nbsp;email us at luis.gasco@bsc.es</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

Scholarly articles written by women extracted from Indexes of Archaeological Papers (1891-1907) and Gomme's Index of Archaeological Papers 1665-1890

<p>The dataset `List-of-Women-in-Archaeological-Indexes_cleaned.tsv` contains scholarly articles written by women extracted from annual <em>Indexes of Archaeological Papers</em> published between 1891 and 1907 inclusive and George Laurence Gomme&#39;s <a href="https://archive.org/details/indexofarchaeolo00gommrich/page/n5/mode/2up"><em>Index of Archaeological Papers 1665-1890</em></a>, referred to hereafter as the source datasets. These Indexes were published in London, initially by the Congress of Archaeological Societies directly, and from 1898 by Archibald Constable &amp; Co. The Indexes were sent to Societies subscribing to the Congress, but could also be acquired separately. A list of indexes consulted in available on our <a href="https://www.zotero.org/groups/4509023/beyond_notability/collections/FDKSUL8J">Zotero library</a>.</p> <p>The dataset is published in .tsv and .xslx formats.</p> <p>This is v2 of the dataset, including some cleaned publication titles and the addition of the socities that published each journal (v2.1 fixes a faulty dataset export in v2).</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Index / koronavírus

<p>This object has been created as a part of the web harvesting project of the E&ouml;tv&ouml;s Lor&aacute;nd University Department of Digital Humanities <a href="http://elte-dh.hu/en">ELTE DH</a>. Learn more about the workflow <a href="https://www.aclweb.org/anthology/2020.wac-1.5/">HERE about the software used </a><a href="https://github.com/ELTE-DH/WebArticleCurator/">HERE</a>.The aim of the project is to make online news articles and their metadata suitable for research purposes. The archiving workflow is designed to prevent modification or manipulation of the downloaded content. The current version of the curated content with normalized formatting in standard <a href="https://tei-c.org/">TEI XML</a> format with Schema.org encoded metadata is available <a href="https://doi.org/10.5281/zenodo.5833023">HERE</a>. The detailed description of the raw content is the following:</p><ul><li>The portal&#39;s archived content (from 2013-03-28 to 2021-01-30) in WARC format available <a href="https://doi.org/10.5281/zenodo.4899569">HERE</a> (crawled: 2021-01-30 13:01:33.164201 - 2021-01-31 19:47:27.191660).</li><li>The portal&#39;s archived content (from 2021-01-31 to 2021-05-24) in WARC format available <a href="https://doi.org/10.5281/zenodo.4899579">HERE</a> (crawled: 2021-05-25 08:47:51.634068 - 2021-05-25 17:04:43.254877).</li></ul><br>Please fill in the following form before requesting access to this dataset:<a href="https://forms.office.com/r/VnP05QYTGV">ACCES FORM</a>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Qualitative coding of brief videos that teach about the h-index

<p><strong>Dataset of qualitative coding of 31 Youtube videos on the h-index.&nbsp;</strong>The study aimed to characterize educational videos about the h-index to understand available resources and provide recommendations for future educational initiatives.</p> <p><em>Data.csv</em>: contains the metadata and qualitative coding for 31&nbsp;videos.</p> <p><em>ReadMe.csv</em>: contains the codebook including a description of variables.</p> <p><strong>Abstract. </strong>The authors analyzed videos on the h-index posted to YouTube. Videos were identified by searching YouTube and were screened by two authors. To code the videos the authors created a coding sheet, which assessed content and presentation style with a focus on the videos&rsquo; educational quality based on Cognitive Load Theory. Two authors coded each video independently with discrepancies resolved by group consensus.&nbsp;Thirty-one videos met inclusion criteria. Twenty-one videos (68%) were screencasts and seven used a &ldquo;talking head&rdquo; approach. Twenty-six videos defined the h-index (83%) and provided examples of how to calculate and find it. The importance of the h-index in high-stakes decisions was raised in 14 (45%) videos. Sixteen videos (52%) described caveats about using the h-index, with potential disadvantages to early researchers the most prevalent (n=7; 23%). All videos incorporated various educational approaches with potential impact on viewer cognitive load.&nbsp;Most videos (n=21; 68%) displayed amateurish production quality.&nbsp;The videos featured content with potential to enhance viewers&rsquo; metrics literacies such that many defined the h-index and described its calculation, providing viewers with skills to recognize and interpret the metric. However, less than half described the h-index as an author quality indicator, which has been contested, and caveats about h-index use were inconsistently presented, suggesting room for improvement. While most videos integrated practices to facilitate balancing viewers&rsquo; cognitive load, few (32%) were of professional production quality. Some videos missed opportunities to adopt particular practices that could benefit learning.&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Indicators and socio-spatial vulnerability index

<p>This table contains all variables used to compute the socio-spatial vulnerability index and the values of this index, at the dristrict scale, on the coastal zone of Bangladesh (16 districts).</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Habitat Protection Indexes - new monitoring measures for the conservation of threatened marine habitats - Datasets and supporting files

<p>The supporting datasets, scripts, and supplementary information for the manuscript, &quot;Habitat Protection Indexes -&nbsp;new monitoring measures for the conservation of threatened marine habitats,&quot; are available within this repository.</p> <p>We conduct an analysis on the coverage of protected areas that cover six threatened marine and coastal and developed two indexes, the Local Proportion of Habitat Protected&nbsp;Index and the Global Proportion of Habitat Protected Index, describing the protection of these habitats locally and globally. The habitats considered are the following: cold corals, warm water corals, knolls and seamounts, mangroves, saltmarshes, and seagrasses.</p> <p>The index scores of each jurisdiction are made available for download in the dataset: <em>habitat_protection_indexes_average.csv</em></p> <p>The habitat specific index scores for each jurisdiction are made available for download in the dataset: <em>habitat_protection_indexes.csv.&nbsp;</em></p> <p>Column name descriptions are available in the text file: <em>Column_name_descriptions_20220301</em></p> <p>The scripts used to run the workflow to calculate the indexes, create figures, and calculate statistics for the manuscript are also included. The script <em>01_Workflow sources</em> the first 9 scripts in the <em>scripts</em> folder to calculate the indexes which relies on the functions script within the functions folder. The rest of the scripts in the folder create the figures and calculate the statistics for the manuscript.</p> <p>A readme pdf file is included here to ease with reproducing the workflow, but we strongly suggest to please visit our github (<a href="https://github.com/jkumagai96/Marine_Habitat_protection">https://github.com/jkumagai96/Marine_Habitat_protection</a>) to reproduce the entire calculation where we provide detailed information on how to run the workflow and package management.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Geopotential-based Multivariate MJO Index (GMM Index)

<p>This repository contains the data of the Geopotential-based Multivariate MJO Index (GMM Index).<br> &nbsp;</p> <p><strong>About the GMM Index</strong></p> <ul> <li>The GMM index is an RMM-like index, derived from outgoing longwave radiation (OLR), 200hPa and 850hPa zonal wind, that can be extended to the pre-satellite era where satellite-based OLR data is unavailable.</li> <li>The short record of OLR observation limits the data length of the RMM index (<a href="https://journals.ametsoc.org/view/journals/mwre/132/8/1520-0493_2004_132_1917_aarmmi_2.0.co_2.xml">Wheeler and Hendon, 2004</a>), hindering the research of long-term variability of the Madden-Julian Oscillation (MJO) in the past century. And, the GMM index, which can be extended to the early 20th century (and even to the 19th century), is proposed to solve the problem.</li> <li>The construction method of the GMM index is the same as the RMM index, except that: <ul> <li>the OLR input is derived from upper-tropospheric geopotential data, based on the intrinsic relationship between MJO convection and upper-level geopotential (see&nbsp;<a href="https://doi.org/10.1007/s00382-016-3431-x">Leung and Qian, 2017</a>);</li> <li>the OLR, 200hPa and 850hPa zonal wind input of the GMM index are bandpass-filtered anomalies.</li> </ul> </li> <li>Please refer to&nbsp;<a href="https://doi.org/10.1007/s00382-022-06142-2">Leung et al. (2022)</a>&nbsp;for more detail about the calculation procedure and the theory behind it.</li> </ul> <p>&nbsp;</p> <p><strong>Files</strong></p> <ul> <li>GMM index derived based on the ERA-Interim reanalysis <ul> <li>Location:&nbsp;<a href="https://github.com/jeremychleung/Geopotential-based-Multivariate-MJO-Index/blob/main/data/gmm_index_erai.csv">data/gmm_index_erai.csv</a></li> <li>Temporal coverage: 1979&ndash;2013</li> </ul> </li> <li>GMM index derived based on the ERA-20C reanalysis <ul> <li>Location:&nbsp;<a href="https://github.com/jeremychleung/Geopotential-based-Multivariate-MJO-Index/blob/main/data/gmm_index_era20c.csv">data/gmm_index_era20c.csv</a></li> <li>Temporal coverage: 1900&ndash;2010</li> </ul> </li> <li>GMM index derived based on the NOAA-20CRv3 reanalysis <ul> <li>Link: (to be uploaded)</li> <li>Temporal coverage: 1836&ndash;2015</li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>If you use the GMM index in a publication or for any other purposes, please cite</p> <ul> <li>Leung, J.CH., Qian, W., Zhang, P. et al. Geopotential-based Multivariate MJO Index: extending RMM-like indices to pre-satellite era. Clim Dyn (2022).&nbsp;<a href="https://doi.org/10.1007/s00382-022-06142-2">https://doi.org/10.1007/s00382-022-06142-2</a></li> <li>Zenodo archive:&nbsp;<a href="https://doi.org/10.5281/zenodo.6331379">https://doi.org/10.5281/zenodo.6331379</a></li> </ul> <p>&nbsp;</p> <p><strong>References</strong></p> <ul> <li>Leung, J.CH., Qian, W., Zhang, P. et al. Geopotential-based Multivariate MJO Index: extending RMM-like indices to pre-satellite era. Clim Dyn (2022).&nbsp;<a href="https://doi.org/10.1007/s00382-022-06142-2">https://doi.org/10.1007/s00382-022-06142-2</a></li> <li>Leung, J.CH., Qian, W. Monitoring the Madden&ndash;Julian oscillation with geopotential height. Clim Dyn 49, 1981&ndash;2006 (2017).&nbsp;<a href="https://doi.org/10.1007/s00382-016-3431-x">https://doi.org/10.1007/s00382-016-3431-x</a></li> <li>Wheeler, M.C., Hendon, H.H. An All-Season Real-Time Multivariate MJO Index: Development of an Index for Monitoring and Prediction. Mon Weather Rev 132:1917&ndash;1932 (2004).&nbsp;<a href="https://journals.ametsoc.org/view/journals/mwre/132/8/1520-0493_2004_132_1917_aarmmi_2.0.co_2.xml">https://journals.ametsoc.org/view/journals/mwre/132/8/1520-0493_2004_132_1917_aarmmi_2.0.co_2.xml</a></li> </ul> <p>&nbsp;</p> <p><strong>Contact</strong></p> <p>If you have any questions about the data, feel free to contact Dr. Jeremy Leung&nbsp;(<a href="mailto:chleung@pku.edu.cn">chleung@pku.edu.cn</a>&nbsp;or&nbsp;<a href="mailto:liangzx@gd121.cn">liangzx@gd121.cn</a>).</p>

openother-openMar 2022View details →
zenodo44/100

Air stagnation index

<p>Daily air stagnation index datasets in China (2010-2021).</p> <p>Please refer to the study &quot;An air stagnation index to qualify extreme haze events in northern China&quot;(10.1175/JAS-D-17-0354.1).</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

1-km forest tree height, cover, plant area index, and foliage height diversity for the CONUS

<p>Consistent and spatially explicit periodic monitoring of forest structure is essential for estimating forest-related carbon emissions, analyzing forest degradation, and supporting sustainable forest management policies.&nbsp; To date, few products are available that allow for continental to global operational monitoring of changes in canopy structure.&nbsp; In this study, we explored the synergy between the NASA&rsquo;s spaceborne Global Ecosystem Dynamics Investigation (GEDI) waveform LiDAR and the Visible Infrared Imaging Radiometer Suite (VIIRS) data to produce spatially explicit and consistent annual maps of canopy height (CH), percent canopy cover (PCC), plant area index (PAI), and foliage height diversity (FHD) across the conterminous United States (CONUS) at 1-km resolution for 2013-2020.&nbsp; The accuracies of the annual maps were assessed using forest structure attribute derived from airborne laser scanning (ALS) data acquired between 2013 and 2020 for the 48 National Ecological Observatory Network (NEON) field sites distributed across the CONUS.&nbsp; The root mean square error (RMSE) values of the annual canopy height maps as compared with the ALS reference data varied from a minimum of 3.31-m for 2020 to a maximum of 4.19-m for 2017.&nbsp; Similarly, the RMSE values for PCC ranged between 8% (2020) and 11% (all other years).&nbsp; Qualitative evaluations of the annual maps using time series of very high-resolution images further suggested that the VIIRS-derived products could capture both large and &ldquo;more&rdquo; subtle changes in forest structure associated with partial harvesting, wind damage, wildfires, and other environmental stresses.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Compressed COBS indexes (XZ) - HQ 661k - Part 1

<p><strong>Compressed COBS indexes (XZ) - HQ 661k - v1</strong>&nbsp;&ndash;&nbsp;<a href="https://doi.org/10.5281/zenodo.6845083">part1</a>,&nbsp;<a href="https://doi.org/10.5281/zenodo.6849657">part2</a></p> <p>Part 1 of the compressed COBS indexes built from the high-quality assemblies&nbsp;of the 661k dataset.&nbsp;More information can be found on&nbsp;<a href="https://brinda.eu/mof/">https://brinda.eu/mof/</a>.</p> <p><strong>Citation:</strong></p> <blockquote> <p>K. Břinda, L. Lima, S. Pignotti, N. Quinones-Olvera, K. Salikhov, R. Chikhi, G. Kucherov, Z. Iqbal, and M. Baym,&nbsp;<a href="https://doi.org/10.1101/2023.04.15.536996"><strong>Efficient and robust search of microbial genomes via phylogenetic compression</strong></a>,&nbsp;<em>bioRxiv</em>&nbsp;2023.04.15.536996, 2023.&nbsp;<strong><a href="https://doi.org/10.1101/2023.04.15.536996">https://doi.org/10.1101/2023.04.15.536996</a></strong></p> </blockquote>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Compressed COBS indexes (XZ) - HQ 661k - Part 2

<p><strong>Compressed COBS indexes (XZ) - HQ 661k - v1</strong>&nbsp;&ndash;&nbsp;<a href="https://doi.org/10.5281/zenodo.6845083">part1</a>,&nbsp;<a href="https://doi.org/10.5281/zenodo.6849657">part2</a></p> <p>Part 2&nbsp;of the compressed COBS indexes built from the high-quality assemblies&nbsp;of the 661k dataset.&nbsp;More information can be found on&nbsp;<a href="https://brinda.eu/mof/">https://brinda.eu/mof/</a>.</p> <p><strong>Citation:</strong></p> <blockquote> <p>K. Břinda, L. Lima, S. Pignotti, N. Quinones-Olvera, K. Salikhov, R. Chikhi, G. Kucherov, Z. Iqbal, and M. Baym,&nbsp;<a href="https://doi.org/10.1101/2023.04.15.536996"><strong>Efficient and robust search of microbial genomes via phylogenetic compression</strong></a>,&nbsp;<em>bioRxiv</em>&nbsp;2023.04.15.536996, 2023.&nbsp;<strong><a href="https://doi.org/10.1101/2023.04.15.536996">https://doi.org/10.1101/2023.04.15.536996</a></strong></p> </blockquote>

opencc-by-4.0Jul 2022View details →
zenodo44/100

A termite genome reference and its Bowtie2 index

<p>This dataset contains a fasta file and its Bowtie2 index. The fasta file includes publicly available genomes of 5 termite species, namely <em>Zootermopsis nevadensis </em>(Terrapon, N., Li, C., Robertson, H. M., Ji, L., Meng, X., Booth, W., ... &amp; Liebig, J. (2014). Molecular traces of alternative social organization in a termite genome. Nature communications, 5(1), 1-12.), <em>Cryptotermes secundus</em> (Harrison, M. C., Jongepier, E., Robertson, H. M., Arning, N., Bitard-Feildel, T., Chao, H., ... &amp; Bornberg-Bauer, E. (2018). Hemimetabolous genomes reveal molecular basis of termite eusociality. Nature ecology &amp; evolution, 2(3), 557-566.), <em>Macrotermes natalensis</em> (Poulsen, M., Hu, H., Li, C., Chen, Z., Xu, L., Otani, S., ... &amp; Zhang, G. (2014). Complementary symbiont contributions to plant decomposition in a fungus-farming termite. Proceedings of the National Academy of Sciences, 111(40), 14500-14505.), <em>Coptotermes formosanus</em> (Draft genome sequence of the termite, Coptotermes formosanus: Genetic insights into the pyruvate dehydrogenase complex of the termite) and <em>Reticulitermes speratus </em>(Shigenobu, S., Hayashi, Y., Watanabe, D., Tokuda, G., Hojo, M. Y., Toga, K., Saiki, R., Yaguchi, H., Masuoka, Y., Suzuki, R., Suzuki, S., Kimura, M., Matsunami, M., Sugime, Y., Oguchi, K., Niimi, T., Gotoh, H., Hojo, M. K., Miyazaki, S., &hellip; Maekawa, K. (2022). Genomic and transcriptomic analyses of the subterranean termite Reticulitermes speratus: Gene duplication facilitates social evolution. Proceedings of the National Academy of Sciences, 119(3), e2110361119.). These genomic sequences have been classified with Kraken 2 v2.1.2 (Wood, D. E., Lu, J., &amp; Langmead, B. (2019). Improved metagenomic analysis with Kraken 2. Genome Biology, 20(1), 1&ndash;13 and Wood, D. E., &amp; Salzberg, S. L. (2014). Kraken: Ultrafast metagenomic sequence classification using exact alignments. Genome Biology, 15(3).) to remove all microbial sequences. This cleaned fasta file was indexed using the bowtie2-build command from Bowtie2 (Langmead, B., &amp; Salzberg, S. L. (2012). Fast gapped-read alignment with Bowtie 2. Nature Methods, 9(4), 357&ndash;359.) and can be used to perform &nbsp;alignments.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Numerical refractive index correction for the stitching procedure in tomographic quantitative phase imaging – dataset

<p>Raw volumetric data used in the work &quot;Numerical refractive index correction for the stitching procedure in tomographic quantitative phase imaging&quot; (<a href="http://doi.org/10.1364/BOE.466403">doi.org/10.1364/BOE.466403</a>). The data is packaged using the FIJI BigStitcher into HDF5 file. The file is split into 89 parts in ZIP format. Additionally we provide XML file needed for opening the data with BigStitcher and the TXT file with the nominal locations of the volumes based on the readings from the X-Y translation stage. The volumes inside the HDF5 file are already registered for stitching using the BigStitcher pairwise registration and global optimization procedure. Using the BigStitcher option &quot;Resave to TIFF&quot; one can access the raw data that we processed in the work. The processing code which operates on TIFF files is available here: <a href="https://github.com/biopto/QPI-stitching-2D-3D">https://github.com/biopto/QPI-stitching-2D-3D</a>.</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

diFUME Leaf Area Index V0.1

<p>Description:</p> <p>The Level 2A (L2A) product by the Theia Land Data Centre of CNES (Centre national d&#39;&eacute;tudes spatiales) is used, which provides georeferenced and orthorectified surface reflectance (SR), water vapor content (WVC), aerosol optical thickness (AOT), cloud and geophysical masks, processed by the MAJA atmospheric processing software. The 10 m SR bands in red (SRred : 665 nm) and near-infrared (SRNIR : 842 nm) are used to compute NDVI (Normalized Difference Vegetation Index) as (SRNIR &ndash; SRred)/(SRNIR + SRred) and the product is masked for clouds, cloud shadows and snow according to the L2A product flags. NDVI is converted to Leaf Area Index (LAI) values by applying an empirical exponential formula and is then resampled from 10 m to 5 m resolution, enhancing the initial LAI values, using the vegetation fraction derived by the 1 m Land Cover product.</p> <p>&nbsp;</p> <p>Data specifications:</p> <p>CRS: EPSG:32632 - WGS 84 / UTM zone 32N - Projected</p> <p>Spatial Extent: 392120.0,5266860.0 : 395160.0,5269840.0</p> <p>Temporal Extent: 2018 - 2020</p> <p>Units: meters</p> <p>Width: 608</p> <p>Height: 596</p> <p>Bands: 1</p> <p>Pixel Size: 5,-5</p> <p>Data type: Float32 - Thirty two bit floating point</p> <p>GDAL Driver Description: GTiff</p> <p>GDAL Driver Metadata: GeoTIFF</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Hourly LC impacts - Soil Quality Index - current mix and future scenarios, average demand

<p>Dataset on LCA results of electricity generation and supply&nbsp;in Italy for 2018, 2019 and 2020 (current mix) and two future scenarios (2030) - Soil Quality Index, average demand perspective.</p> <p>Modelling materials and methods are described in the paper &quot;Life-cycle assessment of current and future electricity supply in Italy: addressing average and marginal hourly demand&quot;.</p>

opencc-by-4.0Oct 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record