Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4,481

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

4,481 results for “list”

Learn how ShareScore rates datasets ↗
edi60/100

Fish and Amphibians species list of the Andrews Experimental Forest, 1987 to present

This is a compilation of fish species currently known to be present within the H.J. Andrews Experimental Forest. Species were extracted from several different aquatic vertebrate studies over time. Taxonomy is updated as needed.

openCC (other)Jan 2025View details →
edi60/100

Plant species list for Niwot Ridge and Green Lakes Valley, 1970 - ongoing.

A plant species list was created for Niwot Ridge and Green Lakes Valley from species identified in those areas by NWT scientists, working primarily at the Saddle and Martinelli sites. Additions to this list included species identified by Komarkova (1979) in the Indian Peaks Wilderness area but not on Niwot Ridge or in the Green Lakes Valley because of the likelihood that those species might exist within the LTER research area. Additions to the list were also provided by Terry Theodose, Leeanne Lestak, Teresa Nettleton, Susan Sherrod, Laura Mujica-Crapanzano (2004), Hope Humphries (2006), and Jane G. Smith (2019-2025). The list was revised to remove duplicate entries, correct typos, and resolve synonymy problems. Species and non-species categories received USDA PLANTS database names and codes.

openCC (other)Jul 2025View details →
zenodo56/100

S71 | CECSCREEN | HBM4EU CECscreen: Screening List for Chemicals of Emerging Concern Plus Metadata and Predicted Phase 1 Metabolites

<p>This is the collection associated with list S71 CECSCREEN HBM4EU CECscreen: Screening List for Chemicals of Emerging Concern Plus Metadata and Predicted Phase 1 Metabolites<strong> </strong>on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>CECScreen is part of the HBM4EU project (coord. UBA) &gt; WP16 &quot;emerging chemicals&quot; (lead INRA, JP Antignac/L Debrauwer) &gt; Task 16.1 (lead IRAS, J Vlanderen / R Vermeulen) &gt; Main contributor (J Meijer) &gt; Involved Partners (M Lamoree, T Hamers, S Hutinet, A, Covaci, C Huber, M Krauss, DI Walker, EL Schymanski). Further details in Meijer et al (2021) DOI: <a href="https://doi.org/10.1016/j.envint.2021.106511">10.1016/j.envint.2021.106511</a>. Dataset DOI: <a href="https://doi.org/10.5281/zenodo.3956586">10.5281/zenodo.3956586</a>.</p> <p>Update 23/7/2020 (v0.1.1): updated MetFrag files to remove elements causing errors (Os, Pd, Ag, Be). Update 8 Nov 2022 (v0.1.2) removed new lines in several synonyms as detected at BioHackEU22.</p>

opencc-by-4.0Dec 2019View details →
zenodo52/100

ABG-IMRHT and Ames-2000K IR line lists for N2O as reported in "Accurate N2O IR Line Lists with Consistent Empirical Line Positions: ABG-IMRHT and Ames-2000K"

<p><strong>[2025-06-17, v2.0</strong>, updated from v1.1 (<a href="https://doi.org/10.5281/zenodo.14834513">10.5281/zenodo.14834513</a>)]<br>(1). This upgraded version is recommended for analysis involving B1b or ABG(-IMRHT) related line lists, or E' &gt; 10,000 cm<sup>-1</sup>.&nbsp; &nbsp;<br>(2). Energy level lists computed on the B1b PES and related 296K ABG-IMRHT line intensity and line lists, and B1b-based 1000-3000K line list files have been updated, along with n2olist.f90.v1.4.&nbsp;<br>(3). The full 296 K IR line lists of&nbsp;<strong>all 12 stable isotopologues</strong> are provided in "n2o.296K.ABG-IMRHT.20250602.with.broadening.dat.v5corrected.100pct-abundance.xz", in which the line intensities assume 100% abundance for every isotopologue.<br>(4). Please note that the B1b PES-based energy levels in "empirical.corrections.for.ABG-IMRHT.zip" are still valid, but the B1b PES based energy levels are not updated in the files inside that .zip.&nbsp;</p> <p>[<strong>Link to <a href="https://www.sciencedirect.com/science/article/pii/S0022407325001645">the N2O paper</a> </strong>published at JQSRT (2025) 343, 109502, part of VSI:HITRAN2024, doi:<a href="https://doi.org/10.1016/j.jqsrt.2025.109502">10.1016/j.jqsrt.2025.109502</a>]</p> <p>[Updated from v1.0 (<a href="https://doi.org/10.5281/zenodo.14174307">10.5281/zenodo.14174307</a>). The J=151-210 transitions of <sup>14</sup>N<sub>2</sub><sup>16</sup>O were missing from v1.0&nbsp; files ]</p> <p><strong>1. Second generation of Ames-296K IR line list for "natural" Nitrous Oxide (N<sub>2</sub>O), denoted ABG-IMRHT</strong>, which was computed from Ames-B1b PES refinement (using Benjamin Schr&ouml;der's ab initio PES Comp I, doi:10.1515/zpch-2015-0622) and 2023 dmsG-10K accurately fitted from CCSD(T)/aug-cc-pV(T,Q,5)Z dipoles. In this major upgrade to Ames-296K N<sub>2</sub>O line list (10.1080/00268976.2023.2232892 and 10.5281/zenodo.7888194),&nbsp;rovibrational energy levels computed from NOSL-296 EH model were adopted to match and replace ~100,000 <sup>14</sup>N<sub>2</sub><sup>16</sup>O levels.&nbsp; More consistent empirical corrections are determined for multiple isotopologues from comparison with RITZ (IAO), MARVEL (ExoMol), HITRAN, and JPL datasets.&nbsp; The ABG-IMRHT line list provides the most reliable and consistent IR intensity predictions and accurate line positions in the range of 0 - 10,000 cm<sup>-1</sup>. All 12 stable isotopologues are included.&nbsp;</p> <p><strong>2. Ames-2000K</strong> IR line list provides complete, reliable and consistent IR predictions in the 0 - 15,000 cm<sup>-1</sup> range. It includes transitions of 12 isotopologues. Their intensities are scaled by corresponding terrestrial "natural" abundances. Ames-2000K is a composite list, including 4 component lists, to achieve better accuracy and reliability needs at both shorter and longer wavelengths:&nbsp;</p> <ul> <li>#1. a hot line list of&nbsp;<sup>14</sup>N<sub>2</sub><sup>16</sup>O, computed on Ames-1 PES and 2023 dmsG-wgt2d, J'&lt;210, E'&lt;25,000 cm<sup>-1</sup>, T=1000 / 1500 / 2000 / 3000 K.&nbsp; It occupies 24 GB in compressed .xz format.&nbsp;&nbsp;</li> <li>#2. 1000 K line lists of #2-#12 minor isotopologues, computed on Ames-1 PES and DMS, J'&lt;150, E'&lt;0.125 au - zpe (iso 2-6) or 0.08 -0.10 au - zpe (iso 7-12), S<sub>1000K </sub>&gt; 10<sup>-34</sup> cm/molecule, size-reduction with 99.9% intensity conservation in cm<sup>-1</sup> bins.&nbsp;</li> <li>#3. a hot line list of&nbsp;<sup>14</sup>N<sub>2</sub><sup>16</sup>O, computed on Ames-B1b PES and 2023 dmsG-10Kcm<sup>-1</sup>, J'&lt;150, E'&lt;16,000 cm<sup>-1</sup>, T=1000 / 1500 / 2000 / 3000 K.</li> <li>#4. ABG-IMRHT IR line list at 296 K, J'&lt;150, E'&lt;16,000 cm<sup>-1</sup>, with best empirical line positions and highly consistent intensity predictions up to 10,000 cm<sup>-1</sup>.&nbsp; Coverage beyond 10,000 cm<sup>-1 </sup>is limited to strong lines.</li> </ul> <p><strong>3. List of files:</strong>&nbsp; (decompress .xz files first, "xz -dkf -T0 file.xz")</p> <ul> <li>IAO_N2O_levels.tar.xz:&nbsp; rovibrational energy levels of 6 N<sub>2</sub>O isotopologues, as computed using global Effective Hamiltonian (EH) models developed by Dr. Sergey Tashkun from IAO (Institute of Atmospheric Optics, Tomsk, Russia, <a href="https://www.iao.ru/">https://www.iao.ru/</a>).&nbsp;<br><br></li> <li>N2O.Ames-B1b.PES.and.Ames-2023.DMS.zip:&nbsp; &nbsp;Ames-B1b PES subroutine &amp; coefficient file, see the note in n2opes2.f90; Ames 2023 dmsG subroutine and coefficients for fits up to 10K/12K/15K/17K/20K/25K cm<sup>-1</sup>, along with fitting residuals and compared to Ames-1 style (dmsC) coeffs and residuals.&nbsp; The geometry set has ~80 points in each 100 cm<sup>-1</sup>.&nbsp;<br><br></li> <li>n2olist.f90.v1.4:&nbsp; the main Fortran program for Ames-2000K generation, customizable, see the note at its beginning.&nbsp;<br>[ default Ames-2000K = Ames-1 (12 iso) + [B1b (446) + ABG-IMRHT (12 iso)] if (E'&lt;15,000 cm-1 &amp; J&lt;=150) ]<br><br></li> <li>empirical.corrections.for.ABG-IMRHT.zip: subroutine and paired/corrected energy level lists used in the empirical correction of energy levels computed on B1b PES [<em>this file is not updated from v1.1 to v2.0</em>]<br><br></li> <li>n2o.partition.1-4000K.iso.1-12.scaled: partition sum for 12 isotopologues, input file required by n2olist.f90, excluding g_n<br>n2o.partition.1-4000K.iso.1-12.scaled.with.degeneracy.txt:&nbsp; same as above, g_n included, see the note inside<br><br></li> <li>list.of.n2o.xz.files :&nbsp; list of .xz files to read/regenerate from, under subdir "xz", input file required by n2olist.f90<br>ames.n2o.xx000-yy000.cm-1.xz:&nbsp; compressed data files for Ames-2000K (line list component #1+#2)<br><br></li> <li>n2o.iso1-12.levels.Ames-1.dat.xz:&nbsp; energy levels of 12 N<sub>2</sub>O isotopologues on Ames-1 PES, input file required by n2olist.f90<br><br></li> <li>n2o.446.B1b-PES.levels.xz:&nbsp; energy levels of <sup>14</sup>N<sub>2</sub><sup>16</sup>O computed on Ames-B1b PES, input file to n2olist.f90<br><br></li> <li>n2o.446.B1b-dmsG10K.1000K-3000K.Eup16K.0-10Kcm-1.compressed.xz: data file for&nbsp;<sup>14</sup>N<sub>2</sub><sup>16</sup>O hot list on Ames-B1b and dms 2023-G10K (component #3), input file required by n2olist.f90<br>n2o.446.B1b-dmsG10K.1000K-3000K.Eup16K.0-10Kcm-1.dat.xz:&nbsp; independent line list (component #3)<br><br></li> <li>n2o.iso1-12.levels.ABG-IMRHT.dat.xz:&nbsp; energy levels in ABG-IMRHT list, with empirical corrections included, input file required by n2olist.f90</li> <li>n2o.296K.ABG-IMRHT.20250602.dat.v5.corrected.xz:&nbsp; Latest Ames-296K line list with empirical energy level corrections (component #4), input file for n2olist.f90<br>n2o.296K.ABG-IMRHT.20250602.with.broadening.dat.v5.corrected.xz:&nbsp; same as above, in HITRAN format, including line-broadening parameters<br><br></li> <li>n2o.296K.ABG-IMRHT.20250602.dat.v5corrected.100pct-abundance.xz: Latest Ames-296K line list with empirical energy level corrections, including the full set of IR transitions for <strong>12 stable isotopologues, assuming 100% abundannce </strong>for every isotopologue, and 1E-31 cm/molecule intensity cut-off at 296 K.</li> <li>n2o.296K.ABG-IMRHT.20250602.with.broadening.dat.v5corrected.100pct-abundance.xz: &nbsp;same as above, in HITRAN format, including line-broadening parameters<br><br></li> <li>ames.n2o.intensity.xz:&nbsp; line count and intensity sum (original, selected, iso #1 and iso #2-12)&nbsp; in each 0.01 cm<sup>-1</sup> bins at 296 K, 1000 K, 1500 K, 2000 K, and 3000 K, for reference and statistics, optional input file to n2olist.f90<br><br></li> <li>ames.n2o.sint.reductions.xz: A check-point file during Ames-2000K file generation, for reference only. including # of lines (total &amp; selected), intensity sum (total, iso #1, iso #2-12), and intensity retention ratio in each 0.01 cm<sup>-1</sup> bins for total, iso #1, and iso #2-12.<br><br></li> </ul> <p><strong>4. List of 12 isotopologues</strong>, abundances adopted, and number of lines in ABG-IMRHT and Ames-2000K line list,&nbsp;<br>[updated 2025-06-17 v2.0: the #lines in ABG-IMRHT are updated to S(296K) cut-off at 1E-31 cm/molecule, instead of 5E-31]</p> <table> <tbody> <tr> <td>#</td> <td>ISO</td> <td>abundance</td> <td># lines in ABG-IMRHT</td> <td># lines in Ames-2000K</td> </tr> <tr> <td>1</td> <td>446</td> <td>0.990333</td> <td>1388017</td> <td>3014643103</td> </tr> <tr> <td>2</td> <td>456</td> <td>3.64093E-3</td> <td>375541</td> <td>97932924</td> </tr> <tr> <td>3</td> <td>546</td> <td>3.64093E-3</td> <td>411353</td> <td>118432059</td> </tr> <tr> <td>4</td> <td>448</td> <td>1.98582E-3</td> <td>377552</td> <td>88169721</td> </tr> <tr> <td>5</td> <td>447</td> <td>3.69280E-4</td> <td>238921</td> <td>40212389</td> </tr> <tr> <td>6</td> <td>556</td> <td>1.33858E-5</td> <td>93933</td> <td>7534736</td> </tr> <tr> <td>7</td> <td>548</td> <td>7.30080E-6</td> <td>93674</td> <td>5583963</td> </tr> <tr> <td>8</td> <td>458</td> <td>7.30080E-6</td> <td>86446</td> <td>4995485</td> </tr> <tr> <td>9</td> <td>547</td> <td>1.35765E-6</td> <td>55360</td> <td>2588347</td> </tr> <tr> <td>10</td> <td>457</td> <td>1.35765E-6</td> <td>50754</td> <td>2275332</td> </tr> <tr> <td>11</td> <td>558</td> <td>2.68412E-8</td> <td>15814</td> <td>500190</td> </tr> <tr> <td>12</td> <td>557</td> <td>4.99134E-9</td> <td>8537</td> <td>274623</td> </tr> <tr> <td>&nbsp;</td> <td>Total</td> <td>1.00000069</td> <td>3195902</td> <td>3383142872</td> </tr> </tbody> </table> <p>5. This project is funded by NASA Grant 18-APRA18-0013 through NASA/SETI Institute Co-operative Agreement 80NSSC20K1358. &nbsp;Resources supporting this work were provided by the NASA High-End Computing (HEC) Program through the NASA Advanced Supercomputing (NAS) Division at Ames Research Center.</p>

opencc-by-4.0Dec 2024View details →
zenodo52/100

S0 | SUSDAT | Merged NORMAN Suspect List: SusDat

<p>This is the collection associated with list S0 SUSDAT on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a> and <a href="https://www.norman-network.com/nds/susdat/">https://www.norman-network.com/nds/susdat/</a></p> <p>S0 SUSDAT <strong>Merged NORMAN Suspect List: SusDat</strong></p> <p><a href="https://www.norman-network.com/nds/susdat/">Interactive Data table</a> (csv of latest version included in dataset here)</p> <p>UPDATED <strong><em>Jun 3, 2025 to latest version</em></strong>. Compiled by Reza Aalizadeh, University of Athens, including RTI and toxicity values, support by Nikiforos Alygizakis, EI. <em>Work in progress ... please report any issues!</em></p> <p>For an explanation of column headers, please see metadata files (xlsx and csv). For curation notes, see "SusDat_curation_notes.txt".&nbsp;</p>

opencc-by-4.0Jan 2024View details →
zenodo52/100

List of TEI rolename annotations in the ISicily EpiDoc corpus

<p>This CSV file details every instance of a 'roleName' tag in the I.Sicily (sicily.classics.ox.ac.uk) EpiDoc TEI files, reporting the ID number of the file in which it appears, and the value of the @type and @subtype attributes in each case - as such it serves as an index of roleName attestations in the I.Sicily dataset (also recoverable directly from the EpiDoc files). The file will be updated in future.</p>

opencc-by-4.0Apr 2024View details →
zenodo52/100

Digitised, searchable Holle List in Stokhof (1980)

<p>This repository contains the digitised Holle List in Stokhof (<a href="https://core.ac.uk/reader/159464813">1980</a>). Details and the interactive web version of the list can be accessed via <a title="Digitised, searchable Holle List" href="https://engganolang.github.io/digitised-holle-list/" target="_blank" rel="noopener">https://engganolang.github.io/digitised-holle-list/</a> (Rajeg 2023).</p> <p>The work in this repository is part of the <a href="https://gtr.ukri.org/projects?ref=AH%2FW007290%2F1">AHRC-funded research</a> on <a title="Lexical resources for Enggano" href="https://portal.sds.ox.ac.uk/Lexical_resources_for_Enggano" target="_blank" rel="noopener"><em>Lexical resources for Enggano, a threatened language of Indonesia</em></a> (central webpage of the Enggano project: <a title="Enggano research" href="https://enggano.ling-phil.ox.ac.uk/" target="_blank" rel="noopener">https://enggano.ling-phil.ox.ac.uk/</a>)</p> <h2>Updates in version 1.4.1</h2> <ul> <li>Adding the Transcription table (Stokhof 1980: 75-77, &sect;6.2) on the interactive webpage version (see <a title="Transcription Symbols" href="https://engganolang.github.io/digitised-holle-list/#:~:text=Transcription%20Symbols" target="_blank" rel="noopener">Table 4 "Transcription Symbols"</a>)</li> <li>Adding an update on the potential for the list to be included in the <em>Concepticon</em> (cf. the note&nbsp;<a href="https://github.com/concepticon/concepticon-data/issues/1324">here</a>)</li> <li>Adding reference to <em>EnoLEX</em>, a diachronic lexical database for the Enggano language (cf. <a href="https://doi.org/10.25446/oxford.28282169">here</a>)</li> </ul> <h2>References</h2> <p>Forkel, Robert, Johann-Mattis List, Simon J. Greenhill, Christoph Rzymski, Sebastian Bank, Michael Cysouw, Harald Hammarstr&ouml;m, Martin Haspelmath, Gereon A. Kaiping &amp; Russell D. Gray. 2018. Cross-Linguistic Data Formats, advancing data sharing and re-use in comparative linguistics. Scientific Data. Nature Publishing Group 5(1). 180205. <a href="https://doi.org/10.1038/sdata.2018.205" target="_blank" rel="noopener">https://doi.org/10.1038/sdata.2018.205</a>.</p> <p>Krau&szlig;e, Daniel; Rajeg, Gede Primahadi Wijaya; Pramartha, Cokorda Rai Adi; Zobel, Erik; Nothofer, Bernd; Hemmings, Charlotte; et al. (2024). EnoLEX: A diachronic lexical database for the Enggano language. University of Oxford. Online database. <a title="EnoLEX metadata record" href="https://doi.org/10.25446/oxford.28282169" target="_blank" rel="noopener">https://doi.org/10.25446/oxford.28282169</a>.</p> <p>Rajeg, Gede Primahadi Wijaya. 2023. Digitised, searchable Holle List in Stokhof (1980). Dataset. University of Oxford. <a href="https://doi.org/10.25446/oxford.23205173" target="_blank" rel="noopener">https://doi.org/10.25446/oxford.23205173</a></p> <p>Stokhof, W. A. L. (ed.). 1980. Holle lists, vocabularies in languages of Indonesia, vol. 1: Introductory volume. Vol. Materials in Languages of Indonesia. Canberra, A.C.T., Australia: Dept. of Linguistics, Research School of Pacific Studies, The Australian National University. <a title="Source Holle List in PDF" href="https://core.ac.uk/reader/159464813" target="_blank" rel="noopener">https://core.ac.uk/reader/159464813</a>.</p>

opencc-by-sa-4.0Dec 2022View details →
zenodo52/100

CLDF dataset of the Enggano word list from 1895 in Stokhof and Almanar's (1987) Holle List

<p>The repository for the digitised Enggano word list from 1895 (see Stokhof and Almanar 1987 for the original source) that has been matched with the <a href="https://engganolang.github.io/digitised-holle-list/">digitised Holle List</a> (Rajeg 2023a; cf. Stokhof 1980), providing the English and Indonesian glosses for the Enggano forms. The data set <a href="https://github.com/engganolang/holle-list-enggano-1895/actions/workflows/cldf-validation.yml">conforms</a> to the Wordlist module of the Cross-Linguistic Data Format (<a href="https://cldf.clld.org/">CLDF</a>) (Forkel et al. 2018).</p> <p><em>The work in this repository is part of the <a href="https://gtr.ukri.org/project/8AB0C3DC-F1C9-4CFA-BB4D-5BE748213372">AHRC-funded research</a> on <strong>Lexical resources for Enggano, a threatened language of Indonesia</strong> (visit the <a href="https://enggano.ling-phil.ox.ac.uk/">central webpage of the Enggano research</a> and the specific repository of the <a href="https://portal.sds.ox.ac.uk/Lexical_resources_for_Enggano">Lexical Resources for Enggano</a> project as well the <a href="https://portal.sds.ox.ac.uk/Enggano/groups">main Enggano repository</a> on the University of Oxford's Sustainable Digital Scholarship (SDS))</em></p> <h1>Updates in version 2.0.0</h1> <p>The following items summarise the major updates in version 2.0.0:</p> <ul> <li> <p><strong>Adding <a href="https://github.com/engganolang/holle-list-enggano-1895/blob/main/cldf/media.csv">MediaTable</a></strong> to accommodate <a href="https://github.com/engganolang/holle-list-enggano-1895/tree/main/img">images</a> in/for note ID &lt;26&gt; (commits <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/dab95401f128bd4203a81294f2e9f4620d45b145">dab9540</a> &amp; <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/a0040038e2577cb37ff8ad3ac68e8bbdedc26291">a004003</a> <a href="https://github.com/engganolang/holle-list-enggano-1895/blob/a0040038e2577cb37ff8ad3ac68e8bbdedc26291/code/Enggano-Holle-List-with-NBL.R#L229-L232">at this line</a> and <a href="https://github.com/engganolang/holle-list-enggano-1895/blob/2aab3ab385fc82c0613de2aacb5947ca4981b417/code/Enggano-Holle-List-with-NBL.R#L370-L397">these lines</a>)</p> </li> <li> <p><strong>Splitting multiple forms in a cell</strong> into their own rows, both for the original list and the forms in the Notes (commit <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/39cdc663843b265aa8f3c5bbdcb11628fcc17b5e">39cdc66</a> at <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/39cdc663843b265aa8f3c5bbdcb11628fcc17b5e#diff-ac46f8a3edb85868970d77f55bb86c5f0449feb25c37295895c4a9e560564301R83">this line</a> and <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/39cdc663843b265aa8f3c5bbdcb11628fcc17b5e#diff-ac46f8a3edb85868970d77f55bb86c5f0449feb25c37295895c4a9e560564301R83">this line</a>, and commit <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/a0040038e2577cb37ff8ad3ac68e8bbdedc26291">a004003</a> at <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/a0040038e2577cb37ff8ad3ac68e8bbdedc26291#diff-5b59e4a74b953f80c8867a60c422bd8405c9d6236bd61239a0d3f20b0d582b78R69">this line</a>)</p> </li> <li> <p><strong>Orthography transliteration</strong> into Enggano's common orthography and IPA (across several commits and [closed] issues [#1 #3 #4 #5 #7], but see <a href="https://github.com/engganolang/holle-list-enggano-1895/blob/2aab3ab385fc82c0613de2aacb5947ca4981b417/code/Enggano-Holle-List-with-NBL.R#L32-L93">these lines</a> for retrieving the existing orthography profile and doing the editing, and <a href="https://github.com/engganolang/holle-list-enggano-1895/blob/2aab3ab385fc82c0613de2aacb5947ca4981b417/code/Enggano-Holle-List-with-NBL.R#L220-L306">these lines</a> for running the transliteration using the <a href="https://cran.r-project.org/web/packages/qlcData/index.html">qlcData</a> R package [Moran &amp; Cysouw 2018; Cysouw 2024])</p> <ul> <li>In the <a href="https://github.com/engganolang/holle-list-enggano-1895/blob/main/cldf/forms.csv">FormTable</a>, the <code>Form</code> column contains the Enggano forms in their common orthography; the <code>Value</code> column contains their original transcription/orthography, with their tokenised/segmented formats available under the <code>Graphemes</code> column; the <code>Segments</code> column, finally, contains the segmented IPA transliteration of the Enggano forms (cf. #6 ). The <code>Comment</code> column is derived from the contents of the Notes. It includes, if any, Enggano forms in their original transcription followed by their segmented/tokenised forms in IPA in square brackets, their glosses in English (<strong>EN</strong>) and/or Indonesian (<strong>ID</strong>) inside the bracket, and finally the ID of the Notes in the original document inside angular brackets. The <code>English</code> and <code>Indonesian</code> columns respectively are glosses of the given language from the master/main Holle List (Stokhof 1980) that has been digitised (Rajeg 2023a).</li> <li>The output files of the orthography profiling and transliteration (commit <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/2aab3ab385fc82c0613de2aacb5947ca4981b417">2aab3ab</a>) are available in <a href="https://github.com/engganolang/holle-list-enggano-1895/tree/main/data-raw">data-raw</a> with the file names prefixed with <code>ortho-...</code>.</li> </ul> </li> </ul> <h2>References</h2> <p>Cysouw, Michael. 2024. qlcData: Processing Data for Quantitative Language Comparison. https://cran.r-project.org/web/packages/qlcData/index.html. (25 December, 2024). Version 0.3</p> <p>Forkel, Robert, Johann-Mattis List, Simon J. Greenhill, Christoph Rzymski, Sebastian Bank, Michael Cysouw, Harald Hammarstr&ouml;m, Martin Haspelmath, Gereon A. Kaiping &amp; Russell D. Gray. 2018. Cross-Linguistic Data Formats, advancing data sharing and re-use in comparative linguistics. Scientific Data. Nature Publishing Group 5(1). 180205. https://doi.org/10.1038/sdata.2018.205.</p> <p>Moran, Steven &amp; Michael Cysouw. 2018. The Unicode cookbook for linguists: Managing writing systems using orthography profiles (Translation and Multilingual Natural Language Processing 10). Berlin: Language Science Press. https://doi.org/10.5281/zenodo.1296780.</p> <p>Rajeg, Gede Primahadi Wijaya. 2023a. Digitised, Searchable Holle List in Stokhof (1980) [Data set]. (1.3.0). Zenodo. https://doi.org/10.5281/ZENODO.7972273. https://engganolang.github.io/digitised-holle-list/. https://ora.ox.ac.uk/objects/uuid:a511951b-86fb-4019-94d4-280efa83de02</p> <p>Rajeg, Gede Primahadi Wijaya. 2023b. CLDF dataset of the Enggano word list from 1895 in Stokhof and Almanar's (1987) Holle List [Data set]. https://github.com/engganolang/holle-list-enggano-1895 https://doi.org/10.25446/oxford.23515788</p> <p>Stokhof, W. A. L., ed. 1980. Holle Lists, Vocabularies in Languages of Indonesia, Vol. 1: Introductory Volume. Vol. Materials in Languages of Indonesia. Canberra, A.C.T., Australia: Dept. of Linguistics, Research School of Pacific Studies, The Australian National University. https://core.ac.uk/reader/159464813.</p> <p>Stokhof, W. A. L., and Alma E. Almanar. 1987. Holle Lists, Vocabularies in Languages of Indonesia, Vol. 10/3: Islands Off the West Coast of Sumatra. Vol. Materials in Languages of Indonesia. Pacific Linguistics (Series d) 76. Canberra, A.C.T., Australia: Dept. of Linguistics, Research School of Pacific Studies, The Australian National University. http://hdl.handle.net/1885/144589.</p>

opencc-by-sa-4.0Dec 2022View details →
zenodo52/100

S21 | UATHTARGETS | University of Athens Target List

<p>This is the collection associated with list S21 UATHTARGETS on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>S21 | UATHTARGETS | <strong>University of Athens Target List </strong></p> <p>Update 22/3/2020: added InChIKey file. Update 8/2/2022: new files from Dec 2021 with new compounds, NORMAN ID, classification and comments (provided by Maristina Nika). (v0.2.1 - attempted fix of CSV headers; v0.2.2 many small fixes flagged via PubChem deposit)</p> <p>Additional grant acknowledgement: Aristeia-Excellence: Transformation products of emerging pollutants in the aquatic environment (TREMEPOL project), 2012-2015, &nbsp;European Social Fund-Ministry of Education, <a href="http://tremepol.chem.uoa.gr/">http://tremepol.chem.uoa.gr/</a></p>

opencc-by-4.0Mar 2018View details →
zenodo52/100

List of the structures of S-protein in complex with ligands deposited in the Protein Data Bank until the 1st January 2021.

<p>All 131 structures of SARS-CoV-2 S-protein in complex with a ligand released on the PDB until the 1<sup>st</sup> January 2021 were categorised by ligand type: hACE2, antibody Fab fragments, VHH antibody fragments or <em>de novo</em> designed peptide scaffolds. The ligands&rsquo; amino acid sequences, the method by which the structures were determined and their resolution were retrieved from the PDB. Information regarding the ligands&#39; production method, dissociation constants (K<sub>D</sub>), S-protein segment against which the K<sub>D</sub> were measured and the determination methods were retrieved from the respective references. The categorisation of ligands by S-protein binding site and listing of S-protein conformation in each structure were achieved by visual analysis of all the structures using molecular visualisation software PyMOL.</p>

opencc-by-4.0Sep 2021View details →
zenodo48/100

Curated list of HAR datasets

<p>A curated list of <em>preprocessed</em> &amp; <em>ready to use under a minute</em> Human Activity Recognition datasets.</p> <p>All the datasets are preprocessed in <a href="https://www.hdfgroup.org/solutions/hdf5/">HDF5</a> format, created using the <a href="http://www.h5py.org">h5py</a> python library. Scripts used for data preprocessing are provided as well (Load.ipynb and load_jordao.py)</p> <p>Each HDF5 file contains at least the keys:</p> <ul> <li><code>x</code> a single array of size <code>[sample count, temporal length, sensor channel count]</code>, contains the actual sensor data. Metadata contains the names of individual sensor channel count. All samples are zero-padded for constant length in the file, original lengths before padding available under the <code>meta</code> keys.</li> <li><code>y</code> a single array of size <code>[sample count]</code> with integer values for target classes (zero-based). Metadata contains the names of the target classes.</li> <li><code>meta</code> contain various metadata, depends on the dataset (original length before padding, subject no., trial no., etc.)</li> </ul> <p>Usage example</p> <pre><code>import h5py with h5py.File(f'data/waveglove_multi.h5', 'r') as h5f: x = h5f['x'] y = h5f['y']['class'] print(f'WaveGlove-multi: {x.shape[0]} samples') print(f'Sensor channels: {h5f["x"].attrs["channels"]}') print(f'Target classes: {h5f["y"].attrs["labels"]}') first_sample = x[0] # Output: # WaveGlove-multi: 10044 samples # Sensor channels: ['acc1-x' 'acc1-y' 'acc1-z' 'gyro1-x' 'gyro1-y' 'gyro1-z' 'acc2-x' # 'acc2-y' 'acc2-z' 'gyro2-x' 'gyro2-y' 'gyro2-z' 'acc3-x' 'acc3-y' # 'acc3-z' 'gyro3-x' 'gyro3-y' 'gyro3-z' 'acc4-x' 'acc4-y' 'acc4-z' # 'gyro4-x' 'gyro4-y' 'gyro4-z' 'acc5-x' 'acc5-y' 'acc5-z' 'gyro5-x' # 'gyro5-y' 'gyro5-z'] # Target classes: ['null' 'hand swipe left' 'hand swipe right' 'pinch in' 'pinch out' # 'thumb double tap' 'grab' 'ungrab' 'page flip' 'peace' 'metal'] </code></pre> <p>Current list of datasets:</p> <ul> <li>WaveGlove-single (waveglove_single.h5)</li> <li>WaveGlove-multi (waveglove_multi.h5)</li> <li>uWave (uwave.h5)</li> <li>OPPORTUNITY (opportunity.h5)</li> <li>PAMAP2 (pamap2.h5)</li> <li>SKODA (skoda.h5)</li> <li>MHEALTH (non overlapping windows) (mhealth.h5)</li> <li>Six datasets with all four predefined train/test folds<br> as preprocessed by Jordao et al. originally in <a href="https://github.com/arturjordao/WearableSensorData">WearableSensorData</a><br> (FNOW, LOSO, LOTO and SNOW prefixed .h5 files)</li> </ul>

opencc-by-4.0May 2020View details →
zenodo48/100

S70 | EISUSGCEIMS | Environmental Institute GC-EI-MS suspect list

<p>This is the collection associated with list S70 EISUSGCEIMS Environmental Institute GC-EI-MS suspect list on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>GC-EI-MS suspect list of Environmental Institute. Provided by Peter Oswald, Nikiforos Alygizakis, Martina Oswaldova, Jaroslav Slobodnik. Dataset DOI: <a href="https://doi.org/10.5281/zenodo.3894827">10.5281/zenodo.3894827</a>.</p>

opencc-by-4.0Jun 2020View details →
zenodo48/100

Listed values of Ibex 35 companies from January 2020 to November 2020

<p>The data included in the&nbsp;dataset are the listed values of the Ibex 35 companies from January 2020 to November 2020. The dataset is a file for each Ibex 35 company.</p>

opencc-by-4.0Nov 2020View details →
zenodo48/100

S100 | PFASREACH | List of PFAS identified in REACH 2019

<p>This is the collection associated with list S100 PFASREACH&nbsp;List of PFAS identified in REACH 2019 on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>A list of 437 PFAS identified in Registration, Evaluation, Authorisation and Restriction of Chemicals <a href="https://eur-lex.europa.eu/eli/reg/2008/1272">(REACH) Reg. (EC) No 1272/2008</a> in September 2019. Of these, 17 are produced at &gt;1000 tonnes/year and 84 at &gt;10 tonnes per year for use in Europe. Collaborative effort between Hans Peter Arp (NGI) and Emma Schymanski (LCSB) within <a href="https://zenodo.org/communities/zeropm-h2020?page=1&amp;size=20">ZeroPM</a> (EU H2020 grant 101036756).</p> <p>Updates: 14 Dec 2022: added 5 new CIDs following PubChem deposition. 16 April 2025: removed CID <a href="https://pubchem.ncbi.nlm.nih.gov/compound/10996402">10996402</a> and <a href="https://pubchem.ncbi.nlm.nih.gov/compound/21967041">21967041</a> as they are not PFAS, due to external report via PubChem.&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo48/100

S49 | CPPDBLISTB | Database of Chemicals possibly (List B) associated with Plastic Packaging (CPPdb)

<p>This is the collection associated with list S49 CPPDBLISTB on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>S49 | CPPDBLISTB | <strong>Database of Chemicals associated with Plastic Packaging (CPPdb)</strong></p> <p>A database of chemicals likely (List A, 903 - in another upload) and possibly (List B, 3353 - this upload) associated with plastic packaging, with hazard data, from Groh et al 2019 DOI: <a href="https://doi.org/10.1016/j.scitotenv.2018.10.015">10.1016/j.scitotenv.2018.10.015</a>. Mapped to structures by CAS/Name by K. Groh &amp; E. Schymanski. 2025: added new CSV file with duplicate headers renamed.&nbsp;</p> <p>Latest version of original data (last update Oct 2018): DOI: <a href="http://doi.org/10.5281/zenodo.1287773">10.5281/zenodo.1287773</a></p>

opencc-by-4.0Mar 2019View details →
zenodo48/100

Taxonomic list of Brazilian fruit-bearing plants for human use

<h3>Lista taxon&ocirc;mica de plantas frut&iacute;feras para consumo humano, com curadoria da equipe do projeto <a href="https://www.inaturalist.org/projects/pomar-urbano">Pomar Urbano</a>.&nbsp;</h3> <p><em>[see English description below]</em></p> <p><br>As planilhas est&atilde;o organizadas da seguinte forma:</p> <p><strong>PT_lista_especies_aceitas_v.3.0</strong>: cont&eacute;m os nomes de todas as esp&eacute;cies atualmente indexadas na base de dados do <a href="https://www.inaturalist.org/projects/pomar-urbano">Pomar Urbano</a>.</p> <p><strong>PT_lista_especies_adicionadas_v.3.0</strong>: cont&eacute;m os nomes das novas esp&eacute;cies que passam a integrar a base de dados do Pomar Urbano a partir da vers&atilde;o 3.0.</p> <p><strong>PT_lista_especies_removidas_v3.0</strong>: cont&eacute;m os nomes das esp&eacute;cies removidas da vers&atilde;o 3.0 da lista, e que portanto n&atilde;o fazem mais parte do banco de dados do projeto.&nbsp;</p> <p>&nbsp;</p> <p><strong>Metadados usados nas planilhas:</strong></p> <ul> <li><em>Nome cient&iacute;fico</em>: O nome cient&iacute;fico completo, com autoria e data, se conhecidos.</li> <li><em>Fam&iacute;lia</em>: O nome cient&iacute;fico completo da fam&iacute;lia.</li> <li><em>Nome vernacular</em>: nome comum, popular.</li> <li><em>Origem</em><strong>: </strong>Declara&ccedil;&atilde;o sobre se um organismo foi introduzido em um local e tempo espec&iacute;ficos por meio da atividade direta ou indireta dos seres humanos modernos.</li> <li><em>Distribui&ccedil;&atilde;o geogr&aacute;fica</em>: &aacute;rea geogr&aacute;fica ou regi&atilde;o onde uma esp&eacute;cie ocorre no Brasil. Foram considerados como valores v&aacute;lidos para este campo apenas as macrorregi&otilde;es do Brazil, a saber: S = Sul, SE = Sudeste, CO = Centro-Oeste, NE = Nordeste, N = Norte.</li> <li><em>&Uacute;ltima atualiza&ccedil;&atilde;o</em>: A data mais recente em que a entrada no cat&aacute;logo foi alterada, atualizada ou modificada.</li> </ul> <h3>--------------------------------------------------------------------------------------------------------------------------------------<br><br>Taxonomic list of fruit-bearing plants for human consumption, curated by the <a href="https://www.inaturalist.org/projects/pomar-urbano">Pomar Urbano project</a></h3> <p><em>[Vernacular names are presented only in Portuguese; for properly processing in data management tools, downloading a Portuguese language package might be necessary]</em></p> <p>The spreadsheets are organized as follows:</p> <p>EN_list_accepted_species_v.3.0: contains the names of all species currently indexed in the <a href="https://www.inaturalist.org/projects/pomar-urbano">Pomar Urbano</a> database.</p> <p>EN_new_added_species_v.3.0: contains the names of new species that are included in the Pomar Urbano database starting from version 3.0.</p> <p>EN_removed_species_v.3.0: contains the names of species that were present in the version 2.0 of the list and are therefore no longer part of the version 3.</p> <p>&nbsp;</p> <p><strong>Metadata used in the spreadsheets</strong>:</p> <p><em>Scientific Name</em>: The complete scientific name, including authorship and date, if known. <em>ExactMatch</em>: <a href="http://rs.tdwg.org/dwc/terms/scientificName">dwc:scientificName</a>.&nbsp;</p> <p><em>Family</em>: The full scientific name of the family. <em>ExactMatch</em>: <a href="http://rs.tdwg.org/dwc/terms/family">dwc:family.</a></p> <p><em>Vernacular Name</em>: Common or popular name. <em>ExactMatch</em>: <a href="http://rs.tdwg.org/dwc/terms/vernacularName">dwc:vernacularName</a></p> <p><em>Establishment Means</em>: Statement about whether an organism has been introduced to a specific place and time through the direct or indirect activity of modern humans. <em>ExactMatch</em>: <a href="http://rs.tdwg.org/dwc/terms/establishmentMeans">dwc:establishmentMeans</a></p> <p><em>Higher geography</em>: The geographical area or region where a species occurs in Brazil. Only the macroregions of Brazil are considered valid values for this field within this dataset, namely: S = South, SE = Southeast, CO = Central-West, NE = Northeast, N = North. <em>CloseMacth</em>: <a href="http://rs.tdwg.org/dwc/terms/higherGeography">dwc:higherGeography</a></p> <p><em>Last Update</em>: The most recent date on which the catalog entry was changed, updated, or modified. <em>ExactMatch</em>: <a href="http://purl.org/dc/terms/modified">dct:modified</a></p> <p>&nbsp;</p>

opencc-zeroNov 2023View details →
zenodo48/100

List of capacity building resources for combating climate mis/disinformation created by EU-funded projects

<p>This dataset is the result of collaborative work for Deliverable 1.3 (WP1; T1.3) of the AGORA project. It compiles resources from projects funded by the European Commission under the last two Framework Programmes (Horizon 2020 and Horizon Europe) and focused on combating climate change misinformation and disinformation. The resources identified and analysed include training materials, guidelines and interactive digital platforms designed for various target groups.</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

List of capacity building resources for climate change adaptation created by EU-funded projects

<p>This dataset is the result of collaborative work for Deliverable 1.3 (WP1; T1.3) of the AGORA project. It compiles resources from projects funded by the European Commission under the last two Framework Programmes (Horizon 2020 and Horizon Europe) and focused on climate change adaptation. The resources identified and analysed include training materials, guidelines and interactive digital platforms designed for various target groups.</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

S65 | UATHTARGETSGC | University of Athens GC-APCI-HRMS Target List

<p>This is the collection associated with list S65 UATHTARGETSGC - University of Athens GC-APCI-HRMS Target List on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>GC-APCI-HRMS target list of University of Athens. Provided by the research group of Prof. Nikolaos Thomaidis (<a href="http://trams.chem.uoa.gr/">http://trams.chem.uoa.gr/</a>) and hosted on the NORMAN Suspect List Exchange (<a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a>). DOI: 10.5281/zenodo.3753372.</p> <p>Updated Feb 2024 following feedback from Peter Oswald, EI.</p>

opencc-by-4.0Apr 2020View details →
zenodo48/100

Construction Industry Steel Ordering Lists (CISOL) Dataset

<p>The Construction Industry Steel Ordering Lists (CISOL) dataset comprises table-centric, real-world documents from the construction industry, annotated to facilitate the testing and training of deep learning models for table detection (TD) and table structure recognition (TSR).&nbsp;</p> <p>CISOL Key Features:</p> <ul> <li>Steel ordering lists from 24 construction projects carried out between 2015-2023, contributed by 10 distinct German structural engineering firms.</li> <li>Anonymized images to ensure the unrecognizability of specific project or creator information.</li> <li>A total of 3280 images, with 844 annotated following the CISOL annotation guidelines.</li> </ul> <p>CISOL is structured into two tracks:</p> <ul> <li><strong>Track A: TD-TSR&nbsp;</strong>version for end-to-end table detection and table structure recognition tasks.</li> <li><strong>Track B: TSR-</strong>only version for table structure recognition tasks, featuring images cropped to the actual table areas with accordingly adjusted annotations.</li> </ul> <p>The dataset is developed in accordance with the FAIR Principles, ensuring that it is Findable, Accessible, Interoperable, and Reusable. The CISOL dataset permits expansion following the established annotation guideline.</p> <p>Access to the CISOL Leaderboard will be provided at <a href="https://eval.ai/web/challenges/challenge-page/2257" target="_blank" rel="noopener">EvalAI.</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record