Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,481
datasets available to search
ShareScore release 0.9.0
Dataset results
4,481 results for “list”
CLDF dataset derived from Aaley and Bodt's "New Kusunda data: A list of 250 concepts" from 2020
<p>Cite the source of the dataset as:</p> <blockquote> <p>Uday Raj Aaley and Timotheus A. Bodt (2020): New Kusunda data: A list of 250 concepts. Computer-Assisted Language Comparison in Practice 3.4 (08/04/2020), URL: https://calc.hypotheses.org/2414.</p> </blockquote>
Updates of listings of published versions and references relating to Rubáiyát of Omar Khayyám
<p>These files provide updates of listings of published versions of the <em>Rubáiyát </em>of Omar Khayyám and references relating to <em>Rubáiyát</em>. They form part of an archive of research data relating to the spread and influence of the <em>Rubáiyát</em> of Omar Khayyám. The data have been compiled by independent researchers W H (Bill) Martin and Sandra Mason; their contact details are in the README file. The files comprise a number of searchable listings relating to the poem and its different manifestations. </p> <p>This section of the archive contains two database tables updating earlier material to September 2018<em>.</em> The first of these, ROKlist2018, provides a listing of over 1750 editions or reprints of the poem, in different languages and translations, which we know either to exist or to have been identified by other compilers of <em>Rubáiyát </em>bibliographies. The second database ROKref2018 lists more than 500 books, articles and other material relating to the <em>Rubáiyát of Omar Khayyám</em>, which we have identified in the course of our research. Further details of sources, and of the fields and codings used in the databases, are given in the accompanying README document.</p>
S8 | ATHENSSUS | University of Athens Surfactants and Suspects List
<p>This is the collection associated with list S8 ATHENSSUS on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/?q=suspect-list-exchange">https://www.norman-network.com/?q=suspect-list-exchange</a></p> <p>S8ATHENSSUS<strong>University of Athens Surfactants and Suspects List </strong></p> <p>Gago Ferrero <em>et al </em><a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/031017Update/GagoFerrero_etal_2015_SuspectsNontargets_wDTXSIDs.csv">CSV</a>, <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/031017Update/GagoFerrero_etal_2015_SuspectsNontargets_wDTXSIDs.xlsx">XLSX</a> (3/10/2017)</p> <p>CompTox <a href="https://comptox.epa.gov/dashboard/chemical_lists/athenssus">ATHENSSUS List</a></p> <p><a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/UniAthens_SuspectAndSurfactants_InChIKeys.txt">UniAthens InChIKeys</a> (28/01/2016)</p> <p>Gago-Ferrero <em>et </em><em>al</em>. 2015. DOI: <a href="http://pubs.acs.org/doi/abs/10.1021/acs.est.5b03454">10.1021/acs.est.5b03454</a></p>
S16 | FRENCHLIST | French Monitoring List
<p>This is the collection associated with list S16 FRENCHLIST on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/?q=suspect-list-exchange">https://www.norman-network.com/?q=suspect-list-exchange</a></p> <p>S16</p> <p>FRENCHLIST</p> <p><strong>French Monitoring List</strong></p> <p>French List <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/French_List_08052017.csv">CSV</a> (8/05/2017)</p> <p>CompTox <a href="https://comptox.epa.gov/dashboard/chemical_lists/frenchlist">French Monitoring List</a></p> <p>Further curation in progress…</p> <p><a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/FrenchList_UniqueInChIKeys_08052017.txt">French List Unique InChIKeys</a> (8/05/2017)</p> <p>Provided by Valeria Dulio, curated by Reza Aalizadeh, University of Athens.</p>
S23 | EIUBASURF | Surfactant Suspect List from EI and UBA
<p>This is the collection associated with list S23 EIUBASURF on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/?q=suspect-list-exchange">https://www.norman-network.com/?q=suspect-list-exchange</a></p> <p>S23</p> <p>EIUBASURF</p> <p><strong>Surfactant Suspect List from EI and UBA</strong></p> <p>Surfactant List <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/190618Update/SurfactantSuspects_EI_UBA_15032018_wDTXSIDs.xlsx">XLSX</a> (19/06/2018)<br> Surfactant List <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/190618Update/SurfactantSuspects_EI_UBA_15032018_wDTXSIDs.csv">CSV</a> (19/06/2018)<br> CompTox <a href="https://comptox.epa.gov/dashboard/chemical_lists/eiubasurf">EIUBASURF List</a> </p> <p>EI UBA Surfactant <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/190618Update/SurfactantSuspects_EI_UBA_15032018_InChIKeys.txt">InChIKeys</a> (19/06/2018)</p> <p>A compiled list of eco-labeled surfactants from Environmental Institute (EI, SK) and the German Federal Environmental Agency (UBA, DE) assigning chemical structures to UVCB chemicals based on names and prior knowledge. Provided by Nikiforos Alygizakis, EI.</p>
S51 | WRIGCHRMS | GC-HRMS target list of WRI
<p>This is the collection associated with list S51 WRIGCHRMS on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/?q=suspect-list-exchange">https://www.norman-network.com/?q=suspect-list-exchange</a></p> <p>S51</p> <p>WRIGCHRMS</p> <p><strong>GC-HRMS target list of WRI</strong></p> <p>WRIGCHRMS <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/080419Update/WRIGCHRMS_04042019.xlsx">XLSX</a>, <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/080419Update/WRIGCHRMS_04042019.csv">CSV</a> (04/04/2019)</p> <p>WRIGCHRMS <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/080419Update/WRIGCHRMS_InChIKeys_04042019.txt">InChIKeys</a> (04/04/2019)</p> <p>GC-HRMS target list of WRI. The method was established by Agilent. The list was provided by Michal Kirchner (Slovak Water Research Institute, WRI) and curated by Nikiforos Alygizakis (EI/UoA)</p>
List of (tentatively) identified non-target structures
<p>This is a list of candidate structures of organic contaminants (tentatively) identified in a riverbank filtration system in The Netherlands by non-target screening of high-resolution mass spectrometry data (associated with ACS Publication <a href="https://doi.org/10.1021/acs.est.9b01750">10.1021/acs.est.9b01750</a>).</p>
List.MID: A MIDI-Based Benchmark for Evaluating RDF Lists
<p>Linked lists represent a countable number of ordered values, and are among the most important abstract data types in computer science. With the advent of RDF as a highly expressive knowledge representation language for the Web, various implementations for RDF lists have been proposed. Yet, there is no benchmark so far dedicated to evaluate the performance of triple stores and SPARQL query engines on dealing with ordered linked data. Moreover, essential tasks for evaluating RDF lists, like generating datasets containing RDF lists of various sizes, or generating the same RDF list using different modelling choices, are cumbersome and unprincipled. In this paper, we propose List.MID, a systematic benchmark for evaluating systems serving RDF lists. List.MID consists of a dataset generator, which creates RDF list data in various models and of different sizes; and a set of SPARQL queries. The RDF list data is coherently generated from a large, community-curated base collection of Web MIDI files, rich in lists of musical events of arbitrary length. We describe the List.MID benchmark, and discuss its impact and adoption, reusability, design, and availability.</p>
Kusunda 250 Word List Audio Files
<p>This data set contains 2 zip files and 3 pdf files containing the cut sound files of the elicited 250-concept word list (plus numerous additional concepts and lexical items) recorded with the last two Kusunda speakers from Nepal, Gyani Maiya Sen Kusunda and Kamala Khatri (Sen Kusunda), on in late July and early August 2019, in Kathmandu, Nepal.</p> <p>As per our knowledge, these are the first recordings of Kusunda that are available in the public domain.</p> <p>The zip files, when unpacked, will reveal several hundreds of cut sound files. At least the concepts from the list are triple repeated. There are also numerous additional concepts on recording. The pdf files contain explanations regarding the elicited concepts as well as the way in which the forms of the two speakers were used to come to an underlying form, a 'reconstruction' of some sorts. </p> <p>The file numbers of the individual sound files refer to both the speakers (GM, K), to the date of recording and the exact origin file number. More information regarding a certain concept or pronunciation can be found in these respective origin sound files.</p> <p>Please note that the actual transcriptions in IPA may differ between the file names of the individual sound recordings, the individual speaker's pdf file, and the aggregated pdf file. We request users of these data to take the transcriptions in the aggregated pdf file as the most recent ones, or alternatively contact us for the most recent, updated transcriptions.</p> <p>This research was funded by a 2,000 USD grant from the Endangered Language Fund (<a href="http://www.endangeredlanguagefund.org/">http://www.endangeredlanguagefund.org/</a>), a 700 euro contribution by the European Research Council Starting Grant 715618 “Computer-Assisted Language Comparison” (<a href="http://calc.digling.org/">http://calc.digling.org</a>), and a total of 2,320 euros raised through crowdfunding at GoFundMe (<a href="https://www.gofundme.com/f/saving-the-kusunda-language-in-nepal">https://www.gofundme.com/f/saving-the-kusunda-language-in-nepal</a>). Many thanks to all generous contributors.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <strong><em>of any kind</em></strong><em>, which includes conversion into commercial audio-visual media (documentaries etc.), (paid) public screening other than for educational purposes, storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Create Commons Attribution 4.0 International / Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>We greatly value feedback, suggestions, advice, analysis etc. based on this material which will help in the description of the Kusunda language, especially any comments and suggestions that will enable the revitalisation of the language, including a standardisation of the phonology and a phonologically consistent but also practical orthography in both देवनागरी Devanāgarī and Roman script.</p> <p>Uday Raj Aaley: aarambhkhabar (at) gmail (dot) com</p> <p>Tim Bodt: bodttim (at) gmail (dot) com</p>
S39 | KEMIWWSUS | Wastewater Suspect List based on Swedish Product Data
<p>This is the collection associated with list S39 KEMIWWSUS on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/?q=suspect-list-exchange">https://www.norman-network.com/?q=suspect-list-exchange</a></p> <p>S39</p> <p>KEMIWWSUS</p> <p><strong>Wastewater Suspect List based on Swedish Product Data</strong></p> <p>Wastewater Suspect List <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/120219Update/Suspects_WasteWater_Sweden_KEMI20190212.xlsx">Original File with Mapped DTXSIDs</a> (12/02/2019)</p> <p>KEMIWWSUS <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/120219Update/KEMIWWSUS_InChIKeys_12022019.txt">InChIKeys</a> (12/02/2019)</p> <p>A prioritized list of 1,123 substances relevant for wastewater based on Swedish product registry data, including scores. Provided by Stellan Fischer, KEMI.</p> <p>Update 14/11/2019: added CSV version and resulting updated XLSX.</p> <p> </p>
S14 | KEMIPFAS | PFAS Highly Fluorinated Substances List: KEMI
<p>This is the collection associated with list S14 KEMIPFAS on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>Appendix 2 from Swedish Chemicals Agency <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/report-7-15-occurrence-and-use-of-highly-fluorinated-substances-and-alternatives.pdf">KEMI Report 7/15</a>. Provided by Stellan Fischer, KEMI. Registration and mapping to <a href="https://comptox.epa.gov/dashboard/">CompTox Dashboard</a> by Antony Williams, US EPA. Nov 17 2019 update: added CSV files for PubChem upload.</p>
CLDF Dataset derived from List's "Sample Size and Cognate Detection" from 2014
<p>Cite the source of the dataset as:</p> <blockquote> <p>List, Johann-Mattis (2014): Investigating the impact of sample size on cognate detection. Journal of Language Relationship. 11. 91-102. DOI: https://doi.org/10.31826/jlr-2014-110111</p> </blockquote>
IOC World Bird List (IOC) - active: IOC World Bird List with higherClassification
This is a Darwin Core Archive version of data from the [IOC World Bird List](<p></p>http://www.worldbirdnames.org/ "<p></p>http://www.worldbirdnames.org"). The IOC World Bird List is an open access resource of the international community of ornithologists. Their goal is to facilitate worldwide communication in ornithology and conservation based on an up-to-date classification of world birds and a set of English names that follows explicit guidelines for spelling and construction (Gill & Wright 2006). Citation: Gill, F & D Donsker (Eds). 2017. IOC World Bird List (v 7.1). doi: [10.14344/IOC.ML.7.1] (<p></p>http://doi.org/10.14344/IOC.ML.7.1) [More...](<p></p>http://www.worldbirdnames.org/ "<p></p>http://www.worldbirdnames.org")<p></p> The Encyclopedia of Life (EOL, eol.org) aggregates biodiversity information from more than 400 sources and provides access to the data through taxon pages, visual query and application programming interfaces. Scientific names are essential elements of the data integration infrastructure, but their shortcomings as key identifiers are well documented (Patterson et al., 2016). Complex automated workflows and continuous manual curation are required to address idiosyncrasies of source taxonomies, variation in data quality, and conflicting taxonomic opinions. To achieve a harmonized taxonomic view of EOL content, names from data sources are mapped to a dynamic reference hierarchy ([see current version here](<p></p>https://opendata.eol.org/dataset/tram-807-808-809-810-dh-v1-1/resource/00adb47b-57ed-4f6b-8f66-83bfdb5120e8)) using an algorithm that leverages canonical name strings, hierarchical information (ancestry, descendants), taxonomic ranks, synonym data, and author strings. Names that cannot be associated with a reference taxon are still accessible, but their unmapped status excludes them and any associated content from certain core EOL functions. For more information about the EOL taxonomy, see [EOL Dynamic Hierarchy](<p></p>https://eol.org/docs/eol-dynamic-hierarchy)
IOC World Bird List (IOC) - active: IOC World Bird List
This is a Darwin Core Archive version of data from the [IOC World Bird List](<p></p>http://www.worldbirdnames.org/ "<p></p>http://www.worldbirdnames.org"). The IOC World Bird List is an open access resource of the international community of ornithologists. Their goal is to facilitate worldwide communication in ornithology and conservation based on an up-to-date classification of world birds and a set of English names that follows explicit guidelines for spelling and construction (Gill & Wright 2006). Citation: Gill, F & D Donsker (Eds). 2017. IOC World Bird List (v 7.1). doi: [10.14344/IOC.ML.7.1] (<p></p>http://doi.org/10.14344/IOC.ML.7.1) [More...](<p></p>http://www.worldbirdnames.org/ "<p></p>http://www.worldbirdnames.org")<p></p>Simple version without higher classification. The Encyclopedia of Life (EOL, eol.org) aggregates biodiversity information from more than 400 sources and provides access to the data through taxon pages, visual query and application programming interfaces. Scientific names are essential elements of the data integration infrastructure, but their shortcomings as key identifiers are well documented (Patterson et al., 2016). Complex automated workflows and continuous manual curation are required to address idiosyncrasies of source taxonomies, variation in data quality, and conflicting taxonomic opinions. To achieve a harmonized taxonomic view of EOL content, names from data sources are mapped to a dynamic reference hierarchy ([see current version here](<p></p>https://opendata.eol.org/dataset/tram-807-808-809-810-dh-v1-1/resource/00adb47b-57ed-4f6b-8f66-83bfdb5120e8)) using an algorithm that leverages canonical name strings, hierarchical information (ancestry, descendants), taxonomic ranks, synonym data, and author strings. Names that cannot be associated with a reference taxon are still accessible, but their unmapped status excludes them and any associated content from certain core EOL functions. For more information about the EOL taxonomy, see [EOL Dynamic Hierarchy](<p></p>https://eol.org/docs/eol-dynamic-hierarchy)
National Checklists 2017: Guatemala Species List
Lists of taxa for each country and a few other administrative zones harvested from effechecka using simplified versions of geonames polygons. See <p></p>https://github.com/diatomsRcool/checklists for details<p></p>A list of species from Guatemala collected using effechecka and geonames polygons
National Checklists 2017: South America Species List
Lists of taxa for each country and a few other administrative zones harvested from effechecka using simplified versions of geonames polygons. See <p></p>https://github.com/diatomsRcool/checklists for details<p></p>List of species from South America inferred from individual country lists that were derived from effechecka and modified geonames polygons
National Checklists 2017: Oceania Species List
Lists of taxa for each country and a few other administrative zones harvested from effechecka using simplified versions of geonames polygons. See <p></p>https://github.com/diatomsRcool/checklists for details<p></p>List of species from Oceania inferred from individual country lists that were derived from effechecka and modified geonames polygons
National Checklists 2017: North America Species List
Lists of taxa for each country and a few other administrative zones harvested from effechecka using simplified versions of geonames polygons. See <p></p>https://github.com/diatomsRcool/checklists for details<p></p>List of species from North America inferred from individual country lists that were derived from effechecka and modified geonames polygons
National Checklists 2017: Asia Species List
Lists of taxa for each country and a few other administrative zones harvested from effechecka using simplified versions of geonames polygons. See <p></p>https://github.com/diatomsRcool/checklists for details<p></p>List of species from Asia inferred from individual country lists that were derived from effechecka and modified geonames polygons
National Checklists 2017: Europe Species List
Lists of taxa for each country and a few other administrative zones harvested from effechecka using simplified versions of geonames polygons. See <p></p>https://github.com/diatomsRcool/checklists for details<p></p>List of species from Europe inferred from individual country lists that were derived from effechecka and modified geonames polygons
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.