Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
188
datasets available to search
ShareScore release 0.9.0
Dataset results
188 results for “Items”
New items activated in the SLUB catalog in 2022
<p>The data set is available as a gzip-compressed, line-delimited JSON file and contains the 18,790 documents with new items activated in the <a href="https://katalog.slub-dresden.de">SLUB catalog</a> in 2022 with the fields id (document identifier), de14_new_item_act_date (timestamp) and de14_new_item_act_date_mv (timestamps). The id consists of a source id and a record id according to the scheme {source_id}-{record_id} and enables the detailed view of a document to be retrieved using the following URL scheme: https://katalog.slub-dresden.de/id/{id}. Example: <a href="https://katalog.slub-dresden.de/id/0-173837243X">https://katalog.slub-dresden.de/id/0-173837243X</a>. Since the documents in this data set originate exclusively from the Southwest German Library Network (SWB), the source id is always 0 and the record id corresponds to the Pica Production Number (PPN) of the K10plus. The values for the fields de14_new_item_act_date and de14_new_item_act_date_mv are based on the NewItemActDate field of the corresponding items in the Libero Library Management System. While the multi-valued field de14_new_item_act_date_mv contains all dates from the year 2022 that could be assigned to a document, only the most recent date is specified in the field de14_new_item_act_date. Both fields are dynamic fields according to VuFind's Solr index schema. Until January 2021, the date values were used to display a list of new acquisitions on the SLUB website.</p>
Newly purchased items in the SLUB catalog in 2022
<p>The data set is available as a gzip-compressed, line-delimited JSON file and contains the 39,013 documents with newly purchased items in the <a href="https://katalog.slub-dresden.de">SLUB catalog</a> in 2022 with the fields id (document identifier), de14_purchase_date (timestamp), de14_purchase_date_mv (timestamps) and facet_de14_acquisition_code (parameters). The id consists of a source id and a record id according to the scheme {source_id}-{record_id} and enables the detailed view of a document to be retrieved using the following URL scheme: https://katalog.slub-dresden.de/id/{id}. Example: <a href="https://katalog.slub-dresden.de/id/0-173837243X">https://katalog.slub-dresden.de/id/0-173837243X</a>. Since the documents in this data set originate exclusively from the Southwest German Library Network (SWB), the source id is always 0 and the record id corresponds to the Pica Production Number (PPN) of the K10plus. The values for the fields de14_purchase_date and de14_purchase_date_mv are based on the DatePurchased field of the corresponding items in the Libero Library Management System. While the multi-valued field de14_purchase_date_mv contains all dates from the year 2022 that could be assigned to a document, only the most recent date is specified in the field de14_purchase_date. Both fields are dynamic fields according to VuFind's Solr index schema. The acquisition parameters specified in the field facet_de14_acquisition_code are based on the AcquisitionType field of the corresponding items in the Libero Library Management System. The multi-valued facet field is a dynamic field according to finc's Solr index schema. A complete list of acquisition types that appear in the data set is available as a CSV file. In addition to the codes, their translations can be found there.</p>
Food items matching
<p>It contains Slovenian food names from receipts linked to items in the Slovenian food composition database (FCDB). They are also annotated with the NAct ontology.<strong> </strong>Since food names on receipts are often abbreviated and vary from those in the FCDB, this dataset is crucial for future research in food, nutrition, data science, and AI. It can be used to develop food recommender systems commonly found in food applications.</p> <p> </p>
A list of items in the FAIRsFAIR training library
<p>A file containing a list of items, with basic metadata, that were included in the <a href="https://www.fairsfair.eu/competence-centre/training-library">FAIRsFAIR training library</a> </p> <p> </p>
Adelie penguin diet composition, secondary prey items, 1991-2020
The fundamental long-term objective of the seabird component of the Palmer LTER (PAL) has been to identify and understand the mechanistic processes that regulate the mean fitness (population growth rate) of regional penguin populations. Since the inception of PAL, Adélie penguin populations have effectively collapsed, gentoo penguin populations have increased dramatically and chinstrap penguin populations have remained relatively stable. These trends are spatially and temporally coherent with regional warming and decreasing sea ice duration. Adélie penguins are an ice-obligate polar species whose life history is intimately linked to the presence of sea ice, while chinstrap and gentoo penguins are ice-intolerant species whose life histories evolved in the sub-Antarctic, where sea ice is a less permanent feature of the marine ecosystem. The PAL study region includes five main islands on which Adélie penguin colonies have historically occurred, with each island containing a different number of spatially segregated sub-colonies. These colonies are censused to determine the total number of nests and chicks produced each year, and breeding success. Diet samples are acquired to understand diet composition (e.g., krill, fish) and krill length-frequencies. In general, krill constitute the most important component of the summer diets by mass of these three penguin species, but changes in PAL krill abundances have exhibited no long-term trends and thus far, have failed to explain the divergent patterns in penguin populations evident in our time series. Chick fledging masses are recorded as a cumulative measure of climate, weather, diet, and parental influences on chick health at the end of the breeding season. These data have provided valuable insights into the marine and terrestrial factors that influence Adélie penguin population fitness. No data were collected during the 2021-2022 season due to the Palmer Station pier rebuild.
Neural Overlap in Item Representations Across Episodes Impairs Context Memory
Open the record for dataset details and reuse information.
A Dataset of Work Items
This is a dataset containing mined work items (commits that logically belong together). See the README.md file for more details.
Taxon item properties
<p>Diagram showing examples of Wikidata properties that can be used on a Wikidata item for a taxon. </p>
Prioritization of semantic over visuo- perceptual aspects in multi-item working memory
<p>All data and code supporting Prioritization of semantic over visuo- perceptual aspects in multi-item working memory</p>
Metadata, Title Pages, and Network Graph of the Digitized Content of the Berlin State Library (146,000 items)
<p>The data set has been downloaded via the OAI-PMH endpoint of the Berlin State Library/Staatsbibliothek zu Berlin’s Digitized Collections (<a href="https://digital.staatsbibliothek-berlin.de/oai">https://digital.staatsbibliothek-berlin.de/oai</a>) on March 1<sup>st</sup> 2019 and converted into common tabular formats on the basis of the provided Dublin Core metadata. It contains 146,000 records.</p> <p>In addition to the bibliographic metadata, representative images of the works have been downloaded, resized to a 512 pixel maximum thumbnail image and saved in JPEG format. The image data is split into title pages and first pages. Title pages have been derived from structural metadata created by scan operators and librarians. If this information was not available, first pages of the media have been downloaded. In case of multi-volume media, title pages are not available.</p> <p>In total, 141,206 images title/first pages are available.</p> <p> </p> <p>Furthermore, the tabular data has been cleaned and extended with geo-spatial coordinates provided by the OpenStreetMap project (<a href="https://www.openstreetmap.org">https://www.openstreetmap.org</a>). The actual data processing steps are summarized in the next section. For the sake of transparency and reproducibility, the original data taken from the OAI-PMH endpoint is still present in the table.</p> <p> </p> <p>To conclude with, various graphs in GML file format are available that can be loaded directly into graph analysis tools such as Gephi (<a href="https://gephi.org/">https://gephi.org/</a>).</p> <p> </p> <p>The implementation of the data processing steps (incl. graph creation) are available as a Jupyter notebook provided at <a href="https://github.com/elektrobohemian/SBBrowse2018/blob/master/DataProcessing.ipynb">https://github.com/elektrobohemian/SBBrowse2018/blob/master/DataProcessing.ipynb</a>.</p> <p> </p> <p>Tabular Metadata</p> <p> </p> <p>The metadata is available in Excel (cleanedData.xlsx) and CSV (cleanedData.csv) file formats with equal content.</p> <p>The table contains the following columns. Italique columns have not been processed.</p> <p>· <em>title</em> The title of the medium</p> <p>· <em>creator</em> Its creator (family name, first name)</p> <p>· <em>subject</em> A collection’s name as provided by the library</p> <p>· <em>type</em> The type of medium</p> <p>· <em>format</em> A MIME type for full metadata download</p> <p>· <em>identifier</em> An additional identifier (most often the PPN)</p> <p>· <em>language</em> A 3-letter language code of the medium</p> <p>· <em>date</em> The date of creation/publication or a time span</p> <p>· <em>relation</em> A relation to a project or collection a medium has been digitized for.</p> <p>· <em>coverage</em> The location of publication or origin (ranging from cities to continents)</p> <p>· <em>publisher</em> The publisher of the medium.</p> <p>· <em>rights</em> Copyright information.</p> <p>· <em>PPN</em> The unique identifier that can be used to find more information about the current medium in all information systems of Berlin State Library/Staatsbibliothek zu Berlin.</p> <p>· spatialClean In case of multiple entries in coverage, only the first place of origin has been extracted. Additionally, characters such as question marks, brackets, or the like have been removed. The entries have been normalized regarding whitespaces and writing variants with the help of regular expressions.</p> <p>· dateClean As the original date may contain various format variants to indicate unclear creation dates (e.g., time spans or question marks), this field contains a mapping to a certain point in time.</p> <p>· spatialCluster The cluster ID determined with the help of the Jaro-Winkler distance on the spatialClean string. This step is needed because the spatialClean fields still contain a huge amount of orthographic variants and latinizations of geographic names.</p> <p>· spatialClusterName A verbal cluster name (controlled manually).</p> <p>· latitude The latitude provided by OpenStreetMap of the spatialClusterName if the location could be found.</p> <p>· longitude The longitude provided by OpenStreetMap of the spatialClusterName if the location could be found.</p> <p>· century A century derived from the date.</p> <p>· textCluster A text cluster ID on the basis of a k-means clustering relying on the title field with a vocabulary size of 125,000 using the tf*idf model and k=5,000.</p> <p>· creatorCluster A text cluster ID based on the creator field with k=20,000.</p> <p>· titleImage The path to the first/title page relative to the img/ subdirectory or None in case of a multi-volume work.</p> <p>Other Data</p> <p> </p> <p><em>graphs.zip</em></p> <p> </p> <p>Various pre-computed graphs.</p> <p><em> </em></p> <p><em>img.zip</em></p> <p> </p> <p>First and title pages in JPEG format.</p> <p> </p> <p><em>json.zip</em></p> <p> </p> <p>JSON files for each record in the following format:</p> <p> </p> <p>ppn "PPN57346250X"</p> <p>dateClean "1625"</p> <p>title "M. Georgii Gutkii, Gymnasii Berlinensis Rectoris Habitus Primorum Principiorum, Seu Intelligentia; Annexae Sunt Appendicis loco Disputationes super eodem habitu tum in Academia Wittebergensi, tum in Gymnasio Berlinensi ventilatae"</p> <p>creator "Gutke, Georg"</p> <p>spatialClusterName "Berlin"</p> <p>spatialClean "Berolini"</p> <p>spatialRaw "Berolini"</p> <p>mediatype "monograph"</p> <p>subject "Historische Drucke"</p> <p>publisher "Kallius"</p> <p>lat "52.5170365"</p> <p>lng "13.3888599"</p> <p>textCluster "45"</p> <p>creatorCluster "5040"</p> <p>titleImage "titlepages/PPN57346250X.jpg"</p>
HELP Study Data Dictionary / Catalog of Items
<p>The <a href="https://doi.org/10.1136/bmjopen-2019-033391">HELP study</a> was a clinical trial in the form of a multicenter interventional randomized controlled trial (RCT) at five German university hospitals, conducted from 2020 to 2022. It aimed to enhance the clinical management of Staphylococcus bacteremia and used data both from Electronic Health Records (EHR), provided by German university hospital's Data Integration Centers in the HL7 FHIR format (German profiles of the Medical Informatics Initiative Core Data Set – <a href="https://www.medizininformatik-initiative.de/en/medical-informatics-initiatives-core-data-set">MII CDS</a>), as well as data from Electronic Case Report Forms (eCRF) used in the study for data which was too unstructered or not availabe in the EHR (hybrid data collection approach).<br>This dataset is a tabular listing and description of the data items used in the study and <em>serves (primarily) as a template for information and data modeling.</em></p>
The International Transport Energy Modeling (iTEM) Open Data & Harmonized Transport Database
<p>This dataset and documentation contains detailed information of the iTEM Open Database, a harmonized transport data set of historical values, 1970 - present. It aims to create transparency through two key features:</p> <ul> <li>Open-Data: Assembling a comprehensive collection of publicly-available transportation data</li> <li>Open-Code: All code and documentation will be publicly accessible and open for modification and extension. <a href="https://github.com/transportenergy">https://github.com/transportenergy</a></li> </ul> <p>The iTEM Open Database is comprised of individual datasets collected from public sources. Each dataset is downloaded, cleaned, and harmonised to the common region and technology definitions defined by the iTEM consortium https://transportenergy.org. For each dataset, we describe the name of the dataset, the web link to the original source, the web link to the cleaning script (in python), variables, and explain the data cleaning steps (which explains the data cleaning script in plain English).</p> <p>Shall you find any problems with the dataset, please report the issues here <a href="https://github.com/transportenergy/database/issues">https://github.com/transportenergy/database/issues</a>. </p> <p> </p>
Wikidata Taxon Items in JSON Lines Format hash://sha256/13ffa9679bae381aa5914d810638fb5a0c75d71f5f7d47f38b3c00d750c88b9c hash://md5/bdcc99bfedfd34abdfdd3802182f225c
<p>Wikidata contains information about taxonomic names, and these taxonomic names are key to integrating biodiversity datasets across different platforms, datasets and institutions. </p> <h2>Content</h2> <table> <tbody> <tr> <td><strong>filename/alias</strong></td> <td><strong>content ids</strong></td> </tr> <tr> <td>wikidata-taxon.json.bz2</td> <td> <p><a href="https://linker.bio/hash://sha256/a7592b72c9013d67d655b6ea5d1f4f67f2057dc5b5ee52578a07f58fea835580">hash://sha256/</a><a href="https://zenodo.org/api/records/13920038/draft/files/701a1382e304a6b1bb38fe828d82f7b8b562c77f918f33097966e38bacf0b2e7/content" target="_blank" rel="noopener noreferrer">701a1382e304a6b1bb38fe828d82f7b8b562c77f918f33097966e38bacf0b2e7</a></p> <p><a href="https://linker.bio/hash://md5/d5bad3553470506f3bde383566a5dea3">hash://md5/d5bad3553470506f3bde383566a5dea3</a></p> </td> </tr> <tr> <td>wikidata-taxa.sh</td> <td> <p><a href="https://linker.bio/hash://sha256/6f4fac44054d54ec3006d091ba702f872b3f4d013628add98fcca08a3b768962">hash://sha256/6f4fac44054d54ec3006d091ba702f872b3f4d013628add98fcca08a3b768962</a></p> <p><a href="https://linker.bio/hash://md5/1d80083d498d61c8b63fbd46d51f7c5c">hash://md5/1d80083d498d61c8b63fbd46d51f7c5c</a></p> </td> </tr> <tr> <td>Q140.json (example)</td> <td> <p><a href="https://linker.bio/hash://md5/44ab0031091fb96caa063e3fe41a85f2">hash://md5/44ab0031091fb96caa063e3fe41a85f2</a></p> </td> </tr> </tbody> </table> <h2>Provenance</h2> <h3>for humans</h3> <p>This dataset contains a subset of Wikidata items referencing the taxonomic name concept https://www.wikidata.org/wiki/Q16521 and is expressed in JSON Lines format. </p> <p>An example of such item is https://wikidata.org/wiki/Q140, an item that describes the taxonomic name associated with <em>Panthera leo</em>, commonly known as Lion (English), León (Spanish), or 狮子 (Chinese). You can find a "pretty" printed example of Q140 in the file "Q140.json" included in this publication. The first 10 lines of "Q140.json" are shown below: </p> <pre><code>{ "type": "item", "id": "Q140", "labels": { "fr": { "language": "fr", "value": "lion" }, "it": { "language": "it", ...</code></pre> <p> </p> <p>The reason for creating a wikidata subset is because all of wikidata (~85G) didn't fit in Zenodo. </p> <h3>for machines</h3> <p>This dataset was generated using the script below with content id <a href="https://linker.bio/hash://sha256/6f4fac44054d54ec3006d091ba702f872b3f4d013628add98fcca08a3b768962">hash://sha256/6f4fac44054d54ec3006d091ba702f872b3f4d013628add98fcca08a3b768962</a> or <a href="https://linker.bio/hash://md5/1d80083d498d61c8b63fbd46d51f7c5c">hash://md5/1d80083d498d61c8b63fbd46d51f7c5c</a></p> <pre><code> 1 #!/bin/bash 2 # 3 # streams Wikidata taxon items (or items containing https://www.wikidata.org/wiki/Q16521) 4 # from latest data dump in line json (one json object per line) 5 # 6 curl --silent "https://dumps.wikimedia.org/wikidatawiki/entities/latest-all.json.bz2"\ 7 | bunzip2\ 8 | grep -E "Q16521[^0-9]"\ 9 | sed 's/,$//g'\ 10 | bzip2 </code></pre> <p>The script first downloads a recent copy of all wikidata entities in bzip2 compressed format (line 6), decompresses them (line 7), selects only lines containing "Q16521" (line 8), removes any trailing commas (line 9), and recompresses the output. With this, the output contains wikidata items/entities as described earlier.</p> <p>Preston, a biodiversity data tracker, was used to (a) track the script, as well as (b) recording a script execution and (c) tracking the outcome by running :</p> <pre><code>#!/bin/bash # # run the script with id hash://sha256/13ff... # preston bash\ --remote https://linker.bio\ -c "hash://sha256/13ffa9679bae381aa5914d810638fb5a0c75d71f5f7d47f38b3c00d750c88b9c" </code></pre> <p>The recording of this process is identified with hash://sha256/13ffa9679bae381aa5914d810638fb5a0c75d71f5f7d47f38b3c00d750c88b9c and hash://md5/bdcc99bfedfd34abdfdd3802182f225c , and can be reconstructed using </p> <pre><code>preston ls\ --remote https://linker.bio/,https://zenodo.org/records/13920038/files\ --anchor hash://sha256/13ffa9679bae381aa5914d810638fb5a0c75d71f5f7d47f38b3c00d750c88b9c</code></pre> <p>Which is expected to produce:</p> <pre><code><https://preston.guoda.bio> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://www.w3.org/ns/prov#SoftwareAgent> <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><https://preston.guoda.bio> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://www.w3.org/ns/prov#Agent> <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><https://preston.guoda.bio> <http://purl.org/dc/terms/description> "Preston is a software program that finds, archives and provides access to biodiversity datasets."@en <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://www.w3.org/ns/prov#Activity> <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> <http://purl.org/dc/terms/description> "Executes script and captures stdout"@en <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> <http://www.w3.org/ns/prov#startedAtTime> "2024-10-10T16:37:36.659Z"^^<http://www.w3.org/2001/XMLSchema#dateTime> <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> <http://www.w3.org/ns/prov#wasStartedBy> <https://preston.guoda.bio> <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><https://doi.org/10.5281/zenodo.1410543> <http://www.w3.org/ns/prov#usedBy> <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><https://doi.org/10.5281/zenodo.1410543> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://purl.org/dc/dcmitype/Software> <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><https://doi.org/10.5281/zenodo.1410543> <http://purl.org/dc/terms/bibliographicCitation> "Jorrit Poelen, Icaro Alzuru, & Michael Elliott. 2018-2024. Preston: a biodiversity dataset tracker (Version 0.9.9-SNAPSHOT) [Software]. Zenodo. https://doi.org/10.5281/zenodo.1410543"@en <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><urn:uuid:0659a54f-b713-4f86-a917-5be166a14110> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://www.w3.org/ns/prov#Entity> <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><urn:uuid:0659a54f-b713-4f86-a917-5be166a14110> <http://purl.org/dc/terms/description> "A biodiversity dataset graph archive."@en <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><hash://sha256/e76276c283090381fc4b3efe28fc61c28f5bf03db0f3743f7178b999ebccada2> <http://www.w3.org/ns/prov#usedBy> <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><hash://sha256/6f4fac44054d54ec3006d091ba702f872b3f4d013628add98fcca08a3b768962> <http://purl.org/dc/elements/1.1/format> "text/x-shellscript" .<br><urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> <http://www.w3.org/ns/prov#used> <hash://sha256/6f4fac44054d54ec3006d091ba702f872b3f4d013628add98fcca08a3b768962> .<br><urn:uuid:7ecf1c84-0438-4224-8909-0804028cf3f6> <http://www.w3.org/ns/prov#wasGeneratedBy> <urn:uuid:fdc316b0-457d-4d22-85a1-d2ce65c2e440> .<br><hash://sha256/701a1382e304a6b1bb38fe828d82f7b8b562c77f918f33097966e38bacf0b2e7> <http://www.w3.org/ns/prov#wasGeneratedBy> <urn:uuid:bb57ae4b-1ba1-4188-8e95-3f3f6cdcab6b> <urn:uuid:bb57ae4b-1ba1-4188-8e95-3f3f6cdcab6b> .<br><hash://sha256/701a1382e304a6b1bb38fe828d82f7b8b562c77f918f33097966e38bacf0b2e7> <http://www.w3.org/ns/prov#qualifiedGeneration> <urn:uuid:bb57ae4b-1ba1-4188-8e95-3f3f6cdcab6b> <urn:uuid:bb57ae4b-1ba1-4188-8e95-3f3f6cdcab6b> .<br><urn:uuid:bb57ae4b-1ba1-4188-8e95-3f3f6cdcab6b> <http://www.w3.org/ns/prov#generatedAtTime> "2024-10-11T02:08:03.739Z"^^<http://www.w3.org/2001/XMLSchema#dateTime> <urn:uuid:bb57ae4b-1ba1-4188-8e95-3f3f6cdcab6b> .<br><urn:uuid:bb57ae4b-1ba1-4188-8e95-3f3f6cdcab6b> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://www.w3.org/ns/prov#Generation> <urn:uuid:bb57ae4b-1ba1-4188-8e95-3f3f6cdcab6b> .<br><urn:uuid:bb57ae4b-1ba1-4188-8e95-3f3f6cdcab6b> <http://www.w3.org/ns/prov#used> <urn:uuid:7ecf1c84-0438-4224-8909-0804028cf3f6> <urn:uuid:bb57ae4b-1ba1-4188-8e95-3f3f6cdcab6b> .<br><urn:uuid:7ecf1c84-0438-4224-8909-0804028cf3f6> <http://purl.org/pav/hasVersion> <hash://sha256/701a1382e304a6b1bb38fe828d82f7b8b562c77f918f33097966e38bacf0b2e7> <urn:uuid:bb57ae4b-1ba1-4188-8e95-3f3f6cdcab6b> .<br><https://preston.guoda.bio> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://www.w3.org/ns/prov#SoftwareAgent> <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> . <https://preston.guoda.bio> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://www.w3.org/ns/prov#Agent> <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> . <https://preston.guoda.bio> <http://purl.org/dc/terms/description> "Preston is a software program that finds, archives and provides access to biodiversity datasets."@en <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> . <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://www.w3.org/ns/prov#Activity> <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> . <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> <http://purl.org/dc/terms/description> "Executes script and captures stdout"@en <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> . <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> <http://www.w3.org/ns/prov#startedAtTime> "2024-06-22T10:40:12.016Z"^^<http://www.w3.org/2001/XMLSchema#dateTime> <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> . <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> <http://www.w3.org/ns/prov#wasStartedBy> <https://preston.guoda.bio> <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> . <https://doi.org/10.5281/zenodo.1410543> <http://www.w3.org/ns/prov#usedBy> <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> . <https://doi.org/10.5281/zenodo.1410543> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://purl.org/dc/dcmitype/Software> <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> . <https://doi.org/10.5281/zenodo.1410543> <http://purl.org/dc/terms/bibliographicCitation> "Jorrit Poelen, Icaro Alzuru, & Michael Elliott. 2021. Preston: a biodiversity dataset tracker (Version 0.8.4) [Software]. Zenodo. https://doi.org/10.5281/zenodo.1410543"@en <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> . <urn:uuid:0659a54f-b713-4f86-a917-5be166a14110> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://www.w3.org/ns/prov#Entity> <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> . <urn:uuid:0659a54f-b713-4f86-a917-5be166a14110> <http://purl.org/dc/terms/description> "A biodiversity dataset graph archive."@en <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> . <hash://sha256/6f4fac44054d54ec3006d091ba702f872b3f4d013628add98fcca08a3b768962> <http://purl.org/dc/elements/1.1/format> "text/x-shellscript" . <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> <http://www.w3.org/ns/prov#used> <hash://sha256/6f4fac44054d54ec3006d091ba702f872b3f4d013628add98fcca08a3b768962> . <urn:uuid:6fa51a99-a137-4387-90b1-23589d7b60ae> <http://www.w3.org/ns/prov#wasGeneratedBy> <urn:uuid:096ba92f-9d5c-4cb1-9a3d-7a95c5228758> . <hash://sha256/a7592b72c9013d67d655b6ea5d1f4f67f2057dc5b5ee52578a07f58fea835580> <http://www.w3.org/ns/prov#wasGeneratedBy> <urn:uuid:e0d76a06-5241-4dfe-8429-46164190ab0e> <urn:uuid:e0d76a06-5241-4dfe-8429-46164190ab0e> . <hash://sha256/a7592b72c9013d67d655b6ea5d1f4f67f2057dc5b5ee52578a07f58fea835580> <http://www.w3.org/ns/prov#qualifiedGeneration> <urn:uuid:e0d76a06-5241-4dfe-8429-46164190ab0e> <urn:uuid:e0d76a06-5241-4dfe-8429-46164190ab0e> . <urn:uuid:e0d76a06-5241-4dfe-8429-46164190ab0e> <http://www.w3.org/ns/prov#generatedAtTime> "2024-06-22T19:01:55.863Z"^^<http://www.w3.org/2001/XMLSchema#dateTime> <urn:uuid:e0d76a06-5241-4dfe-8429-46164190ab0e> . <urn:uuid:e0d76a06-5241-4dfe-8429-46164190ab0e> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://www.w3.org/ns/prov#Generation> <urn:uuid:e0d76a06-5241-4dfe-8429-46164190ab0e> . <urn:uuid:e0d76a06-5241-4dfe-8429-46164190ab0e> <http://www.w3.org/ns/prov#used> <urn:uuid:6fa51a99-a137-4387-90b1-23589d7b60ae> <urn:uuid:e0d76a06-5241-4dfe-8429-46164190ab0e> . <urn:uuid:6fa51a99-a137-4387-90b1-23589d7b60ae> <http://purl.org/pav/hasVersion> <hash://sha256/a7592b72c9013d67d655b6ea5d1f4f67f2057dc5b5ee52578a07f58fea835580> <urn:uuid:e0d76a06-5241-4dfe-8429-46164190ab0e> . </code></pre>
Nicobarese 100 item wordlist for phylogenetic analyses
<p>The data set is based on a modified Swadesh 100 list, intended to provide indications of the internal branching of the Nicobarese languages. To date little work has been done on the classification of the small Nicobarese group, which appears to consisted of approximately seven distinct languages spoken across an island chain. Only two of the languages are have extensive dictionaries and grammatical descriptions, while the others are only partially documented, and the materials can be highly problematic to work with. The excel includes the author's nexus file used for input to phylogentic software, such as SplitsTree.<br> The data supports the author's paper for the 9th ICAAL meeting, Novemer 2021, Lund, Sweden and subsequent published versions.</p>
Datasets from the RecSys 2023 article "Ex2Vec: Characterizing Users and Items from the Mere Exposure Effect".
<p>We have publicly released the anonymized "new_release_stream.csv" dataset from the music streaming platform Deezer. This dataset is described in detail in the article titled "Ex2Vec: Characterizing Users and Items from the Mere Exposure Effect", which was published in the proceedings of the 17th ACM Conference on Recommender Systems (RecSys 2023).</p> <p>Each row in the dataset contains an anonymized user and item identifier, a reference timestamp in seconds (measured from the first consumption in the dataset), and a binary value "y". This "y" value indicates whether a song was listened to for more than 80% of its duration (y = 1) or not (y = 0).</p> <p>You can find this dataset in the GitHub repository <a href="https://github.com/deezer/ex2vec">deezer/ex2vec</a>, where it is used to reproduce experiments discussed in the article.</p> <p>If you plan to use our code or data in your work, please make sure to cite our paper accordingly.</p> <p> </p> <pre><code>@inproceedings{sguerra2023ex2vec, title={Ex2Vec: Characterizing Users and Items from the Mere Exposure Effect}, author={Sguerra, Bruno and Tran, Viet-Anh and Hennequin, Romain}, booktitle = {Proceedings of the 17th ACM Conference on Recommender Systems}, year = {2023} }</code></pre> <p> </p> <p> </p> <p> </p>
PRISMA-P (Preferred Reporting Items for Systematic Review and Meta-Analysis Protocols) of the research entitled "Development of Competences for the Fashion Designer: a Scope Review
<p>PRISMA-P (Preferred Reporting Items for Systematic review and Meta-Analysis Protocols) 2015 checklist: recommended items to address in a systematic review protocol and Check list CAPSI - Critical analysis of the articles related to the specific objective: map the current themes that permeate the competencies of fashion design professionals through a scoping review.</p>
Table 4: The means of the items comprising the test anxiety questionnaire
<p>The purpose of the study was twin: to investigate the test-taking anxiety of ESP students of<br> Engineering taking a course in general English and to shed light on the relationship between the<br> students' test-taking anxiety and their performance on a general English test. To this end, the first<br> phase of the study was devoted to the reliability of the two instruments employed to address the<br> research question: (a) the anxiety questionnaire (TAS), and (b) the general English test. The second<br> or main phase of the study was concerned with the four research question. In this section, the results<br> of analyses related to the two phases are presented.</p> <p>Table 4 also shows the mean of each of 37 items comprising the TAS.</p>
Figure 3 in Seasonal analysis of food items and feeding habits of endangered riverine catfish Rita rita (Hamilton, 1822)
Figure 3. Seasonal variation in frequency of food items assessed by non-metric multidimensional scaling (nMDS) analysis in R. rita sampled from Padma River.
Figure 2 in Seasonal analysis of food items and feeding habits of endangered riverine catfish Rita rita (Hamilton, 1822)
Figure 2. Fullness index of fish stomach in different seasons (a) and size groups (b) of R. rita sampled from Padma River.
Figure 6 in Seasonal analysis of food items and feeding habits of endangered riverine catfish Rita rita (Hamilton, 1822)
Figure 6. Canonical correspondence analysis of food items and morphometric measures of R. rita sampled from Padma river (TL = Total length; BW = Body weight; HG = horizontal mouth gape; VG = vertical mouth gape; MA = mouth area)
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.