Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
243
datasets available to search
ShareScore release 0.9.0
Dataset results
243 results for “categories”
A visual summary of the three categories of Brutus's drumming communication
<p>[This figure appears in DOI: 10.31234/osf.io/kgqy9 (Version 3). The figure was not included in the published version of the manuscript DOI: 10.1007/s10071-021-01584-3.].</p>
Dataset, statistical analysis code, and supplementary material of juvenile ravens' responses towards acoustic cues of different social categories
<p>Social competence i.e., defined as the ability to adjust the expression of social behaviour to the available social information, is known to be influenced by early-life conditions. Brood size might be one of the factors determining such early conditions, particularly in species with extended parental care. We here tested in ravens, whether growing up in families of different sizes affects the chicks' responsiveness to social information. We experimentally manipulated the brood size of 20 captive raven families, creating either small or large families. Simulating dispersal, juveniles were separated from their parents and temporarily housed in one of two captive non-breeder groups. After five weeks of socialization, each raven was individually tested in a playback setting with food-associated calls from three social categories: sibling, familiar unrelated raven they were housed with, and unfamiliar unrelated raven from the other non-breeder aviary. We found that individuals reared in small families were more attentive than birds from large families, in particular towards the familiar unrelated peer. These results indicate that variation in family size during upbringing can affect how juvenile ravens value social information. Whether the observed attention patterns translate into behavioural preferences under daily life conditions remains to be tested in future studies.</p>
Model output data for Smith et al., "Effects of increasing the category resolution of the sea ice thickness distribution in a coupled climate model on Arctic and Antarctic sea ice"
<p>Model output data for Smith et al., "Effects of increasing the category resolution of the sea ice thickness distribution in a coupled climate model on Arctic and Antarctic sea ice", in review in Journal of Geophysical Research-Oceans, 2022. Details on CESM model settings and run setups can be found within the manuscript. </p>
Wikipedia Category Granularity (WikiGrain) data
<p>The "Wikipedia Category Granularity (WikiGrain)" data consists of three files that contain information about articles of the English-language version of Wikipedia (https://en.wikipedia.org).</p> <p>The data has been generated from the database dump dated 20 October 2016 provided by the Wikimedia foundation licensed under the GNU Free Documentation License (GFDL) and the Creative Commons Attribution-Share-Alike 3.0 License. </p> <p>WikiGrain provides information on all 5,006,601 Wikipedia articles (that is, pages in Namespace 0 that are not redirects) that are assigned to at least one category.</p> <p>The WikiGrain Data is analyzed in the paper</p> <p>Jürgen Lerner and Alessandro Lomi: <a href="http://journals.plos.org/plosone/article?id=10.1371/journal.pone.0190674"><strong>Knowledge categorization affects popularity and quality of Wikipedia articles</strong></a>. <em>PLoS ONE</em>, 13(1):e0190674, 2018.</p> <p>===============================================================<br> Individual files (tables in comma-separated-values-format):</p> <p>---------------------------------------------------------------<br> * article_info.csv contains the following variables:</p> <p>- "id"<br> (integer) Unique identifier for articles; identical with the page_id in the Wikipedia database.</p> <p>- "granularity"<br> (decimal) The granularity of an article A is defined to be the average (mean) granularity of the categories of A, where the granularity of a category C is the shortest path distance in the parent-child subcategory network from the root category (Category:Articles) to C. Higher granularity values indicate articles whose topics are less general, narrower, more specific. </p> <p>- "is.FA"<br> (boolean) True ('1') if the article is a featured article; false ('0') else.</p> <p>- "is.FA.or.GA"<br> (boolean) True ('1') if the article is a featured article or a good article; false ('0') else.</p> <p>- "is.top.importance"<br> (boolean) True ('1') if the article is listed as a top importance article by at least one WikiProject; false ('0') else.</p> <p>- "number.of.revisions"<br> (integer) Number of times a new version of the article has been uploaded.</p> <p><br> ---------------------------------------------------------------<br> * article_to_tlc.csv<br> is a list of links from articles to the closest top-level categories (TLC) they are contained in. We say that an article A is a member of a TLC C if A is in a category that is a descendant of C and the distance from C to A (measured by the number of parent-child category links) is minimal over all TLC. An article can thus be member of several TLC.<br> The file contains the following variables:</p> <p>- "id"<br> (integer) Unique identifier for articles; identical with the page_id in the Wikipedia database.</p> <p>- "id.of.tlc"<br> (integer) Unique identifier for TLC in which the article is contained; identical with the page_id in the Wikipedia database. </p> <p>- "title.of.tlc"<br> (string) Title of the TLC in which the article is contained.</p> <p>---------------------------------------------------------------<br> * article_info_normalized.csv<br> contains more variables associated with articles than article_info.csv. All variables, except "id" and "is.FA" are normalized to standard deviation equal to one. Variables whose name has prefix "log1p." have been transformed by the mapping x --> log(1+x) to make distributions that are skewed to the right 'more normal'. <br> The file contains the following variables:</p> <p>- "id"<br> Article id.</p> <p>- "is.FA"<br> Boolean indicator for whether the article is featured.</p> <p>- "log1p.length"<br> Length measured by the number of bytes.</p> <p>- "age"<br> Age measured by the time since the first edit.</p> <p>- "log1p.number.of.edits"<br> Number of times a new version of the article has been uploaded.</p> <p>- "log1p.number.of.reverts"<br> Number of times a revision has been reverted to a previous one.</p> <p>- "log1p.number.of.contributors"<br> Number of unique contributors to the article.</p> <p>- "number.of.characters.per.word"<br> Average number of characters per word (one component of 'reading complexity').</p> <p>- "number.of.words.per.sentence"<br> Average number of words per sentence (second component of 'reading complexity').</p> <p>- "number.of.level.1.sections"<br> Number of first level sections in the article.</p> <p>- "number.of.level.2.sections"<br> Number of second level sections in the article.</p> <p>- "number.of.categories"<br> Number of categories the article is in.</p> <p>- "log1p.average.size.of.categories"<br> Average size of the categories the article is in.</p> <p>- "log1p.number.of.intra.wiki.links"<br> Number of links to pages in the English-language version of Wikipedia.</p> <p>- "log1p.number.of.external.references"<br> Number of external references given in the article.</p> <p>- "log1p.number.of.images"<br> Number of images in the article.</p> <p>- "log1p.number.of.templates"<br> Number of templates that the article uses.</p> <p>- "log1p.number.of.inter.language.links"<br> Number of links to articles in different language edition of Wikipedia.</p> <p>- "granularity"<br> As in article_info.csv (but normalized to standard deviation one).<br> </p>
SSP1 and SSP3 population projections with demographic categories
<p>Demographic projections for the Shared Socioeconomic Pathways SSP1 and SSP3 scenarios with demographic breakdown for "young", "adult" and "old" populations, defined as <15 years old, 15-65 years old, and >65 years old, at 10 year intervals. Projections are calculated combining SSP population projections, UN WPP country wide demographic trends, and SEDAC demographic spatial distributions.</p> <p> </p> <p>The "interp" version of the files includes linear interpolationed values for each grid cell for every year.</p>
Scripts and analysis files for categorization of PKZILLA matching proteomic peptides into protein-unique, protein-multimatch & exon-unique, exon-multimatch categories.
<p>A .zip file containing the source data files & Jupyter notebook for analysis of the <em>Prymnesium parvum</em> 12B1 PKZILLA-detecting proteomic results (<a href="https://doi.org/10.5281/zenodo.10023441">https://doi.org/10.5281/zenodo.10023441</a>), and the resulting files from the workflow. See "Analysis of proteomic results" section of the manuscript Materials and Methods for further detail. </p> <p><strong>Key files:</strong></p> <ul> <li>'PKZILLA-1_classify_peptides.txt' - A plaintext report of the # of classified peptides for PKZILLA-1</li> <li>'PKZILLA-2_classify_peptides.txt' - A plaintext report of the # of classified peptides for PKZILLA-2</li> <li>'./hierarchical_classified_xlsx/' - Excel spreadsheets with the classified peptides for PKZILLA-1 and PKZILLA-2</li> <li>'./Process_into_polypeptide_coordinates/' - Workflow, results, and plots for back-alignment of peptides back to PKZILLA-1 and PKZILLA-2 genomic loci</li> </ul>
Fig. 5 in Evaluating categories of resistance in soybean genotypes from the United States and Brazil to Aphis glycines (Hemiptera: Aphididae)
Fig. 5. Mortality (%) of Aphis glycines on 7 soybean genotypes at 5, 7, and 10 d afer infestation (23 ± 3 °C; 60 ± 10% RH; 16:8 h L:D photoperiod).
Fig. 4 in Evaluating categories of resistance in soybean genotypes from the United States and Brazil to Aphis glycines (Hemiptera: Aphididae)
Fig. 4. Cumulative aphid-days (CAD) for soybean genotypes infested with Aphis glycines at V1 and V3 stages (23 ± 3 °C; 60 ± 10% RH; 16:8 h L:D photoperiod).
Fig. 3 in Evaluating categories of resistance in soybean genotypes from the United States and Brazil to Aphis glycines (Hemiptera: Aphididae)
Fig. 3. Number (mean ± SE) of Aphis glycines individuals on 7 soybean genotypes 24 h afer infestation (23 ± 3 °C; 60 ± 10% RH; 16:8 h L:D photoperiod). Means with the same lower case letter do not differ by Fisher's LSD test (P> 0.05). (F = 1.74; df = 6; P = 0.0110).
Fig. 2 in Evaluating categories of resistance in soybean genotypes from the United States and Brazil to Aphis glycines (Hemiptera: Aphididae)
Fig. 2. Number (mean ± SE) of Aphis glycines individuals on KS4202 plants 24 h afer infestation (23 ± 3 °C; 60 ± 10% RH; 16:8 h L:D photoperiod). Means with the same lower case letter do not differ by Fisher's LSD test (P> 0.05). (F = 1.09; df = 6; P = 0.3897).
Embedded HuffPost New Category Dataset
<p>This dataset was created by embedding the concatenation of title and short description of each entry in the <a href="https://arxiv.org/pdf/2209.11429">HuffPost news category</a> dataset, ordered by timestamp, using OpenAI's t<a href="https://openai.com/index/new-embedding-models-and-api-updates/">ext-embedding-3-small embedding</a>. Usage is subject to Huffington Post's <a href="https://www.huffpost.com/static/user-agreement">user agreement</a> and OpenAI's <a href="https://openai.com/policies/row-terms-of-use/">term of use</a>.</p>
BRAIN Journal-An Enhancement over Texture Feature Based Multiclass Image Classification under Unknown Noise-Figure 1. Sample Images of sixteen categories
<p>An image is often corrupted by noise in its acquisition or transmission. Noise is any<br> undesired information that degrades the image and appears in images from a variety of sources.</p> <p>Basically, there are three standard noise models [17], which model the types of noise<br> encountered in most images; they are additive noise, multiplicative noise and impulse noise. In this<br> work we have considered the occurrence of additive noise. An image function is given by f (x, y)<br> where (x, y) is spatial coordinate and f is intensity at point(x, y). Let f (x, y) be the original image,<br> g(x, y) be the noisy version and η(x, y) be the noise function, which returns random values coming<br> from an arbitrary distribution.</p>
BRAIN Journal-Pros and Cons Gamification and Gaming in Classroom-Figure 2. Course categories within UVAB University Moodle platform, (portal-eifr.ub.ro)
<p>A study was conducted aiming to assess the impact of introduction of ranking block plugin as a gamification element within Moodle learning management system (Ranking block Moodle, 2017). We mention that the Moodle platform, version 3.2 is dedicated to extramural and distance learning. It supports various gamification elements such as avatars, badges, leaderboard, levels, displaying quiz results or progress bars. The ranking block plugin was introduced and configured to be available for procedural programming course activities, at the beginning of the first semester of 2016, which starts in October. It displays a course leaderboard visible to all users as a way of obtaining recognition from other users. It is based on points instead of badges and it can monitor included activities based on accumulated points. The experiment involved first year bachelor students in computer science (32 students, extramural education) from UVAB University (www.ub.ro) who are using the Moodle platform in their tutorial based activities. The main page of UVAB Moodle platform is presented in Figure 2. </p>
Categories of Datasets Offered by Current Open Data Portals
<p>This dataset lists categories of open datasets from 40 European open data catalogs. The 40 European data catalogues were taken from four countries (France, Germany, Spain, and the United Kingdom) at a rate of 10 per country. These categories are an indicator of the topics of interests in the respective countries (back in 2016), the types of questions data publishers assume users will ask, and ultimately the types of questions citizens can ask.</p>
COG_Functional_Category_Abundances_and_GTDB_Taxonomy
<p><strong>Dataset S1:</strong></p> <p><strong>Individual rows correspond to individual genomes (excluding the top row which are column headers). Columns 1 through 25 correspond to raw abundances for each COG functional category. Column 26 corresponds to the total number of COGs in a genome. Columns 27, 28, 29, 30, 31, 32, and 33, correspond to the GTDB domain, phylum, class, order, family, genus, and species classification, respectively. Column 34 corresponds to the culture-status. Column 35 is the genomes size in base pairs. Column 36 corresponds to the accession number for each genome. Accessions starting with GCF and GCA are from Refseq and Genbank, respectively. Accessions that are numbers only correspond to IMG/G. Column 37 corresponds to the total number of open reading frames in the genome.</strong></p>
Text-fig. 1. Bottom view of the spoke bone of the excavated greyhound-like dog (above) and the spoke bone of a dog of the same size category (below). The gracility of the spoke bone of the greyhound-like dog is clearly visible. Photo by D. Nývlt. in Genetic Analysis Of Possibly The Oldest Greyhound Remains Within The Territory Of The Czech Republic As Proof Of A Local Elite Presence At Chotěbuz-Podobora Hillfort In The 8 -9 Century Ad
Text-fig. 1. Bottom view of the spoke bone of the excavated greyhound-like dog (above) and the spoke bone of a dog of the same size category (below). The gracility of the spoke bone of the greyhound-like dog is clearly visible. Photo by D. Nývlt.
Use and sharing of raw data in the Journal Citation Reports' Emergency Medicine Category: Metrics and Journals including supplementary material classification sorted by quartile of the JCR emergency medicine category.
<p>Raw data belonged to the study of use and sharing of raw research data in the Journal Citation Reports' Emergency Medicine Category.</p>
Text-fig. 51. Number of specimens and number of species for the five categories of angiosperms distinguished from the Catefica mesofossil flora. in The Early Cretaceous Mesofossil Flora Of Catefica, Portugal: Angiosperms
Text-fig. 51. Number of specimens and number of species for the five categories of angiosperms distinguished from the Catefica mesofossil flora.
Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore (Main Study Data)
<p>Datasets from Main study of <strong>Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore</strong> (project linked here: https://doi.org/10.17605/OSF.IO/S6VDN).</p>
Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore (Pilot Study Data)
<p>Datasets from Pilot study of <strong>Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore</strong> (project linked here: https://doi.org/10.17605/OSF.IO/S6VDN).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.