Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25
datasets available to search
ShareScore release 0.9.0
Dataset results
25 results for “denoting”
Herbarium specimen image of Marsupella sullivantii (Denot) Evans, part of the collection of Royal Botanic Garden Edinburgh
Part of a training dataset of scanned herbarium specimens. The data paper and a summary landing page will be published on Zenodo as it gets published.<br><br>Content of this deposition:<br><br>- A JSON-LD datafile listing the label data associated with this herbarium specimen. The Darwin and Dublin Core data standards are used for most values.<br>- A JPEG image file of the scanned herbarium sheet.<br>- A lossless TIFF image from which the JPEG image has been derived.
nal; E', F', internal, G'-J', Pal. 1341 (peripheral 8); G', H', external; I', J', internal; K'-N', Pal. 1343 (peripheral 9); K', L', external; M', N', internal; O'-R', Pal. 1345 (peripheral 10); O', P', external; Q', R', internal; S'-V', Pal. 1348 (peripheral 11); S', T', external; U', V', internal views; W', reconstruction of carapace Thick lines correspond to scute sulci, dotted lines denote plate sutures, and oblique lines indicate missing plate portions. Abbreviations: Ce, cervical; co, costal; Ma, marginal; ne, neural; nu, nuchal; per, peripheral; Pl, pleural; py, pygal; sp, suprapygal; Spr, supracaudal; Ve, vertebral. Scale bars: A-V', 1 cm; W', 2.5 cm. in Fossil turtles from the early Miocene localities of Mokrá-Quarry (Burdigalian, MN4), South Moravian Region, Czech Republic
nal; E', F', internal, G'-J', Pal. 1341 (peripheral 8); G', H', external; I', J', internal; K'-N', Pal. 1343 (peripheral 9); K', L', external; M', N', internal; O'-R', Pal. 1345 (peripheral 10); O', P', external; Q', R', internal; S'-V', Pal. 1348 (peripheral 11); S', T', external; U', V', internal views; W', reconstruction of carapace Thick lines correspond to scute sulci, dotted lines denote plate sutures, and oblique lines indicate missing plate portions. Abbreviations: Ce, cervical; co, costal; Ma, marginal; ne, neural; nu, nuchal; per, peripheral; Pl, pleural; py, pygal; sp, suprapygal; Spr, supracaudal; Ve, vertebral. Scale bars: A-V', 1 cm; W', 2.5 cm.
sal; W, visceral; X-Z, Pal. 1308 (costal 5); X, Y, dorsal; Z, visceral; A'-C', Pal. 1309 (costal 6); A', B', dorsal; C', visceral; D'-F', Pal. 1310 (costal 8); D', E', dorsal; F', visceral; G'-J', Pal. 1312 (peripheral 1); G', H', dorsal; I', J', visceral; K'-N', Pal. 1313 (peripheral 7); K', L', dorsal; M', N', visceral; O'-R', Pal. 1314 (peripheral 8); O', P', dorsal; Q', R', visceral views; S', reconstruction of carapace. Thick lines indicate to scute sulci, dotted lines sutures and oblique lines denote missing plate portions. Abbreviations: Ce, cervical; co, costal; Ma, marginal; ne, neural; nu, nuchal; per, peripheral; Pl, pleural; py, pygal; sp, suprapygal; Ve, vertebral. Scale bars: 1 cm. in Fossil turtles from the early Miocene localities of Mokrá-Quarry (Burdigalian, MN4), South Moravian Region, Czech Republic
sal; W, visceral; X-Z, Pal. 1308 (costal 5); X, Y, dorsal; Z, visceral; A'-C', Pal. 1309 (costal 6); A', B', dorsal; C', visceral; D'-F', Pal. 1310 (costal 8); D', E', dorsal; F', visceral; G'-J', Pal. 1312 (peripheral 1); G', H', dorsal; I', J', visceral; K'-N', Pal. 1313 (peripheral 7); K', L', dorsal; M', N', visceral; O'-R', Pal. 1314 (peripheral 8); O', P', dorsal; Q', R', visceral views; S', reconstruction of carapace. Thick lines indicate to scute sulci, dotted lines sutures and oblique lines denote missing plate portions. Abbreviations: Ce, cervical; co, costal; Ma, marginal; ne, neural; nu, nuchal; per, peripheral; Pl, pleural; py, pygal; sp, suprapygal; Ve, vertebral. Scale bars: 1 cm.
Text-fig. 3. a: Pterigophycos sp., whole-plant specimen, coll. No. 22.116, larger blades denoted B1–4 (for details, see text), scale bar = 2 cm; b: Close-up of blades B1 and B2, which resemble P. spectabilis A.MASSAL. and P. canossae A.MASSAL., respectively; c: Closeup of blade B3 resembling P. canossae; d: Close-up of blade B4 resembling P. gazolanus A.MASSAL. or Laminarites irideaephyllus A.MASSAL. Scale bars = 1 cm unless otherwise stated. in A Whole-Plant Specimen Of The Marine Macroalga Pterigophycos From The Eocene Of Bolca (Veneto, N-Italy)
Text-fig. 3. a: Pterigophycos sp., whole-plant specimen, coll. No. 22.116, larger blades denoted B1–4 (for details, see text), scale bar = 2 cm; b: Close-up of blades B1 and B2, which resemble P. spectabilis A.MASSAL. and P. canossae A.MASSAL., respectively; c: Closeup of blade B3 resembling P. canossae; d: Close-up of blade B4 resembling P. gazolanus A.MASSAL. or Laminarites irideaephyllus A.MASSAL. Scale bars = 1 cm unless otherwise stated.
◂Fig. 14 Scanning electron micrographs (SEM) showing radular ribbon form, middle adhesive zone (az) and rows of dentition (rd) of Dinaride and Iberian individuals (notation denotes aspects on one Dinaride Zospeum and one Iberozospeum ribbon); (a) Z. exiguum (NMBE 553384), Križna jama, Slovenia (45.7452, 14.4673), long and narrow, tapered anterior end (tae), short adhesive zone (az), bottom furled with narrow obtuse or straight base (nosb); (b) Z. pretneri, (NMBE 553290), Gornja Cerovačka pećina, Croatia (44.2701, 15.8855), ibid., with straight base; (c) I. vasconicum, (AJC 1848), Cueva Ermita de Sandaili (42.9994, -2.4381), moderately long and broad, tapered anterior end (tae), prominent adhesive zone (az), straight base (sb); (d) I. zaldivarae, (AJC 1876), Cueva de Las Paúles (43.1282, -2.7362), ibid.; (e) Iberozospeum sp. (RMNH.MOL. 234,109), Cueva de la Foz, long and broad, ibid; (f) Iberozospeum sp., (RMNH.MOL. 234,144), Cueva de Rales, very long and broad, ibid; (g) Iberozospeum sp., (RMNH.MOL. 234,116), Cueva a Sul, long and broad, ibid; (h) Iberozospeum sp., (RMNH.MOL. 234,108), Cueva de Torcona, very long and broad, ibid. — Magnification varies for each perspective, see scale bars; all Figs imaged by M. Ruppel, (ret.) Goethe University Frankfurt am Main in Molecular investigation and description of Iberozospeum n. gen., including the description of one new species (Eupulmonata, Ellobioidea, Carychiidae)
◂Fig. 14 Scanning electron micrographs (SEM) showing radular ribbon form, middle adhesive zone (az) and rows of dentition (rd) of Dinaride and Iberian individuals (notation denotes aspects on one Dinaride Zospeum and one Iberozospeum ribbon); (a) Z. exiguum (NMBE 553384), Križna jama, Slovenia (45.7452, 14.4673), long and narrow, tapered anterior end (tae), short adhesive zone (az), bottom furled with narrow obtuse or straight base (nosb); (b) Z. pretneri, (NMBE 553290), Gornja Cerovačka pećina, Croatia (44.2701, 15.8855), ibid., with straight base; (c) I. vasconicum, (AJC 1848), Cueva Ermita de Sandaili (42.9994, -2.4381), moderately long and broad, tapered anterior end (tae), prominent adhesive zone (az), straight base (sb); (d) I. zaldivarae, (AJC 1876), Cueva de Las Paúles (43.1282, -2.7362), ibid.; (e) Iberozospeum sp. (RMNH.MOL. 234,109), Cueva de la Foz, long and broad, ibid; (f) Iberozospeum sp., (RMNH.MOL. 234,144), Cueva de Rales, very long and broad, ibid; (g) Iberozospeum sp., (RMNH.MOL. 234,116), Cueva a Sul, long and broad, ibid; (h) Iberozospeum sp., (RMNH.MOL. 234,108), Cueva de Torcona, very long and broad, ibid. — Magnification varies for each perspective, see scale bars; all Figs imaged by M. Ruppel, (ret.) Goethe University Frankfurt am Main
Figure. Mean pre-adult development time (in days) values for all strains. Vertical bars denote 0.95 confidence intervals. in Effects of artificial migration of susceptible individuals on resistance and fitness of a fenitrothion-resistant strain of Musca domestica (L.) Diptera
Figure. Mean pre-adult development time (in days) values for all strains. Vertical bars denote 0.95 confidence intervals.
Figure 1. Pandirodesmus rutherfordi habitus. The arrow denotes a in The enigmatic milliped genus Pandirodesmus Silvestri 1932 and description of a new species from Tobago represented by males (Polydesmida: Leptodesmidea: Chelodesmidae: Chelodesminae: Pandirodesmini)
Figure 1. Pandirodesmus rutherfordi habitus. The arrow denotes a rounded accumulation of sand on the right anterior leg on segment 15.
A combination of HLA-DP α and β chain polymorphisms paired with a SNP in the DPB1 3' UTR region, denoting expression levels, are associated with Atopic Dermatitis
<p>The publication "A combination of HLA-DP α and β chain polymorphisms paired with a SNP in the DPB1 3’ UTR region, denoting expression levels, are associated with Atopic Dermatitis" contains analysis from two different cohorts: Genetics in Atopic Dermatitis (GAD), which is the main dataset, and Pediatric Eczema Elective Registry (PEER), which is the replication cohort. Included herein are the HLA Class II genotypes for both the GAD and PEER cohorts at 2-field resolution, which forms the basis for the analysis included in the publication. (DOI: 10.3389/fgene.2023.1004138)</p>
Animacy and non-animacy denoting nouns
<p>Universität Zürich<br> Institut für Computerlinguistik<br> Manfred Klenner</p> <p>Die Daten sind im Rahmen eines Projekts, das vom Schweizer Nationalfond gefördert wurde (Nr. 105215-179302), entstanden.<br> ------------------------------------------------------------------------------------------------------------------------<br> **License**: Creative Commons Attribution-ShareAlike 4.0 International Public License</p> <p><br> Repository: animacy data for animcay classification</p> <p>This is the training data for an animacy classifier (see References LREC)</p> <p>1) gold_actor: 7468 nouns denoting animate entities<br> 2) gold_nonactor 5511 nouns denoting non-animate entities</p> <p>subsets of 1:</p> <p>gold_direct 6897 nouns directly denoting animate entities<br> gold_metonym 587 metonymy trigger nouns</p> <p><br> Format: just lists</p> <p>Note: although some person names are in the data, a separate NER for person names should be used .</p> <p>References:</p> <p>@inproceedings{LREC,<br> month = {Juni},<br> author = {Manfred Klenner and Anne G{\"o}hring},<br> booktitle = {Proceedings of the Language Resources and Evaluation Conference},<br> address = {Marseille, France},<br> title = {Animacy Denoting {G}erman Nouns: Annotation and Classification},<br> publisher = {European Language Resources Association},<br> pages = {1360--1364},<br> year = {2022},<br> language = {english},<br> url = {https://doi.org/10.5167/uzh-219148},<br> abstract = {In this paper, we introduce a gold standard for animacy detection comprising almost 14,500 German nouns that might be used to denote either animate entities or non-animate entities. W<br> e present inter-annotator agreement of our crowd-sourced seed annotations (9,000 nouns) and discuss the results of machine learning models applied to this data.}<br> }<br> </p>
Figure 1. - Bayesian phylogeny of Euptychia based on one mitochondrial (COI) and one nuclear (EF1-a) gene. Posterior probabilities are listed above and bootstrap values below branches. A dash denotes bootstrap support lower than 50%. (Euptychiaattenboroughi is not included in the analysis – see text for details.)
Figure 1. - Bayesian phylogeny of Euptychia based on one mitochondrial (COI) and one nuclear (EF1-a) gene. Posterior probabilities are listed above and bootstrap values below branches. A dash denotes bootstrap support lower than 50%. (Euptychiaattenboroughi is not included in the analysis – see text for details.)
Data for manuscript "The Prevalence of Terms Denoting Far-right and Far-left Political Extremism in U.S. and U.K. News Media"
<p>This data set belongs to an academic manuscript examining longitudinally (2000-2019) the prevalence of terms denoting far-right and far-left political extremism in a large corpus of more than 32 million written news and opinion articles from 54 news media outlets popular in the United States and the United Kingdom.</p> <p>The textual content of news and opinion articles from the 54 outlets listed in the main manuscript is available in the outlet's online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We used derived word frequency counts from these sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific xpath expressions. Tokens were lowercased prior to estimating frequency counts. To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1 million words of article content from an outlet. This threshold was chosen to maximize inclusion in our analysis of outlets with sparse amounts of articles text per year. </p> <p>Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript </p> <p>-articlesContainingTargetWords.rar contains counts of target words in outlets articles as well as total counts of words in articles</p> <p> </p> <p>Usage Notes</p> <p>In a small percentage of articles, outlet specific XPath expressions failed to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles are not precise. In a random sample of articles and outlets, manual estimation of target words counts overlapped with the automatically derived counts for over 90% of the articles.</p> <p>Most of the incorrect frequency counts were minor deviations from the actual counts such as for instance counting the word "Facebook" in an article footnote encouraging article readers to follow the journalist’s Facebook profile and that the XPath expression mistakenly included as the content of the article main text. Some additional outlet-specific inaccuracies that we could identify occurred in "The Hill" and "Newsmax" news outlets where XPath expressions had some shortfalls at precisely capturing articles’ content. For "The Hill", in years 2007-2009, XPath expressions failed to capture the complete text of the article in about 40% of the articles. This does not necessarily result in incorrect frequency counts for that outlet but in a sample of articles’ words that is about 40% smaller than the total population of articles words for those three years. In the case of "NewsMax", the issue was that for some articles, XPath expressions captured the entire text of the article twice. Notice that this does not result in incorrect frequency counts. If a word appears x times in an article with a total of y words, the same frequency count will still be derived when our scripts count the word 2x times in the version of the article with a total of 2y words.</p> <p>To conclude, in a data analysis of 32 million articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 1 in the main manuscript for illustration of the accuracy of the frequency counts).</p>
Data set in the form a relational database (sql) to denote a network of service providers, service clients and recommenders
<p>This data-set pertains to a network (i.e. graph) represented in the form of a relational data-base of service providers (nodes), service clients (nodes), service recommenders (nodes) and relationaships between then (i.e. a client used a provider, a recommender recommended a service to another client), along with some initial values of the QoS level perceived by any client whi have used a service and the reputation of a recommender. The data-set can be used for developing a reputation-based trust system. </p>
Fig. 1. Map denoting 3 in Marine algal flora of Oho-ri, Gosung-gun, Gangwon-do, Korea
Fig. 1. Map denoting 3 islets and its diving stations where the survey was conducted.
FIGURE 4. Discriminant function analysis depicting morphological differentiation within the C. vittatus-hansenae group. Close grey circle indicates C. vittatus Group II. Open circle denotes C. hansenae Group I. Closed black circle represents C. hansenae Group II in Re-evaluating the taxonomic status of Chiromantis in Thailand using multiple lines of evidence (Amphibia: Anura: Rhacophoridae)
FIGURE 4. Discriminant function analysis depicting morphological differentiation within the C. vittatus-hansenae group. Close grey circle indicates C. vittatus Group II. Open circle denotes C. hansenae Group I. Closed black circle represents C. hansenae Group II.
Figure 6: Tin price evolution from 01/2003 to 01/2022. The timing of the release from the MV Weserland cargo matches the peak of the tin price (2021) in US Dollars. Left axis denotes tin price (US dollars per ton) and right axis price variation since 01/2003. Source: https://markets.businessinsider.com/commodities/tin-price
Open the record for dataset details and reuse information.
Data for manuscript: "The Prevalence of Prejudice Denoting Terms in Spanish Newspapers"
<p>This data set contains frequency counts of target words in 5 million news and opinion articles from 3 popular newspapers in Spain: El País, El Mundo and ABC. The target words are listed in the associated manuscript and are mostly words that denote some type of prejudice. A few additional words not denoting prejudice are also available since they are used in the manuscript for illustration purposes.</p> <p>The textual content of news and opinion articles from the outlets listed in Figure 1 of the main manuscript is available in the outlet's online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We used derived word frequency counts from original sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific xpath expressions. Tokens were lowercased prior to estimating frequency counts. To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1 million words of article content from an outlet. </p> <p>Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript </p> <p>-targetWordsInArticlesCounts.rar contains counts of target words in outlets articles as well as total counts of words in articles</p> <p>Usage Notes</p> <p>In a small percentage of articles, outlet specific XPath expressions can fail to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles might not be precise. </p> <p>To conclude, in a data analysis of millions of news articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 2 of main manuscript for supporting evidence).</p>
FIGURE 2. Discretized continuous characters. Each point denotes a in Revision and phylogenetic analysis of the orb-weaving spider genus Glenognatha Simon, 1887 (Araneae, Tetragnathidae)
FIGURE 2. Discretized continuous characters. Each point denotes a measured individual (up to three for each species) and the horizontal lines are the limits between states (Mann-Whitney U=0.0, p<0.001). PDP: paracymbium distal portion. PBP: paracymbium basal portion. Ret: Retromarginal tooth.
the area between the left femur and left manual digits. Arrows denote the positions of the samples. Scale bars, 2 µm. in A new Jurassic scansoriopterygid and the loss of membranous wings in theropod dinosaurs
the area between the left femur and left manual digits. Arrows denote the positions of the samples. Scale bars, 2 µm.
FIG UR E 3 (a) Dated phylogeny of the genus Theodoxus constructed in BEAST based on COI, 16S and ATPα. Node labels denote divergence times in millions of years ago (Ma); node bars indicate the 95% credibility interval around these dates. Small squares at nodes indicate significant support of divergence events found with BEAST and other phylogenetic analyses (see Figures S2.1 and S2.2), as explained through the key. Where MOTUs (A–R) show conspecifics among a number of morphospecies, species names are given in order of their year of description. Morphospecies, incorporated from GenBank, where determination was potentially dubious are highlighted by an asterisk. Clades (C) and subclades (SC) are demarcated by dashed lines between MOTUs. (b) LTT plots indicating the build‐up of lineages in Theodoxus over geological time. Dashed lines surrounding the solid LTT lines indicate the 95% confidence intervals. Where intra‐ and interspecific diversity diverge, interspecific diversity is highlighted in blue and intraspecific diversity in red. Transitions in geological ages are highlighted by narrow grey lines, while the grey bar marks the period of pronounced glacial cycles (last 900 kyr) [Colour figure can be viewed at wileyonlinelibrary.com] in Contributions of biogeographical functions to species accumulation may change over time in refugial regions
FIG UR E 3 (a) Dated phylogeny of the genus Theodoxus constructed in BEAST based on COI, 16S and ATPα. Node labels denote divergence times in millions of years ago (Ma); node bars indicate the 95% credibility interval around these dates. Small squares at nodes indicate significant support of divergence events found with BEAST and other phylogenetic analyses (see Figures S2.1 and S2.2), as explained through the key. Where MOTUs (A–R) show conspecifics among a number of morphospecies, species names are given in order of their year of description. Morphospecies, incorporated from GenBank, where determination was potentially dubious are highlighted by an asterisk. Clades (C) and subclades (SC) are demarcated by dashed lines between MOTUs. (b) LTT plots indicating the build‐up of lineages in Theodoxus over geological time. Dashed lines surrounding the solid LTT lines indicate the 95% confidence intervals. Where intra‐ and interspecific diversity diverge, interspecific diversity is highlighted in blue and intraspecific diversity in red. Transitions in geological ages are highlighted by narrow grey lines, while the grey bar marks the period of pronounced glacial cycles (last 900 kyr) [Colour figure can be viewed at wileyonlinelibrary.com]
Data for manuscript: "Prevalence of prejudice denoting words in news media discourse: a chronological analysis"
<p>This data set contains frequency counts of target words in 27 million news and opinion articles from 47 popular news media outlets in the United States. The target words are listed in the associated manuscript and are mostly words that denote some type of prejudice. A few additional words not denoting prejudice are also available since they are used in the manuscript for illustration purposes.</p> <p>The textual content of news and opinion articles from the outlets listed in Figure 4 of the main manuscript is available in the outlet's online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We used derived word frequency counts from these sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific xpath expressions. Tokens were lowercased prior to estimating frequency counts. To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1.25 million words of article content from an outlet. This threshold was chosen to maximize inclusion in our analysis of outlets with sparse amounts of articles text per year such as Reason, Alternet or The American Spectator. </p> <p>Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript </p> <p>-articlesContainingTargetWords.rar contains counts of target words in outlets articles as well as total counts of words in articles</p> <p>-cableNews.rar contains prevalence of target words in TV cable news. Data is from Stanford Cable TV News Analyzer (https://tvnews.stanford.edu/)</p> <p>-surveyData.rar contains longitudinal survey data used in the manuscript and links to original sources</p> <p>Usage Notes</p> <p>In a small percentage of articles, outlet specific XPath expressions failed to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles are not precise. In a random sample of articles and outlets, manual estimation of target words counts overlapped with the automatically derived counts for over 90% of the articles.</p> <p>Most of the incorrect frequency counts were minor deviations from the actual counts such as for instance counting the word "Facebook" in an article footnote encouraging article readers to follow the journalist’s Facebook profile and that the XPath expression mistakenly included as the content of the article main text. Some additional outlet-specific inaccuracies that we could identify occurred in "The Hill" and "Newsmax" news outlets where XPath expressions had some shortfalls at precisely capturing articles’ content. For "The Hill", in years 2007-2009, XPath expressions failed to capture the complete text of the article in about 40% of the articles. This does not necessarily result in incorrect frequency counts for that outlet but in a sample of articles’ words that is about 40% smaller than the total population of articles words for those three years. In the case of "NewsMax", the issue was that for some articles, XPath expressions captured the entire text of the article twice. Notice that this does not result in incorrect frequency counts. If a word appears x times in an article with a total of y words, the same frequency count will still be derived when our scripts count the word 2x times in the version of the article with a total of 2y words. To conclude, in a data analysis of 27 million articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 1 and Figure 2 of main manuscript for supporting evidence).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.