Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,582
datasets available to search
ShareScore release 0.9.0
Dataset results
1,582 results for “manuscript”
Data for manuscript: "The Prevalence of Prejudice Denoting Terms in Spanish Newspapers"
<p>This data set contains frequency counts of target words in 5 million news and opinion articles from 3 popular newspapers in Spain: El País, El Mundo and ABC. The target words are listed in the associated manuscript and are mostly words that denote some type of prejudice. A few additional words not denoting prejudice are also available since they are used in the manuscript for illustration purposes.</p> <p>The textual content of news and opinion articles from the outlets listed in Figure 1 of the main manuscript is available in the outlet's online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We used derived word frequency counts from original sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific xpath expressions. Tokens were lowercased prior to estimating frequency counts. To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1 million words of article content from an outlet. </p> <p>Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript </p> <p>-targetWordsInArticlesCounts.rar contains counts of target words in outlets articles as well as total counts of words in articles</p> <p>Usage Notes</p> <p>In a small percentage of articles, outlet specific XPath expressions can fail to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles might not be precise. </p> <p>To conclude, in a data analysis of millions of news articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 2 of main manuscript for supporting evidence).</p>
wln510/Data_for_SWCs: Data for SWCs (for revision manuscript).
<p>This is the data for revision manuscript.</p>
Datasets to reproduce the figures of the manuscript "Hamming Distance and the onset of quantum criticality"
<p>These files, separated by folders, contain the datasets and the Python scripts necessary to reproduce the Figures of the manuscript "Hamming Distance and the onset of quantum criticality". Figures are also attached. An updated version of Matplotlib is required to plot the data.</p>
Supporting data for manuscript: Beyond Bulk
<p>We used soil density fraction data from The International Soil Radiocarbon Database (ISRaD v. 1.1.2 Lawrence et al., 2020; <a href="http://www.soilradiocarbon.org">www.soilradiocarbon.org</a>). ISRaD is an online repository for environmental radiocarbon data with a specific emphasis on soils and soil fractions. We utilized a subset of ISRaD data comprising measurements of radiocarbon (persistence), organic C concentration (abundance), or the proportion of organic C in the mineral-associated fraction (distribution) made on soil density fractions for the current analysis. Radiocarbon data are reported in units of Δ<sup>14</sup>C (‰) normalized to account for the year of sampling (Shi et al., 2020) (see below). In studies that employed sequential density separation (isolation of multiple free light, occluded light, and heavy fractions for the same sample), the multiple fractions were combined by taking a mass-weighted average for C abundance and C-weighted average for Δ<sup>14</sup>C values. C distribution among density fractions was normalized to sum to 100%. Overall, our meta-analysis included data from 52 studies. In addition to C measurements, ISRaD compiles ancillary data regarding site and sample characteristics that were either provided directly in the associated published works or provided as supplementary information from manuscript authors. When variables of interest were not available directly through ISRaD, these variables were populated through utilization of geolocated databases (see supplemental materials in associated published manuscript).</p>
Data for the manuscript: Historical biogeography of Pomaderris (Rhamnaceae): continental vicariance in Australia and repeated independent dispersals to New Zealand
<p>Gondwanan biogeographic patterns include a combination of old vicariance events following the breakup of the supercontinent, and more recent long-distance dispersals across the southern landmasses. Floristic relationships between Australia and New Zealand have mostly been attributed to recent dispersal events rather than vicariance. We assessed the biogeographic history of Pomaderris (Rhamnaceae), which occurs in both Australia and New Zealand, by constructing a time-calibrated molecular phylogeny to infer (1) phylogenetic relationships and (2) the relative contributions of vicariance and dispersal events in the biogeographic history of the genus. Using hybrid capture and high throughput sequencing, we generated nuclear and plastid data sets to estimate phylogenetic relationships and fossil calibrated divergence time estimates for Pomaderris . BioGeoBEARS and biogeographical stochastic mapping (BSM) were used to assess the ancestral area of the genus and the relative contributions of vicariance vs dispersal, and the directionality of dispersal events. Our analyses indicate that Pomaderris originated in the Oligocene and had a widespread Australian distribution. Vicariance of western and eastern Australian clades coincides with the uplift of the Nullarbor Plain c. 14 Ma, followed by subsequent in-situ and within-biome diversification with little exchange across regions. A rapid radiation of southeastern Australian taxa beginning c. 10 Ma was the source for at least six independent long-distance dispersal events to New Zealand during the Pliocene–Pleistocene. Our study demonstrates the importance of dispersal in explaining not only the current cross-Tasman distributions of Pomaderris, but for the New Zealand flora more broadly. The pattern of multiple independent long-distance dispersal events for Pomaderris , without significant radiation within New Zealand, is congruent with other lowland plant groups, suggesting that this biome has a different evolutionary history compared with the younger alpine flora of New Zealand, which exhibits extensive radiations often following single long distance dispersal events.</p>
Raw data of manuscript "Downregulated dual-specificity protein phosphatase 1 in ovarian carcinoma: a comprehensive study with multiple methods" submitted to PeerJ
<p>Some raw data and original calculating results of this manuscript.</p>
Study data supporting the manuscript titled: Effect of buprenorphine on fentanyl-induced respiratory depression
<p><b>Background: </b>Opioid-induced respiratory depression driven by ligand binding to mu-opioid receptors is a leading cause of opioid-related fatalities. Buprenorphine, a partial agonist, binds with high affinity to mu-opioid receptors but displays partial respiratory depression effects. The authors examined whether sustained buprenorphine plasma concentrations similar to those achieved with some extended-release injections used to treat opioid use disorder could reduce the frequency and magnitude of fentanyl-induced respiratory depression.</p> <p><b>Methods:</b> In this two-period crossover, single-centre study, 14 healthy volunteers (single-blind, randomized) and eight opioid-tolerant (OT) patients taking daily opioid doses ≥90 mg oral morphine equivalents (open-label) received continuous intravenous buprenorphine or placebo for 360 minutes, targeting buprenorphine plasma concentrations of 0.2 or 0.5 ng/mL in healthy volunteers and 1.0, 2.0 or 5.0 ng/mL in OT patients. Upon reaching target concentrations, participants received up to four escalating intravenous doses of fentanyl. The primary endpoint was change in isohypercapnic minute ventilation (V<sub>E</sub>). Additionally, occurrence of apnea was recorded.</p> <p><b>Results:</b> Fentanyl-induced changes in V<sub>E</sub> were smaller at higher buprenorphine plasma concentrations. In healthy volunteers, at target buprenorphine concentration of 0.5 ng/mL, the first and second fentanyl boluses reduced V<sub>E</sub> by [LSmean (95% CI)] 26% (13-40%) and 47% (37-59%) compared to 51% (38-64%) and 79% (69-89%) during placebo infusion (<i>p</i>=0.001 and <.001, respectively). Discontinuations for apnea limited treatment comparisons beyond the second fentanyl injection. In OT patients, fentanyl reduced V<sub>E</sub> up to 49% (21-76%) during buprenorphine infusion (all concentration groups combined) versus up to 100% (68-132%) during placebo infusion (<i>p</i>=0.006). In OT patients, the risk of experiencing apnea requiring verbal stimulation following fentanyl boluses was lower with buprenorphine than with placebo (odds ratio: 0.07; 95% CI: 0.0 to 0.3; <i>p</i>=0.001).</p> <p><b>Interpretation:</b> Results from this proof-of-principle study provide the first clinical evidence that high sustained plasma concentrations of buprenorphine may protect against respiratory depression induced by potent opioids like fentanyl.</p>
Data Associated with Manuscript Titled "Geochemistry and provenance of springs in a Baja California Sur mountain catchment"
<p>This dataset provides the data, data summaries, and modeling constraints used for the analyses in the manuscript titled "Geochemistry and provenance of springs in a Baja California Sur mountain catchment".</p>
Data Associated with Manuscript Titled "Development of a graphical resilience framework to understand a coupled human-natural system in a remote arid highland of Baja California Sur"
<p>This is a dataset in support of analyses of water chemistry and social networks described and interpreted in the associated manuscript titled "Development of a graphical resilience framework to understand a coupled human-natural system in a remote arid highland of Baja California Sur"</p>
Data and scripts for manuscript "Full-field modeling of heat transfer in asteroid regolith 2: Effects of porosity"
<p>Includes summary spreadsheet, solution files (zipped) in vtk format, and scripts. Vtk filenames are the same as those listed in the tables in the supplement to the JGR:Planets paper. </p>
Data and analysis for the manuscript VISUAL NOISE EFFECT ON READING IN THREE DEVELOPMENTAL DISORDERS: ASD, ADHD, AND SLD
<p>These are raw data and their analysis for the manuscript VISUAL NOISE EFFECT ON READING IN THREE DEVELOPMENTAL DISORDERS: ASD, ADHD, AND SLD</p>
data for manuscript 'Two remarkable characteristics of near-inertial wave propagation in the subtropical northwestern Pacific'
<p>Subsurface mooring observational data of Typhoon Sunvn in western Pacific</p>
Data for manuscript on objective phenotyping of alfalfa roots
<p>These are images for the manuscript, "Objective phenotyping of root system architecture using image augmentation and machine learning in alfalfa."</p> <p>Alfalfa_root_ML_segmented.zip contains the root images segmented from the background using RootPainter and that were used as input for RhizoVision Explorer.</p> <p>Alfalfa_root_ML_color_photographs.zip contains the original color photographs of the roots, except that the circular scale and ID tag have been cropped out and filled with black to aid image analysis. </p>
The EWAS Catalog manuscript: Extended data
<p>Epigenome-wide association studies (EWAS) of 387 traits were conducted within the Accessible Resource for Integrated Epigenomic Studies (ARIES). These results were uploaded to The EWAS Catalog: http://www.ewascatalog.org/. The table presented here represents the extended data in The EWAS Catalog manuscript. </p>
The EWAS Catalog manuscript: Underlying data
<p>Epigenome-wide association studies (EWAS) of 40 traits were conducted using publically available data from the Gene Expression Omnibus (GEO) online repository. These results were uploaded to The EWAS Catalog: http://www.ewascatalog.org/. The table presented here represents the underlying data in The EWAS Catalog manuscript and contains the traits along with the corresponding GEO accession numbers and PubMed IDs.</p>
Videos belonging to Manuscript:"Prior experience of captivity affects behavioural responses to 'novel' environments"
<p>Three one minute video clips to show the different behaviours measured on wild caught great tits (<em>Parus major</em>):</p> <p>1) in the exploration room (RoomExploration_1minclip_bird676935V)</p> <p>2) in the exploration cage (Exploration_cage_1minclip_G373)</p> <p>3) the social response behaviour towards a mirror (Mirror_cage_1minclip_G373)</p> <p> </p> <p><em>Ethics</em></p> <p>All experiments complied with Finnish law on animal experimentations. Permits for the capture and use of great tits in experiments were granted by the Central Finland Centre for Economic Development Transport and Environment (ELY; VARELY/294/2015) and licensed from the National Animal Experiment Board (ESAVI/9114/04.10.07/2014).</p>
Dataset associated to the "ADHERENT: Learning Human-like Trajectory Generators for Whole-body Control of Humanoid Robots" paper (manuscript DOI: 10.1109/LRA.2022.3141658)
<pre><code class="language-markdown">This dataset contains data accompanying the work: @ARTICLE{9676410, author={Viceconte, Paolo Maria and Camoriano, Raffaello and Romualdi, Giulio and Ferigo, Diego and Dafarra, Stefano and Traversaro, Silvio and Oriolo, Giuseppe and Rosasco, Lorenzo and Pucci, Daniele}, journal={IEEE Robotics and Automation Letters}, title={ADHERENT: Learning Human-like Trajectory Generators for Whole-body Control of Humanoid Robots}, year={2022}, volume={7}, number={2}, pages={2779-2886}, doi={10.1109/LRA.2022.3141658}} The dataset is organized in folders, whose content can be summarized as follows: - mocap: motion capture data collected from human motion - retargeted_mocap: motion capture data retargeted on the robot - IO_features: input and output features extracted from the retargeted mocap data to train the trajectory generator - training_D2_D3_subsampled_mirrored_4ew_98%: training data - inference: data collected while generating trajectories - trajectory_control_simulation: data collected while controlling trajectories in simulation - trajectory_control_real_robot: data collected while controlling trajectories on the real robot - additional_figures: additional data to reproduce some figures in the paper and portions of the supplementary video A more detailed description of the content of each folder is provided in the README.txt file included in the dataset.</code></pre>
Datasets for manuscript: The nature, representation and measure of genetic information
<p><span>Current studies in genetics very often refer to notions from information science. The concept of genetic information is still disputed because it attributes semantic traits to what seem to be regular biochemical entities. Some researchers maintain that the use of information in biology is just metaphorical and maybe even misleading. In this paper, we offer an analysis of the nature and characteristics of the use of information in proteins, protein families, and their sequences. It is argued that the foundation of the metaphorical view is relatively weak given the current findings in bioinformatics, and it is shown that the present understanding of genetics fits well into the context of the modern philosophy of information. Here, we propose an extension of Floridi's conceptual model of information to include genetic information better. In addition, we discuss how to understand the qualitative aspects of genetic information and how to measure its quantitative aspects and present a joint statistical model including qualitative genetics, where the nominal genetic function is represented jointly with its metric self-information. The functional information of protein families in the <strong><em>Cath</em></strong> and <strong><em>Pfam</em></strong> databases are analysed. The paper concludes that scientific work may place information firmly as one of the fundamental components of molecular biology.</span></p> <p><span>The protein alignment files from the </span><strong><em>Cath</em></strong> and <strong><em>Pfam</em></strong> databases are in FASTA format.</p>
Daiber Collection Database of Arabic Manuscripts.
<p><a href="http://ricasdb.ioc.u-tokyo.ac.jp/daiber/db_index_eng.html">Daiber Collection Database of Arabic Manuscripts in the Institute of Oriental Culture, University of Tokyo</a>. </p>
Waveform data for the manuscript "Varying Shear Wave Splitting Parameters Suggest Interaction between Lithosphere and Asthenosphere in Arxan-Chaihe Volcanic Field, NE China"
<p>The folder contains the seismic waveform data (in SAC format) used for shear wave splitting measurements in this study, which has been filtered with corner frequencies of 0.02–1.00 Hz. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.