Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
794
datasets available to search
ShareScore release 0.7.1
Dataset results
794 results for “publishing”
IBD segments of previously published Eurasian individuals from "Accurate detection of identity-by-descent segments in human ancient DNA""
<p>IBD segments among 4,248 previously published Eurasian individuals. For details please refer to our manuscript. </p>
Fig. 16 in Consortium of European TaXonomic Facilities (CETAF) best practices in electronic publishing in taXonomy
Fig. 16. Scheme extracted from Roarmap showing the different policies regarding Open Access.
Phylogenetic data for construction of bryophyte tree using published sanger data
Open the record for dataset details and reuse information.
Multiple sequence alignments of newly reconstructed and published cervid and human mtDNA
Open the record for dataset details and reuse information.
Scholarly publishing in a developing nation
Open the record for dataset details and reuse information.
Published correlational effect sizes in social and developmental psychology
Open the record for dataset details and reuse information.
Survey of software engineering in code used in published papers
Open the record for dataset details and reuse information.
Conceptualizing relationships among hyporheic exchange, storage, and water age: data represented in published figures
Open the record for dataset details and reuse information.
Data from: Landscape composition and life-history traits influence bat movement and space use: analysis of 30 years of published telemetry data
Open the record for dataset details and reuse information.
For Articles Published in 1995: Evolution Of The Top 50 Most Cited
<p>ascii data files (cleaned, calibrated, processed) and journal-level figure figure (in vector graphics format) shown in the the youtube video "The Top 50 Most Cited Articles - The Short and Long of It" at <a href="https://www.youtube.com/watch?v=4wyy80QZ0lg">https://www.youtube.com/watch?v=4wyy80QZ0lg</a></p>
Data showcase papers published in the Mining Software Repositories (MSR) conference (v2.2)
<p>Data regarding data showcase papers published in the Mining Software Repositories (MSR) conference.</p> <p>The following data files are included.</p> <ul> <li>citing_dp_dois_citations.txt: Strong and weak citations of (strong and weak) citation papers</li> <li>data_paper_clustering.csv: The clustering process of MSR data papers</li> <li>data_paper_clusters.csv: Clusters of MSR data papers</li> <li>data_papers.bib: Bibliographic details of MSR data papers, along with their assigned clusters (field 'cluster') and strong citations (field 'usedby')</li> <li>dp_dois_citations.txt: Strong and weak citations of MSR data papers</li> <li>msr-all: Bibliographic details of all MSR (data and non-data) papers</li> <li>ndp_dois_citations.txt: Strong and weak citations of MSR non-data papers</li> <li>ndp_rand_dois_citations.txt: Strong and weak citations of a randomly chosen MSR non-data paper weighted sample</li> <li>self-citations.txt: Strong citations of MSR data papers by their authors</li> <li>strong_citation_classification.csv: The classification process of strong citation papers according to the SWEBOK knowledge areas</li> <li>strong_citation_fields.csv: SWEBOK knowledge areas of strong citation papers</li> <li>strong_citations.bib: Bibliographic details of strong citation papers</li> <li>survey_questionnaire.pdf: The final survey questionnaire</li> <li>survey_responses.csv: Anonymized responses of the final survey questionnaire (Email addresses have been excluded for privacy reasons.)</li> <li>weak_citations_notes.bib: Weak citations of MSR data papers and the use they make</li> </ul>
Figure 7. The new F1000Workspace that provides a in The five deadly sins of science publishing
Figure 7. The new F1000Workspace that provides a set of tools to enable researchers to collect references, write research articles and grant applications, and collaborate with co-authors and colleagues.
Figure 4 in The five deadly sins of science publishing
Figure 4. An example of an F1000Prime recommendation showing the Faculty Members who made the article recommendation, the rating they gave the article, and the associated comment as to why they felt that article was so interesting.
Figure 3 in The five deadly sins of science publishing
Figure 3. Faculty of 1000, which now comprises 3 core services: F1000Prime, F1000Workspace and F1000Research, and is overseen by the F1000 Faculty comprising over 11,000 members.
Quality of advertisements for prescription drugs in family practice medical journals published in Australia, Canada, and the United States with different regulatory controls: a cross-sectional study
<p><span><span><span><span><span><span><span><span><span><span><span><b>OBJECTIVE: </b>To assess if different forms of regulation lead to differences in the quality of journal advertisements.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>DESIGN: </b>Cross-sectional study.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>PARTICIPANTS: </b>Thirty advertisements from family practice journals published from 2013-2015 were extracted for three countries with distinct regulatory pharmaceutical promotion systems: Australia, Canada, and the United States (US). </span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>PRIMARY AND SECONDARY OUTCOME MEASURES: </b>Advertisements under each regulatory system were compared concerning three domains: information included in the advertisement, references to scientific evidence, and pictorial appeals and portrayals. An overall ranking for advertisement quality among countries was determined using the first two domains as the information assessed has been associated with more appropriate prescribing. </span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>RESULTS: </b>Advertisements varied significantly for number of claims with quantitative benefit (Australia: 0.0 (0.0-3.0); Canada: 0.0 (0.0-5.0); US: 1.0 (0.0-6.0); p=0.01); statistical method used in reporting benefit (RRR, ARR, and NNT) (Australia: 6.7%, n=2; Canada: 10.0%, n=3; US: 36.6%, n=11; p=0.02); mention of adverse effects, warnings, or contraindications (Australia: 13.3%, n=4; Canada: 23.3%, n=7; US: 53.3%, n=16; p=0.002); equal prominence between safety and benefit information (Australia: 25.0%, n=1; Canada: 28.6%, n=2; US: 75.0%, n=12; p=0.04); and methodologic quality of references score (Australia: 0.4150 (0.25-0.70); Canada: 0.25 (0.00-0.63); US: 0.25 (0.00-0.75); p<0.001). The US ranked first, Canada second, and Australia third for overall quality of journal advertisements. Significant differences for humor appeals (Australia: 3.3%, n=1; Canada: 13.3%, n=4; US: 26.7%, n=8; p=0.04), positive emotional appeals (Australia: 26.7%, n=8; Canada: 60.0%, n=18; US: 50.0%, n=15; p=0.03), social approval portrayals (Australia: 0.0%, n=0; Canada: 0.0%, n=0; US: 10.0%, n=3; p=0.04), and lifestyle or work portrayals (Australia: 43.3%, n=13; Canada: 50.0%, n=15; US: 76.7%, n=23; p=0.02) were found among countries.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>CONCLUSIONS:</b> <a name="_Hlk34742358">Different regulatory systems influence journal advertisement quality concerning all measured domains. </a>However, differences may also be attributed to other regulatory, legal, cultural, or health system factors unique to each country.</span></span></span></span></span></span></span></span></span></span></span></p>
New data on the publishing productivity of American sociologists
<p><br> <strong>OVERVIEW</strong><br> <br> This data file, compiled from multiple online sources, presents 2013–2017 publication counts—articles, articles in high-impact journals, books, and books from high-impact publishers—for 2,132 professors and associate professors in 426 U.S. departments of sociology. It also includes information on institutional characteristics (e.g., institution type, highest sociology degree offered, department size) and individual characteristics (e.g., academic rank, gender, PhD year, PhD institution).<br> <br> The data may be useful for investigations of scholarly productivity, the correlates of scholarly productivity, and the contributions of particular individuals and institutions. Complete population data are presented for the top 26 doctoral programs, doctoral institutions other than R1 universities, the top liberal arts colleges, and other bachelor's institutions. Sample data are presented for Carnegie R1 universities (other than the top 26) and master's institutions.<br> <br> <strong>USER NOTES</strong><br> <br> Please see our paper in <em>Scholarly Assessment Reports</em>, freely available at <a href="https://doi.org/10.29024/sar.36">https://doi.org/10.29024/sar.36</a> , for full information about the data set and the methods used in its compilation. The section numbers used here refer to the Appendix of that paper. See the References, below, for other papers that have made use of these data.<br> <br> The data file is a single Excel file with five worksheets: <em>Sampling</em>, <em>Articles</em>, <em>Books</em>, <em>Individuals</em>, and <em>Departments</em>. Each worksheet has a simple rectangular format, and the cells include just text and values—no formulas or links. A few general notes apply to all five worksheets.<br> <br> • The yellow column headings represent institutional (departmental) data. The blue column headings represent data for individual faculty.<br> <br> • <em>iType</em> is institution type, as described in section A.2—TopR (top research universities), R1 (other R1 universities), OD (other doctoral universities), M (master's institutions), TopLA (top liberal arts colleges), or B (other bachelor's institutions). <em>nType</em> provides the same information, but as a single-digit code that is more useful for sorting the rows; 1=TopR, 2=R1, 3=OD, 4=M, 5=TopLA, and 6=B.<br> <br> • <em>Inst</em> is a four-digit institution code. The first digit corresponds to <em>nType</em>, and the last three digits allow for alphabetical sorting by institution name. <em>Indiv</em> is a one- or two-digit code that can be used to sort the individuals by name within each department. The <em>Inst</em>, <em>nType</em>, and <em>Indiv</em> codes are consistent across the five worksheets.<br> <br> • For binary variables such as <em>Full professor</em> and <em>Female</em>, 1 indicates <em>yes</em> (full professor or female) and 0 indicates <em>no</em> (associate professor or male).<br> <br> The five worksheets represent five distinct stages in the data compilation process. First, the <em>Sampling</em> worksheet lists the 1,530 base-population institutions (see section A.3) and presents the characteristics of the faculty included in the data file. Each row with an entry in the <em>Individual</em> column represents a faculty member at one of the 426 institutions included in the data set. Each row without an entry in the <em>Individual</em> column represents an institution that either (a) did not meet the criteria for inclusion (section A.1) or (b) was not needed to attain the desired sample size for the R1 or M groups (section A.3).<br> <br> The <em>Articles</em> worksheet includes the data compiled from SocINDEX, as described in section A.6. Each row with an entry in the <em>Journal</em> column represents an article written by one of the 2,132 faculty included in the data. Each row without an entry in the <em>Journal</em> column represents a faculty member without any article listings in SocINDEX for the 2013–2017 period. (Note that SocINDEX items other than peer-reviewed articles—editorials, letters, etc.—may be listed in the <em>Journal</em> column but assigned a value of 1 in the <em>Excluded</em> column and a value of 0 in the <em>Article credit</em> and <em>HI article credit</em> columns. We assigned no credit for items such as editorial and letters, but other researchers may wish to include them.) The <em>N</em> and <em>i</em> columns represent, for each article, the number of authors (<em>N</em>) and the faculty member's place in the byline (<em>i</em>), as described in section A.8. The <em>CiteScore</em> and <em>Highest percentile</em> columns were used to identify high-impact journals, as indicated in the <em>HI journal</em> column. The <em>Article credit</em> and <em>HI article credit</em> columns are article counts, adjusted for co-authorship.<br> <br> The <em>Books</em> worksheet includes data compiled from Amazon and other sources, as described in section A.7. Each row with an entry in the <em>Book</em> column represents a book written by one of the 2,132 faculty. Each row without an entry in the <em>Book</em> column represents a faculty member without any book listings in Amazon during the 2013–2017 period. The publication counts in the <em>Books</em> worksheet—<em>Book credit</em> and <em>HI book credit</em>—follow the same format as those in the <em>Articles</em> worksheet.<br> <br> The <em>Individuals</em> worksheet consolidates information from the <em>Articles</em> and <em>Books</em> worksheets so that each of the 2,132 individuals is represented by a single row. The worksheet also includes several categorical variables calculated or otherwise derived from the raw data—<em>Years since PhD</em>, for instance, and the three corresponding binary variables. We suspect that many data users will be most interested in the <em>Individuals</em> worksheet.<br> <br> The <em>Departments</em> worksheet collapses the individual data so that each of the 426 institutions (departments) is represented by a single row. Individual characteristics such as <em>Female</em> and <em>Years since PhD</em> are presented as percentages or averages—<em>% Female</em> and <em>Avg years since PhD</em>, for instance. Each of the four productivity measures is represented by a departmental total, an average (the total divided by the number of full and associate professors), a departmental standard deviation, and a departmental median.</p>
Replication package for the paper accepted at Springer's EMSE Journal: Publish or Perish - But do not Forget your Software Artifacts
<p>This is the replication package for the paper "Publish or Perish - But do not Forget your Software Artifacts", accepted at Springer's EMSE Journal in June 2020.</p> <p>It contains:</p> <ul> <li>A ReadMe file with instructions on how to use the replication scripts</li> <li>The complete, labeled dataset of 792 ICSE papers as CSV</li> <li>All the python scripts that we used for data-acquisition and -preparation</li> <li>The jupyter notebook that we used for the evaluation. This also contains some additional analyses which are not included in the paper.</li> <li>An html export of the notebook for quick reference</li> <li>A folder containing all of the diagrams in pdf form</li> </ul>
Newly discovered cichlid fish biodiversity threatened by hybridization with non-native species - Data supporting published version
<p><span><a name="_Hlk503794553"><span>Invasive freshwater fish systems are known to readily hybridize with indigenous congeneric species, driving loss of unique and irreplaceable genetic resources. Here we reveal that newly discovered (2013-2016) evolutionarily significant populations of Korogwe tilapia (<i>Oreochromis korogwe</i>) from southern Tanzania are threatened by hybridization with the larger invasive Nile tilapia (<i>Oreochromis niloticus</i>). We use a combination of morphology, microsatellite allele frequencies and whole genome sequences to show that <i>O. korogwe</i> from southern lakes (Nambawala, Rutamba and Mitupa) are distinct from geographically-disjunct populations in northern Tanzania (Zigi River and Mlingano Dam). We also provide genetic evidence of <i>O. korogwe</i> x <i>niloticus</i> hybrids in three southern lakes and </span></a><span>demonstrate heterogeneity in the extent of admixture across the genome. Finally, using the least admixed genomic regions we estimate that the northern and southern <i>O. korogwe</i> populations most plausibly diverged approximately 140,000 years ago, suggesting that the geographical separation of the northern and southern groups is not a result of a recent translocation, and instead these populations represent independent evolutionarily significant units. We conclude that these newly-discovered and phenotypically unique </span><span>cichlid populations are already threatened by hybridization with an invasive species, and propose that these irreplaceable genetic resources would benefit from conservation interventions.</span></span></p>
Rapid publishing for public health books against COVID-19
<p>A YouTube version of the video is <a href="https://www.youtube.com/watch?v=v6WUUTv-GIc&t=3s&ab_channel=SimonWorthington">available here</a>.</p> <p>Presented are two case studies covering the barriers to be overcome to fully automate the production workflow for Open Access multi-format books, to produce and distribute the following – ebook, print-on-demand, screen PDF, webbook, website, and an interoperable source.</p> <p>The first case study involves producing eight book sprints for <a href="https://github.com/akademie-oeffentliches-gesundheitswesen">training manuals</a>, some with MOOC modules, for the Academy of Public Health in Dusseldorf (Germany) which was run as a <a href="https://github.com/TIBHannover/Rapid-Collaborative-Health-Publishing">research cooperation</a> with the Open Science Lab, TIB – German National Library of Science and Technology. </p> <p>The second case study involves converting the existing reports of <a href="https://www.independentsage.org/">Independent SAGE</a> (UK) – as Open Access, multi-format, enabling multi-channel distribution, and deposing in academic repositories. </p> <p>The indie_SAGE project involved creating a volunteer academic working group to carry out the work. Here it is important to add that I am acting as a private individual. The Independent Science Advisory Group for Emergencies (indie_SAGE) was formed in May 2020 by the former chief Scientific Adviser to the UK government Sir David King, quote, ‘on how to minimise deaths and support Britain’s recovery from the COVID-19 crisis’. </p> <p>The <a href="https://github.com/Independent-SAGE/Technical-Publishing-Working-Group">working group</a> is newly formed and welcomes help and volunteers! </p>
Data from: Reporting guidelines for community-based participatory research (CBPR) did not improve the reporting quality of published studies: a systematic review of studies on smoking cessation
<p>Although a guideline for reporting the results of community-based participatory research (CBPR) was published in 2010, the impact on the quality of reporting a CBPR on smoking cessation is unknown. Here we provide the raw data of a systematic review that assessed the impact of a 2010 community-based participatory research reporting guideline. on the quality of reporting a CBPR on smoking cessation. Specifically, we searched the MEDLINE, Embase, the Cochrane Central Register for Controlled Trials (CENTRAL), PsycINFO, and CINAHL databases and included articles published up to October 2018 (PROSPERO: CRD42019111668). We assessed reporting quality using a 13-item checklist.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.