Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
78
datasets available to search
ShareScore release 0.9.0
Dataset results
78 results for “preprints”
Sambutan pembukaan acara "Preprint dan manajemen data riset" UGM 26.08.19
<p>Pidato sambutan ini dibuat untuk acara lokakarya "Preprint dan manajemen data riset" di Perpustakaan Fakultas Teknik Universitas Gadjah Mada (FT UGM) hari Senin tanggal 26.08.19. Lokakarya ini diinisiasi oleh Mas Eric Kunto Wibowo bersama Tim Pustakawan FT UGM. Semoga acara tersebut berjalan lancar dan menginspirasi para pesertanya, karena jelas menginspirasi saya. Berikut alamat <a href="https://www.erickunto.com/">blog Pak Eric Kunto</a>.</p>
dataset for bioRxiv preprint titled 'Evolution of drug resistance drives progressive destabilizations in functionally conserved molecular dynamics of the flap region of the HIV-1 protease'
<p>This data supports the Figures in the preprint titled</p> <p><strong>Evolution of drug resistance drives progressive destabilizations in functionally conserved molecular dynamics of the flap region of the HIV-1 protease</strong></p> <p><strong>working abstract</strong></p> <p>The HIV-1 protease is one of several common key targets of combination drug therapies for human immunodeficiency virus infection and acquired immunodeficiency syndrome (HIV/AIDS). During the progression of the disease, some individual patients acquire -drug resistance due to mutational hotspots on the viral proteins targeted by combination drug therapies. It has recently been discovered that drug-resistant mutations accumulate on the ‘flap region’ of the HIV-1 protease, which is a critical dynamic region involved in non-specific polypeptide binding during invasion and infection of the host cell. In this study, we utilize machine learning assisted comparative molecular dynamics, conducted at single amino acid site resolution, to investigate the dynamic changes that occur during functional dimerization and polypeptide binding of the main protease. We use a multi-agent machine learning model to identify conserved dynamics of the HIV-1 main protease that are preserved across simian and feline protease orthologs (SIV and FIV). We also investigate changes in dynamics due to common drug-resistant mutations in many patients. We find that a key functional site in the flap region, a solvent-exposed isoleucine (ILE50) and surrounding sites that control flap dynamics is often targeted by drug-resistance mutations, likely leading to malfunctional molecular dynamics affecting the overall flexibility of the flap region. We conclude that better long term patient outcomes may be achieved by designing drugs that target protease regions which are less dependent upon single sites with large functional binding effects.</p>
All data for the preprint Population genetics of Glossina palpalis gambiensis in the sleeping sickness focus of Boffa (Guinea) before and after eight years of vector control: no effect of control despite a significant decrease of human exposure to the disease
<p>Data set for the paper titled "Population genetics of <em>Glossina palpalis gambiensis</em> in the sleeping sickness focus of Boffa (Guinea) before and after eight years of vector control: no effect of control despite a significant decrease of human exposure to the disease"</p>
To Preprint or Not to Preprint: A Global Researcher Survey
<p>This data set contains the results of a survey about researchers’ adoption of and attitudes toward preprinting. The survey respondents are authors of journal articles published in 2021 and early 2022 and indexed in the Web of Science database. The questions in the survey are grouped into three sections: adoption of preprinting, attitudes toward preprinting, and demographic questions.</p> <p>The Word document contains the survey form.</p> <p>The Excel spreadsheet contains the raw survey data. Free-text responses are not included in the spreadsheet because they may reveal sensitive information.</p> <p>The "heatmap_Rdata.xlsx" and "heatmap_Rcode.txt" contain data and R code in Figs. 8, 9, 10. </p> <p>For more information about this data set, please see the paper "To Preprint or Not to Preprint: A Global Researcher Survey" by Rong Ni and Ludo Waltman. The paper is available at <a href="https://doi.org/10.31235/osf.io/k7reb">https://doi.org/10.31235/osf.io/k7reb</a>.</p>
Data from: International authorship and collaboration across bioRxiv preprints
<p>Data and supplementary tables for <a href="https://doi.org/10.7554/eLife.58496">"International authorship and collaboration across bioRxiv preprints,"</a> a paper first posted to <a href="https://doi.org/10.1101/2020.04.25.060756">bioRxiv</a> and now published in <em>eLife</em>.</p> <ul> <li><strong>"reproduce.md"</strong> includes all R code used to generate figures and perform analyses described in the paper.</li> <li><strong>"biorxiv_countries.postgres.backup"</strong> is a database snapshot that can be loaded into a PostgreSQL database to access all data collected and used in the study.</li> <li><strong>"schema.pdf"</strong> describes each field in each table of the database.</li> <li><strong>"manual_edits.sql"</strong> describes all corrections made to the automated inference of the country-level affiliations inferred for all authors.</li> <li><strong>"affiliation_corrections.csv"</strong> lists every unique affiliation string that was re-categorized after institutional corrections. The consequences of the corrections described in "manual_edits.sql."</li> <li><strong>"institution_corrections_summary.csv"</strong> summarizes "affiliation_corrections.csv" by listing each "before" and "after" correction one time. It is important to note that each before/after pair does not necessarily indicate that <em>every</em> affiliation string from the "before" institution was reassigned to the "after" institution, just that at least one affiliation string was switched from one to the other. <ul> <li>Note that the final two "corrections" files describe steps taken to correct the institution-level associations between authors and countries. The final set of corrections assigned authors to countries <strong>using heuristics that did not take institution-level accuracy into account</strong>.</li> </ul> </li> </ul> <p><strong>Version history:</strong></p> <ul> <li><strong>1.0.0: </strong>New files uploaded reflecting substantial corrections to the data, mostly linked to classification of authors and preprints previously without a country classification. (26 Jun 2020)</li> <li><strong>0.2.1:</strong> Added "schema.pdf" file, previously only in the manuscript.</li> <li><strong>0.2.0:</strong> Added new files "affiliation_corrections.csv" and "institution_corrections_summary.csv"</li> <li><strong>0.1.1: </strong>Database snapshot added.</li> <li><strong>0.1.0: </strong>First version with supplementary tables added.</li> </ul>
cweibel2018/More_adapted_species_have_higher_SD: Preprint code
<p>Data and code for <a href="https://www.biorxiv.org/content/10.1101/2020.10.15.341313v1">https://www.biorxiv.org/content/10.1101/2020.10.15.341313v1</a></p> <p>Includes raw data subsets from Phylostratigraphy Database published by <a href="https://www.biorxiv.org/content/10.1101/2020.03.26.010728v1">https://www.biorxiv.org/content/10.1101/2020.03.26.010728v1</a></p>
Dataset for article: Changes in evidence for studies assessing interventions for COVID-19 reported in preprints: meta-research study. BMC Med 18, 402 (2020).
<p>This dataset was used in the analyses reported in Oikonomidi, T., Boutron, I., Pierre, O. et al. Changes in evidence for studies assessing interventions for COVID-19 reported in preprints: meta-research study. BMC Med 18, 402 (2020). https://doi.org/10.1186/s12916-020-01880-8. </p> <p>This project is ancillary to the COVID-NMA Living systematic review and network meta-analysis of Covid-19 trials: https://covid-nma.com/</p> <p>The first spreadsheet entitled "key" includes the full description of the dataset.</p>
IMP HTML reports (preprint)
<p>ZIP file contains the HTML files which are referred to as the "Additional file 1" in the IMP manuscript.<br> </p>
Preprints from arXiv.org in cs and q-bio
<p>arXiv Quantitative Biology (q-bio)</p> <p>This dataset contains all preprints with the label “q-bio” from 2003 (when the section was introduced) to 2014. Downloaded on 10 June, 2016.</p> <p>arXiv CS (cs)</p> <p>This dataset contains all preprints with the label “cs” from 2003 to 2014. Downloaded on 10 June, 2016.</p>
ASAPbio Preprint Service Technical Workshop
<p>Video recording of the ASAPbio Technical Workshop held on August 30, 2016 at the American Academy of Arts and Sciences in Cambridge, MA. More information can be found at http://asapbio.org/asapbio-technical-workshop</p>
Preprint reviews per month
<p>Growth of preprint review over time. Preprints reviewed per month on Sciety as of 2023-10-30, excluding reviews conducted by automated tools (ScreenIT) and reviews by journals posted after publication of the journal version. The original Google Sheets file may be viewed here: https://docs.google.com/spreadsheets/d/1rjH2weFpeXjY5cGAU3s7wDzjKPuAkf_opjpcEo-_4wg/edit?usp=sharing</p>
Crossref relationships between preprints and journal articles
<p>This dataset contains preprint-journal article relationships deposited by publishers with Crossref and/or discovered by an automated preprint matching strategy. It includes preprints and journal articles deposited until the end of February 2025. The following fields are included:</p> <ul> <li>preprint DOI (string)</li> <li>journal article DOI (string)</li> <li>whether the publisher of the journal article deposited this relationship (boolean)</li> <li>whether the publisher of the preprint deposited this relationship (boolean)</li> <li>the confidence score returned by the strategy (float, empty if the strategy did not discover this relationship)</li> </ul> <p>The code of the preprint matching strategy is available <a href="https://gitlab.com/crossref/labs/marple/-/tree/main/strategies_available/preprint_sbmv">here</a>.</p>
Data for preprint: "Non-Telecentric 2P microscopy for 3D random access mesoscale imaging "
<p>Numerical data used in latest version of preprint: "Non-Telecentric 2P microscopy for 3D random access mesoscale imaging ", https://www.researchsquare.com/article/rs-121292/v1</p>
Publication times for COVID-19 articles and associated preprints
<p>These are two datasets of COVID-19 peer-reviewed journal and review articles and associated with them medRxiv and bioRxiv preprints. Datasets contain DOI of preprint and its published version, and publication times associated with each step: pre-submission time, review time, production stage time, and the elapsed time. One dataset was collected on 4 May 2021 and it includes a total of 4,031 deduplicated journal article-preprint pairs. Another dataset was collected on 19 October 2020 and it includes 1,099 journal article-preprint pairs. These files are associated with the submission of a manuscript "Publication Practices during the COVID-19 Pandemic: Expedited publishing or simply an early bird effect?" to the journal Learned Publishing in 2022.</p>
Data supporting the preprint: Soot and charcoal as reservoirs of extracellular DNA
<p>- Adsorption isotherms and kinetic data for adsorption of short and long DNA strands at soot and charcoal particles as a function of solution composition, pH and presence of competing compounds such as phosphates and alcohols.</p> <p>- Material characterisation analyses: XRD, XPS, Raman, water adsorption, zeta potential, mass titration</p>
Forecasting the publication and citation outcomes of Covid-19 preprints
<p>The scientific community reacted quickly to the <em>Covid-19</em> pandemic in 2020, generating an unprecedented increase in publications. Many of these publications were released on preprint servers such as <em>medRxiv</em> and <em>bioRxiv</em>. It is unknown however how reliable these preprints are, and if they will eventually be published in scientific journals. In this study, we use crowdsourced human forecasts to predict publication outcomes and future citation counts for a sample of 400 preprints with high <em>Altmetric</em> scores. Most of these preprints were published within one year of upload on a preprint server (70%), and 46% of the published preprints appeared in a high-impact journal with a Journal Impact Factor of at least 10. On average, the preprints received 162 citations within the first year. We found that forecasters can predict if preprints will be published after one year and if the publishing journal has high impact. Forecasts are also informative with respect to preprints' rankings in terms of <em>Google</em> <em>Scholar</em> citations within one year of upload on a preprint server. For both types of assessment, we found statistically significant positive correlations between forecasts and observed outcomes. While the forecasts can help to provide a preliminary assessment of preprints at a faster pace than the traditional peer-review process, it remains to be investigated if such an assessment is suited to identify methodological problems in pre-prints. </p>
Preprint Infrastructure
<p>Diagram showing preprint infrastructure, including relationships between objects, actors, tools/services, and policies/practices.</p>
Data and R code used in the preprint "Parasite intensity is driven by temperature in a wild bird" (doi 10.1101/323311)
<p>Data (as text files) and R code used in the preprint entitled "Parasite intensity is driven by temperature in a wild bird", recommended by <em>Peer Community In Ecology </em>(doi 10.1101/323311). See preprint and supplementary material; some explanations are also included in the R code. </p>
Model sets and data used in the preprint "Evaluating functional dispersal and its eco-epidemiological implications in a nest ectoparasite"
<p>Model sets and data used in the preprint "Evaluating functional dispersal and its eco-epidemiological implications in a nest ectoparasite", reviewed and recommended by Peer Community In Ecology (https://dx.doi.org/10.24072/pci.ecology.100013). See preprint and supplementary materials.</p>
bioRxiv preprint and publication details, 2014-2023
<p>bioRxiv preprint and publication details, 2014-2023</p> <p>Details at: <a href="https://blog.stephenturner.us/p/exploring-the-biorxiv-api-with-r-httr2-rvest-tidytext-datawrapper" target="_blank" rel="noopener">https://blog.stephenturner.us/p/exploring-the-biorxiv-api-with-r-httr2-rvest-tidytext-datawrapper</a></p> <p>Code at: <a href="https://gist.github.com/stephenturner/e1487c90a98e6d5805a3211f0140e198" target="_blank" rel="noopener">https://gist.github.com/stephenturner/e1487c90a98e6d5805a3211f0140e198</a></p> <p>Original data pulled from the bioRxiv API: <a href="https://api.biorxiv.org/" target="_blank" rel="noopener">https://api.biorxiv.org/</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.