Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
26
datasets available to search
ShareScore release 0.9.0
Dataset results
26 results for “Web Access”
Annual Article Processing Charges (APCs) and number of gold and hybrid open access articles in Web of Science indexed journals published by Elsevier, Sage, Springer-Nature, Taylor & Francis and Wiley 2015-2018
<p><strong>Dataset of annual Article Processing Charges (APCs) for 6,252 journals from 2015 to 2018. </strong>The dataset contains annual APCs for journals indexed in the Web of Science (WoS) and published by the oligopoly of academic publishers (Elsevier, Sage, Springer-Nature, Taylor & Francis, Wiley). It also includes an estimate of the total APCs paid by the academic community based on the number of gold and hybrid articles published between 2015 and 2018. The dataset was created using publication data from WoS, OA status from Unpaywall and annual APC prices from open datasets (<a href="https://doi.org/10.5281/ZENODO.3841568">Matthias, 2020</a>; <a href="https://doi.org/10.5683/SP2/84PNSG">Morrison, 2021</a>) and historical fees retrieved via the Internet Archive Wayback Machine. </p> <p>Detailed methods and findings are reported in the following journal article</p> <p>Butler, L.-A., Matthias, L., Simard, M.-A., Mongeon, P., & Haustein, S. (2023). The Oligopoly's Shift to Open Access. How the Big Five Academic Publishers Profit from Article Processing Charges. <em>Quantitative Science Studies</em>. Preprint: <a href="https://doi.org/10.5281/zenodo.8322555">https://doi.org/10.5281/zenodo.8322555</a></p> <p><strong>Description of included files (v1):</strong></p> <p><em>APCs.csv: </em>contains the annual APCs for gold and hybrid OA journals indexed in Web of Science published by the oligopoly of academic publishers (Elsevier, Sage, Springer-Nature, Taylor & Francis, Wiley) between 2015 and 2018 including the total estimate of APCs paid per journal per year. It contains APC data for 18,846 journal-year-OA status combinations.</p> <p><em>countries.csv</em>: contains the fractionalized number of annual gold and hybrid OA articles by oligopoly publishers between 2015 and 2018 and the total estimate of fractionalized APCs paid per country per journal per year.</p> <p><em>oecd.csv</em>: contains the fractionalized number of annual gold and hybrid OA articles by oligopoly publishers between 2015 and 2018 and the total estimate of fractionalized APCs per discipline per journal per year.</p> <p><em>ReadMe.csv</em>: contains a description of the variables used in <em>APCs.csv</em>, <em>countries.csv</em> and <em>oecd.csv</em>.</p> <p> </p>
Ecological data for: Subsidy accessibility drives asymmetric food web responses
<p>Global change is fundamentally altering flows of natural and anthropogenic subsidies across space and time. After a pointed call for research on subsidies in the 1990s, an industry of empirical work has documented the ubiquitous role subsidies play in ecosystem structure, stability and function. Here, we argue that physical constraints (e.g., water temperature) and species traits can govern a species' accessibility to resource subsidies, which has been largely overlooked in the subsidy literature. We examined the input of a high quality, point-source anthropogenic subsidy (aquaculture feed) into a recipient freshwater lake food web. By using a combined bio-tracer approach, we detect a gradient in accessibility of the anthropogenic subsidy within the surrounding food web driven by the thermal preferences of three constituent species, effectively rewiring the recipient lake food web. Since aquaculture is predicted to increase significantly in coming decades to support growing human populations, and global change is altering temperature regimes, then this form of food web alteration may be expected to occur frequently. We argue that subsidy accessibility is a key characteristic of recipient food web interactions that must be considered when trying to understand the impacts of subsidies on ecosystem stability and function under continued global change.</p>
Data from: Open access levels: a quantitative exploration using Web of Science and oaDOI data
<p>This is the raw data behind the publication (on PeerJ Preprints):</p> <p><strong>Open access levels: a quantitative exploration using Web of Science and oaDOI data</strong></p> <p>Across the world there is growing interest in open access publishing among researchers, institutions, funders and publishers alike. It is assumed that open access levels are growing, but hitherto the exact levels and patterns of open access have been hard to determine and detailed quantitative studies are scarce. Using newly available open access status data from oaDOI in Web of Science we are now able to explore year-on-year open access levels across research fields, languages, countries, institutions, funders and topics, and try to relate the resulting patterns to disciplinary, national and institutional contexts. With data from the oaDOI API we also look at the detailed breakdown of open access by types of gold open access (pure gold, hybrid and bronze), using universities in the Netherlands as an example. There is huge diversity in open access levels on all dimensions, with unexpected levels for e.g. Portuguese as language, Astronomy & Astrophysics as research field, countries like Tanzania, Peru and Latvia, and Zika as topic. We explore methodological issues and offer suggestions to improve conditions for tracking open access status of research output. Finally, we suggest potential future applications for research and policy development. We have shared all data and code openly.</p>
Prototyping 3D Virtual Learning Environments with X3D-based Content and Visualization Tools-Figure 11. Online accessible repository of digital data on cultural heritage with X3D models (STARC Web Repository, 2017, © Copyright 2017, STARC, Cyprus Institute. Used with permission)
<p>Prototyping can also include the development of toolkits for automatic content generation simulator, but in the case of an architectural environment, the components are too complex to be automatically generated. Furniture elements or the learning artifacts (i.e. content created by learners) can be converted to be viewed in X3D compatible browsers or included in online galleries (Figure 11). After functional and 3D content prototyping, certain components of the virtual campus can be easily modified and adapted as needed.</p>
The mOTUs online database provides web-accessible genomic context to taxonomic profiling of microbial communities - Supplementary Tables
<p><strong>Supplementary Table 1:</strong></p> <p>A map between each of the genomes in mOTUs-db (3’747’151), the associated study and its metagenomic sample (in case of MAGs).</p> <p>Columns:</p> <p><code> GENOME → Unique mOTUs-db name of the genome</code><br><code> STUDY → Unique mOTUs-db name of the study</code><br><code> IS_MAG → True if genome is a MAG, otherwise False </code><br><code> METAGENOMIC_SAMPLE → Unique name of the metagenomic sample or NA in case of non-MAG genome</code></p> <p>Example:</p> <p><code> GENOME STUDY IS_MAG METAGENOMIC_SAMPLE</code><br><code> ---------------------------------------------------------------------------------------------</code><br><code> ACIN21-1_SAMN05421555_MAG_00000001 ACIN21-1 True ACIN21-1_SAMN05421555_METAG</code><br><code> RSGB23-1_GCA-006096615-V1_GENO_10000001 RSGB23-1 False NA</code></p> <p><strong>Supplementary Table 2:</strong></p> <p>A map between all non-MAG genomes (919’090) and their source (e.g. Refseq or JGI).</p> <p>Columns:</p> <p><code> GENOME → Unique mOTUs-db name of the genome</code><br><code> SOURCE_SAMPLE_LINK → Link to the original location of this genome</code></p> <p>Example:</p> <p><code> #GENOME SOURCE_SAMPLE_LINK</code><br><code> --------------------------------------------------------------------------------------------------------</code><br><code> JGIG23-1_GA0055041_GENO_10000001 https://gold.jgi.doe.gov/analysis_project?id=Ga0055041</code><br><code> RSGB23-1_GCA-006717865-V1_GENO_10000001 https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_006717865.1</code></p> <p><strong>Supplementary Table 3:</strong></p> <p>A list of all metagenomic studies processed for the mOTUs-db, their number of samples, the number of reconstructed MAGs and the associated publication.</p> <p>Columns:</p> <p><code> STUDY --> Unique mOTUs-db study identifier</code><br><code> BIOPROJECT --> Public identifier (NCBI/JGI) of metagenomic sequencing project</code><br><code> SAMPLES --> Number of metagenomic samples</code><br><code> MAGs --> Number of reconstructed MAGs</code><br><code> PUBLICATION --> Link to publication</code></p> <p>Example:</p> <p><code> STUDY BIOPROJECT SAMPLES MAGs PUBLICATION</code><br><code> -------------------------------------------------------------------------------------------------</code><br><code> ACIN21-1 PRJEB44456 58 1,110 https://www.nature.com/articles/s42003-021-02112-2</code></p> <p><strong>Supplementary Table 4:</strong></p> <p>Mapping between mOTUs-db sample identifier, the associated biosample and the environment.</p> <p>Columns:</p> <p><code> SAMPLE --> Unique mOTUS-db sample identifier</code><br><code> BIOSAMPLE --> Public identifier (NCBI/JGI) of metagenomic sample</code><br><code> STUDY --> Unique mOTUs-db study identifier</code><br><code> ENVIRONMENT --> Environment of metagenomic sample</code><br><code> SOURCE_SAMPLE_LINK --> Link to the original location of this sample</code></p> <p>Example:</p> <p><code> #SAMPLE BIOSAMPLE STUDY ENVIRONMENT SOURCE_SAMPLE_LINK</code><br><code> ---------------------------------------------------------------------------------------------------------------------</code><br><code> ACIN21-1_SAMN05421555_METAG SAMN05421555 ACIN21-1 marine https://www.ncbi.nlm.nih.gov/biosample/SAMN05421555/</code></p> <p><strong>Supplementary Table 5:</strong></p> <p>A list of environments covered in the mOTUs-db mapped to the respective NCBI taxonomy (if possible)</p> <p>Columns:</p> <p><code> TERM --> Unique environment name</code><br><code> NCBI TAXONOMY ID --> Link to the NCBI taxonomy</code></p> <p>Example:</p> <p><code> TERM NCBI TAXONOMY ID</code><br><code> ----------------------------------------------</code><br><code> activated sludge metagenome NCBI:txid942017</code><br><code> air metagenome NCBI:txid655179</code></p>
Ecological data for: Subsidy accessibility drives asymmetric food web responses
Open the record for dataset details and reuse information.
Interviews on Current Practices for Describing and Providing Access to UK Public Sector Web Archives
<p>This dataset contains qualitative interview data which investigated current practice for describing and providing access to UK Public Sector Web Archives. Participants included staff responsible for the management and curation of the following web archives:</p><ul><li>UK Web Archive (four of the six Legal Deposit libraries: the British Library, Bodleian Libraries, Cambridge University Library, and the National Library of Scotland)</li><li>UK Government Web Archive (The National Archives)</li><li>UK Parliament Web Archive (Parliamentary Archives)</li><li>NRS Web Archive (National Records of Scotland) and</li><li>PRONI Web Archive (Public Record Office Northern Ireland).</li></ul><p>Available to the public are the University of Dundee (UoD) ethics application for this study, including the research data management plan and information provided to organisations before participating in the study. The report of interview codes and code groups (the 'Codebook') demonstrates the connections made across responses. This is supplemented by a redacted report of quotations by code, organised by code group and document.</p><p>This qualitative interview data, and subsequent analysis, forms the basis of the Masters thesis 'Web Archives for All? Towards Equitable Access to UK Public Sector Web Archives' submitted as part of the MLitt Archives and Records Management at the University of Dundee. </p>
Web Accessibility Evaluation Data - March to September 2021
<p>This dataset contains the dump of a PostgreSQL database with the contents of automated web accessibility evaluation of nearly 3 million webpages obtained with QualWeb between March and September 2021.</p>
Nipah virus: Analysis of the scientific production in Open Access on the Web of Science, 2000 - 2020
<p>Data used in the analysis of the scientific production of Nipah in open access in the period 2000 to 2020</p>
Figure 5 in Actionable, long-term stable and semantic web compatible identifiers for access to biological collection objects
Figure 5. The Wallich Catalogue. Screenshot of Wallich Catalogue hosted by Royal Botanic Garden Edinburgh showing popup for stable URI containing information hosted at Botanic Garden and Botanical Museum Berlin-Dahlem.
Figure 4 in Actionable, long-term stable and semantic web compatible identifiers for access to biological collection objects
Figure 4. The CETAF Specimen URI Tester provides for any given Specimen URI an overview of the redirection process as well as a preview of machine-readable and human-readable data associated with the URI.
Figure 3 in Actionable, long-term stable and semantic web compatible identifiers for access to biological collection objects
Figure 3. CETAF stable HTTP URIs in the GBIF data portal. The Global Biodiversity Information Facility (GBIF) publishes CETAF stable HTTP URIs via their data portal.
Figure 2 in Actionable, long-term stable and semantic web compatible identifiers for access to biological collection objects
Figure 2. Basic redirection mechanisms. Human users are redirected to a human-readable web-representation of the specimens. Software systems are re-directed to a machine-readable metadata record.
Accessibility Issues in Ad-Driven Web Appliactions
Open the record for dataset details and reuse information.
Ditte Laursen: Developing a legal agreement for research-access to a web archive
<p>Ditte Laursen: Developing a legal agreement for research-access to a web archive</p> <p>WARcnet Luxembourg meeting Thursday 5 November 2020</p>
Web Accessible Population Pharmacokinetics Service - Hemophilia: Sources of Variability
ClinicalTrials.gov study NCT03533504. IPD Sharing: NO. Countries: 1. Publications: 2.
Primary Care-based Facilitated Access to a Web Based Brief Intervention to Reduce Alcohol Consumption
ClinicalTrials.gov study NCT02082990. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Freely Accessible eJournals web archive collection derivatives
<p>Web archive derivatives of the <a href="https://archive-it.org/collections/5921">Freely Accessible eJournals</a> collection from <a href="https://archive-it.org/home/Columbia">Columbia University Libraries</a>. The derivatives were created with the <a href="https://github.com/archivesunleashed/aut/">Archives Unleashed Toolkit</a> and <a href="https://cloud.archivesunleashed.org/">Archives Unleashed Cloud</a>.</p> <p>The <strong>cul-5921-parquet.tar.gz</strong> derivatives are in the <a href="https://parquet.apache.org/">Apache Parquet format</a>, which is a <a href="http://en.wikipedia.org/wiki/Column-oriented_DBMS">columnar storage</a> format. These derivatives are generally small enough to work with on your local machine, and can be easily converted to Pandas DataFrames. See <a href="https://github.com/archivesunleashed/notebooks/blob/master/datathon-nyc/parquet_pandas_stonewall.ipynb">this</a> notebook for examples.</p> <p><strong>Domains</strong></p> <pre><code class="language-java">.webpages().groupBy(ExtractDomainDF($"url").alias("url")).count().sort($"count".desc)</code></pre> <p>Produces a DataFrame with the following columns:</p> <ul> <li>domain</li> <li>count</li> </ul> <p><strong>Web Pages</strong></p> <pre><code class="language-java">.webpages().select($"crawl_date", $"url", $"mime_type_web_server", $"mime_type_tika", RemoveHTMLDF(RemoveHTTPHeaderDF(($"content"))).alias("content"))</code></pre> <p>Produces a DataFrame with the following columns:</p> <ul> <li>crawl_date</li> <li>url</li> <li>mime_type_web_server</li> <li>mime_type_tika</li> <li>content</li> </ul> <p><strong>Web Graph</strong></p> <pre><code class="language-java">.webgraph()</code></pre> <p>Produces a DataFrame with the following columns:</p> <ul> <li>crawl_date</li> <li>src</li> <li>dest</li> <li>anchor</li> </ul> <p><strong>Image Links</strong></p> <pre><code class="language-java">.imageLinks()</code></pre> <p>Produces a DataFrame with the following columns:</p> <ul> <li>src</li> <li>image_url</li> </ul> <p><a href="https://github.com/archivesunleashed/aut-docs/blob/master/current/binary-analysis.md#binary-analysis"><strong>Binary Analysis</strong></a></p> <ul> <li>Audio</li> <li>Images</li> <li>PDFs</li> <li>Presentation program files</li> <li>Spreadsheets</li> <li>Text files</li> <li>Word processor files<br> </li> </ul> <p>The <strong>cul-12143-auk.tar.gz </strong>derivatives<strong> </strong>are the <a href="https://cloud.archivesunleashed.org/derivatives">standard set of web archive derivatives</a> produced by the Archives Unleashed Cloud.</p> <ul> <li><strong>Gephi </strong>file, which can be loaded into <a href="https://gephi.org/">Gephi</a>. It will have basic characteristics already computed and a basic layout.</li> <li><strong>Raw Network</strong> file, which can also be loaded into <a href="https://gephi.org/">Gephi</a>. You will have to use that network program to lay it out yourself.</li> <li><strong>Full text</strong> file. In it, each website within the web archive collection will have its full text presented on one line, along with information around when it was crawled, the name of the domain, and the full URL of the content.</li> <li><strong>Domains count</strong> file. A text file containing the frequency count of domains captured within your web archive.</li> </ul>
Data from: IKey+: a new single-access key generation web service
Single-access keys are a major tool for biologists who need to identify specimens. The construction process of these keys is particularly complex (especially if the input dataset is large) so having an automatic single-access key generation tool is essential. As part of the European project ViBRANT, our aim was to develop such a tool as a web service, thus allowing end-users to integrate it directly into their workflow. IKey+generates single-access keys on demand, for single users or research institutions. It receives user input data (using the standard SDD format), accepts several key-generation parameters (affecting the key topology and representation), and supports several output formats. IKey+ is freely available (sources and binary packages) at www.identificationkey.fr. Furthermore, it is deployed on our server and can be queried (for testing purposes) via a simple web client also available at www.identificationkey.fr. Finally a client plugin will be integrated to the Scratchpads biodiversity networking tool (scratchpads.eu).
Figure 1 in Actionable, long-term stable and semantic web compatible identifiers for access to biological collection objects
Figure 1. Example specimen. Example physical herbarium object and its stable HTTP URI identifier.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.