Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

45

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

45 results for “ORF”

Learn how ShareScore rates datasets ↗
zenodo40/100

Similarity of ORF genes grouped by chromosome without chromosome arm correction

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo36/100

Average fold coverage and annotation results of those ORF that were increasing over treatment time

<p>Table listing those ORF that presented an increased frequency across the treatment time</p> <p>This a supplementary material for the doctoral thesis entitle: <em><strong>Understanding microbiome supra-metabolism responses under strong selective pressures as a resource for designing synthetic gene arrangements encoding key co-selected functions for environmental biotechnology applications. </strong></em>Villegas-Plazas M, 2020. Universidad del Valle. Cali, Colombia</p>

opencc-by-4.0May 2020View details →
zenodo36/100

ProTInSeq: transposon insertion tracking by ultra-deep DNA sequencing applied to identify small and large translated ORFs

<p>ProTInSeq is a novel -omics technique designed to characterize proteomes by using DNA ultra-deep sequencing. The technique is based on transposons engineered to have a positive or negative protein selection marker expressed when the transposon is inserted in-frame into a protein-coding gene. In the genome-reduced bacterium Mycoplasma pneumoniae, ProTInSeq identifies 80% of known expressed proteins, as well as 5 new open reading frames (ORFs; &gt;100 amino acids); and 153 novel small ORF-encoded proteins (SEPs; &le;100 aa) that represent up to 18% of this bacterium&rsquo;s proteome. ProTInSeq can be used to detect translational noise, for protein quantification and to provide insight into functional protein aspects such as relative half-life, stability, and membrane topology. Herein, we describe a methodology that can be easily implemented in any living system and allows the deep understanding of proteomes and more importantly the identification of small proteins by DNA ultra-sequencing.</p> <p>We include the following files:</p> <p>- processed_inscalling.zip: output obtain after running FASTQINS transposon calling tool over the datasets found at&nbsp;<a href="https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-10380?key=5f54209d-ce59-490a-9bcd-7084e9c619ee">https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-10380?key=5f54209d-ce59-490a-9bcd-7084e9c619ee</a>. This include every genome position in <em>M. pneumoniae</em>, the number of times an insertion has been mapped to that position and the total read count value.</p> <p>- separated_library_metrics.zip: insertion and read count processed from&nbsp;processed_inscalling files associated to every ORF&nbsp; and intergenic region in&nbsp;<em>M.&nbsp;penumoniae. </em>Columns include frame measured (0 - whole gene, 1 - in-frame, 2 and 3 for following positions) and metric. Metrics account for number insertions in-frame (<em>I</em>), read count (<em>R</em>), linear density from non-coding regions used in the Poisson evaluation (<em>rNC</em>), probability measured (<em>sfNC</em>) and a binary for prediction (<em>pred</em>; 0 - no significant, 1 - significant).</p> <p>- allmetrics.xlsx: merged table with the combination of results from separated_library_metrics.zip tab-delimited files.</p> <p>- selective_metrics_allannotations.xlsx:&nbsp; this table includes all the available information about the 30,112 sequences that could encode for a coding sequence in <em>M. pneumoniae</em>. For each identifier (column B), we include coordinates information and nucleotide and amino acid length information (columns C-H). Column I includes the gene name when the entry is found annotated in <em>M. pneumoniae</em>. Localization and function are described in columns J and K. Column L includes the operon number in which the annotation would be expressed. We also included transcription-related information average expression (column M; as log2(gene read count/gene length) and estimated average RNA copies per cell (column N) considering 4 RNA sequencing samples covering different growth times (6, 24 and 48 hours, ArrayExpress identifier E‐MTAB‐6203). Column O accounts for the number of mass spectrometry experiments detecting that entry (to a maximum of 116) and column P accounts for the total number of unique tryptic peptides detected. This comes <a href="https://paperpile.com/c/BImj5N/eljz">[5]</a>, available for 12,426 sequences that present an amino acid length &ge;19 (from 116 mass spectrometry experiments, ID PRIDE: PXD008243). Columns Q to T recapitulate protein copies per cell under different conditions (overall, extracting with urea, extracting with SDS and mean, respectively). Column U includes half-lives of the proteins. Columns V and W describe the reference density of insertion and essentiality assigned in previous studies. Column X-AA includes the predicted RanSEPs score, ribosome binding site presence, homology into seven groups: 0&mdash;no hits passed the thresholds defined; 1&mdash;conserved with an annotated function; 2&mdash;conserved as an annotated SEP in NCBI but no associated function; 3&mdash;conserved in a different species but target and homologous sequence not found in NCBI; 4&mdash;sequence is completely or partially (&gt; 75%) repeated &ge; 3 times in the reference genome; 5&mdash;potential pseudogene; and 6&mdash;to depict those annotations that are found in the reference NCBI annotation file. , and function expected by homology, respectively. Columns AB to AD cover the output provided by Phobius, including the number of transmembrane segments, presence of signal peptide and transmembrane topology predicted by TM-HMM. Column AE includes the complex information where 1 implies that entry is functional as a monomer, 2 as dimer, and so on. Finally, columns AF-AH will be 1 if the protein is a Lon protease target, a lipoprotein, and/or a truncated gene or pseudogene, respectively, 0 otherwise.&nbsp;Following columns include for every sample presenting selective insertion rates in-frame using the following identifiers separated by underscores: marker (BarnB, Cm or Ery), type (control-AC or selection-BD, antibiotic concentration, sample replicate, frame measured, metric. Metrics account for number insertions in-frame (<em>I</em>), read count (<em>R</em>), linear density from non-coding regions used in the Poisson evaluation (<em>rNC</em>), probability measured (<em>sfNC</em>) and a binary for prediction (<em>pred</em>; 0 - no significant, 1 - significant). Last columns combine the number of samples each annotation has been identified. Notice for barnase library the results need to be interpreted considering it is a negative selection marker inverting the 0 and 1 meaning.</p>

openOct 2023View details →
dryad36/100

Metatranscriptomic analysis uncovers prevalent viral ORFs compatible with mitochondrial translation

Open the record for dataset details and reuse information.

publicMay 2025View details →
dryad36/100

Upstream ORF-Encoded ASDURF Is a Novel Prefoldin-like Subunit of the PAQosome

Open the record for dataset details and reuse information.

publicNov 2019View details →
zenodo32/100

Phenotypically active ORF and CRISPR consensus profiles

This hosts the files on phenotypically active genes for JUMP, derived from https://zenodo.org/records/14025602 after filtering samples with no matches above a given threshold (see filenames for thresholds).

opencc-zeroNov 2024View details →
zenodo32/100

Structure prediction from SARS-CoV-2 accessory proteins ORF-6

<p>Structure prediction made with Collabfold for SARS-CoV-2 accessory protein ORF-6.</p> <p>The archive contains both the structure and</p>

opencc-by-4.0Nov 2022View details →
zenodo32/100

Structure prediction from SARS-CoV-2 accessory proteins ORF-7B

<p>Structure prediction made with Collabfold for SARS-CoV-2 accessory protein ORF-7B.</p> <p>The archive contains both the structure and the logs from the prediction.</p>

opencc-by-4.0Nov 2022View details →
ClinicalTrials.gov32/100

A Study Assessing the Safety, Tolerability, Immunogenicity of COVID-19 Vaccine Candidate PRIME-2-CoV_Beta, Orf Virus Expressing SARS-CoV_2 Spike and Nucleocapsid Proteins

ClinicalTrials.gov study NCT05367843. IPD Sharing: Not stated. Countries: 2. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo28/100

Phenotypically active ORF and CRISPR consensus profiles

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
ClinicalTrials.gov28/100

Pharmacokinetics and Safety of ORF Tablets in Pediatric Patients

ClinicalTrials.gov study NCT01160614. IPD Sharing: Not stated. Countries: 4. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
geo24/100

Pan-viral ORFs discovery using massively parallel ribosome profiling

GEO Series GSE272406. Homo sapiens; synthetic construct. 7 samples. Type: Other.

openGEO-OpenJun 2025View details →
geo24/100

smORFer: a modular algorithm to detect small ORFs in prokaryotes

GEO Series GSE150601. Staphylococcus aureus. 2 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMay 2021View details →
geo24/100

DAP5 enables main ORF translation on mRNAs with structured and uORF-containing 5’ leaders

GEO Series GSE155854. Homo sapiens. 8 samples. Type: Expression profiling by high throughput sequencing; Other.

openGEO-OpenNov 2022View details →
geo24/100

Telomeric ORFs (TLOs) in Candida spp. encode Mediator subunits that regulate distinct virulence traits

GEO Series GSE60173. Candida dubliniensis. 3 samples. Type: Genome binding/occupancy profiling by genome tiling array.

openGEO-OpenAug 2014View details →
geo24/100

Developmental regulation of Canonical and small ORF translation from mRNAs

GEO Series GSE147619. Drosophila melanogaster. 13 samples. Type: Other.

openGEO-OpenMar 2020View details →
geo24/100

Exploring the molecular mechanisms of positive and negative regulation of apoptosis post Orf virus infection in sheep

GEO Series GSE95203. Ovis aries. 8 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenFeb 2018View details →
geo24/100

Identification of small ORFs in vertebrates using ribosome footprinting and evolutionary conservation

GEO Series GSE53693. Danio rerio. 30 samples. Type: Expression profiling by high throughput sequencing; Non-coding RNA profiling by high throughput sequencing.

openGEO-OpenApr 2014View details →
ClinicalTrials.gov24/100

A Study of BBP-711 (ORF-229) in Healthy Adult Volunteers

ClinicalTrials.gov study NCT04876924. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
geo20/100

Genomewide demarcation of RNA PolII transcription units by physical fractionation of chromatin- ORF enrichment

GEO Series GSE5649. Saccharomyces cerevisiae. 5 samples. Type: Other.

openGEO-OpenAug 2006View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record