Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
45
datasets available to search
ShareScore release 0.9.0
Dataset results
45 results for “ORF”
Similarity of ORF genes grouped by chromosome without chromosome arm correction
Open the record for dataset details and reuse information.
Average fold coverage and annotation results of those ORF that were increasing over treatment time
<p>Table listing those ORF that presented an increased frequency across the treatment time</p> <p>This a supplementary material for the doctoral thesis entitle: <em><strong>Understanding microbiome supra-metabolism responses under strong selective pressures as a resource for designing synthetic gene arrangements encoding key co-selected functions for environmental biotechnology applications. </strong></em>Villegas-Plazas M, 2020. Universidad del Valle. Cali, Colombia</p>
ProTInSeq: transposon insertion tracking by ultra-deep DNA sequencing applied to identify small and large translated ORFs
<p>ProTInSeq is a novel -omics technique designed to characterize proteomes by using DNA ultra-deep sequencing. The technique is based on transposons engineered to have a positive or negative protein selection marker expressed when the transposon is inserted in-frame into a protein-coding gene. In the genome-reduced bacterium Mycoplasma pneumoniae, ProTInSeq identifies 80% of known expressed proteins, as well as 5 new open reading frames (ORFs; >100 amino acids); and 153 novel small ORF-encoded proteins (SEPs; ≤100 aa) that represent up to 18% of this bacterium’s proteome. ProTInSeq can be used to detect translational noise, for protein quantification and to provide insight into functional protein aspects such as relative half-life, stability, and membrane topology. Herein, we describe a methodology that can be easily implemented in any living system and allows the deep understanding of proteomes and more importantly the identification of small proteins by DNA ultra-sequencing.</p> <p>We include the following files:</p> <p>- processed_inscalling.zip: output obtain after running FASTQINS transposon calling tool over the datasets found at <a href="https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-10380?key=5f54209d-ce59-490a-9bcd-7084e9c619ee">https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-10380?key=5f54209d-ce59-490a-9bcd-7084e9c619ee</a>. This include every genome position in <em>M. pneumoniae</em>, the number of times an insertion has been mapped to that position and the total read count value.</p> <p>- separated_library_metrics.zip: insertion and read count processed from processed_inscalling files associated to every ORF and intergenic region in <em>M. penumoniae. </em>Columns include frame measured (0 - whole gene, 1 - in-frame, 2 and 3 for following positions) and metric. Metrics account for number insertions in-frame (<em>I</em>), read count (<em>R</em>), linear density from non-coding regions used in the Poisson evaluation (<em>rNC</em>), probability measured (<em>sfNC</em>) and a binary for prediction (<em>pred</em>; 0 - no significant, 1 - significant).</p> <p>- allmetrics.xlsx: merged table with the combination of results from separated_library_metrics.zip tab-delimited files.</p> <p>- selective_metrics_allannotations.xlsx: this table includes all the available information about the 30,112 sequences that could encode for a coding sequence in <em>M. pneumoniae</em>. For each identifier (column B), we include coordinates information and nucleotide and amino acid length information (columns C-H). Column I includes the gene name when the entry is found annotated in <em>M. pneumoniae</em>. Localization and function are described in columns J and K. Column L includes the operon number in which the annotation would be expressed. We also included transcription-related information average expression (column M; as log2(gene read count/gene length) and estimated average RNA copies per cell (column N) considering 4 RNA sequencing samples covering different growth times (6, 24 and 48 hours, ArrayExpress identifier E‐MTAB‐6203). Column O accounts for the number of mass spectrometry experiments detecting that entry (to a maximum of 116) and column P accounts for the total number of unique tryptic peptides detected. This comes <a href="https://paperpile.com/c/BImj5N/eljz">[5]</a>, available for 12,426 sequences that present an amino acid length ≥19 (from 116 mass spectrometry experiments, ID PRIDE: PXD008243). Columns Q to T recapitulate protein copies per cell under different conditions (overall, extracting with urea, extracting with SDS and mean, respectively). Column U includes half-lives of the proteins. Columns V and W describe the reference density of insertion and essentiality assigned in previous studies. Column X-AA includes the predicted RanSEPs score, ribosome binding site presence, homology into seven groups: 0—no hits passed the thresholds defined; 1—conserved with an annotated function; 2—conserved as an annotated SEP in NCBI but no associated function; 3—conserved in a different species but target and homologous sequence not found in NCBI; 4—sequence is completely or partially (> 75%) repeated ≥ 3 times in the reference genome; 5—potential pseudogene; and 6—to depict those annotations that are found in the reference NCBI annotation file. , and function expected by homology, respectively. Columns AB to AD cover the output provided by Phobius, including the number of transmembrane segments, presence of signal peptide and transmembrane topology predicted by TM-HMM. Column AE includes the complex information where 1 implies that entry is functional as a monomer, 2 as dimer, and so on. Finally, columns AF-AH will be 1 if the protein is a Lon protease target, a lipoprotein, and/or a truncated gene or pseudogene, respectively, 0 otherwise. Following columns include for every sample presenting selective insertion rates in-frame using the following identifiers separated by underscores: marker (BarnB, Cm or Ery), type (control-AC or selection-BD, antibiotic concentration, sample replicate, frame measured, metric. Metrics account for number insertions in-frame (<em>I</em>), read count (<em>R</em>), linear density from non-coding regions used in the Poisson evaluation (<em>rNC</em>), probability measured (<em>sfNC</em>) and a binary for prediction (<em>pred</em>; 0 - no significant, 1 - significant). Last columns combine the number of samples each annotation has been identified. Notice for barnase library the results need to be interpreted considering it is a negative selection marker inverting the 0 and 1 meaning.</p>
Metatranscriptomic analysis uncovers prevalent viral ORFs compatible with mitochondrial translation
Open the record for dataset details and reuse information.
Upstream ORF-Encoded ASDURF Is a Novel Prefoldin-like Subunit of the PAQosome
Open the record for dataset details and reuse information.
Phenotypically active ORF and CRISPR consensus profiles
This hosts the files on phenotypically active genes for JUMP, derived from https://zenodo.org/records/14025602 after filtering samples with no matches above a given threshold (see filenames for thresholds).
Structure prediction from SARS-CoV-2 accessory proteins ORF-6
<p>Structure prediction made with Collabfold for SARS-CoV-2 accessory protein ORF-6.</p> <p>The archive contains both the structure and</p>
Structure prediction from SARS-CoV-2 accessory proteins ORF-7B
<p>Structure prediction made with Collabfold for SARS-CoV-2 accessory protein ORF-7B.</p> <p>The archive contains both the structure and the logs from the prediction.</p>
A Study Assessing the Safety, Tolerability, Immunogenicity of COVID-19 Vaccine Candidate PRIME-2-CoV_Beta, Orf Virus Expressing SARS-CoV_2 Spike and Nucleocapsid Proteins
ClinicalTrials.gov study NCT05367843. IPD Sharing: Not stated. Countries: 2. Publications: 1.
Phenotypically active ORF and CRISPR consensus profiles
Open the record for dataset details and reuse information.
Pharmacokinetics and Safety of ORF Tablets in Pediatric Patients
ClinicalTrials.gov study NCT01160614. IPD Sharing: Not stated. Countries: 4. Publications: 0.
Pan-viral ORFs discovery using massively parallel ribosome profiling
GEO Series GSE272406. Homo sapiens; synthetic construct. 7 samples. Type: Other.
smORFer: a modular algorithm to detect small ORFs in prokaryotes
GEO Series GSE150601. Staphylococcus aureus. 2 samples. Type: Expression profiling by high throughput sequencing.
DAP5 enables main ORF translation on mRNAs with structured and uORF-containing 5’ leaders
GEO Series GSE155854. Homo sapiens. 8 samples. Type: Expression profiling by high throughput sequencing; Other.
Telomeric ORFs (TLOs) in Candida spp. encode Mediator subunits that regulate distinct virulence traits
GEO Series GSE60173. Candida dubliniensis. 3 samples. Type: Genome binding/occupancy profiling by genome tiling array.
Developmental regulation of Canonical and small ORF translation from mRNAs
GEO Series GSE147619. Drosophila melanogaster. 13 samples. Type: Other.
Exploring the molecular mechanisms of positive and negative regulation of apoptosis post Orf virus infection in sheep
GEO Series GSE95203. Ovis aries. 8 samples. Type: Expression profiling by high throughput sequencing.
Identification of small ORFs in vertebrates using ribosome footprinting and evolutionary conservation
GEO Series GSE53693. Danio rerio. 30 samples. Type: Expression profiling by high throughput sequencing; Non-coding RNA profiling by high throughput sequencing.
A Study of BBP-711 (ORF-229) in Healthy Adult Volunteers
ClinicalTrials.gov study NCT04876924. IPD Sharing: NO. Countries: 1. Publications: 0.
Genomewide demarcation of RNA PolII transcription units by physical fractionation of chromatin- ORF enrichment
GEO Series GSE5649. Saccharomyces cerevisiae. 5 samples. Type: Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.