Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,523
datasets available to search
ShareScore release 0.7.1
Dataset results
7,523 results for “Annotation”
CREMP-CycPeptMPDB: Conformer-rotamer ensembles of macrocyclic peptides for machine learning with permeability annotations
<p>CREMP-CycPeptMPDB: A resource generated for the rapid development and evaluation of machine learning models for permeable macrocyclic peptides. CREMP-CycPeptMPDB contains 3,258 unique macrocyclic peptides and their high-quality structural ensembles generated using the Conformer-Rotamer Ensemble Sampling Tool (CREST). Altogether, this dataset contains nearly 8.7 million unique macrocycle geometries, each annotated with energies derived from semi-empirical tight-binding DFT calculations and with experimental membrane permeability measurements obtained from the <a href="http://cycpeptmpdb.com/" target="_blank" rel="noopener">CycPeptMPDB</a> database. We anticipate that this dataset will enable the development of machine learning models that can improve peptide design and optimization for novel therapeutics.</p> <p>This dataset complements the <a title="CREMP" href="../doi/10.5281/zenodo.7931444" target="_blank" rel="noopener">CREMP dataset</a>, which contains a larger selection of conformer ensembles for homodetic macrocyclic peptides.</p> <p>We provide the data in two available formats, either as Python pickle files, which provide quick read access with RDKit version 2022.09.5 or later, and as text-based SDF files with associated metadata in JSON format. Each file is named based on its amino acid sequence, with residues separated by periods, using standard one-letter codes with lowercase letters representing D-amino acids and "Me" prefixes representing <em>N</em>-methylated amino acids. The sequences are in no particular order, e.g., "C.R.E.M.P" and "R.E.M.P.C" correspond to the same peptide macrocycle. The filename extensions are ".pickle", ".sdf", and ".json".</p> <p>Each file in the “pickle” folder contains a Python dictionary with amino acid sequence, SMILES, CREST metadata, and a single RDKit molecule object containing all conformers. All files in the folder were compressed into a single “pickle.tar.gz” archive. In the “sdf_and_json” folder, each individual SDF file contains all conformers, each associated with its own JSON file that contains CREST metadata. Similarly, all are compressed into another single archive, “sdf_and_json.tar.bz2”. A single summary CSV file is also provided containing ”sequence”, “smiles”, “num_monomers”, “num_atoms”, “num_heavy_atoms”, along with the CREST metadata “totalconfs”, “uniqueconfs”, “lowestenergy”, “poplowestpct”, “temperature”, “ensembleenergy”, “ensembleentropy”, and “ensemblefreeenergy”. The number of unique conformers with different 3D structures is given by “uniqueconfs”, while “totalconfs” includes the number of rotamers in addition.</p> <p>The unzipped sizes of the archives are approximately 13 GB for "pickle.tar.gz" and 84 GB for "sdf_and_json.tar.bz2". If you encounter errors when trying to load the pickle files, please make sure your RDKit version is at least 2022.09.5. If that doesn't work, try other Python versions.</p>
UnientrezDB: Large-scale Gene Ontology Annotation Dataset and Evaluation Benchmarks with Unified Entrez Gene Identifiers
<p>Our work focuses on providing a comprehensive dataset and benchmarks for evaluating gene ontology annotations using a unified system of Entrez Gene Identifiers.</p>
GO Term annotations for five plants species from Phytozome by FANTASIA
<p>This is the GO term annotation made with FANTASIA for five species (Arabidopsis thaliana, Oryza sativa, Zea mays, Populus trichocarpa, and Solanum lycopersicum) from the Phytozome 13 datasets as proof of concept for this tool.</p>
LyBar v2.0 Genome Assembly and Annotation for Lycium barbarum
<p>LyBar v2.0 genome assembly and annotation files for <em>Lycium barbarum</em>.</p>
Gene Annotations of 49 Bacillariophyta Genome Assemblies (Individual gff3 files)
<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div> </div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a> . It is a copy of the data hostet at <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a> , but instead of storing one archive will all gff3 files included, the gff3 files are here hosted, individually. This copy was made upon request from the RDA Working Group "FAIRification of Genomic Annotations – metadata harmonisation at scale".</p> <div> <h2>Files</h2> <div>The following gzip-compressed gff3-files with structural and functional genome annotation are included:</div> <div> </div> <div>Asterionella_formosa.gff3.gz<br>Asterionellopsis_glacialis.gff3.gz<br>Bacterosira_constricta.gff3.gz<br>Chaetoceros_muellerii.gff3.gz<br>concatenated_output.gff3.gz<br>Conticribra_guillardii.gff3.gz<br>Conticribra_weissflogii.gff3.gz<br>Craspedostauros_australis.gff3.gz<br>Cyclostephanos_invisitatus.gff3.gz<br>Cyclostephanos_tholiformis.gff3.gz<br>Cyclotella_atomus.gff3.gz<br>Cyclotella_baltica.gff3.gz<br>Cyclotella_choctawhatcheeana.gff3.gz<br>Cyclotella_cryptica.gff3.gz<br>Cylindrotheca_fusiformis.gff3.gz<br>Detonula_confervacea.gff3.gz<br>Discostella_pseudostelligera.gff3.gz<br>Discostella_stelligera.gff3.gz<br>Discostella_stelligeroides.gff3.gz<br>Epithemia_pelagica.gff3.gz<br>Fistulifera_pelliculosa.gff3.gz<br>Fistulifera_solaris.gff3.gz<br>Fragilaria_radians.gff3.gz<br>Fragilariopsis_cylindrus.gff3.gz<br>Licmophora_abbreviata.gff3.gz<br>Mediolabrus_comicus.gff3.gz<br>Nitzschia_palea.gff3.gz<br>Nitzschia_putrida.gff3.gz<br>Porosira_glacialis.gff3.gz<br>Psammoneis_japonica.gff3.gz<br>Pseudo-nitzschia_multiseries.gff3.gz<br>Pseudo-nitzschia_pungens.gff3.gz<br>Skeletonema_costatum.gff3.gz<br>Skeletonema_marinoi.gff3.gz<br>Skeletonema_menzelii.gff3.gz<br>Skeletonema_potamos.gff3.gz<br>Skeletonema_tropicum.gff3.gz<br>Stephanocyclus_meneghinianus.gff3.gz<br>Stephanodiscus_minutulus.gff3.gz<br>Stephanodiscus_triporus.gff3.gz<br>Thalassiosira_allenii.gff3.gz<br>Thalassiosira_delicatula.gff3.gz<br>Thalassiosira_exigua.gff3.gz<br>Thalassiosira_gravida.gff3.gz<br>Thalassiosira_livingstoniorum.gff3.gz<br>Thalassiosira_mediterranea.gff3.gz<br>Thalassiosira_oceanica.gff3.gz<br>Thalassiosira_ordinaria.gff3.gz<br>Thalassiosira_pacifica.gff3.gz<br>Thalassiosira_profunda.gff3.gz</div> <div> </div> <div>To extract individual files after download execute the following command:</div> <div> </div> <div><code>gunzip *.gff3.gz</code></div> <h2>Genome Assemblies</h2> <p> </p> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <p> </p> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p> </p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p> </p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p> </p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^>/ s/ .*//' genome.fasta > genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <p> </p> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <p> </p> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>The submission and release was made upon request of the RDA working group "FAIRification of Genomic Annotations – metadata harmonisation at scale". The contained data is identical to <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a></p> <h2>License</h2> <p> </p> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div> <p> </p> </div> </div>
Fraxinus pennsylvanica genome assembly and annotation
<p>We report the first chromosome-level assembly for green ash (<em>Fraxinus pennsylvanica</em>) to assist in ash breeding efforts to propagate resistance to the emerald ash borer. The final haploid assembly consists of 23 chromosomes and 87 unplaced scaffolds of 10 kb or more. Over 99% of the bases anchored to the chromosomes. The assembly spans 757 Mb and consists of 49.43% repetitive DNA. Gene annotation yielded 35,470 high-confidence gene models, all located on the chromosomes and assigned to 22,976 Asterid Orthogroups.</p>
ForTrunkDet - Image dataset of visible and thermal annotated images for forest tree trunk detection
<p>Forest dataset composed by visible and thermal images with trunk annotations. The images were acquired in three different portuguese forests and were captured by four different cameras:</p> <ul> <li>GoPro Hero6</li> <li>Allied Mako G-125</li> <li>FLIR M232</li> <li>ZED Stereo</li> </ul> <p>The images and annotations are stored in two zip files:</p> <ul> <li>forest_dataset_original.zip - original dataset</li> <li>forest_dataset_augmented - augmented dataset</li> </ul> <p>Also, the subsets that were used to train, validate and test some deep learning models are available in three .TXT files (train.txt, val.txt and test.txt), where each file line corresponds to an image name.</p> <p> </p>
Bakta Annotation Examples
<p>This data repository provides exemplary bacterial genome annotations conducted with Bakta v1.1 comprising a broad taxonomical range of many pathogenic (all ESKAPE), commensal and environmental genomes from RefSeq.</p> <p>Bakta is a tool for the rapid & standardized local annotation of bacterial genomes & plasmids. It provides <strong>dbxref</strong>-rich and <strong>sORF</strong>-including annotations in machine-readble <code>JSON</code> & bioinformatics standard file formats for automatic downstream analysis: https://github.com/oschwengers/bakta</p>
ArXiV-Entity/Relation annotated dataset
<p>This dataset is a collection of abstracts from the CS section of ArXiV, each annotated with <a href="https://github.com/dwadden/dygiepp">DyGIE++</a> (SciERC model)</p> <p>The dataset can be used to train triple extractors or to cluster triples (in the Computer Science and AI domains).</p> <p>Supersedes the ArXiV-AIKG dataset as these triples are unconstrained (so they don't forcibly appear in AIKG)</p>
Translation Alignment: Ancient Greek to English. Annotation Style Guide and Gold Standard.
<p>This dataset contains guidelines and a gold standard for the alignment of Ancient Greek texts with English translations.</p> <p>The guidelines were used to annotate a diverse dataset including Homeric epic, Attic prose, and Platonic dialogue, and were tested by measuring inter-annotator agreement of 80% or higher. The Ancient Greek texts used are almost entirely available through the Scaife viewer (<a href="https://scaife.perseus.org/">https://scaife.perseus.org/</a>).</p> <p>The datasets used to develop the gold standard were aligned using the Ugarit Translation Alignment Editor for Historical languages (<a href="http://ugarit.ialigner.com/">http://ugarit.ialigner.com/</a>).</p> <p>The materials available here can be used to perform and evaluate alignments of various texts in Ancient Greek, to create new gold standard corpora, and to train automated translation models.</p> <p>The guidelines can also be further adapted to address similar language pairs including an inflected and a synthetic language, such as Latin and English, or can provide a structure for the alignment of other historical texts against modern translations. However, the guidelines are not project-specific: they were specifically intended for the creation of a Gold Standard in the scenario of machine translation. Different scenarios, such as language research or pedagogy, may need further tweaking to these guidelines to make them more compatible with different underlying principles.</p> <p>For further information on Ugarit and translation alignment of historical languages, see <a href="http://ugarit.ialigner.com/bib.php">http://ugarit.ialigner.com/bib.php</a> and follow us on Twitter (@ugarit_ty).</p>
Mycobacteroides abscessus subp. bolletii strain associated with a persistent infection (genome assembly and annotation dataset)
<p>This dataset includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) of a <strong><em>Mycobacteroides abscessus subp. bolletti </em></strong>strain associated with a persistente infection. (genome anotation was performed using Bakta v1.2.2 https://github.com/oschwengers/bakta)</p> <p>The raw sequence reads were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB57933; Run Accession: ERR10554471).</p>
Wikidata CiTO intention annotations
<p>Dump using three SPARQL queries without provenance of the CiTO intention annotations collected in Wikidata.</p>
A collection of fully-annotated soundscape recordings from the southern Sierra Nevada mountain range
<p>This collection contains 100 soundscape recordings of 10 minutes duration, which have been annotated with 10,296 bounding box labels for 21 different bird species from the Western United States. The data were recorded in 2015 in the southern end of the Sierra Nevada mountain range in California, USA. This collection has been featured as test data in the 2020 BirdCLEF and Kaggle Birdcall Identification competition and can primarily be used for training and evaluation of machine learning algorithms.</p> <p><strong>Data collection</strong></p> <p>The recordings were made in Sequoia and Kings Canyon National Parks, two contiguous national parks in the southern Sierra Nevada mountain range in California, USA. The focus of the acoustic study was the high-elevation region of the Parks; specifically, the headwater lake basins above 3,000 km in elevation. The original intent of the study was to monitor seasonal activity of birds and bats at lakes containing trout and lakes without trout, because the cascading impacts of trout on the adjacent terrestrial zone remain poorly understood. Soundscapes were recorded for 24 h continuously at 10 lakes (5 fishless, 5 fish-containing) throughout Sequoia and Kings Canyon National Parks during June-September 2015. Song Meter SM2+ units (Wildlife Acoustics, USA) powered by custom-made solar panels were used to obviate the need to swap batteries, due to the recording locations being extremely difficult to access. Song Meters continuously recorded mono-channel, 16-bits uncompressed WAVE files at 48 kHz sampling rate. For this collection, recordings were resampled at 32 kHz and converted to FLAC.</p> <p><strong>Sampling and annotation protocol</strong></p> <p>A total of 100 10-minute segments of audio between July 9 and 12, 2015 from morning hours (06:10-09:10 PDT) from all 10 sites were selected at random. Annotators were asked to box every bird call they could recognize, ignoring those that are too faint or unidentifiable. Every sound that could not be confidently assigned an identity was reviewed with 1-2 other experts in bird identification. To minimize observer bias, all identifying information about the location, date and time of the recordings was hidden from the annotator. Raven Pro software was used to annotate the data. Provided labels contain full bird calls that are boxed in time and frequency. In this collection, we use eBird species codes as labels, following the 2021 eBird taxonomy (Clements list). Unidentifiable calls have been marked with “????” and were added as bounding box labels to the ground truth annotations. Parts of this dataset have previously been used in the 2020 BirdCLEF and Kaggle Birdcall Identification competition.</p> <p><strong>Files in this collection</strong></p> <p>Audio recordings can be accessed by downloading and extracting the “soundscape_data.zip” file. Soundscape recording filenames contain a sequential file ID, recording date and timestamp in PDT (UTC-7). As an example, the file “HSN_001_20150708_061805.flac” has sequential ID 001 and was recorded on July 8th 2015 at 06:18:05 PDT. Ground truth annotations are listed in “annotations.csv” where each line specifies the corresponding filename, start and end time in seconds, low and high frequency in Hertz and an eBird species code. These species codes can be assigned to scientific and common name of a species with the “species.csv” file. The approximate recording location with longitude and latitude can be found in the “recording_location.txt” file.</p> <p><strong>Acknowledgements </strong></p> <p>Compiling this extensive dataset was a major undertaking, and we are very thankful to the domain experts who helped to collect and manually annotate the data for this collection (individual contributors in alphabetic order): Anna Calderón, Thomas Hahn, Ruoshi Huang, Angelly Tovar</p>
Toloker Graph: Interaction of Crowd Annotators
<p>The graph contains 11,758 nodes and 519,000 edges representing interactions between crowd annotators on a project labeled on the <a href="https://toloka.ai/">Toloka</a> crowdsourcing platform (see the <a href="https://toloka.ai/en/docs/guide/concepts/overview">Toloka overview</a> for the details on the used terminology).</p> <p>Each node represents an individual annotator; nodes are provided with four numerical and three categorical features. An edge is drawn between a pair of annotators if they annotated the same task. Also, each node is provided with a label showing whether the annotator was banned on this project, or not.</p> <p><strong>Nodes</strong> are stored in the <a href="https://github.com/Toloka/TolokerGraph/blob/main/nodes.tsv">nodes.tsv</a> file in the TSV format of the following structure:</p> <ul> <li><code>id</code>: unique identifier of the annotator</li> <li><code>approved_rate</code>: percentage of the approved labels of this annotator</li> <li><code>skipped_rate</code>: percentage of the skipped tasks of this annotator</li> <li><code>expired_rate</code>: percentage of the expired tasks of this annotator</li> <li><code>rejected_rate</code>: percentage of the rejected labels of this annotator</li> <li><code>education</code>: level of education as self-reported by this annotator (<code>none</code>, <code>basic</code>, <code>middle</code>, <code>high</code>)</li> <li><code>english_profile</code>: knowledge of English as self-reported by this annotator (<code>0</code> for no, <code>1</code> for yes)</li> <li><code>english_tested</code>: whether the annotator passed the Toloka language test for English (<code>0</code> for no, <code>1</code> for yes)</li> <li><code>banned</code>: whether the annotator was banned on this project (<code>0</code> for no, <code>1</code> for yes)</li> </ul> <p>The <code>*_rate</code> attributes should sum up to 1.</p> <p><strong>Edges</strong> are stored in the <a href="https://github.com/Toloka/TolokerGraph/blob/main/edges.tsv">edges.tsv</a> file in the TSV format of the following structure:</p> <ul> <li><code>source</code>: source identifier of the annotator</li> <li><code>target</code>: target identifier of the annotator</li> </ul> <p>As the graph is undirected, <code>source</code> and <code>target</code> can be interchanged for the given pair of nodes.</p>
Planet Microbe Functional and Taxonomic annotation of Illumina WGS Prokaryotic Fraction for Semantic Web Analysis
<p>Functional and Taxonomic annotations computed from a subset of Illumina Whole-Genome Sequencing samples from the prokaryotic fraction of the <a href="https://www.planetmicrobe.org/">Planet Microbe</a> database. Data was computed using the pipeline available from https://github.com/hurwitzlab/planet-microbe-functional-annotation/, and post processing scripts from https://github.com/hurwitzlab/planet-microbe-semantic-web-analysis. Files contain total annotation counts of Interpro, GO and NCBITaxon annotations, as well as additional sample metadata. See readme.txt file for more information.</p>
(Annotation and metric) Thalassiosirales reference transcriptomes
<p>Here are deposited the public data recompiled for the construction of a Thalassiosirales reference database as part of a Ph.D Thesis "TEMPERATURE ACCLIMATION CAPACITY AND COLD-ADAPTATION MECHANIMS IN THALASSIOSIRALES ANTARCTIC MEMBERS". Data correspond to the annotation of 53 transcriptomes from the MMETSP of Thalassiosirales members.</p>
Line-level Named Entity Recognition annotation for the George Washington and IAM datasets
<p>Line-level Named Entity annotation for the George Washington and IAM datasets. The word-level annotations from Oliver Tüselmann [3] were extended to line-level to enable experimentation with line-level coupled HTR+NER models. We also publish the line-level partition files that result from the partition proposed by [3].</p>
PretoxTM Corpus: a gold standard corpus of preclinical treatment-related findings annotated from toxicology reports
<p>The PretoxTM Corpus is a gold standard corpus of preclinical treatment-related findings annotated from toxicology reports.</p> <p>Example documents annotated by domain experts, also known as a gold standard corpus, are needed in order to develop, train, and validate text mining tools. To this aim we designed and performed an annotation activity for the development of the corpus of treatment-related findings: the PretoxTM corpus.</p> <p>A treatment-related finding expression enclose several named entities; the most relevant one is the abnormal effect detected; which depending on the study domain of the finding, can be given by a measurement, test, or examination named Study Test and an abnormal Manifestation result obtained for that study test; or by an abnormal Finding in study domains where there is no associated test or measurement (e.g., clinical, macroscopic and microscopic). Other related named entities that could be present to complete the treatment-related finding are; the Specimen of the abnormal observation, the Sex of the subject, the Group of subjects in which the observation was detected and the Dose level administration of the compound. Examples of sentences with treatment-related findings are: “The decrease in food consumption and body weight of the animals from the mid dose onwards is regarded as evidence of general toxicity.” and "At dose level 3, absolute and relative liver weights were increased in male rats.”.</p> <p>Contributions: The PretoxTM corpus was developed by BSC, with the contribution of IMIM and a team of experts from the eTRANSAFE EFPIA partners.</p> <p>License: Creative Commons Attribution-ShareAlike 4.0 International (cc by sa 4.0).</p> <p>The PretoxTM resources have been developed as part of the eTRANSAFE project.</p> <p>For more information about PretoxTM please visit:</p> <p>PretoxTM Corpus Gitlab: <a href="https://gitlab.bsc.es/inb/etransafe/preclinical-toxicological-corpus/">https://gitlab.bsc.es/inb/etransafe/preclinical-toxicological-corpus/</a></p> <p>PretoxTM central documentation: <a href="https://pretoxtm.gitlab.io/documentation/">https://pretoxtm.gitlab.io/documentation/</a></p>
A Grammatically Annotated Corpus of the Old Latvian Postil of Georg Mancelius
<p>This grammatically annoted corpus aims at facilitating linguistic research on Old Latvian based on the Postil of Georg Mancelius from the year 1654. The corpus is divided into two subcorpora, "pericopes" and "homilies" to make register related research easier.</p> <p>The pericopes were annotated using SIL Toolbox and converted to be used in the search-tool ANNIS using the conversion tool PEPPER.</p> <p>Three formats are provided in this release: 1. the Toolbox files, 2. the transitional Excel files and 3. a zipped folder to be imported into ANNIS.</p> <p>Created in the project B02, <em>Emergence and change of registers: The case of Lithuanian and Latvian</em> of the CRC 1412 "Register" (funded by the Deutsche Forschungsgemeinschaft: DFG, German Research Foundation: 416591334).</p>
Genome assembly and annotation files for Corylus americana accessions 'Rush' and 'Winkler'
<p>The native shrub American hazelnut (<em>Corylus americana</em>) is currently used in breeding programs that are aiming to develop commercially viable hazelnut varieties for the U.S. Upper Midwestern U.S. This species provides significant ecological benefits as it is a perennial crop and well-adapted to this region. Breeding cycles for perennial species are long, and may benefit from the use of predictive methods such as genomic selection to reduce cycle time and increase the efficiency of field trials.</p> <p>High-quality reference genome assemblies are very useful for the implementation marker-assisted selection and genomic prediction, and we therefore developed the first chromosome-scale reference assemblies for <em>C. americana</em>, using the accessions 'Rush' and 'Winkler'. Initial draft assemblies were created using HiFi PacBio reads and Arima Hi-C sequencing to assemble genomes into 11 pseudomolecules. We then utilized Oxford Nanopore reads and a high-density genetic map in order to perform error correction. N50 scores were calculated to be 31.9 Mb and 35.3 Mb for 'Rush' and 'Winkler', respectively, while 97.1% (for 'Winkler') and 90.2% (for 'Rush') of the total genome was assembled into the 11 pseudomolecules. Gene prediction was performed using both RNAseq libraries as well as protein homology data. 'Rush' had a BUSCO score of 99.0 for its assembly and 99.0 for its annotation, while 'Winkler' had corresponding scores of 96.9 and 96.5, indicating extremely high-quality assemblies.</p> <p>These two independent, de novo assemblies enable unbiased assessment of structural variation across the genome, as well as patterns of syntenic relationships within C. americana and the <em>Corylus</em> genus. These assemblies are also an important first step in providing a resource for using next-generation sequencing data in the improvement of <em>C. americana</em>. We demonstrate this utility through the generation of high-density SNP marker sets from genotyping-by-sequencing data for 1,343 <em>C. americana</em>, <em>C. avellana</em>, and <em>C. americana</em> x <em>C. avellana</em> hybrids, in order to assess population structure in natural and breeding populations. Finally, the transcriptomes of these assemblies, as well as several other recently published <em>Corylus</em> genomes, were utilized to perform phylogenetic analysis of sporophytic self-incompatibility (SSI) in hazelnut, providing further evidence of unique molecular pathways governing self-incompatibility in Corylus not exhibited in other well-studied SSI systems. We hope these assemblies will aide in the application of modern breeding methods to the development of commercially viable hazelnut varieties for the U.S. Upper Midwest.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.