Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
528
datasets available to search
ShareScore release 0.9.0
Dataset results
528 results for “correspondence”
High-quality large curated dataset of protein sequences (1.83 million) and their corresponding Position Specific Scoring Matrices
<p>As part of his master thesis at the Rostlab, which is located at the Technical University of Munich (TUM), Mr. Issar Arab developed the first language model that encodes evolutionary information of proteins explicitly. The pre-training involved the creation of a novel high-quality dataset of protein sequences (around 1.83 million proteins, or ~0.8 Billion amino acids) with their corresponding Position Specific Scoring Matrices (PSSMs). Those matrices reflect the relative frequency of each amino acid at each position in a protein and is derived from evolutionarily related proteins.</p> <p>Mr. Arab makes this work publicly available to help other researchers speed up their work to leverage AI to learn the representation of protein evolutionary information more explicitly. The set of sequences was derived by extracting all PSSMs from the <a href="https://predictprotein.org/">PredictProtein</a> (PP) cache, which were also part o the UniProt Reference Cluster with 50% sequence identity (uniref50 2019_12). The overlap between PP and uniref50 was further filtered to only include high-quality samples, e.g. only multiple sequence alignments with a certain number of aligned sequences were considered. The processing led to a training set of 1.83 Million sequences, a validation set of 879 instances, and a test set of 879 entries. The training data of proteins is reduced to 40% sequence identity, with respect to the validation/test sets, and contains sequences ranging between 18 and 9858 residues in length.</p> <p>Refer to the Jupyter notebook for a detailed description of the files' structure and a Python code snippet to correctly manipulate this data.</p> <p>To access the full original work, please visit the following link: <a href="https://mediatum.ub.tum.de/node?id=1579236">Manuscript</a> <br><br><strong>Note:</strong> The dataset was recently used to fine tune a protein sequence language model (<a href="https://github.com/issararab/PEvoLM">PEvoLM</a>). The work was presented at the CIBCB'23 conference. If you use PEvoLM or this dataset in your work, please cite the following publication:</p> <p>- Issar Arab, <strong>PEvoLM: Protein Sequence Evolutionary Information Language Model</strong>, <em>IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), Eindhoven, Netherlands</em>, (2023), pp. 1-8, doi:<a href="https://ieeexplore.ieee.org/document/10264890">10.1109/CIBCB56990.2023.10264890</a></p>
FIG. 6 in More than just a name: Colonel Messager and his correspondents
FIG. 6. — Typical label with Messager's hand writing: name of species (Plectopylis messageri Gude, 1909), and locality in two lines (Muong-Hum, Tonkin).
FIG. 7 in More than just a name: Colonel Messager and his correspondents
FIG. 7. — Several labels of Messager's Vietnamese Plectopylidae collection have been partly eaten by the live snails.
FIG. 4. — A in More than just a name: Colonel Messager and his correspondents
FIG. 4. — A sample from the former Messager collection in the MNHN: thousands of shells of a Pupina Vignard, 1929 species.
FIG. 5. — A in More than just a name: Colonel Messager and his correspondents
FIG. 5. — A sample from the former Messager collection in the MNHN: operculae of Pollicaria A. Gould, 1856.
FIG. 2 in More than just a name: Colonel Messager and his correspondents
FIG. 2. — Birth certificate of Messager stating his official first names (Archives nationales, Base Léonore, dossier LH/1846/14).
FIG. 3 in More than just a name: Colonel Messager and his correspondents
FIG. 3. — Name card of Marcel Messager found in the MNHN. Interestingly, the name "Marcel" does not match with the official names found of Messager ("Louis Gabriel Martin"). Despite this, we are probably talking about the same person, because on the other side of this name card (lower image) there are names of snail genera written with Messager's handwriting.
FIG. 1 in More than just a name: Colonel Messager and his correspondents
FIG. 1. — Document certifying the ranks of Messager in the 'Légion d'honneur' (Archives nationales, Base Léonore, dossier LH/1846/14).
FIG. 8 in More than just a name: Colonel Messager and his correspondents
FIG. 8. — Letter from Messager to Dautzenberg, 12 July 1904 (Dautzenberg autographs archive, RBINS).
A list of collection codes and corresponding BOLD numbers to sixty new dragonfly and damselfly species from Africa
<p>These files contain the data and accession numbers used in the following publication:</p> <p>Dijkstra, Klaas-Douwe B. et al.. (2015). Sixty new dragonfly and damselfly species from Africa (Odonata). Odonatologica 44(4): 447-678. doi:10.5281/zenodo.35388</p> <p>Contents</p> <p>- Lab (BOLD numbers)<br /> - Vouchers<br /> - Taxonomy<br /> - Specimen details<br /> - CollectionData</p> <p> </p> <p>uploaded for Odonatologica by Plazi</p>
Peak flow identified at selected GRDC stations and the corresponding hydrological and hydrometeorological state variables
Data set contains a list of peak flows at selected GRDC stations as well as the start, peak, and end dates of each selected event. Hydrometeorlogical variables and hydrological state variables simulated by a hydrological model E-HYPE corresponding to each selected event are also listed in the dataset. Further description and content of each data file is available in the included metadata.
Odontocete detections and corresponding values of environmental variables in the Hawaiian Archipelago
<p>This dataset contains detections of echolocation clicks at two sites in the Hawaiian Archipelago. These sites are Hawaii and Manawai (also known as Pearl and Hermes Reef). Echolocation clicks have been labeled using a neural network classifier that was trained and tested on data from the Hawaiian Islands and can successfully identify many species of regionally present odontocetes. During the labeling process, clicks were grouped into five-minute bins and each bin was given a class label. The data provided here is further binned at a daily level, where counts of a given class represent the number of five-minute bins within a given day that were labeled as that class. One file is provided per site, and files are in .csv format that can be read using any desired coding language. </p> <p>In addition to acoustic counts, values for environmental variables considered in the corresponding manuscript are provided in the CSV files. The final file in this dataset contains satellite-derived chlorophyll-a concentration values from NASA MODIS for the Hawaiian region (used to create Supplementary Fig. 1). Details on all variables and how they were accessed can be found in the manuscript and in the README file accompanying this dataset. </p>
Fig. 1. Map showing the localities where Acropora corals were sampled during the first Snellius expedition, numbers correspond with Table 1 in The Acropora Humilis Group (Scleractinia) Of The Snellius Expedition (1929-30)
Fig. 1. Map showing the localities where Acropora corals were sampled during the first Snellius expedition, numbers correspond with Table 1.
Data and code corresponding to the article "Interaction network structure explains species temporal persistence in empirical plant-pollinator communities"
<p>This upload contains the Datasets and code to generate the results of the article "Interaction network structure explains species temporal persistence in empirical plant-pollinator communities".</p><p>The database comprises two files containing the abundances of plants and pollinators, and one containing the interaction networks among plants and pollinators. </p><p>The code folder contains the code to generate the results, and to generate the figures of the manuscript. </p>
Supplementary material for "Mutual predictiveness of sound correspondences for reconstruction and language subgrouping: The case of Gyalrongic preinitials"
<p>This is the supplementary material of the paper "Mutual predictiveness of sound correspondences for reconstruction and language subgrouping: The case of Gyalrongic preinitials", to be published in Diachronica. It includes the datasets, codes and images involved in the research. Publication is still several months away, and I am trying to collect and identify more cognates as well as correct mistakes. </p> <p>Version 2 has included a few more cognates. I have added a file explaning my cognate judgements and annotations, and removed some doubtful forms after verification with native speakers. </p> <div> <div> <div> <div> <p>Comments, corrections and suggestions about cognate judgement, annotation and bibliographical references are welcome. I will correct and update the data from time to time.</p> </div> </div> </div> </div> <p> </p>
Quantitative results of the analysis of human native and bioengineered tissues corresponding to the work "Histological, histochemical and immunohistochemical characterization of NANOULCOR nanostructured fibrin-agarose human cornea substitutes generated by tissue engineering"
<p>Dataset containing the quantitative results of the histochemical and immunohistochemical analysis of the following human tissues:</p> <ul> <li>Control native cornea (CTR-C)</li> <li>Control native limbus (CTR-L)</li> <li>Artificial cornea generated by tissue engineering (HAC)</li> </ul> <p>Each tissue type was subjected to histochemical and immunohistochemical analyses and results were quantified using ImageJ software to determine average intensities and area fractions corresponding to positive staining signal for each marker.</p>
Data from: Annual species' experimental germination responses to light and temperature do not correspond with their microhabitat associations in the field
<p>Annual species have evolved sets of germination cues that are thought to be predictive of the post-germination environment. In naturally patchy environments, germination microsites often vary considerably in the amount of light they receive and in the diurnal temperature fluctuations they experience. However, whether species' differential germination responses to light and temperature are associated with their spatial patterns of occurrence remains largely untested.</p> <p>We surveyed species' occurrences in annual plant communities in 150 quadrats across gradients of canopy cover and litter cover. Nineteen species recorded in this survey were then included in a germination experiment that manipulated (1) Light vs. Dark (12h light or continuous dark) approximating seeds near the soil surface versus those covered by litter and (2) Cold vs. Warm temperature regimes (7/18 °C and 7/24 °C) approximating diurnal fluctuations experienced in shaded versus sun-exposed microsites, respectively.</p> <p>In the germination experiment, six species had highest germination probabilities in the Light treatment (regardless of temperature), five in <em>Cold</em> + <em>Light</em>, one in <em>Warm</em> + <em>Light</em>, two were indifferent to the treatments, and four did not germinate at all. Binomial linear mixed-effects models showed that species' maximum responses to light and temperature did not explain their spatial distributions along canopy cover and litter cover gradients, contrary to theoretical expectations of germination being a strong driver of species' occurrences.</p> <p>Despite variation in species' responses to experimental treatments, no association was found with their field microsite associations. Germination strategies in our system were wider than expected for Mediterranean systems. Our results support that germination cues are not strong drivers of microhabitat associations in this system.</p>
РИС. 4. Места находок Amuranodonta kijaensis в бассейне р. Амур: черные точки – ранее иЗвестные местонахождениЯ, белые квадраты – впервые обнаруженные колонии. Номера локалитетов соответствуют таковым в таблице 1. FIG. 4. Localitions of finds of Amuranodonta kijaensis in the Amur River basin: black dots are previously known locations, white squares are newly discovered colonies. The locality numbers correspond to those in Table 1. in Новые данные об охранЯемом пресноводном двустворчатом моллюске Amuranodonta kijaensis Moskvicheva, 1973 (Unionidae, Anodontinae)
РИС. 4. Места находок Amuranodonta kijaensis в бассейне р. Амур: черные точки – ранее иЗвестные местонахождениЯ, белые квадраты – впервые обнаруженные колонии. Номера локалитетов соответствуют таковым в таблице 1. FIG. 4. Localitions of finds of Amuranodonta kijaensis in the Amur River basin: black dots are previously known locations, white squares are newly discovered colonies. The locality numbers correspond to those in Table 1.
Supplementary data frames, AlphaFold models, Normal Mode Analysis (NMA) Data, and NMA of Corresponding NMR Ensembles in the S2RCI, MD, and S2 Datasets for "Gradations in protein dynamics captured by experimental NMR are not well represented by AlphaFold2 models and other computational metrics"
<h1><strong>Changes applied to V2</strong></h1> <p>In addition to the supplementary dataframes and AlphaFold models from each dataset in V1, V2 includes the additional data outlined below.</p> <p>The <strong>S2RCI</strong> and <strong>MD</strong> datasets include comprehensive analyses of AlphaFold2 models (both before and after truncation). These datasets feature: </p> <ul> <li><strong>AlphaFold2 Models</strong>: Both original and truncated structures. </li> <li><strong>WEBnma Modes</strong>: `modes.txt` files generated from WEBnma analysis, available for both non-truncated and truncated AF2 models. </li> <li><strong>Root-Mean-Square-Fluctuations (RMSF)</strong>: Profiles calculated before and after truncation of AF2 models. </li> <li><strong>NMR Data: Normal Mode Analysis (NMA)</strong>: Performed on corresponding NMR ensembles (see below). </li> </ul> <p> </p> <p>The <strong>NMR Data</strong> of NMA in these datasets includes: </p> <ul> <li>NMR ensembles </li> <li>Individual NMR models extracted from each ensemble </li> <li>STRIDE secondary structure calculations per-individual NMR models</li> <li>RMSF profiles per-individual NMR models</li> </ul> <p>For detailed information, please refer to the `Readme.txt` file within each corresponding folder. </p> <p>The <strong>S2 dataset</strong> includes all the features listed above, except for the NMR analysis.</p>
Our predicted Crinkler (CRN) family effector proteins and the corresponding GFF3 files across 128 Phytophthora isolates
<p>These files contain predicted Crinkler (CRN) family effector proteins and the corresponding GFF3 files across 128 Phytophthora isolates.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.