Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
14,965
datasets available to search
ShareScore release 0.7.1
Dataset results
14,965 results for “evolution”
The evolution of genomic, transcriptomic, and single-cell protein markers of metastatic upper tract urothelial carcinoma
<p>The molecular characteristics of metastatic upper tract urothelial carcinoma (UTUC) are unknown. The genomic and transcriptomic differences between primary and metastatic UTUC is not well described either. We combined whole-exome sequencing, RNA-sequencing, and Imaging Mass Cytometry<sup>TM</sup> (IMC<sup>TM</sup>) of 44 tumor samples from 28 patients with high-grade primary and metastatic UTUC. IMC enables spatially resolved single-cell analyses to examine the evolution of cancer cell, immune cell, and stromal cell markers using mass cytometry with lanthanide metal-conjugated antibodies. We discovered that actionable genomic alterations are frequently discordant between primary and metastatic UTUC tumors in the same patient. In contrast, molecular subtype membership and immune depletion signature were stable across primary and matched metastatic UTUC. Molecular and immune subtypes were consistent between bulk RNA-sequencing and mass cytometry of protein markers from 340,798 single-cells. Molecular subtyping at the single cell level was highly conserved between primary and metastatic UTUC tumors within the same patient.</p>
A multi-locus phylogeny for the Diamesinae (Chironomidae: Diptera) provides new insights into evolution of an amphitropical clade
<p>Aligned FASTA files for each locus, Concatinated Dataset, Input files for MrBayes, IqTree2, PartitionFinder and RASP. Ready trees after MrBayes and BEAST. </p>
Island-specific evolution of a sex-primed autosome in the planarian Schmidtea mediterranea
<p>The sexual strain of the planarian <em>Schmidtea mediterranea </em>is a hermaphrodite indigenous to Tunisia and several Mediterranean islands. Here, we isolated individual chromosomes and used sequencing, Hi-C and linkage mapping to assemble a chromosome-scale genome reference. The linkage map revealed an extremely low rate of recombination on chromosome 1. We confirmed suppression of recombination on chromosome 1 by genotyping of individual sperm and oocytes. We showed that previously identified genomic regions that maintain heterozygosity even after prolonged inbreeding comprise essentially all of chromosome 1. Genome sequencing of individuals isolated in the wild indicated that this phenomenon has evolved specifically in populations from Sardinia and Corsica. We found that most known master regulators of the reproductive system are located on chromosome 1. We used RNA interference to knock down a gene with haplotype-biased expression and observed that this led to the formation of a more pronounced female mating organ. Based on these observations, we propose that chromosome 1 is a sex-primed autosome primed for evolution into a sex chromosome.</p>
Dataset for "The Gobolitide (Al-Jibal) microregion: geography and settlement network evolution from Nabataean to Byzantine times"
<p>Dataset for: Kopij, K. and Bała, S. (2021). The Gobolitide (Al-Jibal) microregion: geography and settlement network evolution from Nabataean to Byzantine times. Polish Archaeology in the Mediterranean 30/2) (pp. 181–201). https://doi.org/10.31338/uw.2083-537X.pam30.2.28</p>
Genomics of extreme ecological specialists: multiple convergent evolution but no genetic divergence between ecotypes of Maculinea alcon butterflies
<p>Biotic interactions are often acknowledged as catalysers of genetic divergence and eventual explanation of processes driving species richness. We address the question, whether extreme ecological specialization is always associated with lineage sorting, by analysing polymorphisms in morphologically similar ecotypes of the myrmecophilous butterfly <em>Maculinea alcon</em>. The ecotypes occur in either hygric or xeric habitats, use different larval host plants and ant species, but no significant distinctive molecular traits have been revealed so far. We apply genome-wide RAD-sequencing to specimens originating from both habitats across Europe in order to get a view of the potential evolutionary processes at work. Our results confirm that genetic variation is mainly structured geographically but not ecologically — specimens from close localities are more related to each other than populations of each ecotype from distant localities. However, we found two loci for which the association with xeric versus hygric habitats is supported by segregating alleles, suggesting convergent evolution of habitat preference. Thus, ecological divergence between the forms probably does not represent an early stage of speciation, but may result from independent recurring adaptations involving few genes. We discuss the implications of these results for conservation and suggest preserving biotic interactions and main genetic clusters.</p>
Tectonic evolution and deep mantle structure of the eastern Tethys since the latest Jurassic
<p>Uploaded by Sabin Zahirovic (sabin.zahirovic@sydney.edu.au)<br>6 August 2018</p> <p>Notes:</p> <p>Plate reconstructions can be downloaded from: <br><a href="https://www.earthbyte.org/webdav/ftp/Data_Collections/Zahirovic_etal_ESR_EasternTethys_Supplement.zip" target="_blank" rel="noopener">https://www.earthbyte.org/webdav/ftp/Data_Collections/Zahirovic_etal_ESR_EasternTethys_Supplement.zip</a></p> <p>The relevant seafloor paleo-agegrid can be downloaded from:<br><a href="https://www.earthbyte.org/webdav/ftp/Data_Collections/Zahirovic_etal_2016_ESR_AgeGrid/" target="_blank" rel="noopener">https://www.earthbyte.org/webdav/ftp/Data_Collections/Zahirovic_etal_2016_ESR_AgeGrid/</a> </p> <p>This is internal revision 888 of the 2015_v2 seafloor age-grid. </p> <p>Citation:<br>Zahirovic, S., Matthews, K. J., Flament, N., Müller, R. D., Hill, K. C., Seton, M., and Gurnis, M., 2016, Tectonic evolution and deep mantle structure of the eastern Tethys since the latest Jurassic. Earth Science Reviews, v. 162, p. 293-337."<br><a href="https://www.sciencedirect.com/science/article/pii/S0012825216302872%22" target="_blank" rel="noopener">https://www.sciencedirect.com/science/article/pii/S0012825216302872%22</a> </p>
Global plate boundary evolution and kinematics since the late Paleozoic
<h3>Global plate boundary evolution and kinematics since the late Paleozoic </h3> <p>Kara J. Matthews*^, Kayla T. Maloney*, Sabin Zahirovic*, Simon E. Williams*, Maria Seton*, R. Dietmar Müller*</p> <p>* EarthByte Group, School of Geosciences, The University of Sydney, Sydney, NSW 2006, Australia<br>^ Present address: Department of Earth Sciences, University of Oxford, South Parks Road, Oxford OX1 3AN, UK</p> <p>Contact: karajmatthews@gmail.com</p> <p>CORRECTION applied for the Pacific plate prior to 83 Ma based on Torsvik et al. (2019)</p> <h3><br>Supplementary Material</h3> <p>We provide a digital plate model files (including rotations and geometries) with this publication. These files allow for the visualisation and/or manipulation of the late Paleozoic to present-day (410-0 Ma) global plate motion model presented in this study. </p> <p>#########################################<br>The digital plate model files are compatible with the open-source GPlates plate reconstruction software (<a href="https://www.gplates.org" target="_blank" rel="noopener">www.gplates.org</a>):</p> <p>(1) Rotations - Global rotation model that contains the reconstruction poles that describe the motions of the continents and oceans.<br>- <strong>Global_EB_250-0Ma_GK07_Matthews_etal.rot</strong> (455 KB)<br>- <strong>Global_EB_410-250Ma_GK07_Matthews_etal.rot</strong> (115 KB) - in the comments 'POLE_RECALCULATED' means that we recalculated that finite pole of rotation such that the moving plate moves relative to a neighbouring plate rather than directly to the absolute reference frame (see Section 2.2.1 of the main text for more details). This process should have a minimal effect on the absolute motion of the plate.</p> <p>(2) Plate polygons and boundary geometries - Topologically closed plate polygons are constructed from the intersection of ridges, transforms, subduction zones and other plate boundary geometries. These 'resolved topologies' are valid at 1 Myr intervals (410-0 Ma). The plate boundary geometries and plate polygons have been assigned plate reconstruction IDs to allow them to be reconstructed using the supplied rotation file.<br>- <strong>Global_Mesozoic-Cenozoic_plate_bounds_Matthews_etal.gpml</strong> (36 MB)<br>- <strong>Global_Paleozoic_plate_bounds_Matthews_etal.gpml</strong> (8.7 MB)<br>- <strong>TopologyBuildingBlocks_Matthews_etal.gpml</strong> (2 MB) - this file has not been modified from Müller et al. (2016)</p> <p>(3) Coastlines - Geometries of the present-day coastlines.<br>- <strong>Global_coastlines_low_res_Matthews_etal.gpml</strong> (25.4 MB)<br>- <strong>Global_coastlines_low_res_Matthews_etal.shp</strong> (2.9 MB inc. auxillary files, datum-WGS 1984)<br>NOTE: From 410 to 320-310 Ma Kazakhstania is represented as one or two ('Internal' and 'External' Kazakhstania - Domeier and Torsvik, 2014) ovate polygons. Kazakhstania is highly deformed following a long and complicated history, and so for simplicity we avoid using their present-day outlines in the earlier part of the model.</p> <p>(4) Static polygons (optional) - Includes ocean isochron and terrane polygon geometries.<br>- <strong>Global_EarthByte_GPlates_PresentDay_StaticPlatePolygons_Matthews_etal.shp</strong> (2.7 MB inc. auxillary files, datum-WGS 1984)</p> <p>(5) Continenal polygons (optional) - Includes continental terrane polygon geometries and excludes oceanic lithosphere.<br>- <strong>Global_EarthByte_GPlates_PresentDay_ContinentalPolygons_Matthews_etal.shp</strong> (804 KB inc. auxillary files, datum-WGS 1984)</p> <p>GPLATES: <br>To view the model load all files in GPlates (either drag and drop files onto the globe OR from the navigation bar at the top of the screen click File -> Open Feature Collection and select files). Both rotation files (1) and each of the three plate geometry files (2) need to be loaded for the model to work properly. It is recommended that coastlines (3) are loaded to see how the continents move, however only one coastline file is necessary (.gpml or .shp). The static polygons (4) and continental polygons (5) are optional. </p> <p>The two rotation files need to be 'connected' in order for the model to run continuously from 410 to 0 Ma. In the GPlates 'Layers' window (opened from the main navigation bar, click 'Window' -> 'Show Layers') the rotation files will be highlighted yellow, yet only one will have a yellow tick next to it to signify it is being used. Click the small black triangle to the left the ticked rotation file. Under 'Inputs' -> 'Reconstruction features' click 'Add new connection' and then select the other rotation file from the list of files that will appear. This will ensure that both rotation files are active. </p> <p>Finally, it is recommended to experiment with geometry visibility in order to make the globe less cluttered. For instance, from the navigation bar click View -> Geometry Visibility and untick 'Show Line Geometries'. Alternatively, files can be toggled on and off using the tick boxes in the Layers window. For more information about using GPlates, a set of user tutorials can be accessed from the GPlates website - http://www.gplates.org/docs.html.</p> <p><br>#########################################<br>We also provide a list of the plate reconstruction IDs used in the model:</p> <p>Plate IDs - A list of all the plate IDs used in the rotation and geometry files and their corresponding plate names.<br>- <strong>EarthByte_Plate_ID_Table_Matthews_etal.txt</strong> (33 KB)</p> <p>#########################################<br>MODEL REFERENCING:<br>When using our model, in addition to citing this publication:</p> <p>Matthews, K.J., Maloney, K.T., Zahirovic, S., Williams, S.E., Seton, M. and Müller, R.D., 2016, Global plate boundary evolution and kinematics since the late Paleozoic, Global and Planetary Change, in press, accepted 3 October 2016.</p> <p>please also consider citing the studies of Domeier and Torsvik (2014) and Müller et al. (2016) which served as the basis for this model in the late Paleozoic and Mesozoic-Cenozoic, respectively, and cite any other study that describes refinements to the plate reconstructions in your region of interest. See Section 2 and Section 3 of the main text for more information on how the present model was constructed.</p> <p>- Domeier, M., & Torsvik, T. H. (2014). Plate tectonics in the late Paleozoic. Geoscience Frontiers, 5(3), 303-350. DOI:<a href="https://doi.org/10.1016/j.gsf.2014.01.002" target="_blank" rel="noopener">10.1016/j.gsf.2014.01.002</a><br>- Müller, R. D., Seton, M., Zahirovic, S., Williams, S. E., Matthews, K. J., Wright, N. M., Shephard, G. E., Maloney, K., Barnett-Moore, N., Hosseinpour, M., Bower, D. J., & Cannon, J. (2016). Ocean Basin Evolution and Global-Scale Plate Reorganization Events Since Pangea Breakup. Annual Review of Earth and Planetary Sciences, 44(1). DOI:<a href="https://doi.org/10.1146/annurev-earth-060115-012211" target="_blank" rel="noopener">10.1146/annurev-earth-060115-012211</a></p> <p>Note: We have recently fixed some issues in this model, namely the motion of the Pacific plate (following Torsvik et al., 2019), and some MOR topologies in the Arctic. The fixes are in the model files included in this folder, but the old (published) version of the model is included in a sub-folder called "_OLD_MODEL_DO_NOT_USE". </p> <p>Torsvik, T. H., B. Steinberger, G. E. Shephard, P. V. Doubrovine, C. Gaina, M. Domeier, C. P. Conrad, and W. W. Sager (2019), Pacific‐Panthalassic reconstructions: Overview, errata and the way forward, Geochemistry, Geophysics, Geosystems, 20(7), 3659-3689.</p> <p> </p>
The tectonic evolution of the Arctic since Pangea breakup: Integrating constraints from surface geology and geophysics with mantle structure
<div>Description of Resources - Shephard et al. (2013)</div> <div> </div> <div>This file provides a detailed description of all of the files that make up the data collection associated with the publication: Shephard, G. E., Müller, R. D., & Seton, M. (2013). The tectonic evolution of the Arctic since Pangea breakup: Integrating constraints from surface geology and geophysics with mantle structure. Earth-Science Reviews, 124(0), 148-183. doi: <a href="https://doi.org/10.1016/j.earscirev.2013.05.012" target="_blank" rel="noopener">10.1016/j.earscirev.2013.05.012</a></div> <div> </div> <div>Note: For information on file formats and what programs to use to interact with various file formats, see "File Formats and Recommended Programs”.</div> <div> </div> <div>Note: This paper is based on a global model (Seton et al., 2012), which should also be referenced if looking globally or regions other than the Arctic or northern Panthalassa.</div> <div> </div> <div>The files that make up the tectonic reconstruction model include:</div> <div>• <strong>Rotations </strong>- This is a global rotation model (based on Seton et al., 2012) that includes the new rotations for the Arctic.</div> <div>* Shephard_etal_ESR2013.rot (373 KB)</div> <div> </div> <div>• <strong>Coastlines </strong>- These are present day coastlines that have been assigned plate reconstruction ids to allow them to be reconstructed using the rotation file.</div> <div>* Shephard_etal_ESR2013_Coastlines.gpml (34.1 MB)</div> <div>* Shephard_etal_ESR2013_Coastlines.txt (3.2 MB)</div> <div>* Shephard_etal_ESR2013_Coastlinesc.kml (6.3 MB; datum - WGS 1984)</div> <div>* Shephard_etal_ESR2013_Coastlines.shp (3.2 MB inc auxiliary files; datum - WGS 1984)</div> <div> </div> <div>• <strong>Static polygons </strong>- These are closed polygons that split present day Earth's surface into regions that can be assigned to a given plate id, and therefore reconstructed back through time using the rotation file. These polygons can be used to cookie-cut and assign plate ids to geometry and raster data (for more information on this feature please visit http://gplates.org or http://earthbyte.org).</div> <div>* Shephard_etal_ESR2013_staticpolygons.gpml (19.4 MB)</div> <div>* Shephard_etal_ESR2013_staticpolygons.txt (2.7 MB)</div> <div>* Shephard_etal_ESR2013_staticpolygons.kml (4.4 MB; datum - WGS 1984)</div> <div>* Shephard_etal_ESR2013_staticpolygons.shp (2.3 MB inc auxiliary files; datum - WGS 1984)</div> <div> </div> <div>• <strong>Plate boundary geometries and resolved topologies</strong> – Resolved topologies comprise ridges, transforms, subduction zones and other plate boundary geometries. These boundaries intersect to form closed plate polygons ('resolved topologies') that are valid at 1 Myr intervals (0-200 Ma). The plate boundary geometries and plate polygons have been assigned plate reconstruction ids to allow them to be reconstructed using the rotation file.</div> <div>* Shephard_etal_ESR2013_platebounds.gpml (27.7 MB) - contains both plate boundaries and resolved topological plate polygons</div> <div>* Resolved topologies:</div> <div>- topology_*.00Ma.txt (20.6 MB)</div> <div>- topology_*.00Ma.shp (12.5 MB inc auxiliary files; datum - WGS 1984)</div> <div> </div> <div> </div> <div>References</div> <div> </div> <div>M. Seton, R.D. Müller, S. Zahirovic, C. Gaina, T.H. Torsvik, G. Shephard, A. Talsma, M. Gurnis, M. Turner, S. Maus, M. Chandler, (2012). Global continental and ocean basin reconstructions since 200 Ma. Earth-Science Reviews, 113(3–4), 212-270. doi:<a href="https://doi.org/10.1016/j.earscirev.2012.03.002" target="_blank" rel="noopener">10.1016/j.earscirev.2012.03.002</a></div>
Ocean Basin Evolution and Global-Scale Plate Reorganization Events Since Pangea Breakup
<p><strong>Abstract </strong></p> <p>We present a revised global plate motion model with continuously closing plate boundaries ranging from the Triassic at 230 Ma to the present day, assess differences between alternative absolute plate motion models, and review global tectonic events. Relatively high mean absolute plate motion rates around 9–10 cm yr-1 between 140 and 120 Ma may be related to transient plate motion accelerations driven by the successive emplacement of a sequence of large igneous provinces during that time. A ~100 Ma event is most clearly expressed in the Indian Ocean and may reflect the initiation of Andean-style subduction along southern continental Eurasia, while an ~80 Ma acceleration of mean rates from 6 to 8 cm yr-1 reflects the initial northward acceleration of India and simultaneous speedups of plates in the Pacific. An event at ~50 Ma expressed in relative, and some absolute plate motion changes around the globe and in a reduction of global mean velocities from about 6 to 4–5 cm yr-1, indicates that an increase in collisional forces (such as the India-Eurasia collision) and ridge subduction events in the Pacific (such as the Izanagi-Pacific Ridge) play a significant role in modulating plate velocities.</p> <p><strong>Muller et al. (2016) AREPS model file versions</strong></p> <p>This model has been maintained for some time after initial publication. There are six versions of the model that we provide, including:</p> <ul> <li>v1.10 – Some minor fixes were made to plate topologies, and so conforms to the originally-published model.</li> <li>v1.11 – A back-arc basin north of Arabia was introduced in the Cretaceous (see note below), and hence slightly diverges from the original model in plate topologies, velocities, and seafloor age-grids for this region.</li> <li>v1.14 – The latest version of the model that has duplicated topology segments cleaned from the evolving polygons, which helps with quantifying plate boundary lengths in the resolved topology output.</li> <li>v1.15 – The correction to the pre-83 Ma Pacific rotations according to Torsvik et al. (2019) has been applied.</li> <li>v1.16 – Some fixes to topologies</li> <li>v1.17 – Major update to the seafloor age-grids and topologies. Age-grids are consistent with v1.15 and 1.16 as well. We strongly recommend you use this version of the model.</li> </ul> <p>Note about the evolution of the western Tethys in this model: The Western Tethys, north of Arabia, is punctuated by ophiolite formation and obduction in Cretaceous times. The first end-member involves applying the central and eastern Tethys analogues of back-arc opening and closure following ophiolite obduction, much like is usually implied in the Kohistan-Ladakh and Greater India collision zone. This scenario makes the Western Tethys north of Arabia consistent with the model of the eastern Tethys. However, a second end-member interpretation for the formation of many of the ophiolites in the region is that they develop when a mid-oceanic ridge inverts to become a subduction zone. Both options are plausible, but we implemented a change in this plate model after it was published to reflect the first end-member scenario in order to link the region to the eastern Tethys in a plausible way. This scenario is based on back-arc opening from ~125 Ma (Jolivet et al., 2016), with subduction of back-arc initiating in Albian times from ~110 Ma (Ghazi et at., 2003; Aygul et al., 2015). Obduction and Arabia collision with an arc occurs at 85 Ma (Jolivet et al., 2016; Jagoutz et al., 2016). The scenario is also consistent with the recent work of Morris et al. (2016) on the Oman Ophiolite.</p> <p> </p> <p>The agegrids associated with this model can be accessed at: <a href="https://repo.gplates.org/webdav/PlateModel_Age_SR_Grids/Muller_etal_2016_AREPS/" target="_blank" rel="noopener">https://repo.gplates.org/webdav/PlateModel_Age_SR_Grids/Muller_etal_2016_AREPS/</a></p>
Data from: Padfield et al. (2016) Rapid evolution of metabolic traits explains thermal adaptation in phytoplankton. Ecology letters.
<p>This repository provides the data from the TPC and logistic growth curves from the paper:</p> <p>Padfield, D., Yvon‐Durocher, G., Buckling, A., Jennings, S., & Yvon‐Durocher, G. (2016). Rapid evolution of metabolic traits explains thermal adaptation in phytoplankton. Ecology letters, 19(2), 133-142.</p> <p>metadata.pdf gives a more detailed explanation of the data.</p>
Evolution of cosmic star formation in the SCUBA-2 Cosmology Legacy Survey
<p>This dataset consists of tabulated data from the figures included in the referenced publication. The following datasets are included:</p> <p>Stacked SFR obscuration (IRX=IR/UV) of UVJ-selected star-forming galaxies:</p> <ul> <li>Weighted mean IRX as a function of Muv & stellar mass (Figure 12): MUV_irx1.dat</li> <li>Weighted mean IRX as a function of beta, over all masses and redshifts: beta_irx.dat</li> <li>Weighted mean IRX as a function of beta, binned by stellar mass (Figure 13): beta_irx_mstar.dat</li> <li>Weighted mean IRX as a function of beta, binned by redshift (Figure 14): beta_irx_z.dat</li> </ul> <p>Cosmic SFR density as a function of redshift for massive galaxies log(Ms/Msol)>10 (Figure 15):</p> <ul> <li>All mass-selected galaxies: sfrd_massive.dat</li> <li>UV-luminous galaxies Muv<M*; log(Ms/Msol)>10: sfrd_hiLUV.dat</li> <li>IR-luminous galaxies detected at 450µm: sfrd_IRdet.dat</li> </ul> <p> </p> <p>Cosmic SFR density as a function of redshift corrected to all stellar masses (Figure 16):</p> <ul> <li>All mass-selected galaxies: sfrd_uvlfcorr.dat</li> <li>UV-luminous galaxies Muv<M*; log(Ms/Msol)>10: sfrd_hiLUV_uvlfcorr.dat</li> </ul> <p>Full details of the binning and stacking methodology are explained in the paper.</p>
Investigating instability architectural smells evolution: an exploratory case study
<p>This is the dataset used in our case study on architectural smells evolution. We tracked smells from 524 versions among 14 open source Java systems. </p> <p>More information can be found in our ICSME'19 paper titled: "Investigating instability architectural smells evolution: an exploratory case study".</p> <p>Additionally, you can find the tool on GitHub: <a href="https://github.com/darius-sas/astracker">https://github.com/darius-sas/astracker</a></p>
Datasets and Jupyter notebook for the structural analysis of protein-RNA interface evolution
<p>The present repository contains data and code related to our manuscript "Structural comparison of protein-RNA homologous interfaces reveals widespread overall conservation contrasted with versatility in polar contacts". In the manuscript, we analyze the evolution of protein-RNA interfaces by building a dataset of protein-RNA interologs (homologous interfaces) and exploring how interface contacts are conserved between homologous interfaces, as well as possible explanations for non-conserved contacts.</p> <p>This repository contains the following files:</p> <ul> <li>DataAnalysisNotebook.ipynb is a Jupyter notebook to reproduce contact conservation analysis and all figures from our manuscript, and to explore data</li> <li>env.yaml is an environment file in order to build a Conda/Mamba environment to run the Jupyter notebook </li> <li>2022-02-21-PDB.csv contains data from the PDB about 3D structures of complexes containing interacting protein and RNA chains (PDB structure identifier, chain identifiers, experimental technique and resolution)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.tsv contains more detailed information about interacting protein and RNA chains from these complexes (PDB and chain identifiers, protein and RNA size, interface size and number of contacts)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.txt.selectXE_2.50_p30_r10_pi5_ri5_rep_bc-100.out_RNAcl_0.99.tsv contains the same detailed information, restricted to the filtered dataset used as a starting point in our interolog search pipeline</li> <li>PDBinterfaceAlign.csv contains information about the structural alignment of pairs of protein-RNA interactions (structural alignment TM-scores, sequence identity and coverage)</li> <li>DataInterologsParam.tsv contains information about a pre-filtered set of 2587 potential interologs (including interface RMSD, sequence identity and coverage and interface size)</li> <li>DataInterologsContactsFixedSASA.tsv contains detailed information about conserved and non-conserved contacts in the final set of 2022 interologs (atomic contacts, apolar contacts, hydrogen bonds, salt bridges and stacking information for aminoacid-nucleotide pairs, as well as information about whether each belongs to the interface, secondary structures, and the aminoacid surface accessibility and evolutionary conservation metrics) - compared to version 1, the calculation of solvent accessibility was fixed for a number of interolog pairs</li> <li>DataCons.csv contains precomputed contact conservation metrics for each of the 2022 interolog pairs, for fast reproduction of manuscript figures</li> <li>DataInterologsContactsResampledMaintainStructSeqId.tsv, DataInterologsContactsShuffled.tsv and DataInterologsShuffled.tsv relate to baselines computed for contact conservation assessment</li> <li>clan.txt, clan_membership.txt, ecod.latest.domains.uniq.txt, rfam_interfaces_977.txt, DataGroupsECOD.tsv, DataGroupesRFAM.tsv, DataGroupsRFAMClan.tsv, DataInterfaceGroupsECOD.tsv and DataInterfaceGroupsRFAM.tsv relate to the ECOD (respectively Rfam) classification of protein domains (respectively RNA) in protein-RNA interfaces from our dataset</li> <li>ListeIntraHbonds.pkl and ListeIntraSaltBridges.pkl are pickle-format data files containing intra-molecular hydrogen bonds and salt bridges (respectively) that are used to analyse scenarii of compensation for non-conserved polar contacts.</li> </ul>
Data set: Morphological evolution and niche conservatism across a continental radiation of Australian blindsnakes
<h1>Repository for "Morphological evolution and niche conservatism across a continental radiation of Australian blindsnakes"</h1> <p>---</p> <p>These data scripts were used to perform analyses included in the research paper "Morphological evolution and niche conservatism across a continental radiation of Australian blindsnakes" </p> <p>Main questions for the study:</p> <p>1. What are the main axes of morphological variation?<br>2. Does variation in morphology among species correlate with their current environments? <br>3. Are lineages that occupy ecologically similar habitats morphologically convergent? <br>4. Is speciation predominantly allopatric or sympatric? <br>5. Do sister species have greater morphological and ecological niche overlap than expected relative to non-sister species pairs?</p> <h2>## Data structure</h2> <p>Contents in the data folder is archived as a zip and can be downloaded from Zenodo (for all versions see https://zenodo.org/doi/10.5281/zenodo.10397830). Once you unzip the zipped files, you will see three folders and some files that are no in any folders. </p> <p>/data/ - files that were manually created and the phylogeny</p> <p>/data/script_generated_data/ - A combination of processed data needed to run the analyses </p> <p>/data/dorsal/ - photographs of the head from the dorsal view. These photos were used for digitising landmarks and semilandmarks. </p> <p>/data/worldclim2_30s/ - cropped and merged annual temperature from WorldClim2 (Fick and Hijmans 2017), soil bulk density from <a href="https://esoil.io/TERNLandscapes/Public/Pages/SLGA/GetData.html">Soil and Landscape Grid of Australia</a>, and Global Aridity Index from Zomer et al. (2022). <br><br>/DREaD/ - contains some files required to replicate DREaD analysis</p> <h2>## Code/Software</h2> <p>All scripts can be run using open source software. Scripts should be run in order to create necessary files that will be saved in /data/script_generated_data/ for further scripts. R is required to run R scripts (.R).</p> <h3>### /Code</h3> <p> - utility/*.R - scripts for custom functions. These are sourced in other scripts.<br> - DREaD/*.R - scripts associated with DREaD analyses<br> - 00_linear_measurement_shaperatio.R - script used to account for sexual dimorphism and calculate conventional PCA. Addresses Q1.<br> - 01_model_fitting.R - script used to address Q2 and plot visualisations.<br> - 02_convergence.R - this script calculates Ct1-4 and C5 scores. Addresses Q3.<br> - 02_convergence_model_fitting.R - this script evaluates fit of different evolutionary models to traits. Addresses Q3.<br> - 02_convergence_test_simulations.R - simulation studies to show that our phylogeny has sufficient power to detect convergence.<br> - 03_niche_enmtools_bias_account.R - calculates ecological niche models (ENMs) for each species using MAXENT. Runs Age-Overlap Correlation tests for geography and ENMs. Partially addresses Q4.<br> - 03_DREaD_Blindsnakes_AS.R - script to run DREaD analysis. <br> - 03_morpho_niche_overlap_plots.R - Runs Age-Overlap Correlation tests for body shape and snout shape. Plots AOCs. Partially addresses Q4. <br> - 04_pairwise_distance_test.R - Binomial tests between sister and non-sister pairs for ENMs and Geographic Range. Partially addresses Q5<br> - 04_morpho_pairwise.R - Binomial tests between sister and non-sister pairs for body shape and snout shape. Partially addresses Q5</p> <h2>## Contact</h2> <p>Should you have questions about these scripts or would like to request raw data, please do not hesitate to contact Sarin Tiatragul (contact information can be found in the paper) or on Github (https://github.com/stiatragul/blindsnakemorphoevo)</p> <h2>## References</h2> <p><a name="ref-fickWorldClim2017"></a>Fick, S. E., and R. J. Hijmans. 2017. <a href="https://doi.org/10.1002/joc.5086">WorldClim 2: New 1-km spatial resolution climate surfaces for global land areas</a>. International Journal of Climatology 37:4302–4315.</p> <p><a name="ref-zomerVersion2022"></a>Zomer, R. J., J. Xu, and A. Trabucco. 2022. <a href="https://doi.org/10.1038/s41597-022-01493-1">Version 3 of the global aridity index and potential evapotranspiration database</a>. Scientific Data 9:409.</p>
LukProt - an animal evolution-centric eukaryotic protein database
<p>LukProt is the EukProt database with additional species added, mostly the undersampled animal and some holozoan taxa. The database is composed of sequences translated from annotated genomes, transcriptomes or ESTs. <strong>The main purposes of the database are to consolidate sequences from undersampled animal taxa</strong> and provide usable search tools. The publication associated with LukProt can be found here: <a href="https://doi.org/10.1093/gbe/evae231">https://doi.org/10.1093/gbe/evae231</a>.</p> <p>The current version of the database (v1.5.1) is based on <a href="https://doi.org/10.24072/pcjournal.173">EukProt v3</a>. The home of all public versions of LukProt is this page (Zenodo).</p> <p>Proteomes that are novel in LukProt are denoted as LPXXXXX and those coming from AniProtDB are called APXXXXX. The sequence IDs from EukProt are conserved in LukProt. This means that each sequence is assigned an ID in the following format:</p> <pre><code>(A/E/L)PXXXXX_Species_epithet_(strain)_PYYYYYY</code></pre> <p>where XXXXX is a number from 00001 to 99999 and YYYYYY is a number from 000001 to 999999. Each sequence is assigned a unique number YYYYYY, and each taxon XXXXXX. All the IDs are compatible with BLAST v5 "-parse_seqids" option and the database can be readily deployed, for example on a server running <a href="https://doi.org/10.1093/molbev/msz185">SequenceServer</a>. Within each of the source fasta files, the source sequence identifier was kept after a blank space, so that it can still be retrieved if needed.</p> <p>A publicly available BLAST server providing LukProt search is available at: <a title="LukProt BLAST server" href="https://lukprot.hirszfeld.pl/" target="_blank" rel="noopener">https://lukprot.hirszfeld.pl/</a>.</p> <p>Comparison of EukProt v2/v3, LukProt 1.4.1 and LukProt v1.5.1 in their main areas of difference:</p> <table> <tbody> <tr> <th>Taxogroup</th> <th>EukProt v2</th> <th>EukProt v3</th> <th>LukProt v1.4.1</th> <th>LukProt v1.5.1</th> </tr> <tr> <th> <p>Holozoa</p> <p>(excluding Metazoa)</p> </th> <td>31</td> <td>40</td> <td>39</td> <td>43</td> </tr> <tr> <th>Ctenophora</th> <td>2</td> <td>2</td> <td>35</td> <td>38</td> </tr> <tr> <th>Porifera</th> <td>4</td> <td>5</td> <td>30</td> <td>47</td> </tr> <tr> <th>Placozoa</th> <td>2</td> <td>2</td> <td>3</td> <td>6</td> </tr> <tr> <th>Cnidaria</th> <td>3</td> <td>5</td> <td>65</td> <td>88</td> </tr> <tr> <th>Bilateria</th> <td>51</td> <td>51</td> <td>94</td> <td>142</td> </tr> </tbody> </table> <p>Included with the database are:</p> <ul> <li>ready to use main database files: <ul> <li><em>LukProt_v1.5.1_single_species_FASTA.7z</em> – a FASTA file with the sequences - <a href="https://en.wikipedia.org/wiki/7z">7-zipped</a>, <strong>uncompressed size: 17.6 GB</strong><br> <ul> <li>to concatenate all into one file, run this in the parent directory: <code>for file in $(find . -type f -name "*.fasta"); do awk 'FNR==1{print ""}1' $file >> LukProt_v1.5.1.fa; done</code>. This will create single FASTA file with all the sequences in the parent directory. <code>awk</code> is used to insert a new line after every file because <code>cat</code> would sometimes merge the last sequence with the header of the first sequence.</li> </ul> </li> <li><em>LukProt_v1.5.1_full_BLAST_db.7z</em> – a preformatted, full BLAST database (NCBI BLAST database format version: v5, masked with segmasker), <strong>uncompressed size: 28.3 GB</strong></li> <li><em>LukProt_v1.5.1_taxogroup_BLAST_db.7z</em> – a collection of BLAST databases where each proteome is one taxogroup and is placed within the eukaryotic tree of life directory structure, <strong>uncompressed size: 26.3 GB</strong></li> <li><em>LukProt_v1.5.1_single_species_BLAST_db.7z</em> – a collection of BLAST databases where each proteome is one BLAST database and is placed within the eukaryotic tree of life directory structure, <strong>uncompressed size: 26.4 GB</strong></li> </ul> </li> <li>auxiliary database files: <ul> <li><em>LukProt_v1.5.1.cdhit70.7z</em> – the full database clustered at 70% identity using CD-HIT with the following command: <code>cd-hit -g 1 -d 0 -T 20 -M 90000 -c 0.7 -uL 0.2 -uS 0.9 -s 0.2</code>, <strong>uncompressed sizes: fasta file - 11 GB, clstr file - 2.5 GB</strong></li> <li><em>LukProt_IDs_mapped.txt.gz</em> – a text file mapping the LukProt IDs to the AniProtDB IDs and EukProt IDs that are different</li> <li><em>BUSCO_tables.ods</em> – a spreadsheet with full result tables generated by BUSCO analysis</li> <li><em>OMAmer_output.zip</em> – a folder with full results of OMAmer analyses (includes per-sequence taxonomy classification)</li> <li><em>OMArk_output.zip</em> – a folder with the results of all OMArk analyses</li> </ul> </li> <li>metadata: <ul> <li><em>README.md</em> – a README file describing the metadata</li> <li><strong><em>LukProt_metadata_sheet.ods</em> – main metadata file. A spreadsheet with information about each proteome (in an open .ods format, most compatible with <a href="https://www.libreoffice.org/">LibreOffice</a>)</strong></li> <li><em>LukProt_metadata_other.zip</em> – an archive with other metadata files, documented in the README. Contents include:<br> <ul> <li>the LukProt taxonomy in various formats</li> <li>supporting scripts for data manipulation and visualization</li> </ul> </li> <li>a recoloring script (modified by LFS, originally by Dr. Celine Petitjean). The script is in <a title="formatFigtree2" href="https://doi.org/10.5281/zenodo.10654583">public domain</a> and reuploaded here only for convenience. </li> <li>other files - see README</li> </ul> </li> <li><em>changelog.md</em> – database changelog</li> </ul> <p>Words of caution:</p> <ul> <li>The database has been synchronized to EukProt v3 in version v1.5.1. This means that identifiers were modified in comparison to LukProt v1.4.1. The convention is not expected to change any more in future updates.</li> <li>Many proteomes, especially those transcriptome-based, may contain contamination from different species. In addition, the translation algorithms often introduce errors (e.g. the transcript may not represent a full length protein). For this reason, to get accurate sequences from each organism, users are directed to source data and to the included OMAmer, OMArk and BUSCO data for details.</li> <li>The taxonomy is different to UniEuk/EukMap, but UniEuk data were integrated where possible.</li> <li>A few NCBI taxids are missing and will be added in due course.</li> <li>Proteomes from NCBI and UniProt will be updated to current versions.</li> <li>A number of proteomes present in some metadata, are unpublished and were held back.</li> <li>While the database contains metadata that present a particular phylogeny of animals, holozoans and other eukaryotes, no particular claims or hypotheses are made by the author(s). However, in the future efforts will be made to name clades officially, once they are more firmly established.</li> </ul> <p><strong>Please report any problems or suggestions to Lukasz Sobala: lukasz.sobala (at) hirszfeld.pl.</strong></p> <p> </p> <p>Acknowledgements:</p> <ul> <li> <p>Andrew E. Allen Lab for creating the original <a href="https://allenlab.ucsd.edu/data/" target="_blank" rel="noopener">PhyloDB</a>.</p> </li> <li> <p>Daniel Richter <em>et al.</em> for creating <a href="https://doi.org/10.6084/m9.figshare.12417881">EukProt</a> and keeping it updated.</p> </li> <li> <p>Members of <a href="https://multicellgenome.com/">the Multicellgenome Lab</a>, especially Michelle Leger (for donating her database), for the bioinformatics support and for doing great science.</p> </li> <li> <p>All the authors of the original data.</p> </li> <li> <p>National Science Centre of Poland for funding of the project 2020/36/C/NZ8/00081, "The role of glycosylation in the emergence of animal multicellularity", which enabled the creation of this database.</p> </li> </ul>
Dataset of "Black Titanium Oxide/Activated TaS2 Flakes Photoelectrode for Plasmon Assisted Hydrogen Evolution at Neutral pH at High Current Density"
<p>Nanotubular structure of black titania with sputtered gold and incorporation of 3R-TaS2 self-activated flakes for high current density and neutral pH usage for hydrogen evolution reaction. Dataset consists of electrochemical data (LSV, EIS, CA), x-ray difractograms, Raman spectra, SEM images with EDX mapping, UV-vis spectra, DEMS records, ICP-MS records, XPS spectra and compositional analysis and BET records.</p>
Modulation of bioelectric cues in the evolution of flying fishes [Data set]
<p>Assembled reference contigs for protein-coding exons and conserved non-coding regions from targeted sequence enrichment of beloniform fishes. </p> <p>Current citation: Daane JM, Blum N, Lanni J, Boldt H, Iovine MK, Johnson SL, Lovejoy NR, and MP Harris. (2021). Novel regulators of growth identified in the evolution of fin proportion in flying fish. <em>bioRxiv. </em>doi: 10.1101/2021.03.05.434157</p> <p>-contigs.tar.gz contains the assembled contigs for each species. Each contig represents a targeted region with the addition of flanking DNA sequence</p> <p>-cnes.tar.gz contains the targeted conserved non-coding regions isolated from the larger contigs in contigs.tar.gz</p> <p>-exons.tar.gz contains the targeted protein coding exons isolated from the larger contigs in contigs.tar.gz</p> <p>-translated_exons.tar.gz contains the translated protein coding exons from exons.tar.gz</p> <p>-Beloniformes.tre is the species tree </p> <p>-medaka_cne_great.txt contains the associations between the assembled CNEs and neighboring protein-coding genes based on the GREAT approach </p>
Assembled chromosomes of the blood fluke Schistosoma mansoni provide insight into the evolution of its ZW sex-determination system
<p><em>Schistosoma mansoni </em>has a diploid genome of approximately 380 MB, organized in 7 pairs of autosomes and 2 sex chromosomes. The original <em>Schistosoma mansoni </em>Genome Project was completed by the Wellcome Sanger Institute in collaboration with The Institute for Genome Research using a Whole Genome Shotgun sequencing strategy. The draft assembly was subsequently improved first by incorporating Illumina reads from a clonal (single-miracidial) infection and more recently by incorporating long PacBio reads, HiC, and optical mapping data.</p> <p>Associated manuscript can be found at https://www.biorxiv.org/content/10.1101/2021.08.13.456314v1</p>
Imaging the footprint of nanoscale electrochemical reactions for assessing synergistic hydrogen evolution
<p>Dataset complementary to supporting information, such as optical movies, COMSOL model, and Python codes to analyze the experimental and simulated data according to the manuscript submitted for publication.<br> The movies correspond to cyclic voltammetry operando monitoring by optical microscopy of the reduction of water + KCl in the presence of NiCl2 at an ITO electrode or NiCl2 or MgCl2 at ITO electrode coated with Pt nanoparticles.</p> <p>The python function was used to extract the halo size around each nanoparticle from optical images, the python routines were used to postprocess the COMSOL simulation and evalaute the simulated halo size.</p>
Source Data for "Transport properties and doping evolution of the Fermi surface in cuprates"
<p>Source data for the publication "Transport properties and doping evolution of the Fermi surface in cuprates", in Scientific Reports (https://doi.org/10.1038/s41598-023-39813-z) and on arxiv (https://doi.org/10.48550/arXiv.2303.05254).</p> <p>This dataset is organized in the following way:</p> <p>For every figure of the manuscript there is a separate folder, which includes the figure itself, as well as one or more additional folders for the individual panels. In those, there are one or more .csv files with the data. Some of the .csv files have two header lines, for example when the temperature and <span class="math-tex">\(n_{\mathrm{H}}\)</span> are recorded for multiple doping levels.</p> <p>Additional comments:</p> <ul> <li>Figure 1 <ul> <li>The generic phase boundaries are not included.</li> <li>The precision of values of <span class="math-tex">\(n_{\mathrm{loc}}\)</span>is increased for presentation purposes</li> </ul> </li> <li>experimental doping values are typically rounded to 2 decimal points, doping errors to 3 decimal points</li> <li>estimated <span class="math-tex">\(n_{\mathrm{eff}}\)</span> are rounded to 5 decimal points</li> <li>experimental <span class="math-tex">\(n_{\mathrm{H}}\)</span> from the literature are rounded to 3 decimal points</li> <li>otherwise, if it exists, experimental values are typically rounded to the error</li> <li>temperature is always given in Kelvin</li> <li><span class="math-tex">\(C_2\)</span>is given in <span class="math-tex">\([\mathrm{TK}^{-2}]\)</span> (i.e. Tesla Kelvin^-2)</li> <li>the unit for <span class="math-tex">\(n_{\mathrm{eff}}\)</span>, <span class="math-tex">\(n_{\mathrm{loc}}\)</span>, <span class="math-tex">\(n_{\mathrm{H}}\)</span> is [per CuO2 unit cell]</li> <li>the unit for the resistivity in figure 4 is described in the methods section of the article</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.