Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

29,145

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

29,145 results for “Association”

Learn how ShareScore rates datasets ↗
edi52/100

Meta-analysis of ecosystem services associated with oyster restoration on the Eastern and Gulf coasts of the US

We conducted a meta-analysis to systematically quantify the success and uncertainty of oyster reef restoration for a suite of biological, biogeochemical, and physical ecosystem services relative to both degraded and natural reference habitats. We focused on the eastern oyster, Crassostrea virginica. To evaluate whether restored eastern oyster reefs enhance ecosystem services relative to unaltered, degraded habitats and whether restored reefs provide ecosystem services equivalent to reference reefs, we synthesized data and calculated log response ratios for 245 restored-degraded reef pairs and 136 restored-reference reef pairs from 106 publications collected along 3500 km of U.S. Gulf of Mexico and Atlantic coastlines.

openCustomMar 2023View details →
zenodo48/100

Fedora and Debian software package dependency networks along with description text associated with nodes

<p>Fedora (version 28) and Debian (version 9.5) software package dependency networks along with description text associated with nodes. Also includes learned vectors by using PCTADW-* as in &quot;Kexuan Sun, Shudan Zhong, and Hong Xu. 2020. Learning Embeddings of Directed Networks with Text-Associated Nodes---with Application in Software Package Dependency Networks. 2020 BigGraphs Workshop at IEEE BigData 2020.&quot;</p>

openmit-licenseSep 2018View details →
zenodo48/100

S49 | CPPDBLISTB | Database of Chemicals possibly (List B) associated with Plastic Packaging (CPPdb)

<p>This is the collection associated with list S49 CPPDBLISTB on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>S49 | CPPDBLISTB | <strong>Database of Chemicals associated with Plastic Packaging (CPPdb)</strong></p> <p>A database of chemicals likely (List A, 903 - in another upload) and possibly (List B, 3353 - this upload) associated with plastic packaging, with hazard data, from Groh et al 2019 DOI: <a href="https://doi.org/10.1016/j.scitotenv.2018.10.015">10.1016/j.scitotenv.2018.10.015</a>. Mapped to structures by CAS/Name by K. Groh &amp; E. Schymanski. 2025: added new CSV file with duplicate headers renamed.&nbsp;</p> <p>Latest version of original data (last update Oct 2018): DOI: <a href="http://doi.org/10.5281/zenodo.1287773">10.5281/zenodo.1287773</a></p>

opencc-by-4.0Mar 2019View details →
zenodo48/100

Results: Predicted cooling effect, deaths prevented and associated economic value from public green spaces in Paris V2

<p>This dataset represents results predicting the cooling effect, deaths prevented and associated economic value&nbsp; for public green spaces in Paris for 40 hot days above the minimum mortality threshold in 2019.&nbsp;</p> <p>This is version 2. The value of a statistical life (VSL) has been corrcted and all values adjusted.&nbsp;</p> <p>The data format is a shapefile with coordinate reference system RGF93 v1 / Lambert-93 (EPSG:2154).</p> <p>Please see the Variable_name csv file for description of the variable names.&nbsp;</p> <p>The (non-reproducible) code is available at https://github.com/j-k-garrett/REGREEN_Paris_heat</p> <p>These results are from the submitted (September 2025) paper entitled:</p> <p><strong><span>Nature-Based Solutions for Urban Heat: Health and Economic Value of Paris&rsquo;s Public Green Spaces</span></strong></p> <p>Authored by:</p> <p>Joanne K. Garrett<sup>1</sup>, David Neil Bird<sup>2</sup>, Timothy J. Taylor<sup>1</sup>, Elizabeth McCarthy<sup>3</sup>, David H. Fletcher<sup>4</sup>, Benedict W. Wheeler<sup>1</sup>, Marianne Zandersen<sup>5</sup>, Laurence Jones<sup>3</sup></p> <p><sup>1</sup>European Centre for Environment and Human Health, University of Exeter, Penryn, Cornwall, UK</p> <p><sup>2 </sup>Institute for Climate, Energy and Society, JOANNEUM RESEARCH, Graz, Austria</p> <p><sup>3</sup> Department of Environmental Studies, Schiller Institute for Integrated Science and Society, Boston College, USA</p> <p><sup>4</sup> UK Centre for Ecology &amp; Hydrology, Environment Centre Wales, Bangor, Gwynedd, Wales, UK</p> <p><sup>5 </sup>Department of Environmental Science, iClimate Interdisciplinary Centre for Climate Change, Aarhus University, Denmark</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo48/100

GWAS Summary Statistics from "Sex and statin-related genetic associations at the PCSK9 gene locus – results of genome-wide association meta-analysis"

<p>GWAMA summary statistics of PCSK9 levels stratified by sex and statin useage in Europeans.</p> <p>When using this data, please cite:</p> <p>Pott, J., Kheirkhah, A., Gadin, J.R.&nbsp;<em>et al.</em> Sex and statin-related genetic associations at the <em>PCSK9</em> gene locus: results of genome-wide association meta-analysis. <em>Biol Sex Differ</em> <strong>15</strong>, 26 (2024). https://doi.org/10.1186/s13293-024-00602-6</p> <p>All txt files contain the following columns:</p> <ul> <li>markername (unique SNP ID)</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>EA (effect allele)</li> <li>OA (other allele)</li> <li>EAF (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (in case of double-stratified data: number of studies; in case of single-stratified data: 2, as it is a meta-analysis of the two double-stratified data sets)</li> <li>beta (effect estimate)</li> <li>SE (standard error)</li> <li>pval (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>invalidAssoc (TRUE/FALSE flag if this variant was excluded in our analysis)</li> <li>reason4exclusion (reason why this SNP was excluded)</li> <li>phenotype (phenotyp setting)</li> </ul>

opencc-by-4.0Jan 2024View details →
zenodo48/100

Data Associated with Chemical Cartography with APOGEE: Two-process Parameters and Residual Abundances for 288,789 Stars from Data Release 17

<p>Stellar abundance measurements are subject to systematic errors that induce extra scatter and artificial correlations in elemental abundance patterns. &nbsp;We derive empirical calibration offsets to remove systematic trends with surface gravity log(g) in 17 elemental abundances of 288,789 evolved stars from the SDSS APOGEE survey. &nbsp;We fit these corrected abundances as the sum of a prompt process tracing core-collapse supernovae and a delayed process tracing Type Ia supernovae, thus recasting each star's measurements into the amplitudes A_cc and A_Ia and the element-by-element residuals from this two-parameter fit. Here we present the log(g)-calibrated abundances, fit parameters, process amplitudes, and element-by-element abundance residuals of 288,789 stars (310,427 spectra) in APOGEE DR17 that accompany <a href="https://arxiv.org/abs/2403.08067" target="_blank" rel="noopener">the paper</a>.</p> <p>calibration_values_final.dat contains all derived calibration offsets, including the grids of log(g) calibration offsets and zero-point offsets for two-process model analysis. The first five rows of this catalog are reproduced in Table 2 of the paper.</p> <p>logg_calib_example.ipynb is a Jupyter notebook containing Python code to load calibration_values_final.dat, extract the log(g) calibration offsets for specific element, and apply calibration offsets to 10 sample stars.</p> <p>2process_residual_abund_catalog_final.fits is the catalog of 310,427 APOGEE DR17 spectra (288,789 unique stars) containing calibrated abundances, two-process fit parameters, and abundance residuals. A full listing of columns in this catalog is given in Table 5 of the paper.</p> <p>catalog_examples.ipynb is a Jupyter notebook containing Python code to load 2process_residual_abund_catalog_final.fits, cross match with other catalogs (using AstroNN and the APOGEE DR17 Globular Cluster Value-Added Catalog as examples), and make some example plots utilizing the cross-matched data.</p>

opencc-by-4.0Mar 2024View details →
zenodo48/100

Associated data from: An end-to-end workflow to study newly synthesized mRNA following rapid protein depletion in Saccharomyces cerevisiae

<p>This dataset includes two custom BED files described in "An end-to-end workflow to study newly synthesized mRNA following rapid protein depletion in&nbsp;<em>Saccharomyces cerevisiae</em>" (Ridenour and Donczew, submitted), which were used to define counting windows for processing SLAM-seq data in SLAM-DUNK (version 0.4.3) [1]. The BED files contain all annotated open reading frames (ORFs) in the<em> Saccharomyces cerevisiae</em> genome or the <em>Schizosaccharomyces</em><em>&nbsp;pombe</em> genome and were created using BEDOPS (version 2.4.3) [2]. All ORFs were then extended 250 bp beyond their stop position to capture 3&prime; untranslated regions (UTRs) using SAMtools (version 1.14) [3] and BEDTools (version 2.30.0) [4]. The reference genome annotations for <em>S. cerevisiae</em> strain S288C (version R64-3-1, RefSeq Assembly GCF_000146045.2) and <em>S. pombe</em> strain 972h- (version ASM294v2, RefSeq Assembly GCF_000002945.1) were retrieved from the NCBI Datasets repository. The <em>S. cerevisiae </em>chromosome names were modified to reflect standard nomenclature (https://www.yeastgenome.org/).</p>

opencc-by-4.0Feb 2024View details →
zenodo48/100

Sex affects transcriptional associations with schizophrenia across the dorsolateral prefrontal cortex, hippocampus, and caudate nucleus

<p>This is supplementary data and source data for the manuscript,&nbsp;<em>"Sex affects transcriptional associations with schizophrenia across the dorsolateral prefrontal cortex, hippocampus, and caudate nucleus"</em>.</p> <p><strong>Abstract</strong>: Schizophrenia is a complex neuropsychiatric disorder with sexually dimorphic features, including differential symptomatology, drug responsiveness, and male incidence rate. Prior large-scale transcriptome analyses for sex differences in schizophrenia have focused on the prefrontal cortex. Analyzing BrainSeq Consortium data (caudate nucleus: n=399, dorsolateral prefrontal cortex: n=377, and hippocampus: n=394), we identified 831 unique genes that exhibit sex differences across brain regions, enriched for immune-related pathways. We observed X-chromosome dosage reduction in the hippocampus of male individuals with schizophrenia. Our sex interaction model revealed 148 junctions dysregulated in a sex-specific manner in schizophrenia. Sex-specific schizophrenia analysis identified dozens of differentially expressed genes, notably enriched in immune-related pathways. Finally, our sex-interacting expression quantitative trait loci analysis revealed 704 unique genes, nine associated with schizophrenia risk. These findings emphasize the importance of sex-informed analysis of sexually dimorphic traits, inform personalized therapeutic strategies in schizophrenia, and highlight the need for increased female samples for schizophrenia analyses.</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

Data associated with the following publication: "Giant thermoelectric response of confined electrolytes with thermally activated charge carrier generation"

<p>Data associated with the following publication: "Giant thermoelectric response of confined electrolytes with thermally activated charge carrier generation" (DOI: <a title="" href="https://doi.org/10.48328/tudatalib-1376">https://doi.org/10.48328/tudatalib-1376</a>)</p>

opencc-by-4.0Jan 2024View details →
zenodo48/100

Occurrences records of Herichthys labridens (Cichliformes: Cichlidae), with associated habitat information, in the Media Luna spring, San Luis Potosí, Mexico

<h2><strong>Introduction</strong></h2> <blockquote> <p>Occurrence records of the endemic cichlid <em>Herichthys labridens</em>, by adult and juvenile life stages, during three summer events (years of 1999, 2009, and 2019), in the Media Luna spring, San Luis Potos&iacute; Mexico.&nbsp;</p> </blockquote> <h2><strong>Material and Methods&nbsp;</strong></h2> <blockquote> <p>The occurrence records, ordered by adult and juvenile life stages, were obtained from two sources. For the summer of 1999, data were downloaded from the literature (Palacio-N&uacute;&ntilde;ez et al., 2010). For subsequent events, we recorded new data from 66 underwater transects distributed among 14 sectors (S1 to S14) in the Media Luna spring. We followed the method of Palacio-N&uacute;&ntilde;ez (2007), which maintained the transect location and sector boundaries of the summer of 1999 (Fig. 1a). The 20 m&sup2; transects were placed transversely to the current, from the edge to the central part of the canal (Fig. 1b). This sampling design was selected to meet two basic assumptions for studies of spatial distribution and habitat suitability: (1) the observations within the area are true and, (2) these observations delimit the initial position of the recorded individuals (Buckland &amp; Elston, 1993). The analysis of the spatial information of the sectors, the underwater transects, and the delimitation of the water surface was performed using the QGIS&reg; software version 3.4.8 (Menke, 2019).</p> <p>&nbsp;</p> <p><strong>Figure 1</strong>. <a href="https://zenodo.org/api/records/14231104/draft/files/Sector%20boundary_Transect%20location%20and%20sampling_Media%20Luna%20spring.jpeg/content" target="_blank" rel="noopener noreferrer">Sector boundary_Transect location and sampling_Media Luna spring.jpeg</a>. (a) Location of the transects in the Media Luna spring, Mexico. (b) Design scheme of the sampling transect; a CPVC pipe was used to give width to the edges of the transect and a nylon rope was attached to each side of the pipes to demarcate the length of the transect. Floating rubber buoys were added to the transects (at the edge towards the center of the canal) to prevent them from sinking into the sediment and to locate them among the vegetation. Transect scheme: Jorge Palacio-N&uacute;&ntilde;ez.</p> <p><br>In the summer events where we worked in field, we recorded the spatial location (i.e., GPS coordinates) of each individual and its life stage by direct observation with snorkel equipment and using a Garmin etrex device. The recorded&nbsp; information&nbsp; included the data of water depth and related underwater coverage. It is important to mention that, to prevent a repeat observation of the same organism or to ommit any individual, the transect was swaped slowly and in one direction only (i.e., from the center of the canal to the shore). We also used underwater cameras to validate the information. In adittion, the characterization of <em>H. labridens</em> individuals by life stage was performed by approximate size. For this purpose, previous studies on the life history and biology of the species were reviewed (Miller et al., 2005; De La Maza-Benignos &amp; Lozano-Vilano, 2013). It is worth mentioning that, during fieldwork, we avoided manipulation, damage, or unnecessary capture of the fish (e.g., Prchalov&aacute; et al., 2009).</p> <p><br>The databases by life stage were organized for each summer event, where, each observation record was included along with the associated habitat conditions. Subsequently, we depurated each database to remove atypical spatial data, data without information, incomplete data, or data with duplicate coordinates (Garc&iacute;a-Rosell&oacute; et al., 2014). Then, we performed spatial filtering of the remaining records to validate those that were within the study area, and to prevent that two or more points were within 0.1 m of each other. These steps of our analysis were performed using the software Qgis&reg; version 3.28.4 and Rstudio&reg; (Rstudio team, 2020). Subsequently, with the data set that included fish records, water depth, and underwater coverage variables, we performed a final environmental filter to rule out atypical records. This exploration was performed in Rstudio &reg; using the outliers function, starting from the lowest and highest quantiles.</p> </blockquote> <h2><strong>Results</strong></h2> <blockquote> <p>The final filtered databases were organized by life stage and summer event:</p> <p><strong>Adult:&nbsp;</strong></p> <table> <tbody> <tr> <td>Summer event</td> <td>Database</td> </tr> <tr> <td>1999</td> <td><a href="https://zenodo.org/api/records/14231104/draft/files/Occurrences_records_H_labridens_Adult_1999_Palacio-N%C3%BA%C3%B1ez%20et%20al.,%202010.csv/content" target="_blank" rel="noopener noreferrer">Occurrences_records_H_labridens_Adult_1999_Palacio-N&uacute;&ntilde;ez et al., 2010.csv</a></td> </tr> <tr> <td>2009</td> <td><a href="https://zenodo.org/api/records/14231104/draft/files/Occurrences_records_H_labridens_Adult_2009_Field%20work.csv/content" target="_blank" rel="noopener noreferrer">Occurrences_records_H_labridens_Adult_2009_Field work.csv</a></td> </tr> <tr> <td>2019</td> <td><a href="https://zenodo.org/api/records/14231104/draft/files/Occurrences_records_H_labridens_Adult_2019_Field%20work.csv/content" target="_blank" rel="noopener noreferrer">Occurrences_records_H_labridens_Adult_2019_Field work.csv</a></td> </tr> </tbody> </table> <p><strong>&nbsp;Juvenile:</strong></p> <table> <tbody> <tr> <td>Summer event</td> <td>Database</td> </tr> <tr> <td>1999</td> <td><a href="https://zenodo.org/api/records/14231104/draft/files/Occurrences_records_H_labridens_Juvenile_1999_Palacio-N%C3%BA%C3%B1ez%20et%20al.,%202010.csv/content" target="_blank" rel="noopener noreferrer">Occurrences_records_H_labridens_Juvenile_1999_Palacio-N&uacute;&ntilde;ez et al., 2010.csv</a></td> </tr> <tr> <td>2009</td> <td><a href="https://zenodo.org/api/records/14231104/draft/files/Occurrences_records_H_labridens_Juvenile_2009_Field%20work.csv/content" target="_blank" rel="noopener noreferrer">Occurrences_records_H_labridens_Juvenile_2009_Field work.csv</a></td> </tr> <tr> <td>2019</td> <td><a href="https://zenodo.org/api/records/14231104/draft/files/Occurrences_records_H_labridens_Juvenile_2019_Field%20work.csv/content" target="_blank" rel="noopener noreferrer">Occurrences_records_H_labridens_Juvenile_2019_Field work.csv</a></td> </tr> </tbody> </table> </blockquote> <p>&nbsp;</p> <blockquote> <p>These occurrence records for <em>H. labridens </em>are ready to be used in ecological niche modeling and spatial distribution studies. Also, these records can be used for other ecological and spatial studies, because each record (i.e., individual) included geoespatial coordinates, sector, location, and transect number. Also, we recorded information about the conditions of underwater coverage and water depth, which were asociated to each ocurrence record.</p> <p>For more information about several R codes where the previous databases can be used, visit the following repository URL: <a href="https://doi.org/10.5281/zenodo.7603557">https://doi.org/10.5281/zenodo.7603557</a>.</p> <p>Also, to download the UC and WDp variables to run the spatial and ecological modeling, visit the following repository URL:&nbsp;<a href="https://doi.org/10.5281/zenodo.7603890">https://doi.org/10.5281/zenodo.7603890</a>.</p> </blockquote>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Genome, repeat, and functional annotation associated with the naked mole-rat genome assembly, mHetGlaV3 (GCA_964261345.1)

<p>The naked mole-rat (NMR; Heterocephalus glaber) is a eusocial subterranean rodent with a highly unusual set of physiological traits, such as extreme longevity, that has attracted great interest amongst the scientific community. However, the genetic basis of most of these traits has not been elucidated. To facilitate our understanding of the molecular mechanisms underlying NMR physiology and behaviour, we generated a long-read chromosomal-level genome assembly of the NMR. This genome, mHetGlaV2, was subsequently annotated and incorporated into a &ldquo;91 eutherian mammals&rdquo; multiple whole genome alignment in Ensembl.&nbsp;</p> <p>We identified intra-chromosomal misassemblies within mHetGlaV2. We fixed these misassemblies by comparing syntenic blocks between this assembly and the Canadian Porcupine (EreDor) genome assembly (https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_028451465.1/) and a FISH-Karyotype of the naked mole-rat completed by Romanenko et al., 2023 (PMID: 380307020) to address any misassemblies and place centromeres. Chromosome numbering was identified from a composite karyogram of karyotypes from over 350 cells.&nbsp;This scaffold-corrected assembly is labelled mHetGlaV3 (https://www.ebi.ac.uk/ena/browser/view/GCA_964261345.1).</p> <p>This repository stores the repeat, genome, and epigenome annotations for HetGlaV3.</p> <p>mHetGlaV3.primary.gtf.gz. Gene structures and gene symbols are transferred from ENSEMBL annotations of mHetGlaV2 using liftOff with default parameters. Additional gene symbols were identified using TOGA and manual curation.</p> <p>mHetGlaV3.primary.gtf.gz. Simple repetitive regions and transposable elements were annotated using EarlGrey (https://github.com/TobyBaril/EarlGrey) using "Rodentia" annotations for RepeatMasker.</p> <p>mHetGlaV3.primary.genesymbol_table.txt.txt.gz. A tab-delimited file where rows are gene IDs and columns are gene symbols generated with each method. "Consensus" shows the best matching gene symbol for each gene ID.</p> <p>mHetGlaV3.primary_annotated_blacklist.bed.gz. Provides an assembly "blacklist" for mHetGlaV3. This blacklist is a bed file annotating assembly breakpoints between HetGlaV2 and HetGlaV3. This blacklist contains additional columns (e.g., closest gene, overlapping TE etc.) and should therefore be filtered to the first column before being incorporated into traditional genomic pipelines.</p> <p>mHetGlaV3.primary_hypothalamus_ABC_enhancer.bedpe.gz. Activity-By-Contact enhancers (https://github.com/broadinstitute/ABC-Enhancer-Gene-Prediction) generated in the female subordinate naked mole-rat hypothalamus using Hi-C-seq, ChIP-seq of H3K27Ac data, ATAC-seq, and RNA-seq information.</p> <p>mHetGlaV3.primary_hypothalamus_chromHMM.bed.gz. Chromatin states (using Chromhmm) annotating the female subordinate naked mole-rat hypothalamus using H3K4me3 (promoter), H4K4me2 (promoter-enhancer), H3K27Ac (active enhancer), H3K36me3 (elongated), H3K27me3 (polycomb repressed), H3K9me3 (heterochromatin), and CTCF (whole brain) ChIP-seq data, as well as ATAC-seq and RNA-seq data.</p> <p>mHetGlaV3.primary.fa.gz. Genome assembly fasta file for the naked mole-rat (V3, primary assembly). This assembly matches the primary assembly stored on ENA, however the chromosome names match these files, rather than have chromosome names processed by ENA (e.g. chr 1 instead of "OZ179169.1 Heterocephalus glaber genome assembly, chromosome: 1").</p> <p>&nbsp;</p> <p>UPDATES:</p> <p>* The 1.2 update fixed unscaffolded contig names from those used in-lab to those compatible with ENA.</p> <p>* The 1.3 update added small (50~100kbp) contigs onto mHetGlaV3.primary.fa.gz that were filtered before the ENA submission.</p> <p>* The 1.4 update fixed a small chromosome naming inconsistency spotted in the 1.3 update.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

DEM and associated kinematic GPS coordinates of September 2009 survey of the salar de Uyuni, Bolivia

<p>This dataset consists of two parts: &nbsp;1) the post-processed kinematic GPS coordinates of a September 2009 survey of a 45 x 54 km region of the salar de Uyuni, Bolivia. &nbsp;2) a digital elevation model (DEM) of the salar de Uyuni surface derived from those kinematic GPS data.</p> <p>Details of the survey design are identical to that from an earlier survey in 2002 and can be found in the manuscript, "Topography of the salar de Uyuni, Bolivia from kinematic GPS" (doi: 10.1111/j.1365-246X.2007.03604.x). &nbsp;The DEM is described in the manuscript "A Terrestrial Validation of ICESat Elevation Measurements and Implications for Gloval Reanalysis" (doi: 10.1109/TGRS.2019.2909739). The DEM was generated from fitting two-dimensional Fourier basis set with parameters: L_x = L_y = 70000 meters, m = n = 10. &nbsp;This results in a basis set with a nominal resolution of 7 km.</p> <p>The attached "salar_de_uyuni_2009_dem" files duplicate Figure 1 from the authors' "A terrestrial validation of ICESat elevation measurements and implications for global reanalyses," whose caption is:&nbsp;</p> <p>Landsat image of the salar de Uyuni, showing ICESat tracks 85, 241, 360 and 1320 (red) and the GPS-derived DEM from 2009 (color-coded with&nbsp;respect to mean elevation). The portion of each track plotted in Figure 2 is&nbsp;boxed in black. Total relief on the GPS DEM is less than 1 m over 50 km.</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

Data from 'Local Regions Associated With Interdecadal Global Temperature Variability in the Last Millennium Reanalysis and CMIP5 Models'

<p><strong>Abstract from &#39;<em>Local Regions Associated With Interdecadal Global Temperature Variability in the Last Millennium Reanalysis and CMIP5 Models</em>&#39;:</strong></p> <p>Despite the importance of interdecadal climate variability, we have a limited understanding of which geographic regions are associated with global temperature variability at these timescales. The instrumental record tends to be too short to develop sample statistics to study interdecadal climate variability, and Coupled Model Intercomparison Project, Phase 5 (CMIP5) climate models tend to disagree about which locations most strongly influence global mean interdecadal temperature variability. Here we use a new paleoclimate data assimilation product, the Last Millennium Reanalysis (LMR), to examine where local variability is associated with global mean temperature variability at interdecadal timescales. The LMR framework uses an ensemble Kalman filter data assimilation approach to combine the latest paleoclimate data and state-of-the-art model data to generate annually resolved field reconstructions of surface temperature, which allow us to explore the timing and dynamics of preinstrumental climate variability in new ways. The LMR consistently shows that the middle- to high-latitude north Pacific and the high-latitude North Atlantic tend to lead global temperature variability on interdecadal timescales. These findings have important implications for understanding the dynamics of low-frequency climate variability in the preindustrial era.</p>

opencc-by-4.0Aug 2019View details →
zenodo48/100

Dataset of "TWIST1 expression is associated with high-risk neuroblastoma and promotes primary and metastatic tumor growth"

<p>The embryonic transcription factors TWIST1/2 are frequently overexpressed in cancer, acting as multifunctional oncogenes. Here we investigate their role in neuroblastoma (NB), a heterogeneous childhood malignancy ranging from spontaneous regression to dismal outcomes despite multimodal therapy. We first reveal the association of TWIST1 expression with poor survival and metastasis in primary NB, while TWIST2 correlates with good prognosis. Secondly, suppression of TWIST1 by CRISPR/Cas9 results in a reduction of tumor growth and metastasis in immunocompromised mice. Moreover, TWIST1 knock-out tumors displays a less aggressive cellular morphology and a reduced disruption of the extracellular matrix (ECM) reticulin network. Additionally, we identify a TWIST1-mediated transcriptional program associated with dismal outcome in NB and involved in the control of pathways mainly linked to the signaling, migration, adhesion, the organization of the ECM, and the tumor cells versus tumor stroma crosstalk. Taken together, our findings identified TWIST1 as novel therapeutic target in NB.</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

Dataset for "TWIST1 expression is associated with high-risk neuroblastoma and promotes primary and metastatic tumor growth"

<p>The embryonic transcription factors TWIST1/2 are frequently overexpressed in cancer, acting as multifunctional oncogenes. Here we investigate their role in neuroblastoma (NB), a heterogeneous childhood malignancy ranging from spontaneous regression to dismal outcomes despite multimodal therapy. We first reveal the association of TWIST1 expression with poor survival and metastasis in primary NB, while TWIST2 correlates with good prognosis. Secondly, suppression of TWIST1 by CRISPR/Cas9 results in a reduction of tumor growth and metastasis in immunocompromised mice. Moreover, TWIST1 knock-out tumors displays a less aggressive cellular morphology and a reduced disruption of the extracellular matrix (ECM) reticulin network. Additionally, we identify a TWIST1-mediated transcriptional program associated with dismal outcome in NB and involved in the control of pathways mainly linked to the signaling, migration, adhesion, the organization of the ECM, and the tumor cells versus tumor stroma crosstalk. Taken together, our findings confirm&nbsp;TWIST1 as promising therapeutic target in NB.</p> <p>This dataset comprise images&nbsp;&nbsp;(.ndpi files) of anti-F4/80 IHC staining used for the quantification of macrophages in subcutaneous and orthotopic neuroblastoma xenografts derived from&nbsp;SK-N-Be2c cells expressing TWIST1 or knocked out for TWIST1 through CRISR/Cas9.</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

DATA SET: Peripheral microcirculatory alterations are associated with the severity of acute respiratory distress syndrome in COVID-19 patients admitted to intermediate respiratory and intensive care units

<p>This repository contains the data sets of the article:</p> <p>Mesquida, J., Caballer, A., Cortese, L.&nbsp;<em>et al.</em>&nbsp;Peripheral microcirculatory alterations are associated with the severity of acute respiratory distress syndrome in COVID-19 patients admitted to intermediate respiratory and intensive care units.&nbsp;<em>Crit Care</em>&nbsp;<strong>25,&nbsp;</strong>381 (2021). https://doi.org/10.1186/s13054-021-03803-2</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

Citation data of arXiv eprints and the associated quantitatively-and-temporally normalised impact metrics

<p><strong>Data collection</strong></p> <p>This dataset contains information on the eprints posted on arXiv from its launch in 1991 until the end of 2019 (1,589,006 unique eprints), plus the data on their citations and the associated impact metrics. Here, eprints include preprints, conference proceedings, book chapters, data sets and commentary, i.e. every electronic material that has been posted on arXiv.&nbsp;</p> <p>The content and metadata of the arXiv eprints were retrieved from the arXiv API (https://arxiv.org/help/api/) as of 21st January 2020, where the metadata included data of the eprint&rsquo;s title, author, abstract, subject category and the arXiv ID (the arXiv&rsquo;s original eprint identifier). In addition, the associated citation data were derived from the Semantic Scholar API (https://api.semanticscholar.org/) from 24th January 2020 to 7th February 2020, containing the citation information in and out of the arXiv eprints and their published versions (if applicable). Here, whether an eprint has been published in a journal or other means is assumed to be inferrable, albeit indirectly, from the status of the digital object identifier (DOI) assignment. It is also assumed that if an arXiv eprint received&nbsp;<em>c</em><sub>pre</sub>&nbsp;and&nbsp;<em>c</em><sub>pub</sub>&nbsp;citations until the data retrieval date (7th February 2020) before and after it is assigned a DOI, respectively, then the citation count of this eprint is recorded in the Semantic Scholar dataset as&nbsp;<em>c</em><sub>pre</sub>&nbsp;+&nbsp;<em>c</em><sub>pub</sub>. Both the arXiv API and the Semantic Scholar datasets contained the arXiv ID as metadata, which served as a key variable to merge the two datasets.</p> <p>The classification of research disciplines is based on that described in the arXiv.org website (https://arxiv.org/help/stats/2020_by_area/). There, the arXiv subject categories are aggregated into several disciplines, of which we restrict our attention to the following six disciplines: Astrophysics (&lsquo;astro-ph&rsquo;), Computer Science (&lsquo;comp-sci&rsquo;), Condensed Matter Physics (&lsquo;cond-mat&rsquo;), High Energy Physics (&lsquo;hep&rsquo;), Mathematics (&lsquo;math&rsquo;) and Other Physics (&lsquo;oth-phys&rsquo;), which collectively accounted for 98% of all the eprints. Those eprints&nbsp;tagged to multiple arXiv disciplines were counted independently for each discipline. Due to this overlapping feature, the current dataset contains a cumulative total of 2,011,216 eprints.&nbsp;</p> <p>Some general statistics and visualisations per research discipline are provided in the original article (Okamura, 2022), where the validity and limitations associated with the dataset are also discussed.</p> <p>&nbsp;</p> <p><strong>Description of columns (variables)</strong></p> <ul> <li><strong>arxiv_id</strong> :&nbsp;arXiv ID</li> <li><strong>category</strong> :&nbsp;Research discipline</li> <li><strong>pre_year</strong> :&nbsp;Year of posting v1 on arXiv</li> <li><strong>pub_year</strong> :&nbsp;Year of DOI acquisition</li> <li><strong>c_tot</strong> :&nbsp;No. of citations acquired during 1991&ndash;2019</li> <li><strong>c_pre</strong> :&nbsp;No. of citations acquired before and including the year of DOI acquisition</li> <li><strong>c_pub</strong> :&nbsp;No. of citations acquired after the year of DOI acquisition</li> <li><strong>c_<em>yyyy</em></strong>&nbsp;(<em>yyyy</em>&nbsp;= 1991, &hellip;, 2019) :&nbsp;No. of citations acquired in the year&nbsp;<em>yyyy</em>&nbsp;(with &lsquo;<em>yyyy</em>&rsquo; running from 1991 to 2019)</li> <li><strong>gamma</strong> :&nbsp;The quantitatively-and-temporally normalised citation index</li> <li><strong>gamma_star</strong> :&nbsp;The quantitatively-and-temporally standardised citation index</li> </ul> <p><em>Note:</em> The definition of the quantitatively-and-temporally normalised citation index (&gamma;; &lsquo;gamma&rsquo;) and that of the standardised citation index (&gamma;*; &lsquo;gamma_star&rsquo;) are provided in the original article (Okamura, 2022). Both indices can be used to compare the citational impact of papers/eprints published in different research disciplines at different times.&nbsp;</p> <p>&nbsp;</p> <p><strong>Data files</strong></p> <p>A comma-separated values file (&lsquo;<strong>arXiv_impact.csv</strong>&rsquo;) and a Stata file (&lsquo;<strong>arXiv_impact.dta</strong>&rsquo;) are provided, both containing the same information.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Genome- and transcriptome-wide association summary statistics for outcome from traumatic brain injury

<p>The dataset contains summary statistics for the genome- and transcriptome-wide association studies (GWAS, TWAS) of genetic effects on outcome in traumatic brain injury (TBI). The study participants attended hospital within 24 hours of TBI, and underwent head computed tomography imaging.</p> <p><strong>Study participants</strong></p> <p>European ancestry data set contains 4710 individuals; multi-ethnic cohort 5268 individuals, including Europeans (n = 4710), Africans (n = 245) and Admixed Americans (n = 313).</p> <p>The largest European population contribution was from CENTER-TBI (Collaborative European NeuroTrauma Effectiveness Research, https://www.center-tbi.eu), where each participating center (60 centers from 20 countries in Europe) recruited patients between December 2013 and December 2017. The patients recruited in CENTER-TBI were supplemented by subjects from cohorts recruited at two European centres (Cambridge, UK, and Turku, Finland).</p> <p>The majority of patients in the US cohort were recruited between 2014 and 2018 to TRACK-TBI (Transforming Research and Clinical Knowledge in TBI, https://tracktbi.ucsf.edu) by the 18 US participant sites. The subjects recruited to the US cohort from TRACK-TBI were supplemented by patients recruited to an institutional research initiative at Mass General Brigham (MGB).</p> <p><strong>Outcome definition</strong></p> <p>Outcomes were measured using the extended Glasgow Outcome Scale (GOSE), ranging from 1 (dead) to 8 (upper good recovery), measured 6 months post-TBI. TBI severity was specified using the Glasgow Coma Score (GCS), with TBI classified as mild (GCS 13-15), moderate (GCS 9-12), or severe (GCS 3-8).</p> <p>To account for the effect of injury severity on outcome, sliding dichotomization was used to categorize outcome as favourable or unfavourable. A GOSE &le; 4 was used to define an unfavourable outcome for patients with either moderate (GCS 9-12) or severe (GCS 3-8) TBI, while the unfavourable group was extended to patients with GOSE &le; 7 if they had mild (GCS 13-15) TBI.</p> <p><strong>Genotype data and imputation</strong></p> <p>Genotyping was completed at FIMM Technology Center for CENTER-TBI, Cambridge, Turku patients and the Broad Institute for TRACK-TBI, using the Illumina Global Screening Array (GSA-24v2-0 + Multi-Disease). The MGB cohort were genotyped using Illumina&rsquo;s Multi-Ethnic Global array (MEGA) and the pre-releases forms, including MEGA and MEGA-Ex arrays at Illumina at the MGB Translational Genomics Core.</p> <p>A unified quality control procedure was applied for each study cohort and the array-based genotypes were imputed using the Haplotype Reference Consortium panel. Autosomal chromosomes were considered, post-imputation data was filtered by imputation quality (INFO &gt; 0.4 for CENTER-TBI, Cambridge and Turku;&nbsp;R2 &gt; 0.4 for TRACK-TBI and MGB) and MAF &gt; 1%.</p> <p><strong>Genome-wide association analysis and meta-analysis</strong></p> <p>Genome-wide single-marker scans were performed using a penalized likelihood-based Firth logistic regression, and implemented in PLINK v2.0. Using favourable outcome as reference, models were fitted on the basis of imputed allelic dosages. Age, sex, major extracranial injury, pupillary reactivity, and the first 10 principal components were included as covariates. Study cohort (CENTER-TBI, Cambridge, Turku) was an additional covariate in the CENTER-TBI GWAS.</p> <p>Fixed-effects meta-analysis of the three European ancestry GWAS was performed using METAL. For trans-ethnic meta-analysis, summary statistics of five GWASs in patients of European, African and Admixed Americans were aggregated via MR-MEGA.</p> <p><strong>Transcriptome-wide association study</strong></p> <p>Genetically regulated gene expression (GREx) was imputed using a regression model fitted on a separate gene expression database. Elastic net models provided by PrediXcan for all available GTEx brain tissues and whole blood were used. For TWAS, the same sliding dichotomy model for outcome with the same set of covariates as in the GWAS, but PCA components were replaced with the top five principal components of the respective gene expression data.&nbsp;</p> <p><strong>Column headers - GWAS</strong></p> <p>rsID: variant rsID<br> Chrom: chromosome<br> Pos: position (build GRCh38)<br> A1: effect allele<br> A2: reference allele<br> EAF: allele frequency of effect allele<br> Effect: effect size of effect allele<br> StdErr: standard error of effect size<br> P: p value of association (with genomic correction)<br> N: sample size</p> <p>Note. &#39;Effect&#39; and &#39;StdErr&#39; are only available for the European ancestry meta-analysis.</p> <p><br> <strong>Column headers - TWAS</strong></p> <p>tissue: GTEx tissue type<br> id: ensembl gene id<br> coef: model coefficient<br> se: model standard error for coefficient<br> p: model-based p value<br> symbol: gene symbol<br> name: gene name written out<br> chr: chromosome<br> start: gene start position (build GRCh38)</p>

opencc-by-4.0Jan 2022View details →
zenodo48/100

A blood atlas of COVID-19 defines hallmarks of disease severity and specificity: Associated data

<p>This dataset contains&nbsp;raw and processed data&nbsp;from the COvid-19 Multi-omics Blood&nbsp;ATlas&nbsp;(COMBAT) consortium.&nbsp;Data are divided into 26 datasets&nbsp;representing&nbsp;anonymised&nbsp;raw and processed data from&nbsp;deep immune phenotyping of peripheral blood from COVID-19 patients.&nbsp;</p> <p>In addition to the data listed below, some datasets&nbsp;are&nbsp;available through other repositories:&nbsp;</p> <ul> <li> <p>Proteomics data&nbsp;(CBD-KEY-PROTEOMICS)&nbsp;is available at PRIDE</p> <ul> <li> <p>Accession number: PDX023175</p> </li> <li> <p>Contact: Roman&nbsp;Fischer</p> </li> </ul> </li> </ul> <ul> <li> <p>Genetic data and detailed clinical information&nbsp;are&nbsp;available via a data access&nbsp;agreement through&nbsp;EGA</p> <ul> <li> <p>Study accession: EGAS00001005493&nbsp;</p> </li> </ul> </li> </ul> <p>For further information regarding specific datasets, please contact the individuals listed in Dataset_descriptions.pdf through&nbsp;<a href="mailto:contact@combat.ox.ac.uk">contact@combat.ox.ac.uk</a>.&nbsp;</p>

opencc-by-4.0Aug 2021View details →
zenodo48/100

Data associated with "A weakened recurrent circuit in the hippocampus of Rett syndrome mice disrupts long-term memory representations"

<p><strong>Datasets used in <em>A weakened recurrent circuit in the hippocampus of Rett syndrome mice disrupts long-term memory representations.</em></strong></p> <p><strong>Datatypes:</strong></p> <ol> <li>Multi-index pandas dataframe (.pkl)</li> <li>Numpy array (.npy)</li> <li>Collection of numpy arrays (.npz)</li> <li>Python dictionary objects (.pkl)</li> </ol> <p><strong>Datasets:</strong></p> <p><strong>alignments.pkl: A dataframe containing numpy arrays of image displacements for each mouse in each memory context.</strong></p> <p>This multi-index dataframe has rows&nbsp;indexed by&nbsp;genotype (&#39;wt&#39; or &#39;het&#39;) and mouse_id. The columns are&nbsp;[&#39;T&#39;, &#39;F1&#39;, &#39;N1&#39;, &#39;F2&#39;, &#39;N2&#39;] for the training, recall 1-hour, neutral, recall 1-day, neutral day 2 memory contexts respectively. Each element of this dataframe is a numpy array of shape&nbsp; images x 2 that hold&nbsp;x and y image displacements respectively. These alignments are computed after the inscopix software motion correction and are used in Supplemental Figure 2 of the paper.</p> <p><strong>behavior_df.pkl: A dataframe of behavior readouts recorded by a camera positioned above the mice in each context chamber.</strong></p> <p>This multi-index dataframe has rows&nbsp;indexed by&nbsp;genotype (&#39;wt&#39; or &#39;het&#39;) and mouse_id.The columns are; sample times (*_time), freezing boolean arrays (*_freeze), x-positions in context chamber (*_x) and y-positions in the context chamber (*_y) for each context (*) in (&#39;Train&#39;, &#39;Fear&#39;, &#39;Neutral&#39;, &#39;Fear_2&#39;, &#39;Neutral_2&#39;).</p> <p><strong>correlated_pairs_df.pkl: A dataframe containing arrays of neuron indices that have a correlation in activity pattern &gt; 0.3.</strong></p> <p>This multi-index dataframe has rows&nbsp;indexed by&nbsp;genotype (&#39;wt&#39; or &#39;het&#39;) and mouse_id and treatment (&#39;NA&#39;). The columns contain [&#39;Train&#39;, &#39;Fear&#39;, &#39;Neutral&#39;, &#39;Fear_2&#39;, &#39;Neutral_2&#39;] representing each memory context. Each element of the dataframe is a numpy array with three columns. The first two columns are the neuron indices that are correlated and the last column is the strength of the correlation.</p> <p><strong>dredd_freezes_df.pkl: A dataframe containing freezing percentages for SOM-Cre and RTT-SOM-Cre mice treated with DREADDS.</strong></p> <p>This multi-index dataframe has rows&nbsp;indexed by&nbsp;genotype (&#39;wt&#39; or &#39;het&#39;) and mouse_id and treatment (mcherry, hm3d, hm4d). The columns contain one of [&#39;Neutral&#39;, &#39;Fear&#39;, &#39;Fear_2&#39;]. Each element of the dataframe is a freezing percentage for a single mouse. This dataframe is built from reading the dredd_behavior.xlsx excel file. This is used to generate figure 5E of the paper.</p> <p><strong>high_degree_df.pkl: A dataframe containing list of high degree neuron indices.</strong></p> <p>This multi-index dataframe has rows&nbsp;indexed by&nbsp;genotype (&#39;wt&#39; or &#39;het&#39;) and mouse_id and treatment (&#39;NA&#39;=not applicable since no DREADD used). The columns contain [&#39;Train&#39;, &#39;Fear&#39;, &#39;Neutral&#39;, &#39;Fear_2&#39;, &#39;Neutral_2&#39;] representing each memory context. Each element of the dataframe is a list of neuron indices that are high-degree cells.</p> <p><strong>N006_wt_basis.npz: a dict containing three&nbsp;numpy arrays representing the basis images for mouse N006 of genotype wild-type.</strong></p> <p>This dict has three&nbsp;arrays stored under the variable names &#39;U&#39;,&nbsp;&#39;sigma&#39; and &#39;img_shape&#39;. U is a matrix of column vector basis images. Each column is the vector representation of a basis image (row pixels x column pixels). There are 220 basis images (columns) in U. The sigma variable is the singular value associated with each basis image vector in U. img_shape can be used to reshape each basis column vector into a 2-D image for viewing. This data is used in Supplemental Figure 2 of the paper.</p> <p><strong>N006_wt_cxtbasis.pkl: A dictionary containing arrays for basis images and singular values for each context.</strong></p> <p>This dictionary has keys, [&#39;Train&#39;, &#39;Fear&#39;, &#39;Neutral&#39;, &#39;Fear_2&#39;,&nbsp; &#39;Neutral_2&#39;] representing the memory contexts. Each value is a 2 element list containing the U-basis images as column vectors and singular values, one per basis image in U. The shape of the basis images is the same shape stored&nbsp;in N006_wt_basis.pkl. This dataset is used in Supplementary Figure 2 to track cells across contexts of the CFC task (see also N006_wt_cxtsources.pkl)</p> <p><strong>N006_wt_cxtsources.pkl: A dictionary containing the independent component source images computed from the basis images for automatically identifying regions of interest (ROIs).&nbsp;</strong></p> <p>The dictionary is keyed on&nbsp; [&#39;Train&#39;, &#39;Fear&#39;, &#39;Neutral&#39;, &#39;Fear_2&#39;,&nbsp; &#39;Neutral_2&#39;] contexts. Each value in the dictionary at a given key is a 3-D numpy array of shape sources x height x width. These data were used to construct the source images and max intensity projection image of the sources in Supplemental Figure 2F-J&nbsp;of the paper.</p> <p><strong>N006_wt_rois.pkl: A dictionary containing the boundaries and annuli coordinates of all rois for mouse N006 of genotype wild-type.</strong></p> <p>This dictionary is keyed on&nbsp;[&#39;boundaries&#39;, &#39;annuli&#39;] contexts and each value is a 179 element list of&nbsp;arrays of boundary line coordinates or annulus point coordinates one&nbsp; per ROI&nbsp;detected for this mouse.</p> <p><strong>N006_wt_sources.npy: A numpy array containing all source images computed from all contexts of the CFC task for mouse N006 of genotype wild-type.</strong></p> <p>This numpy array has shape n x height x width where n=205 source images, height=517 pixels and width=704 pixels. This data was used to construct Supplemental Figure 3F.</p> <p><strong>N019_wt_basis.npz: a dict containing three&nbsp;numpy arrays representing the basis images for mouse N019&nbsp;of genotype wild-type.</strong></p> <p>This dict has three&nbsp;arrays stored under the variable names &#39;U&#39;,&nbsp;&#39;sigma&#39; and &#39;img_shape&#39;. U is a matrix of column vector basis images. Each column is the vector representation of a basis image (row pixels x column pixels). There are 220 basis images (columns) in U. The sigma variable is the singular value associated with each basis image vector in U. img_shape can be used to reshape each basis column vector into a 2-D image for viewing. This data is used in Figure 1C&nbsp;of the paper.</p> <p><strong>N019_wt_sources.npy: A numpy array containing all source images computed from all contexts of the CFC task for mouse N019&nbsp;of genotype wild-type.</strong></p> <p>This numpy array has shape n x height x width where n=204&nbsp;source images, height=516&nbsp;pixels and width=698&nbsp;pixels. This data was used to construct Figure 1C of the paper.</p> <p><strong>P80_animals.pkl: A pandas multi-index object containing the genotype, mouse_id and treatment of the top 80% behavioral performance animals.</strong></p> <p>In this study, we drop the lowest 20% performing WT and RTT animals based on freezing percentage during the recall contexts. This multi-index is used to filter the data before each computation or plot in this study. So for example Figure 1B contains only the top 80% performing WT and RTT mice.</p> <p><strong>pc_sipscs_amps.pkl: A dictionary containing the amplitudes of spontaneous IPSCs recorded in pyramidal cells of&nbsp;WT and RTT mice.</strong></p> <p>This dictionary is keyed on [&#39;wt&#39;, &#39;mecp2_pos&#39;, &#39;mecp2_neg&#39;] representing whether the pyramidal cell was recorded from a wild-type mouse (&#39;wt&#39;) or is an MeCP2 negative or MeCP2 positive RTT cell. This value&nbsp;under each key is an array of IPSC amplitudes, one per recorded cell. This data was used to construct Figure 4C in the paper.</p> <p><strong>pc_sipscs_freqs.pkl: A dictionary containing the frequencies&nbsp;of spontaneous IPSCs recorded in pyramidal cells of WT and RTT mice.</strong></p> <p>This dictionary is keyed on [&#39;wt&#39;, &#39;mecp2_pos&#39;, &#39;mecp2_neg&#39;] representing whether the pyramidal cell was recorded from a wild-type mouse (&#39;wt&#39;) or is an MeCP2 negative or MeCP2 positive RTT cell. This value&nbsp;under each key is an array of IPSC frequencies, one per recorded cell. This data was used to construct Figure 4C in the paper.</p> <p><strong>rois_df.pkl: A multi-index dataframe containing all ROI information for each non-DREADD treated cell in this study (Figures 1-3).</strong></p> <p>This dataframe index contains the genotype (&#39;wt&#39;, &#39;het&#39;), the mouse_id, the treatment (&#39;NA&#39;=not applicable since no DREADD used), and the cell index starting from 0. The columns are [&#39;centroid&#39;, &#39;cell_boundary&#39;, &#39;annulus_boundary&#39;]. The centroid for each cell is a 2-tuple of row, column pixel centroid coordinates. The cell_boundary is a two-column array of row, col boundary points for each ROI. The annulus_boundary is a two-column array of row, column interior points in the annulus. The annulus region&nbsp; excludes points of overlap with nearby cell bodies (See STAR methods of the paper).</p> <p><strong>signals_df.pkl: A multi-index dataframe containing calcium signals, inferred spikes and metadata for all Non-DREADD experiments used in this study (Figs 1-3).</strong></p> <p>This dataframe index contains the genotype (&#39;wt&#39;, &#39;het&#39;), the mouse_id, the treatment (&#39;NA&#39;=not applicable since no DREADD used), and the cell index starting from 0 and going up to 5771 cells. The columns are&nbsp;[&#39;channels&#39;, &#39;channel&#39;, &#39;num_pages&#39;, &#39;width&#39;, &#39;height&#39;, &#39;bits&#39;, &#39;Train_signals&#39;, &#39;Fear_signals&#39;, &#39;Neutral_signals&#39;, &#39;Cue_signals&#39;, &#39;Fear_2_signals&#39;, &#39;Neutral_2_signals&#39;, &#39;Cue_2_signals&#39;, &#39;Train_spikes&#39;, &#39;Fear_spikes&#39;, &#39;Neutral_spikes&#39;, &#39;Cue_spikes&#39;, &#39;Fear_2_spikes&#39;, &#39;Neutral_2_spikes&#39;, &#39;Cue_2_spikes&#39;, &#39;sample_rate&#39;]. The channels are all the recorded channels, the channels is the channel on which ROIs were detected, the width and height are the image dimensions, the bits is the image bit depth of the calcium movie. The *_signals&#39; are the df/f signals for each cell in each context. Each signal is a numpy array with the first 800 samples have been set to NAN due to settling time of the miniscope.&nbsp;The&nbsp;&#39;*_spikes&#39; are the inferred spikes for each cell stored as an image index. This signal and spike indices&nbsp;can be converted to time using the&nbsp;sample column. This dataframe is used in the construction of Figures 1-3 in the paper.</p> <p><strong>som_behavior_df.pkl: A dataframe of behavior readouts recorded by a camera positioned above the mice in each context chamber.</strong></p> <p>This multi-index dataframe has rows&nbsp;indexed by&nbsp;genotype (&#39;wt&#39; or &#39;het&#39;) and mouse_id. The columns are; sample times (*_time), freezing boolean arrays (*_freeze), x-positions in context chamber (*_x) and y-positions in the context chamber (*_y) for each context in *=(&#39;Train&#39;, &#39;Fear&#39;, &#39;Neutral&#39;, &#39;Fear_2&#39;, &#39;Neutral_2&#39;). This dataframe was not used in the paper but may still be useful for further analysis.</p> <p><strong>som_sepsc_amplitudes:</strong>&nbsp;<strong>A dictionary containing the amplitudes of spontaneous EPSCs recorded in SOM cells of WT and RTT mice with and without MeCP2.</strong></p> <p>A dictionary with keys [&#39;som&#39;, &#39;som_rett_pos&#39;, &#39;som_rett_neg&#39;] for WT SOM and RTT-SOM cells with and without MeCP2 respectively. Each value is a list of sEPSC amplitudes. This data was used&nbsp;in Figure 4E-G.</p> <p><strong>som_sepsc_freqs:</strong>&nbsp;<strong>A dictionary containing the amplitudes of spontaneous EPSCs recorded in SOM cells of WT and RTT mice with and without MeCP2.</strong></p> <p>A dictionary with keys [&#39;som&#39;, &#39;som_rett_pos&#39;, &#39;som_rett_neg&#39;] for WT SOM and RTT-SOM cells with and without MeCP2 respectively. Each value is a list of sEPSC frequencies. This data was used&nbsp;in Figure 4E-G.</p> <p><strong>som_signals_df.pkl:&nbsp;A multi-index dataframe containing calcium signals, inferred spikes and metadata for all Non-DREADD SOM cell recordings used in this study (Figs 5).</strong></p> <p>This dataframe index contains the genotype (&#39;wt&#39;, &#39;het&#39;), the mouse_id, the treatment (&#39;NA&#39;=not applicable since no DREADD used), and the cell index starting from 0 and going up to 710&nbsp;cells. The columns are&nbsp;[&#39;channels&#39;, &#39;channel&#39;, &#39;num_pages&#39;, &#39;width&#39;, &#39;height&#39;, &#39;bits&#39;, &#39;Train_signals&#39;, &#39;Fear_signals&#39;, &#39;Neutral_signals&#39;, &#39;Cue_signals&#39;, &#39;Fear_2_signals&#39;, &#39;Neutral_2_signals&#39;, &#39;Cue_2_signals&#39;, &#39;Train_spikes&#39;, &#39;Fear_spikes&#39;, &#39;Neutral_spikes&#39;, &#39;Cue_spikes&#39;, &#39;Fear_2_spikes&#39;, &#39;Neutral_2_spikes&#39;, &#39;Cue_2_spikes&#39;, &#39;sample_rate&#39;]. The channels are all the recorded channels, the channels is the channel on which ROIs were detected, the width and height are the image dimensions, the bits is the image bit depth of the calcium movie. The *_signals&#39; are the df/f signals for each cell in each context. Each signal is a numpy array with the first 800 samples have been set to NAN due to settling time of the miniscope.&nbsp;The&nbsp;&#39;*_spikes&#39; are the inferred spikes for each cell stored as an image index. This signal and spike indices&nbsp;can be converted to time using the&nbsp;sample column. This data was used to construct Figure 5B-C.</p> <p><strong>ssn33_sstcre_basis.npz:&nbsp;a dict containing three&nbsp;numpy arrays representing the basis images for mouse ssn33&nbsp;of genotype sst-cre.</strong></p> <p>This dict has three&nbsp;arrays stored under the variable names &#39;U&#39;,&nbsp;&#39;sigma&#39; and &#39;img_shape&#39;. U is a matrix of column vector basis images. Each column is the vector representation of a basis image (row pixels x column pixels). There are 100 basis images (columns) in U. The sigma variable is the singular value associated with each basis image vector in U. img_shape can be used to reshape each basis column vector into a 2-D image for viewing. This data is used in Figure 5A&nbsp;of the paper.</p> <p><strong>ssn33_sstcre_sources.npy: A numpy array containing all source images computed from all contexts of the CFC task for mouse N019&nbsp;of genotype wild-type.</strong></p> <p>This numpy array has shape n x height x width where n=86&nbsp;source images, height=516&nbsp;pixels and width=654&nbsp;pixels. This data was used to construct Figure 5A&nbsp;of the paper.</p>

opencc-by-4.0Feb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record