Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,694
datasets available to search
ShareScore release 0.7.1
Dataset results
4,694 results for “data analysis”
Data from the parametric analysis of multi-ring masonry arches using a limit analysis approach for the span of 3m
<p>The dataset is produced by running a series of 360 simulations using the in-house code ALMA (<em>Analisi Limite Murature Attritive</em>) that implements the upper bound approach of limit analysis to detect the collapse multiplier and mechanism for masonry structures. This set of data contains the results of 120 simulations achieved for the multi-ring masonry arches for span of 3m with different level of parameters considered, namely size of blocks, span of the arch, ring number, interlocking coefficient and the friction angle. In order to simplify references to figures an acronym system is used and it follows a sequence of "<em>size_span_ring_interlock_friction</em>". Size takes attributes as <strong>S</strong>-small and <strong>B</strong>-big while span takes number attributes based on the span like <strong>S3</strong>, <strong>S5</strong>, <strong>S7</strong> for spans of <strong>3</strong>, <strong>5</strong> and <strong>7</strong> meters, respectively. Ring number similarly is based on the number of rings as <strong>R2</strong>, <strong>R3</strong>, <strong>R4</strong> and <strong>R5</strong> and interlock takes the following attributes, I<strong>00</strong>, <strong>I15</strong>, <strong>I35</strong> and <strong>I50</strong> for the interlocking percentage considered such that <strong>00</strong>-stacked, <strong>15%</strong>, <strong>35%</strong> and <strong>50%</strong>, respectively. Finally friction takes attributes following the angle of friction such as <strong>F25</strong>—low level, <strong>F30</strong>—medium level, and <strong>F35</strong>—high level. The acronym used for the equivalent one-ring arches is used as simply <strong>EQ</strong>. For example, the acronym “<em><strong>S_S7_R4_I35_F30</strong></em>" refers to the arch with small blocks, span of 7 meters, consisting of 4 rings and by blocks interlocked at 35% with a joint friction of 30<sup>o</sup>.</p> <p>This database contains a <strong>*.txt</strong>, a <strong>*.vtk</strong> and a <strong>*.png</strong> file. In the .txt file the elapsed time and the collapse multiplier of each simulation can be found. The .vtk file contains all the geometry and displacement values of every masonry panel. Finally, the .png file presents the collapse mechanism obtained. </p>
Data from the parametric analysis of multi-ring masonry arches using a limit analysis approach for the span of 7m
<p>The dataset is produced by running a series of 360 simulations using the in-house code ALMA (<em>Analisi Limite Murature Attritive</em>) that implements the upper bound approach of limit analysis to detect the collapse multiplier and mechanism for masonry structures. This set of data contains the results of 120 simulations achieved for the multi-ring masonry arches for span of 7m with different level of parameters considered, namely size of blocks, span of the arch, ring number, interlocking coefficient and the friction angle. In order to simplify references to figures an acronym system is used and it follows a sequence of "<em>size_span_ring_interlock_friction</em>". Size takes attributes as <strong>S</strong>-small and <strong>B</strong>-big while span takes number attributes based on the span like <strong>S3</strong>, <strong>S5</strong>, <strong>S7</strong> for spans of <strong>3</strong>, <strong>5</strong> and <strong>7</strong> meters, respectively. Ring number similarly is based on the number of rings as <strong>R2</strong>, <strong>R3</strong>, <strong>R4</strong> and <strong>R5</strong> and interlock takes the following attributes, I<strong>00</strong>, <strong>I15</strong>, <strong>I35</strong> and <strong>I50</strong> for the interlocking percentage considered such that <strong>00</strong>-stacked, <strong>15%</strong>, <strong>35%</strong> and <strong>50%</strong>, respectively. Finally friction takes attributes following the angle of friction such as <strong>F25</strong>—low level, <strong>F30</strong>—medium level, and <strong>F35</strong>—high level. The acronym used for the equivalent one-ring arches is used as simply <strong>EQ</strong>. For example, the acronym “<em><strong>S_S7_R4_I35_F30</strong></em>" refers to the arch with small blocks, span of 7 meters, consisting of 4 rings and by blocks interlocked at 35% with a joint friction of 30<sup>o</sup>.</p> <p>This database contains a <strong>*.txt</strong>, a <strong>*.vtk</strong> and a <strong>*.png</strong> file. In the .txt file the elapsed time and the collapse multiplier of each simulation can be found. The .vtk file contains all the geometry and displacement values of every masonry panel. Finally, the .png file presents the collapse mechanism obtained. </p>
Data from the parametric analysis of multi-ring masonry arches using a limit analysis approach for the span of 5m
<p>The dataset is produced by running a series of 360 simulations using the in-house code ALMA (<em>Analisi Limite Murature Attritive</em>) that implements the upper bound approach of limit analysis to detect the collapse multiplier and mechanism for masonry structures. This set of data contains the results of 120 simulations achieved for the multi-ring masonry arches for span of 5m with different level of parameters considered, namely size of blocks, span of the arch, ring number, interlocking coefficient and the friction angle. In order to simplify references to figures an acronym system is used and it follows a sequence of "<em>size_span_ring_interlock_friction</em>". Size takes attributes as <strong>S</strong>-small and <strong>B</strong>-big while span takes number attributes based on the span like <strong>S3</strong>, <strong>S5</strong>, <strong>S7</strong> for spans of <strong>3</strong>, <strong>5</strong> and <strong>7</strong> meters, respectively. Ring number similarly is based on the number of rings as <strong>R2</strong>, <strong>R3</strong>, <strong>R4</strong> and <strong>R5</strong> and interlock takes the following attributes, I<strong>00</strong>, <strong>I15</strong>, <strong>I35</strong> and <strong>I50</strong> for the interlocking percentage considered such that <strong>00</strong>-stacked, <strong>15%</strong>, <strong>35%</strong> and <strong>50%</strong>, respectively. Finally friction takes attributes following the angle of friction such as <strong>F25</strong>—low level, <strong>F30</strong>—medium level, and <strong>F35</strong>—high level. The acronym used for the equivalent one-ring arches is used as simply <strong>EQ</strong>. For example, the acronym “<em><strong>S_S7_R4_I35_F30</strong></em>" refers to the arch with small blocks, span of 7 meters, consisting of 4 rings and by blocks interlocked at 35% with a joint friction of 30<sup>o</sup>.</p> <p>This database contains a <strong>*.txt</strong>, a <strong>*.vtk</strong> and a <strong>*.png</strong> file. In the .txt file the elapsed time and the collapse multiplier of each simulation can be found. The .vtk file contains all the geometry and displacement values of every masonry panel. Finally, the .png file presents the collapse mechanism obtained. </p>
Data for Progress on Climate Action: a Multilingual Machine Learning Analysis of the Global Stocktake
<p>Data to go with our submission to Climatic Change titled "Progress on Climate Action: a Multilingual Machine Learning Analysis of the Global Stocktake".</p> <p>Dataset contains the embeddings (.zip with pickles) as well as the associated document items (idem), the most-closely associated keywords and paragraphs per topic in the final model (.xlsx), the reduced 2d embeddings with all selected paragraphs (.csv utf-8 encoded), as well as an overview with the meta-data per source (.csv utf-8 encoded).</p>
The Authorship of Stephen King's Books Written Under the Pseudonym "Richard Bachman": A Stylometric Analysis (data)
<p>This data accompanies a paper for the 2nd Annual Conference for Computational Literary Studies: "The Authorship of Stephen King’s Books Written Under the Pseudonym 'Richard Bachman': A Stylometric Analysis".</p> <p><strong>Abstract</strong>:</p> <p>Between 1977 and 1984, Stephen King published five novels under the pseudonym “Richard Bachman”. Reviewers noted similarities between King’s and Bachman’s writing styles when <em>Thinner </em>(1984) was published, ultimately leading to King’s unmasking. We investigate, using the Juola protocol, whether computational techniques can correctly identify King as the author of the Bachman books out of a selection of contemporary candidate authors – Dean Koontz, Peter Straub, and Thomas Harris. We also perform a post-hoc analysis of the use of pop-culture references and brand names in Bachman, King, Koontz, Straub, and Harris novels, based on comments in reviews of Bachman and King novels. The references extracted from the Bachman books occurred significantly more often in King’s texts than in the others’, showing that attentive readers could have “heard King’s voice” in the Bachman books through what a reviewer denigratingly called King’s “compulsion to list brand-name products and his affinity for pop-cult teenage junk”. These results contribute to the vexed issue of explainability, which is a recurrent challenge in author identification for literary texts.</p> <p> </p> <p>Below is a description of each file in this repository:</p> <p><strong>bachman_segments_features_array_1000token_segments.csv</strong>, <strong>bachman_segments_features_array_5000token_segments.csv</strong>, and <strong>bachman_segments_features_array_10000token_segments.csv</strong> contain the feature spaces created by vectorizing 1,000-, 5,000-, and 10,000-token segments of Bachman, King, Koontz, and Straub books. Each row of the csv files contains the vectorized segment, the segment's author, the book the segment was drawn from, the book's publication date, and the segment number. </p> <p> </p> <p><strong>bachman_segments_author_candidate_cosine_distances_1000token_segments.csv</strong>, <strong>bachman_segments_author_candidate_cosine_distances_5000token_segments.csv</strong>,and <strong>bachman_segments_author_candidate_cosine_distances_10000token_segments.csv </strong>contain the Bachman segment number, bootstrap iteration number (from 0 and 9,999), the distractor author of the randomly-sampled segment, and the cosine distance between the Bachman segment vector and the distractor author's randomly-sampled segment vector (calculated using the data stored in the bachman_segments_features_array_1000token_segments.csv, bachman_segments_features_array_5000token_segments.csv, and bachman_segments_features_array_10000token_segments.csv files).</p> <p> </p> <p><strong>bachman_segments_author_candidate_ranks_1000token_segments.csv</strong>, <strong>bachman_segments_author_candidate_ranks_5000token_segments.csv</strong>,and <strong>bachman_segments_author_candidate_ranks_10000token_segments.csv </strong>contain the same columns as the 3 files described in the previous paragraph, but the cosine distance between Bachman segment and distractor author segment is converted to a ranking. For each bootstrap iteration there are 4 (one for each candidate author) rows containing the distance ranking between the Bachman segment and a candidate author segment. In a particular bootstrap iteration, if a King segment had the smallest cosine distance to a Bachman segment, King has the ranking "1", and if a Koontz segment had the second smallest distance to a Bachman segment, Koontz has the ranking "2", and so on. </p> <p> </p> <p><strong>predicted_author_candidate_raw_counts_1000token_segments.csv</strong>,<strong> predicted_author_candidate_raw_counts_5000token_segments.csv</strong>, and<strong> predicted_author_candidate_raw_counts_10000token_segments.csv </strong>contain the total number of times King, Straub, Harris, and Koontz segments received a certain distance ranking in the files described in the previous paragraph. </p> <p> </p> <p><strong>predicted_author_candidate_proportions_1000token_segments.csv</strong>, <strong>predicted_author_candidate_proportions_5000token_segments.csv</strong>, and and <strong>predicted_author_candidate_proportions_10000token_segments.csv</strong> contain a Bachman book title, and percentage of that book's segments that received the distance rankings 1-4 of each author. For example, in <strong>predicted_author_candidate_proportions_10000token_segments.csv, </strong><em>The Long Walk</em>'s segments were ranked as most similar (rank= "1") to King segments in 73.3% of bootstrap iterations. </p> <p> </p> <p><strong>pop_culture_refs_counts_books_10000token_segments.csv </strong>contains the author and book title of a randomly-sampled 10,000-token segment from the aforementioned book, the iteration (from 0 to 99), and the number of pop culture references found in the segment that match those extracted from Bachman books.</p> <p> </p> <p> </p> <p> </p>
Dissecting glial scar formation by spatial point pattern and topological data analysis
<p>These data were generated by the Laboratory of Neurovascular Interactions (https://elalilab.com/) at University Laval (Quebec, Canada), and reported in "Dissecting glial scar formation by spatial point pattern and topological data analysis". </p> <p>Please refer to the Open Science Framework (OSF) repository (https://osf.io/3vg8j/) or GitHub (https://github.com/elalilab/GlialScar_PPA-TDA_2022) to see the processing pipeline.</p> <p><strong>AUTHORS</strong><br> Manrique-Castano, Daniel; Bhaskar, Dhananjay; ElAli, Ayman</p> <p><strong>KEYWORDS</strong><br> Stroke, cerebral ischemia, brain injury, glial scar, reactive astrocytes, reactive microglia, </p> <p><br> <strong>1. STUDY DESCRIPTION </strong> <br> This research provides a quantitative analysis of reactive glia and glial scar formation in a mouse model of cerebral ischemia. The dataset in this repository consists of raw widefield microscopy images from healthy and ischemic animals. </p> <p><strong>2. EXPERIMENTAL CONDITIONS</strong><br> Six-month-old C57BL/6 mice were subjected to 30 minutes of cerebral ischemia by middle cerebral artery occlusion (MCAO). Brains were harvested at 5, 15, and 30 days post-ischemia (DPI) (see 10.5281/zenodo.3559570). 5 sham animals were included as controls. The full protocol for brain harvesting is available at 10.17504/protocols.io.4r3l27q5pg1y/v1. Brain sections were stained with NeuN, Gfap, and Iba1 antibodies to detect neurons and reactive glia after injury. Full protocol available at 10.17504/protocols.io.yxmvmk94og3p/v1 <br> <br> <strong>3. FILE DESCRIPTION</strong></p> <p><strong>- GT5X_Gfap_Iba1_NeuN.rar: </strong>Contain widefield (5x magnification) .tif images grouped by animals (5-7 images per animal; see research article for further details). The images were taken with the following parameters.</p> <p>Objective: Fluar 5x/0.25 M27<br> Scaling per pixel: 1.300 x 1.300 µm<br> Bit depth: 16 bit </p> <p>Stainings:<br> Neun Channel AF647; Excitation 653; Emission 668; Exposure 3 s<br> IBA1 Channel AFCy3; Excitation 458; Emission 561; Exposure 4 s<br> GFAP Channel AF488; Excitation 493; Emission 517; Exposure 1 s<br> DAPI Channel AF405; Excitation 353; Emission 465; Exposure 50 ms</p> <p>We used a FIJI script to pre-process the original .czi files. The script is shared in the GitHub repository under the name GT_Exp2_5x_GenerateTiffs.jim.</p> <p><strong>- GT10X_Gfap_Iba1_NeuN.rar:</strong> Contain a single widefield (10x magnification) .tif image per animal at the level of the MCA territory (see research article for further details). The images were taken with the following parameters.</p> <p>Objective: ECM paln-NeoFluar 10x/0.30 M27<br> Scaling per pixel: 0.45 x 0.45 µm<br> Bit depth: 16 bit </p> <p>Stainings:<br> Neun Channel AF647; Excitation 653; Emission 668; Exposure 200 ms<br> IBA1 Channel AFCy3; Excitation 458; Emission 561; Exposure 250 ms<br> GFAP Channel AF488; Excitation 493; Emission 517; Exposure 100 ms<br> DAPI Channel AF405; Excitation 353; Emission 465; Exposure 10 ms</p> <p><br> We used a FIJI script to pre-process the original .czi files. The script is shared in the GitHub repository under the name GT_Exp2_10x_GenerateTiffs.jim.<br> <br> For 5x and 10x images, the following naming strings apply:</p> <p>GT5x: Research project identifier indicating the magnification<br> M01(n): Animal ID<br> 5D(n): Days post-ischemia. 0D refers to healthy (naive) animals. <br> Scene1(n): Bregma level. Scene 1 corresponds to the most anterior area sampled, while Scene 6 or 7 is the most posterior.</p> <p><strong>- PointPatterns_10x.rds: </strong>2D point patterns of GFAP, IBA1, and NeuN generated by the r-package <em>spatstat</em>. The observation window comprises a horizontal ROI from the ventricular area to the outer border of the dorsolateral cerebral cortex. The point patterns were generated from the files and coordinates contained in the <strong>QupathProjects_10x.rar</strong> file in this repository. To reproduce the generation of point patterns please refer to the associated GitHub repository (https://github.com/elalilab/Stroke_GlialScar_PPA-TDA). </p> <p><strong>- PointPatterns_5x.rds: </strong>2D point patterns of GFAP, IBA1, and NeuN generated by the r-package <em>spatstat</em>. The observation window comprises the ischemic hemisphere. The point patterns were generated from the files and coordinates contained in the <strong>QupathProjects_5x.rar</strong> file in this repository. To reproduce the generation of point patterns please refer to the associated GitHub repository (https://github.com/elalilab/Stroke_GlialScar_PPA-TDA). </p> <p><strong>- QupathProjects_5x.rar: </strong>QuPath project folder for 5x images (GT5X_Gfap_Iba1_NeuN.rar). Each subfolder (per animal) contains the necessary files to import annotations (alignment to the Allen Brain Atlas) generated by ABBA (https://biop.github.io/ijp-imagetoatlas/). Please see the research article for further details. </p> <p><strong>**NOTE** </strong>Gfap, Iba1, and NeuN folders contain raw .tsv data originated by QuPath (cell counting). These folders are read in the R processing pipeline to extract the coordinates of each cell. Please make sure the whole folder is in the R working directory. The file "project.qpproj" in each folder opens the QuPath project in QuPath and reads the classifiers and data folders. Each folder also contains "_Alignement.json" and "_Registration_json" files generated during the alignment and annotation procedures in ABBA. However, when the route of the source images is changed, the plugin does not allow rerouting, and the files are of no practical use. The issue has been reported to the ABBA Github repository. </p> <p><strong>- QupathProjects_10x.rar:</strong> QuPath project folder for 10x images (GT5X_Gfap_Iba1_NeuN.rar). The folder contains the necessary files to import annotations (Alignment to the Allen Brain Atlas) generated by ABBA (https://biop.github.io/ijp-imagetoatlas/). Please see the research article for further details. </p> <p><strong>**NOTE** </strong>Gfap, Iba1, NeuN, and DAPI folders contain raw .tsv data originated by QuPath (cell counting). These folders are read in the R processing pipeline to extract the coordinates of each cell. Please make sure the whole folder is in the R working directory. The file "project.qpproj" opens the QuPath project in QuPath and reads the classifiers and data folders. </p>
Preprocessed Indonesian Twitter Dataset on UU Perlindungan Data Pribadi for Sentiment Analysis Research
<p>This dataset, titled 'Preprocessed Indonesian Twitter Dataset on UU Perlindungan Data Pribadi for Sentiment Analysis Research,' is curated and prepared for the purpose of conducting sentiment analysis research as outlined in the project 'ANALISIS SENTIMEN MASYARAKAT TERHADAP UU PERLINDUNGAN DATA PRIBADI PADA APLIKASI X DENGAN METODE SUPPORT VECTOR MACHINE' (Sentiment Analysis of the Community Towards the Personal Data Protection Law on Application X Using Support Vector Machine Method).</p>
Vegetation Density Across NYC: Analysis of Land Cover Data (2017) within 200 meter Buffers of Points
<p><strong>Summary:</strong></p><p>This repository contains spatial data files representing the density of vegetation cover within a 200 meter radius of points on a grid across the land area of New York City (NYC), New York, USA based on 2017 six-inch resolution land cover data, as well as SQL code used to carry out the analysis. The 200 meter radius was selected based on a study led by researchers at the NYC Department of Health and Mental Hygiene, which found that for a given point in the city, cooling benefits of vegetation only begin to accrue once the vegetation cover within a 200 meter radius is at least 32% (Johnson et al. 2020). The grid spacing of 100 feet in north/south and east/west directions was intended to provide granular enough detail to offer useful insights at a local scale (e.g., within a neighborhood) while keeping the amount of data needed to be processed for this manageable. </p><p>The contained files were developed by the NY Cities Program of <a href="https://www.nature.org/newyork">The Nature Conservancy</a> and the <a href="https://nyc-eja.org/">NYC Environmental Justice Alliance</a> through the <a href="https://medium.com/gage-nyc/introducing-the-just-nature-nyc-partnership-513612e8c3b4">Just Nature NYC Partnership</a>. Additional context and interpretation of this work is available in a <a href="https://medium.com/gage-nyc/looking-at-cooling-benefits-of-plants-through-nyc-vegetation-data-ccdeb33cbe17">blog post</a>.</p><p> </p><p><i>References:</i></p><p>Johnson, S., Z. Ross, I. Kheirbek, and K. Ito. 2020. Characterization of intra-urban spatial variation in observed summer ambient temperature from the New York City Community Air Survey. <i>Urban Climate</i> 31:100583. <a href="https://doi.org/10.1016/j.uclim.2020.100583">https://doi.org/10.1016/j.uclim.2020.100583</a></p><p> </p><p><strong>Files in this Repository:</strong></p><p>Spatial Data (all data are in the New York State Plane Coordinate System - Long Island Zone, North American Datum 1983, <a href="https://epsg.io/2263">EPSG 2263</a>):</p><p>Points with unique identifiers (<i>fid</i>) and data on proportion tree canopy cover (<i>prop_canopy</i>), proportion grass/shrub cover (<i>prop_grassshrub</i>), and proportion total vegetation cover (<i>prop_veg</i>) within a 200 meter radius (same data made available in two commonly used formats, Esri File GeoDatabase and GeoPackage):</p><p><i>nyc_propveg2017_200mbuffer_100ftgrid_nowater.gdb.zip</i></p><p><i>nyc_propveg2017_200mbuffer_100ftgrid_nowater.gpkg</i> </p><p>Raster Data with the proportion total vegetation within a 200 meter radius of the center of each cell (pixel centers align with the spatial point data)</p><p><i>nyc_propveg2017_200mbuffer_100ftgrid_nowater.tif</i></p><p>Computer Code:</p><p>Code for generating the point data in PostgreSQL/PostGIS, assuming the data sources listed below are already in a PostGIS database.</p><p><i>nyc_point_buffer_vegetation_overlay.sql</i></p><p> </p><p><strong>Data Sources and Methods:</strong></p><p>We used two openly available datasets from the City of New York for this analysis:</p><p>Borough Boundaries (Clipped to Shoreline) for NYC, from the NYC Department of City Planning, available at <a href="https://www.nyc.gov/site/planning/data-maps/open-data/districts-download-metadata.page">https://www.nyc.gov/site/planning/data-maps/open-data/districts-download-metadata.page</a> </p><p>Six-inch resolution land cover data for New York City as of 2017, available at <a href="https://data.cityofnewyork.us/Environment/Land-Cover-Raster-Data-2017-6in-Resolution/he6d-2qns">https://data.cityofnewyork.us/Environment/Land-Cover-Raster-Data-2017-6in-Resolution/he6d-2qns</a> </p><p>All data were used in the New York State Plane Coordinate System, Long Island Zone (<a href="https://epsg.io/2263">EPSG 2263</a>). Land cover data were used in a polygonized form for these analyses.</p><p>The general steps for developing the data available in this repository were as follows:</p><p>Create a grid of points across the city, based on the full extent of the Borough Boundaries dataset, with points 100 feet from one another in east/west and north/south directions</p><p>Delete any points that do not overlap the areas in the Borough Boundaries dataset.</p><p>Create circles centered at each point, with a radius of 200 meters (656.168 feet) in line with the aforementioned paper (Johnson et al. 2020).</p><p>Overlay the circles with the land cover data, and calculate the proportion of the land cover that was grass/shrub and tree canopy land cover types. Note, because the land cover data consistently ended at the boundaries of NYC, for points within 200 meters of Nassau and Westchester Counties, the area with land cover data was smaller than the area of the circles.</p><p>Relate the results from the overlay analysis back to the associated points.</p><p>Create a raster data layer from the point data, with 100 foot by 100 foot resolution, where the center of each pixel is at the location of the respective points. Areas between the Borough Boundary polygons (open water of NY Harbor) are coded as "no data."</p><p>All steps except for the creation of the raster dataset were conducted in PostgreSQL/PostGIS, as documented in <i>nyc_point_buffer_vegetation_overlay.sql</i>. The conversion of the results to a raster dataset was done in QGIS (version 3.28), ultimately using the <a href="https://gdal.org/programs/gdal_rasterize.html">gdal_rasterize</a> function.</p>
The raw microarray data and the differential expression analysis results from "Manipulating the growth environment through co-culture to enhance stress tolerance and viability of probiotic strains in the gastrointestinal tract".
<p>The signal data for each spot were subsequently quantified by using Feature Extraction software (Agilent Technologies).M1.txt to M5.txt: monoculture; C1.txt to C5.txt: co-culture; P1.txt to P5.txt: pH-controlled monoculture. The differential expression analysis results were obtained by using limma.</p>
Data and scripts for the colour analysis from: Gene flow throughout the evolutionary history of a colour polymorphic and generalist clownfish
<p>Even seemingly homogeneous on the surface, the oceans display high environmental heterogeneity across space and time. Indeed, different soft barriers structure the marine environment, which offers an appealing opportunity to study various evolutionary processes such as population differentiation and speciation. Here, we focus on <em>Amphiprion clarkii </em>(Actinopterygii; Perciformes), the most widespread of clownfishes that exhibits the highest colour polymorphism. Clownfishes can only disperse during a short pelagic larval phase before their sedentary adult lifestyle, which might limit connectivity among populations, thus facilitating speciation events. Consequently, the taxonomic status of <em>A. clarkii</em> has been under debate. We used whole-genome resequencing data of 67 <em>A. clarkii</em> specimens spread across the Indian and Pacific Oceans to characterise the species' population structure, demographic history, and colour polymorphism. We found that <em>A. clarkii</em> spread from the Indo-Pacific Ocean to the Pacific and Indian Oceans following a stepping-stone dispersal and that gene flow was pervasive throughout its demographic history. Interestingly, colour patterns differed noticeably among the Indonesian populations and the two populations at the extreme of the sampling distribution (i.e. Maldives and New Caledonia), which exhibited more comparable colour patterns despite their geographic and genetic distances. Our study emphasises how whole-genome studies can uncover the intricate evolutionary past of wide-ranging species with diverse phenotypes, shedding light on the complex nature of the species concept paradigm.</p>
Sample data for analysis of demographic potential of the 15-minute city in northern and southern France
<pre>This upload contains two Geopackage files of raw data used for urban analysis in the outskirts of Lille and Nice, France. <br>The data include building footprints (layer "building"), roads (layer "road"), and administrative boundaries (layer "adm_boundaries")<br>extracted from version 3.3 of the French dataset BD TOPO®3 (IGN, 2023) for the municipalities of Santes, Hallennes-lez-Haubourdin,<br>Haubourdin, and Emmerin in northern France (Geopackage "DPC_59.gpkg") and Drap, Cantaron and La Trinité in southern France <br>(Geopackage "DPC_06.gpkg").</pre> <pre> </pre> <pre>Metadata for these layers is available here: https://geoservices.ign.fr/sites/default/files/2023-01/DC_BDTOPO_3-3.pdf</pre> <pre> </pre> <pre>Additionally, this upload contains the results of the following algorithms available in GitHub (<a href="https://github.com/perezjoan/emc2-WP2?tab=readme-ov-file">https://github.com/perezjoan/emc2-WP2?tab=readme-ov-file</a>)</pre> <pre><code> </code></pre> <pre>1. The<code> </code>identification<code> </code>of<code> </code>main<code> </code>streets using the QGIS plugin Morpheo (layers "road_morpheo" and "buffer_morpheo") <br><a href="https://plugins.qgis.org/plugins/morpheo/">https://plugins.qgis.org/plugins/morpheo/</a> </pre> <pre><code>2. </code>The<code> </code>identification of main streets in local contexts – connectivity locally weighted<code> </code>(layer "road_LocRelCon")</pre> <pre><code>3. </code>Basic morphometry<code> </code>of<code> </code>buildings<code> </code>(layer "building_morpho")</pre> <pre><code>4. </code>Evaluation<code> </code>of<code> </code>the<code> </code>number<code> </code>of<code> </code>dwellings<code> </code>within<code> </code>inhabited<code> </code>buildings<code> </code>(layer "building_dwellings")</pre> <pre>5. Projecting<code> </code>population<code> </code>potential<code> </code>accessible from<code> </code>main<code> </code>streets<code> </code>(layer "road_pop_results")</pre> <pre> </pre> <pre>Project website: <a href="http://emc2-dut.org/">http://emc2-dut.org/</a></pre> <pre> </pre> <pre>Publications using this sample data: <br>Perez, J. and Fusco, G., 2024. Potential of the 15-Minute Peripheral City: Identifying Main Streets and Population Within Walking Distance. In: O. Gervasi, B. Murgante, C. Garau, D. Taniar, A.M.A.C. Rocha and M.N. Faginas Lago, eds. <em>Computational Science and Its Applications – ICCSA 2024 Workshops. ICCSA 2024</em>. Lecture Notes in Computer Science, vol 14817. Cham: Springer, pp.50-60. <a href="https://doi.org/10.1007/978-3-031-65238-7_4" target="_blank" rel="noopener noreferrer">https://doi.org/10.1007/978-3-031-65238-7_4</a>.</pre> <p><strong>Acknowledgement.</strong> <a name="_Hlk162443883"></a>This work is part of the emc2 project, which received the grant ANR-23-DUTP-0003-01 from the French National Research Agency (ANR) within the DUT Partnership.</p>
Data and tools of the landscape and cost analysis of data repositories currently used by the Swiss research community
<p>This file collection is part of the ORD Landscape and Cost Analysis Project (DOI: 10.5281/zenodo.2643460), a study jointly commissioned by the SNSF and swissuniversities in 2018.</p> <p>Please cite this data collection as:<br> von der Heyde, M. (2019). Data and tools of the landscape and cost analysis of data repositories currently used by the Swiss research community. Retrieved from https://doi.org/10.5281/zenodo.2643495</p> <p>Connected data papers are:<br> von der Heyde, M. (2019). Open Data Landscape: Repository Usage of the Swiss Research Community: Description of collection, collected data, and analysis methods [Data paper]. Retrieved from https://doi.org/10.5281/zenodo.2643430<br> von der Heyde, M. (2019). International Open Data Repository Survey: Description of collection, collected data, and analysis methods [Data paper]. Retrieved from https://doi.org/10.5281/zenodo.2643450</p> <p>Connected data sets are:<br> von der Heyde, M. (2019). Data from the Swiss Open Data Repository Landscape survey. Retrieved from https://doi.org/10.5281/zenodo.2643487<br> von der Heyde, M. (2019). Data from the International Open Data Repository Survey. Retrieved from https://doi.org/10.5281/zenodo.2643493</p> <p> </p> <p><strong>Contact</strong></p> <p>Swiss National Science Foundation (SNSF)</p> <p>Open Research Data Group</p> <p>E-mail: <a href="mailto:ord@snf.ch">ord@snf.ch</a></p> <p> </p> <p>swissuniversities</p> <p>Program "Scientific Information"</p> <p>Gabi Schneider</p> <p>E-Mail: <a href="mailto:isci@swissuniversities.ch">isci@swissuniversities.ch</a></p>
Data and analysis scripts for: Co-occurrence patterns at four spatial scales implicate reproductive processes in shaping community assembly in clovers
Open the record for dataset details and reuse information.
Data from: A novel approach to quantifying mammal locomotor repertoires using scoring and cluster analysis
Open the record for dataset details and reuse information.
Reproducible data, example subsets, and analysis pipeline for the extended TAaCGH study of breast cancer genomic and transcriptomic profiles
Open the record for dataset details and reuse information.
Data and code from: Cost-effectiveness Analysis of Alternative Infant and Neonatal Rotavirus Vaccination Schedules in Malawi
Open the record for dataset details and reuse information.
Data from: Broad-scale meta-analysis of drivers mediating adverse impacts of flow regulation on riparian vegetation
Open the record for dataset details and reuse information.
Matlab example for Local Enrichment Analysis (LEA) analysis with real data
Open the record for dataset details and reuse information.
Code for: A century of wild bee sampling: historical data and neural network analysis reveal ecological traits associated with species loss
Open the record for dataset details and reuse information.
Analysis code and data for the morphometrics and kinematics of tube feet
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.