Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
30,813
datasets available to search
ShareScore release 0.7.1
Dataset results
30,813 results for “typing”
Supplementary data to accompany Information flow, cell types and stereotypy in a full olfactory connectome
<p>Supplemental file 1</p> <p>Layers assigned by the probabilistic graph traversal model. bodyId refers to neurons’ unique ID in ne- uPrint. layer mean contains the mean layer after 10,000 iterations of the main model (Figure 2). layer - olf mean and layer th mean contain the mean layers from running the traversal model with ORNs and THN/HRNs, respectively (Figure S2).</p> <p>S1 hemibrain neuron layers.csv</p> <p>Supplemental file 2</p> <p>Sensory meta-information related to each glomerulus. Columns: glomerulus (canonical name for one of the 51 olfactory + 7 thermo/hygrosensory antennal lobe glomeruli), laterality (whether the glomerulus receives bilateral or only unilateral innervation from ALRNs), expected cit (a citation that describes the expected number of RNs in this glomerulus), expected RN female 1h (number of expected RNs in one hemi- sphere), expected RN female SD (standard deviation in the expected number of RNs), missing (qualitative assessment of glomeruli truncation), RN frag (if the RNs in that glomerulus are fragmented), receptor (the OR or IR expressed by cognate ALRNs (Bates et al., 2020; Task et al., 2020)), odour scenes (the general ‘odour scene(s)’ which this glomerulus may help signal (Mansourian and Stensmyr, 2015; Bates et al., 2020)), key ligand(the ligand that excites the cognate ALLRN or receptor the most, based on pooled data from multiple studies (Mu ̈nch and Galizia, 2016)), valence (the presumed valence of this odour chan- nel (Badel et al., 2016)). Exists as hemibrain glomeruli summary in our R package hemibrainr.</p> <p>S2 hemibrain olfactory information.csv</p> <p>Supplemental file 3</p> <p>File listing all identified antennal lobe receptor neurons (ALRNs) in the hemibrain, including information shown in neuPrint. See above for column explanations. Exists as rn.info in our R package hemibrainr.</p> <p>S3 hemibrain ALRN meta.csv</p> <p>Supplemental file 4</p> <p>All the hemibrain neurons we have classed as antennal lobe local neurons (ALLNs). See above for column explanations. Exists as alln.info in our R package hemibrainr.</p> <p>S4 hemibrain ALLN meta.csv</p> <p>Supplemental file 5</p> <p>All the hemibrain neurons we have classed as antennal lobe projection neurons (ALPNs). See above for column explanations. In addition, across dataset cluster refers to the clustering with left and right FAFB PNs; is canonical indicates whether that ALPN is one of the well studied “canonical” uPNs. Exists as pn.info in our R package hemibrainr.</p> <p>40</p> <p>S5 hemibrain ALPN meta.csv</p> <p>Supplemental file 6</p> <p>All the hemibrain neurons we have classed as third-order olfactory neurons (TOONs) including lateral horn neurons (LHNs), as well as wedge projection neurons (WEDPNs), lateral horn centrifugal neurons (LHCENT) and other projection neuron classes (Figure 1). See above for column explanations. Exists as ton.info in our R package hemibrainr.</p> <p>S6 hemibrain TOON meta.csv</p> <p>Supplemental file 7</p> <p>All the hemibrain neurons we have classed as neurons that descend to the ventral nervous system (DNs). See above for column explanations. Exists as dn.info in our R package hemibrainr.</p> <p>S8 hemibrain DN meta.csv</p> <p>Supplemental file 8</p> <p>The root point in hemibrain voxel space, for each hemibrain neuron. This is either the location of the soma, or the tip of a severed cell body fibre tract, where possible. Exists as hemibrain somas in our R package hemibrainr.</p> <p>S8 hemibrain root points.csv</p> <p>Supplemental file 9</p> <p>The start points for different neuron compartments. Nodes downstream of this position in the 3D structure of the neuron indicated with bodyid, belong to the compartment type designated by Label. A product of running flow centrality on hemibrain neurons, exists as hemibrain splitpoints in our R package hemi- brainr.</p> <p>S9 hemibrain compartment startpoints.csv</p> <p>Supplemental file 10</p> <p>3D triangle mesh for the hemibrain surface as a .obj file. This mesh was generated by first merging individual ROI meshes from neuPrint and then filling the gaps in between in a semi-manual process. It also exists as hemibrain.surf in our R package hemibrainr.</p> <p>S10 hemibrain raw.obj</p> <p>Supplemental file 11</p> <p>3D meshes of 51 olfactory + 7 thermo/hygrosensory antennal lobe glomeruli for the hemibrain volume, generated from ALRN presynapses.</p> <p>41</p> <p>Note that hemibrain coordinate system has the anterior-posterior axis aligned with the Y axis (rather than the Z axis, which is more commonly observed).</p> <p>S11 hemibrain AL glomeruli meshes RN-based.zip</p> <p>Supplemental file 12</p> <p>3D meshes of 51 olfactory + 7 thermo/hygrosensory antennal lobe glomeruli for the hemibrain volume, generated from ALPN presynapses.</p> <p>Note that hemibrain coordinate system has the anterior-posterior axis aligned with the Y axis (rather than the Z axis, which is more commonly observed).</p> <p>These meshes are also available as hemibrain al.surf in our R package hemibrainr. S12 hemibrain AL glomeruli meshes PN-based.zip</p>
Data for Cell-type-specific inhibitory circuitry from a connectomic census of mouse visual cortex
<p>Data for the paper: Cell-type-specific inhibitory circuitry from a connectomic census of mouse visual cortex, Nature 640, 2025</p> <p>In brief, this data archive includes information about the skeleton morphology and synaptic features of neurons whose cell bodies fell within a 100 micron by 100 micron column spanning all layers of mouse visual cortex. See <a href="https://www.microns-explorer.org/cortical-mm3">MICrONs-Explorer</a> for a full description of the broader volume and how it was collected.</p> <p>The data here include both data tables of cell locations, neuronal features, synapse lists, and more, as well as files containing morphological descriptions of all neurons used for the analysis in the initial version of the preprint. See the README.md file for more complete information about the individual files.</p> <p>Note: Data has been updated with post-publication files.</p>
Smartbay Marine Types Object Detection Training dataset
<h1>Training Dataset</h1> <p>The SmartBay Observatory in Galway Bay is an important contribution by Ireland to the growing global network of real-time data capture systems deployed within the ocean – technology giving us new insights into the ocean which we have not had before.</p> <p>The observatory was installed on the seafloor 1.5km off the coast of Spiddal, County Galway, Ireland . The observatory uses cameras, probes and sensors to permit continuous and remote live underwater monitoring. This observatory equipment allows ocean researchers unique real-time access to monitor ongoing changes in the marine environment. Data relating to the marine environment at the site is transferred in real-time from the SmartBay Observatory through a fibre optic telecommunications cable to the Marine Institute headquarters and onwards onto the internet. The data includes a live video stream, the depth of the observatory node, the sea temperature and salinity, and estimates of the chlorophyll and turbidity levels in the water which give an indication of the volume of phytoplankton and other particles, such as sediment, in the water.</p> <p>The Smartbay Marine Types Object Detection training Dataset is an initial Bounding Box Annotated image dataset used in attempting to Train a YOLOv8 Object Detection Model to classify the Marine Fauna observed in the Smartbay Observatory Video footage using broad "Marine Type" classes.</p> <p>The imagery used in this training dataset consists of image frame captures from the <a href="https://smartbay.marine.ie">Smartbay</a> video Archive files, CC-BY imagery from the <a href="https://www.minka-sdg.org">www.minka-sdg.org</a> website and images taken by Eva Cullen in the "<a href="https://nationalaquarium.ie/">Galway Atlantaquaria</a>" Aquarium in Galway, Ireland.</p> <p>The imagery were annotated using CVAT, collated on <a href="https://www.roboflow.com/">Roboflow</a> and exported in YOLOv8 trainign dataset format. </p>
Results of Survey on Playertypes by Gamification User Types Hexad Framework in Higher Education
<p>Survey on playertypes via the validated quesitonaire published in Krath, J., von Korflesch, H.F.O. (2021). Player Types and Game Element Preferences: Investigating the Relationship with the Gamification User Types HEXAD Scale. In: Fang, X. (eds) HCI in Games: Experience Design and Game Mechanics. HCII 2021. Lecture Notes in Computer Science(), vol 12789. Springer, Cham. https://doi.org/10.1007/978-3-030-77277-2_18</p> <p>Between 25.01.23 and 08.02.23 students of the University of Lübeck, Germany were invited to fill out an online questionnaire. The acquisition was done by sending an email to the students. No incentive was offered for participation, except to find out at one's own expression at the end of the survey. In addition to the validated questions, this also included questions about gender and study area. All participants agreed to anonymous data collection and publication. Participants were also asked to confirm that they were completing the survey for the first time, otherwise the return was removed from the result set. </p> <p>The result set is formatted as CSV. Questions and identifiers of the data are shown in the first line.</p>
NICHE Flanders: reference values for the (a)biotic requirements of vegetation types in Flanders, Belgium
<p>This dataset contains site requirements/tolerance limits (or "reference values") for 28 vegetation types found in Flanders. It gives the lower and upper limits or the classes within which these vegetation types can occur, for 7 site factors that determine potential vegetation development. These reference values can be used to determine the potential distribution of the different vegetation types with the ecohydrological model NICHE Flanders (<a href="https://purews.inbo.be/ws/portalfiles/portal/5370206/Callebaut_etal_2007_NicheVlaanderen.pdf">Callebaut et al. 2007</a>, in Dutch).</p> <p>See the Technical info (available in English and Dutch) for more information.</p>
A fading radius valley towards M-dwarfs, a persistent density valley across stellar types -- data
Open the record for dataset details and reuse information.
Replication data for: "The hapax / type ratio: an indicator of minimally required sample size in productivity studies?"
<p>The dataset accompanies the scientific article "The hapax / type ratio: an indicator of minimally required sample size in productivity studies?" and can be used to reproduce the findings presented in this article. This dataset consists of two components, namely (i) the corpus data involving the Dutch semi-copular verb "raken" and (ii) an R analysis script to reproduce the computational steps.</p>
Alignment between type of landmark in different sources and the concept in the spatial reference objects ontology
<p>The five datasets represent a manually alignment between the landmark type of five different datasets archived <a href="https://doi.org/10.5281/zenodo.6480986">here</a> and a common vocabulary extracted from an application ontology defined for mountain rescue purposes, named <a href="https://hamac.ign.fr/owa/redir.aspx?C=cjlWje9SCaYsVOTLbxbOoIBLZUCS56nVb248cRSMTEDSENDFzybaCA..&URL=http%3a%2f%2fchoucas.ign.fr%2fdoc%2fontologies%2foor.owl%2f">Ontology of landmarks</a> (OOR).</p> <p>Each file represents the alignment for features belonging to a data source with the same OOR ontology.</p> <p>For example, the type «bivouac» from camptocamp.org source is aligned with the uri <a href="http://purl.org/choucas.ign.fr/oor#abri">http://purl.org/choucas.ign.fr/oor#abri</a> of the corresponding class «Shelter » in the ontology of landmark. The alignments models can be considered as a ground truth data.</p> <p>The alignments results are obtained using an ontology application named <a href="http://choucas.ign.fr/doc/ontologies/index-fr.html">OOR</a>. These specific results are obtained using the version of OOR V1.0.1 which is an improved version and contains new concepts compared to the first release 1.0.0. The new version of OOR (i.e. 1.0.1) will be released by the end of May 31 2022. The new link will be added here.</p> <p>This archive is released for transparency and reproducibility purposes.</p>
Data from calculated radial neutron flux distributions in a KBS-3 type geological repository
<p>Data from calculations of radial distribution of neutron flux per emitted neutron from rods of spent nuclear fuel in a KBS-3 type geological repository. Reference (<em>Jansson, 2022</em>) contain a summary of the calculations and a description of the structure of this data.</p> <p>This data was computed on resources provided by Swedish National Infrastructure for Computing (SNIC) at Uppsala Multidisciplinary Center for Advanced Computational Science (UPPMAX), National Supercomputer Centre at Linköping University (NSC) and the SNIC Cloud, partially funded by the Swedish Research Council through grant agreement no. 2018-05973, under projects SNIC 2021/5-299 and SNIC 2021/18-12.</p>
Supporting data for: Type 1 diabetes risk genes mediate pancreatic beta cell survival in response to proinflammatory cytokines
<p><strong>SUMMARY OF THE STUDY</strong></p> <p>We combined functional genomics and human genetics to investigate processes that affect type 1 diabetes (T1D) risk by mediating beta-cell survival in response to proinflammatory cytokines. We mapped 38,931 cytokine-responsive candidate <em>cis-</em>regulatory elements (cCREs) in beta-cells using ATAC-seq and snATAC-seq and linked them to target genes using co-accessibility and HiChIP. Using a genome-wide CRISPR screen in EndoC-βH1 cells we identified 867 genes affecting cytokine-induced survival, and genes promoting survival and up-regulated in cytokines were enriched at T1D risk loci. Using SNP-SELEX, we identified 2,229 variants in cytokine-responsive cCREs altering transcription factor (TF) binding, and variants altering binding of TFs regulating stress, inflammation and apoptosis were enriched for T1D risk. At the 16p13 locus, a fine-mapped T1D variant altering TF binding in a cytokine-induced cCRE interacted with <em>SOCS1</em>, which promoted survival in cytokine exposure. Our findings reveal processes and genes acting in beta-cells during inflammation that modulate T1D risk.</p> <p><strong>DESCRIPTION OF FILES:</strong></p> <ul> <li>Supplementary Data 1. List of islet cCREs annotated with cell type and cytokine response - also in GSE205853</li> <li>Supplementary Data 2. Coaccessible sites in untreated beta cells and promoter annotations - also in GSE205853</li> <li>Supplementary Data 3. Coaccessible sites in cytokine-treated beta cells and promoter annotations - also in GSE205853</li> <li>Supplementary Data 4. Coaccessible sites in cytokine treated and untreated beta cells and promoter annotations - also in GSE205853</li> <li>Supplementary Data 5. Chromatin interactions in EndoC-BH1 cells - also in GSE205853</li> <li>Supplementary Data 6. Variants selected for SNP-SELEX assay </li> <li>Supplementary Data 7. Variants with TF binding and allelic binding results from SNP-SELEX</li> <li>Supplementary Data 8. snATAC-seq barcodes and metadata - also in GSE205853</li> <li>Supplementary Data 9. CRISPR-KO screen results - also in GSE205853</li> <li>Supplementary Data 10. Bulk ATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 11. Bulk RNA-seq count matrix - also in GSE205853</li> <li>Supplementary Data 12. Alpha cells snATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 13. Acinar cells snATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 14. Beta cells snATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 15. Stellate cells snATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 16. Endothelial cells snATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 17. Delta cells snATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 18. Luciferase assay rs10483809</li> <li>Supplementary Data 19. SOCS1 knockdown qPCR results</li> <li>Supplementary Data 20. SOCS1 knockdown Apotracker (flow-cytometry)results</li> </ul> <p><strong>Raw data deposited at GEO, accessions GSE205853 and GSE118725.</strong></p> <p><em>Please refer to publication and GEO for details on methods.</em></p>
Organic micropollutants and heavy metals in stormwater runoff of five different catchment types in Berlin (Germany)
<p>This dataset includes concentrations of micropollutants (67), heavy metals (8) and standard parameters (9) for stormwater runoff taken from separated sewers of five catchments between 3 and 37 ha in Berlin (Germany). It also includes rain data of analyzed events as separate file. Samples were taken as part of the OgRe research project of Kompetenzzentrum Wasser Berlin (<a href="https://www.kompetenz-wasser.de/en/project/ogre/">www.kompetenz-wasser.de/en/project/ogre/</a>) in 2014 and 2015. Sampling and analytical methods are detailed in "Concentrations of micropollutants in urban stormwater runoff of different land uses" (<a href="https://doi.org/10.3390/w13091312">https://doi.org/10.3390/w13091312</a>). A dataset with concentrations of the urban stream Panke in Berlin during dry and wet weather (samples were taken as part of the same project) is available separately (<a href="https://zenodo.org/record/4633779">https://zenodo.org/record/4633779</a>).</p> <p><strong>Description of fields (concentrations):</strong></p> <ul> <li><strong>SampleID</strong>: unique sample identifier</li> <li><strong>SiteID</strong>: unique site identifier (catchment type) <ul> <li> 1 - OLD: area with typical five-storey perimeter blocks built between 1870 and 1930 (31 ha)</li> <li> 2 - NEW: newer area of 4-8-storey concrete slab buildings built between 1960 and 1980 (16 ha)</li> <li> 3 - STR: 1.3 km of a busy streeat with intersection with traffic lights and bus stops (3 ha)</li> <li> 4 - OFH: a residential area characterized by one-family houses and villas with gardens (17 ha)</li> <li> 5 - COM: a commercial and industrial area of high imperviousness with large flat-roof buildings and yards (37 ha)</li> <li> 6 - PNK: urban stream Panke (characterized by strong stormwater inputs from separate sewer discharges - available in separate dataset)</li> </ul> </li> <li><strong>LocalDateTime</strong>: start time of sampling (local)</li> <li><strong>DateTimeUTC</strong>: start time of sampling (UTC)</li> <li><strong>UTCOffset</strong>: UTC offset to local time in h</li> <li><strong>SampleType</strong>: either "composite" for volume proportional composite sample (all samples from storm sewers) or "single" for grab sample (all stream samples, separate dataset)</li> <li><strong>VariableName</strong>: name of analysed substance/parameter</li> <li><strong>UnitsAbbreviation</strong>: either "ug/L" (microgram per litre) or "mg/L" (milligram per litre)</li> <li><strong>CensorCode</strong>: either "lt" (less than) for concentration below detection limit (value is detection limit) or "nc" (not censored) for concentration above detection limit</li> <li><strong>DataValue</strong>: measured value (if censor code is lt, value indicates detection limit)</li> </ul> <p><strong>Description of fields (rain data):</strong></p> <ul> <li><strong>SampleID</strong>: sample identifier of matching sample (see above)</li> <li><strong>SiteID and SiteName</strong>: unique site identifier and name (catchment type) (see above)</li> <li><strong>tBeg_rain, tEnd_rain</strong>: begin and end of rain event in local time</li> <li><strong>depth.mm</strong>: rain depth of rain event in mm</li> <li><strong>duration_rain.h</strong>: duration of rain event in h</li> <li><strong>intensity_max_10min.mm_h</strong>: maximum rain intensitity of rain event in 10-min interval in mm/h</li> <li><strong>intensity_mean_event.mm_h</strong>: mean rain intensitity of rain event in mm/h</li> <li><strong>ADD.d</strong>: number of antecedent dry days in days</li> </ul> <p>Rain data was collected by rain gauge network of Berlin waterworks (>40 gauges) — gauge with best correlation between rain depth and event volume in storm sewer was chosen (distances to monitoring sites: 2–6 km).</p> <p>Two data files are provided in comma separated format:</p> <ul> <li>"OgRe_drain.csv" contains concentrations of all stormwater runoff samples taken in separate storm sewers</li> <li>"OgRe_rain.csv" contains rain data for all stormwater runoff samples</li> </ul>
StreetSurfaceVis: a dataset of street-level imagery with annotations of road surface type and quality
<h1>StreetSurfaceVis</h1> <p><em>StreetSurfaceVis</em> is an image dataset containing <strong>9,122 street-level images from Germany</strong> with labels on <strong>road surface type and quality.</strong> The CSV file <code>streetSurfaceVis_v1_0.csv</code> contains all image metadata and four folders contain the image files. All images are available in four different sizes, based on the image width, in 256px, 1024px, 2048px and the original size.<br>Folders containing the images are named according to the respective image size. Image files are named based on the <code>mapillary_image_id</code>.</p> <p>You can find the corresponding publication here: <a href="https://www.nature.com/articles/s41597-024-04295-9#citeas">StreetSurfaceVis: a dataset of crowdsourced street-level imagery with semi-automated annotations of road surface type and quality</a></p> <p> </p> <h3>Image metadata</h3> <p>Each CSV record contains information about one street-level image with the following attributes:</p> <ul> <li><code>mapillary_image_id</code>: ID provided by Mapillary (see information below on Mapillary)</li> <li><code>user_id</code>: Mapillary user ID of contributor</li> <li><code>user_name</code>: Mapillary user name of contributor</li> <li><code>captured_at</code>: timestamp, capture time of image</li> <li><code>longitude</code>, <code>latitude</code>: location the image was taken at</li> <li><code>train</code>: Suggestion to split train and test data. `True` for train data and `False` for test data. Test data contains data from 5 cities which are excluded in the training data.</li> <li><code>surface_type</code>: Surface type of the road in the focal area (the center of the lower image half) of the image. Possible values: asphalt, concrete, paving_stones, sett, unpaved</li> <li><code>surface_quality</code>: Surface quality of the road in the focal area of the image. Possible values: (1) excellent, (2) good, (3) intermediate, (4) bad, (5) very bad (see the attached <strong>Labeling Guide document</strong> for details)</li> </ul> <p> </p> <h3>Image source</h3> <p>Images are obtained from <a href="https://www.mapillary.com/">Mapillary</a>, a crowd-sourcing plattform for street-level imagery. More metadata about each image can be obtained via the <a href="https://www.mapillary.com/developer/api-documentation">Mapillary API . </a>User-generated images are shared by Mapillary under the <a href="https://creativecommons.org/licenses/by-sa/4.0/">CC-BY-SA</a> License.</p> <p>For each image, the dataset contains the <code>mapillary_image_id</code> and <code>user_name</code>. <br>You can access user information on the Mapillary website by <code>https://www.mapillary.com/app/user/<USER_NAME> </code><br>and image information by <code>https://www.mapillary.com/app/?focus=photo&pKey=<MAPILLARY_IMAGE_ID></code></p> <p>If you use the provided images, please adhere to the <a href="https://www.mapillary.com/terms">terms of use of Mapillary.</a></p> <p> </p> <h3>Instances per class</h3> <p>Total number of images: 9,122</p> <table> <tbody> <tr> <td> </td> <td><strong>excellent</strong></td> <td><strong>good</strong></td> <td><strong>intermediate</strong></td> <td><strong>bad</strong></td> <td><strong>very bad</strong></td> </tr> <tr> <td><strong>asphalt</strong></td> <td>971</td> <td>1697</td> <td>821</td> <td>246</td> <td>-</td> </tr> <tr> <td><strong>concrete</strong></td> <td>314</td> <td>350</td> <td>250</td> <td>58</td> <td>-</td> </tr> <tr> <td><strong>paving stones</strong></td> <td>385</td> <td>1063</td> <td>519</td> <td>70</td> <td>-</td> </tr> <tr> <td><strong>sett</strong></td> <td>-</td> <td>129</td> <td>694</td> <td>540</td> <td>-</td> </tr> <tr> <td><strong>unpaved</strong></td> <td>-</td> <td>-</td> <td>326</td> <td>387</td> <td>303</td> </tr> </tbody> </table> <p> </p> <p>For modeling, we recommend using a train-test split where the test data includes geospatially distinct areas, thereby ensuring the model's ability to generalize to unseen regions is tested. We propose five cities varying in population size and from different regions in Germany for testing - images are tagged accordingly.</p> <p>Number of test images (train-test split): 776</p> <h3>Inter-rater-reliablility</h3> <p>Three annotators labeled the dataset, such that each image was annotated by one person. Annotators were encouraged to consult each other for a second opinion when uncertain.<br>1,800 images were annotated by all three annotators, resulting in a <em>Krippendorff's alpha</em> of 0.96 for surface type and 0.74 for surface quality.</p> <h3>Recommended image preprocessing</h3> <p>As the focal road located in the bottom center of the street-level image is labeled, it is recommended to crop images to their lower and middle half prior using for classification tasks.</p> <p>This is an exemplary code for recommended image preprocessing in <strong>Python</strong>:</p> <pre><code>from PIL import Image<br></code><code>img = Image.open(image_path)</code><br><code>width, height = img.size</code><br><code>img_cropped = img.crop((0.25 * width, 0.5 * height, 0.75 * width, height))</code></pre> <h3><br><strong>License</strong></h3> <p><a href="https://creativecommons.org/licenses/by-sa/4.0/">CC-BY-SA</a></p> <p> </p> <h3><strong>Citation</strong></h3> <p>If you use this dataset, please cite as: </p> <p> </p> <p>Kapp, A., Hoffmann, E., Weigmann, E. <em>et al.</em> StreetSurfaceVis: a dataset of crowdsourced street-level imagery annotated by road surface type and quality. <em>Sci Data</em> <strong>12</strong>, 92 (2025). https://doi.org/10.1038/s41597-024-04295-9</p> <p> </p> <p><code>@article{kapp_streetsurfacevis_2025,<br> title = {{StreetSurfaceVis}: a dataset of crowdsourced street-level imagery annotated by road surface type and quality},<br> volume = {12},<br> issn = {2052-4463},<br> url = {https://doi.org/10.1038/s41597-024-04295-9},<br> doi = {10.1038/s41597-024-04295-9},<br> pages = {92},<br> number = {1},<br> journaltitle = {Scientific Data},<br> shortjournal = {Scientific Data},<br> author = {Kapp, Alexandra and Hoffmann, Edith and Weigmann, Esther and Mihaljević, Helena},<br> date = {2025-01-16},<br>}</code></p> <p> </p> <p>-----------------------------------------------------------------------------------------------------------------------------------------------------------</p> <p>This is part of the SurfaceAI project at the University of Applied Sciences, HTW Berlin.</p> <p><br>- Prof. Dr. Helena Mihajlević<br>- Alexandra Kapp<br>- Edith Hoffmann<br>- Esther Weigmann</p> <p>Contact: surface-ai@htw-berlin.de</p> <p>https://surfaceai.github.io/surfaceai/</p> <p><strong>Funding</strong>: SurfaceAI is a mFund project funded by the Federal Ministry for Digital and Transportation Germany.</p> <p> </p>
Types, open citations, closed citations, publishers, and participation reports of Crossref entities
<p>This publication contains several datasets that have been used in the paper "Crowdsourcing open citations with CROCI – An analysis of the current status of open citations, and a proposal" submitted to the <a href="https://www.issi2019.org/">17th International Conference on Scientometrics and Bibliometrics (ISSI 2019)</a>, available at <a href="https://opencitations.wordpress.com/2019/02/07/crowdsourcing-open-citations-with-croci/">https://opencitations.wordpress.com/2019/02/07/crowdsourcing-open-citations-with-croci/</a>.</p> <p>Additional information about the analyses described in the paper, including the code and the data we have used to compute all the figures, is available as a Jupyter notebook at <a href="https://github.com/sosgang/pushing-open-citations-issi2019/blob/master/script/croci_nb.ipynb">https://github.com/sosgang/pushing-open-citations-issi2019/blob/master/script/croci_nb.ipynb</a>. The datasets contain the following information.</p> <p><strong>non_open.zip:</strong> it is a zipped (~5 GB unzipped) CSV file containing the numbers of open citations and closed citations received by the entities in the Crossref dump used in our computation, dated October 2018. All the entity types retrieved from Crossref were aligned to one of following five categories: journal, book, proceedings, dataset, other. The open CC0 citation data we used came from the CSV dump of <a href="https://doi.org/10.6084/m9.figshare.6741422.v3">most recent release of COCI dated 12 November 2018</a>. The number of closed citations was calculated by subtracting the number of open citations to each entity available within COCI from the value “is-referenced-by-count” available in the Crossref metadata for that particular cited entity, which reports all the DOI-to-DOI citation links that point to the cited entity from within the whole Crossref database (including those present in the Crossref ‘closed’ dataset).</p> <p>The columns of the CSV file are the following ones:</p> <ul> <li><em>doi:</em> the DOI of the publication in Crossref;</li> <li><em>type:</em> the type of the publication as indicated in Crossref;</li> <li><em>cited_by:</em> the number of open citations received by the publication according to COCI;</li> <li><em>non_open:</em> the number of closed citations received by the publication according to Crossref + COCI.</li> </ul> <p><strong>croci_types.csv:</strong> it is a CSV file that contains the numbers of open citations and closed citations received by the entities in the Crossref dump used in our computation, as collected in the previous CSV file, alligned in five classes depening on the entity types retrieved from Crossref: <em>journal</em> (Crossref types: journal-article, journal-issue, journal-volume, journal), <em>book</em> (Crossref types: book, book-chapter, book-section, monograph, book track, book-part, book-set, reference-book, dissertation, book series, edited book), <em>proceedings</em> (Crossref types: proceedings-article, proceedings, proceedings-series), <em>dataset</em> (Crossref types: dataset), <em>other</em> (Crossref types: other, report, peer review, reference-entry, component, report-series, standard, posted-content, standard-series).</p> <p>The columns of the CSV file are the following ones:</p> <ul> <li><em>type:</em> the type publication between "journal", "book", "proceedings", "dataset", "other";</li> <li><em>label:</em> the label assigned to the type for visualisation purposes;</li> <li><em>coci_open_cit</em>: the number of open citations received by the publication type according to COCI;</li> <li><em>crossref_close_cit:</em> the number of closed citations received by the publication according to Crossref + COCI.</li> </ul> <p><strong>publishers_cits.csv:</strong> it is a CSV file that contains the top twenty publishers that received the greatest number of open citations. The columns of the CSV file are the following ones:</p> <ul> <li><em>publisher</em>: the name of the publisher;</li> <li><em>doi_prefix</em>: the list of DOI prefixes used assigned by the publisher;</li> <li><em>coci_open_cit</em>: the number of open citations received by the publications of the publisher according to COCI;</li> <li><em>crossref_close_cit</em>: the number of closed citations received by the publications of the publishers according to Crossref + COCI;</li> <li><em>total_cit</em>: the total number of citations received by the publications of the publisher (= <em>coci_open_cit</em> + <em>crossref_close_cit</em>).</li> </ul> <p><strong>20publishers_cr.csv: </strong>it is a CSV file that contains the numbers of the contributions to open citations made by the twenty publishers introduced in the previous CSV file as of 24 January 2018, according to the data available through the Crossref API. The counts listed in this file refers to the number of publications for which each publisher has submitted metadata to Crossref that include the publication’s reference list. The categories 'closed', 'limited' and 'open' refer to publications for which the reference lists are not visible to anyone outside the Crossref Cited-by membership, are visible only to them and to Crossref Metadata Plus members, or are visible to all, respectively. In addition, the file also record the total number of publications for which the publisher has submitted metadata to Crossref, whether or not those metadata include the reference lists of those publications.</p> <p>The columns of the CSV file are the following ones:</p> <ul> <li><em>publisher: </em>the name of the publisher;</li> <li><em>open: </em>the number of publications in Crossref with an 'open' visibility for their reference lists;</li> <li><em>limited: </em>the number of publications in Crossref with an 'limited' visibility for their reference lists;</li> <li><em>closed: </em>the number of publications in Crossref with an 'closed' visibility for their reference lists;</li> <li><em>overall_deposited:</em> the overall number of publications for which the publisher has submitted metadata to Crossref.</li> </ul>
Multi-omic Insights into Molecular Mechanism and Therapeutic Targets in Spinocerebellar Ataxia type 7
<p>The molecular mechanism in spinocerebellar ataxia type 7 is currently poorly understood. To provide understandings, a multi-omic study was performed using SCA7266Q/5Q mice. At week 12, entire brain tissue samples were collected and RNA sequencing, methylation analysis, and proteomic analysis were performed. Results were integrated to identify genes with identical trends in expression. Data was also compared with SCA patient serum proteomic analysis, and based on common differentially expressed proteins, a Naïve Bayesian network model was constructed to predict nilotinib treatment response. Data from RNA sequencing and methylation analysis revealed 58 significantly hypomethylated-upregulated genes and 62 hypermethylated-downregulated genes, mostly enriched in GO terms of regulation of axonogenesis, channel activity, and monoamine signaling. In the proteomic analysis, 211 upregulated and 281 downregulated DEPs associated mostly with immune response and cellular mobility were identified. Two genes, Fam107b and Tph2, showed differential expression in both transcriptomic and proteomic analysis. Forty-two overlapping proteins were identified compared with SCA patient serum, and Bayesian network analysis revealed that nilotinib treatment response was associated with the protein expression of CLU, CA2, GLUL, PRDX6, C1QA, PLXNB1, and age. These findings will serve as an important reference for future studies on the pathogenesis and discovery of druggable targets. </p>
Standardized map of habitat types and regionally important biotopes in Flanders
<p>The <code>habitatmap_stdized.gpkg</code> file is a processed version of the <a href="https://www.vlaanderen.be/datavindplaats/catalogus/biologische-waarderingskaart-en-natura-2000-habitatkaart-toestand-2023">Natura 2000 habitat map of Flanders</a> (De Saeger et al., 2023; see also De Saeger et al. 2017). It contains all polygons with Natura 2000 habitat types or regional important biotopes (RIB). This file is used as a basis for designing monitoring schemes in Flanders. </p> <p>In the original habitat map, every polygon can consist of maximum 5 different types (habitat (sub)types and regionally important biotopes). This information is stored in the columns <code>HAB1</code>, <code>HAB2</code>,..., <code>HAB5</code> of the attribute table. The fraction of each type within the polygons is stored in the columns <code>PHAB1</code>, <code>PHAB2</code>, ..., <code>PHAB5</code>.</p> <p>The <code>habitatmap_stdized.gpkg</code> file is a GeoPackage that contains:</p> <ul> <li><code>habitatmap_polygons</code>: a spatial layer with every habitat map polygon that contains a Natura 2000 habitat or RIB type.</li> <li><code>habitatmap_types</code>: a table with information on the habitat and RIB types (HAB1, HAB2,..., HAB5) that occur within each polygon of <code>habitatmap_polygons.</code></li> </ul> <p>The processing of the habitatmap_types table included following adjustments:</p> <ul> <li>For some polygons the type is uncertain, and the type code in the raw habitatmap data source consists of 2 or 3 possible types, separated with a ','. The different possible types are split up and one row is created for each of them, with <code>phab</code> for each new row simply set to the original value of <code>phab</code>. The variable <code>certain</code> will be <code>FALSE</code> if the original type code consists of 2 or 3 possible types, and <code>TRUE</code> if only one type is provided.</li> <li>Some polygons contain both a standing water habitat type and <code>rbbmr</code>: <ul> <li><code>3130_rbbmr</code>,</li> <li><code>3140_rbbmr</code>,</li> <li><code>3150_rbbmr</code>, and</li> <li><code>3160_rbbmr</code>.</li> </ul> </li> <li>Since <code>habitatmap_stdized_2020_v1</code>, the two types <code>31xx</code> and <code>rbbmr</code> are split up and one row is created for each of them, with <code>phab</code> for each new row simply set to the original value of <code>phab</code>. The variable certain in this case will be <code>TRUE</code> for both types.</li> <li>After those steps, a given polygon could contain the same type with the same value for <code>certain</code> repeated several times, e.g. when <code>31xx_rbbmr</code> is present with <code>phab</code> = yy% and <code>31xx</code> is present with <code>phab</code> = zz%. In that case the rows with the same <code>polygon_id</code>, <code>type</code> and <code>certain</code> were gathered into one row and the respective phab values were added up.</li> </ul> <p>The R-code for creating the <code>habitatmap_stdized</code> data source can be found in the GitHub repository <a href="https://github.com/inbo/n2khab-preprocessing/tree/abf596e/src/generate_habitatmap_stdized">'n2khab-preprocessing' at commit abf596e</a>.</p> <p>A reading function to return the data source in a standardized way into the R environment is provided by the R-package <a href="https://github.com/inbo/n2khab">n2khab</a>.</p> <p>Attributes of <code>habitatmap_polygons</code>:</p> <ul> <li><code>polygon_id</code></li> <li><code>description_orig</code>: polygon description based on the original type codes in the raw habitatmap </li> </ul> <p>Attributes of <code>habitatmap_types</code>:</p> <ul> <li><code>polygon_id</code></li> <li><code>type</code>: the interpreted habitat or RIB type</li> <li><code>certain</code>: <code>TRUE</code> when type is certain and <code>FALSE</code> when type is uncertain</li> <li><code>code_orig</code>: original type code in raw habitatmap</li> <li><code>phab</code>: proportion of polygon covered by type, as a percentage.</li> </ul> <p>Since version <code>habitatmap_stdized_2020_v1</code>, rows are unique only by the combination of the <code>polygon_id</code>, <code>type</code> and <code>certain</code> columns.</p>
ManyTypes4Py: A Benchmark Python Dataset for Machine Learning-Based Type Inference
<ul> <li>The dataset is gathered on Sep. 17th 2020 from GitHub.</li> <li>It has <em>clean</em> and <em>complete</em> versions (from v0.7): <ul> <li>The clean version has 5.1K <strong>type-checked </strong>Python repositories and 1.2M type annotations.</li> <li>The complete version has 5.2K Python repositories and 3.3M type annotations.</li> </ul> </li> <li>The dataset's source files are type-checked using <a href="https://mypy.readthedocs.io/">mypy</a> (clean version).</li> <li>The dataset is also de-duplicated using the <a href="https://github.com/saltudelft/CD4Py">CD4Py</a> tool.</li> <li>Check out the <strong>README.MD</strong> file for the description of the dataset.</li> <li>Notable changes to each version of the dataset are documented in <strong>CHANGELOG.md</strong>.</li> <li>The dataset's scripts and utilities are available on <a href="https://github.com/saltudelft/many-types-4-py-dataset">its GitHub repository</a>.</li> </ul>
Data from: Will Current Protected Areas Harbour Refugia for Threatened Arctic Vegetation Types until 2050? A First Assessment
<p>We present predictions of Arctic vegetation for 2050 based on a combination of climate models (namely, EC-Earth3-Veg, IPSL-CM6A-LR, and MRI-ESM2-0), emission scenarios (names, SSP126 and SSP585) and tree dispersal rate scenarios (unrestricted, 20km and 5km) based on the methods of Pearson et al. (2013) and the new raster version of the Circumpolar Arctic Vegetation Map (CAVM) (Raynolds et al. 2019). We additionally present a dataset summarising total areas for each vegetation type in the CAVM and the forecasted models based on the computation of zonal histograms in ArcGIS (zonal_histogram_results.csv), for the total Arctic as well as only within protected areas, defined by the Map of Arctic Protected Areas (CAFF and PAME 2017). We also present a potential map of refugia for what we deem the realistic model (IPSL, SSP585, 20 km tree dispersal) as a raster file. Refugia were identified as regions where the vegetation remained the same between the CAVM and the predictions. Additionally, we present a map of model agreement, showing the degree to which other models agree with the vegetation classification for our refugia.</p> <p>All predictions named according to the tree dispersal rate, climate model, and emissions scenario, preceded by the term "pred". For example: "pred_unres_mri_585" represents the unrestricted tree dispersal, MRI-ESM-0 climate model, and SSP585 scenario-based prediction. The MRI-ESM-0 x SSP585 combination had gaps in data which results in a lack of predictions in some areas; this affects 3 models.</p> <p>Further details and all code associated with these datasets are found <a href="https://github.com/PlekhanovaElena/Arctic_vegetation_prediction">here</a>.</p>
Rare type III responses: data & data methods (v1.0.0)
<p>This repository includes the data (data-rare-type3-responses.csv) for Kalinkat et al. (2023).</p> <p>The data comprises a literature review on type III functional responses between 2002 and 2022 with 12 variables and 107 observations. Please read the accompanying document (Rall_et_al_2023_Zenodo_Rare-type3-responses_data_methods_v1_0_0.pdf) for the methods and a description of the variables. </p> <p><strong>Reference</strong></p> <p>Kalinkat, G. <em>et al.</em> (2023) ‘Empirical evidence of type III functional responses and why it remains rare’, <em>Frontiers in Ecology and Evolution</em>, 11:1033818. Available at: https://doi.org/10.3389/fevo.2023.1033818.</p>
Raw data to "Series expansions in closed and open quantum many-body systems with multiple quasiparticle types"
<p>This collection of data is complementary to the publication "Series expansions in closed and open quantum many-body systems with multiple quasiparticle types", Lea Lenke, Andreas Schellenberger, Kai Phillip Schmidt, <a href="https://arxiv.org/abs/2302.01000">arXiv:2302.01000</a> (<a href="https://arxiv.org/abs/2302.01000">https://arxiv.org/abs/2302.01000</a>).</p> <p>It contains all data used for Figure 2 given in the file `Figure_2_complementary_data.yaml` and all needed data to recalculate the energies of the visualized modes in the files `Figure_2_coefficients_expectation_values.yaml` and `Figure_2_broad_signum_coefficients_expectation_values.yaml`.</p> <p>For the last two files, we used a program to calculate the coefficients. The source code for coefficient calculation is openly available under GitHub (<a href="https://github.com/FAU-kpslab/pcstpp_CoefficientGenerator">https://github.com/FAU-kpslab/pcstpp_CoefficientGenerator</a>) including configuration files to reproduce the coefficients given here.</p> <p>All files are self-consistent, for further information we recommend the comments directly in the files.</p> <p>For further details on the used method pcst<sup>++ </sup>and discussion of the results we refer to the linked publication.</p> <p>If any question may arise, you are highly welcome to contact us (see e.g. contact information on the publication).</p>
Data on different types of green spaces and their accessibility in the seven largest urban regions in Finland
<p>This repository contains data described in the article "Data on different types of green spaces and their accessibility in the seven largest urban regions in Finland" (Heikinheimo et al. 2023) and used in the research article "Associations of neighborhood-level socioeconomic status, accessibility, and quality of green spaces in Finnish urban regions" (Viinikka et al. 2023). <br> <br> This repository contains data on green space quality and path distances to different types of green spaces. The path distances represent green space accessibility using active travel modes (walking, cycling). The path distances were calculated using the pedestrian street network across the seven largest urban regions in Finland. We derived the green space typology from the Urban Atlas Data that is available across functional urban areas in Europe and enhanced it with national data on water bodies, conservation areas and recreational facilities and routes from Finland. We extracted the walkable street network from OpenStreetMap and calculated shortest paths to different types of green spaces using open-source Python programming tools. Network distances were calculated up to ten kilometers from each green space edge and the distances were aggregated into a 250 m x 250 m statistical grid that is interoperable with various statistical data from Finland. The geospatial data files representing the different types of green spaces, network distances across the seven urban regions, as well as the processing and analysis scripts are shared in an open repository. These data offer actionable information about green space accessibility in Finnish city regions and support the integration of green space quality and active travel modes into further research and planning activities.</p> <p> </p> <p><strong>Data description article: </strong></p> <p>Heikinheimo, V., Tiitu, M., & Viinikka, A. (2023). Data on different types of green spaces and their accessibility in the seven largest urban regions in Finland. <em>Data in Brief</em>, <em>50</em>, 109458. <a href="https://doi.org/10.1016/j.dib.2023.109458">https://doi.org/10.1016/j.dib.2023.109458</a></p> <p><strong>Related research article:</strong> </p> <p>Viinikka, A., Tiitu, M., Heikinheimo, V., Halonen, J. I., Nyberg, E., & Vierikko, K. (2023). Associations of neighborhood-level socioeconomic status, accessibility, and quality of green spaces in Finnish urban regions. <em>Applied Geography</em>, <em>157</em>, 102973. <a href="https://doi.org/10.1016/j.apgeog.2023.102973">https://doi.org/10.1016/j.apgeog.2023.102973</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.