Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,709
datasets available to search
ShareScore release 0.7.1
Dataset results
3,709 results for “feature”
Long term response of arctic tussock tundra to thermal erosion features: A modeling analysis. Tussock tundra greenhouse simulation
The Multiple Element Limitation (MEL) model is used to simulate the recovery of Alaskan arctic tussock tundra to thermal erosion features (TEFs) caused by permafrost thaw and mass wasting. TEFs could be significant to regional carbon (C) and nutrient budgets because permafrost soils contain large stocks of soil organic matter (SOM) and TEFs are expected to become more frequent as climate warms. These simulations deal only with recovery following TEF stabilization and do not address initial losses of C and nutrients during TEF formation. To capture the variability among and within TEFs, we simulate a range of post-stabilization conditions by varying the initial size of SOM pools and nutrient supply rates. This file contains the results for 25 years of tussock tundra under greenhouse conditions.
Long term response of arctic tussock tundra to thermal erosion features: A modeling analysis. Tussock tundra nitrogen fertilized simulation
The Multiple Element Limitation (MEL) model is used to simulate the recovery of Alaskan arctic tussock tundra to thermal erosion features (TEFs) caused by permafrost thaw and mass wasting. TEFs could be significant to regional carbon (C) and nutrient budgets because permafrost soils contain large stocks of soil organic matter (SOM) and TEFs are expected to become more frequent as climate warms. These simulations deal only with recovery following TEF stabilization and do not address initial losses of C and nutrients during TEF formation. To capture the variability among and within TEFs, we simulate a range of post-stabilization conditions by varying the initial size of SOM pools and nutrient supply rates. This file contains the results for 25 years of tussock tundra under nitrogen fertilization conditions.
Long term response of arctic tussock tundra to thermal erosion features: A modeling analysis. Tussock tundra nitrogen and phosphorus fertilization simulation
The Multiple Element Limitation (MEL) model is used to simulate the recovery of Alaskan arctic tussock tundra to thermal erosion features (TEFs) caused by permafrost thaw and mass wasting. TEFs could be significant to regional carbon (C) and nutrient budgets because permafrost soils contain large stocks of soil organic matter (SOM) and TEFs are expected to become more frequent as climate warms. These simulations deal only with recovery following TEF stabilization and do not address initial losses of C and nutrients during TEF formation. To capture the variability among and within TEFs, we simulate a range of post-stabilization conditions by varying the initial size of SOM pools and nutrient supply rates. This file contains the results for 25 years of tussock tundra under nitrogen and phosphorus fertilization conditions.
Long term response of arctic tussock tundra to thermal erosion features: A modeling analysis. Tussock tundra phosphorus fertilization simulation
The Multiple Element Limitation (MEL) model is used to simulate the recovery of Alaskan arctic tussock tundra to thermal erosion features (TEFs) caused by permafrost thaw and mass wasting. TEFs could be significant to regional carbon (C) and nutrient budgets because permafrost soils contain large stocks of soil organic matter (SOM) and TEFs are expected to become more frequent as climate warms. These simulations deal only with recovery following TEF stabilization and do not address initial losses of C and nutrients during TEF formation. To capture the variability among and within TEFs, we simulate a range of post-stabilization conditions by varying the initial size of SOM pools and nutrient supply rates. This file contains the results for 25 years of tussock tundra under phosphorus fertilization conditions.
Long term response of arctic tussock tundra to thermal erosion features: A modeling analysis. Tussock tundra shade house simulation
The Multiple Element Limitation (MEL) model is used to simulate the recovery of Alaskan arctic tussock tundra to thermal erosion features (TEFs) caused by permafrost thaw and mass wasting. TEFs could be significant to regional carbon (C) and nutrient budgets because permafrost soils contain large stocks of soil organic matter (SOM) and TEFs are expected to become more frequent as climate warms. These simulations deal only with recovery following TEF stabilization and do not address initial losses of C and nutrients during TEF formation. To capture the variability among and within TEFs, we simulate a range of post-stabilization conditions by varying the initial size of SOM pools and nutrient supply rates. This file contains the results for 25 years of tussock tundra under shade conditions.
Urban habitat features and patterns of snake removals in the greater Phoenix, Arizona (USA) metropolitan area (March 2021 - March 2022)
In urban and suburban areas, wildlife and people are often in close quarters, leading to human-wildlife interactions (HWI). Understanding how wildlife interact with humans and the built environment is critical as urbanization contributes to habitat change and fragmentation globally. In our study, we partnered with a local business that removes and relocates snakes from homes and businesses in the Phoenix area. The most frequently removed were venomous (family Viperidae, e.g., rattlesnakes) and nonvenomous (family Colubridae, e.g., gophersnakes) snakes. Using these records, we investigated taxa-specific habitat trends at two spatial scales. The neighborhood scale focused on front yard measures of cover and vegetation classes and the landscape scale focused on variables related to vegetation indices and degree of urbanization. Both analyses compared areas where snakes were removed to random locations in the city to represent possible habitat available to snakes. At the neighborhood scale (n=60), we found that removals occurred in yards with abundant cover opportunities. At the landscape scale (n=764), we found species-specific differences with nonvenomous snakes removed from areas of higher urbanization compared to venomous snakes. Understanding these distinct habitat patterns in residential yards can identify areas with potential human-snake conflict.
Urban Residential Surface and Subsurface Hydrology: Synergistic Effects of Low-Impact Features at the Parcel Scale
Accurately predicting the hydrologic effects of urbanization requires an understanding of how hydrologic processes are affected by low‐impact development practices. In this study, we explored how growing season surface runoff, deep drainage, and evapotranspiration on a residential parcel are affected by several low‐impact interventions, including three "impervious‐centric" interventions (disconnecting downspouts, disconnecting sidewalks, and adding a transverse slope to the driveway and front walk), two "pervious‐centric" interventions (decompacting soil and adding microtopography), and all possible "holistic" combinations. Results were compared to both a highly and moderately compacted baseline parcel under an average and a dry weather scenario for a temperate climate. We find that under reasonable assumptions for highly compacted soil, pervious areas are a major source of runoff and disconnecting impervious surfaces may be relatively less effective without improving soil conditions. Under both highly and moderately compacted soil conditions, combining efforts to decompact soil with impervious disconnection has a synergistic effect on reducing surface runoff and increasing deep drainage and evapotranspiration. All combinations of interventions enhance infiltration, but the partitioning of additional root zone water between deep drainage and evapotranspiration depends on the weather scenario. Importantly, when all low‐impact interventions are applied together, growing season deep drainage is higher than that from a vacant lot with no impervious surfaces. We infer that ecohydrologic interfaces between impervious and pervious areas are strong controls on urban hydrologic fluxes and that high‐resolution, process‐based models can be used to account for these interfaces and thereby improve predictions of the hydrologic effects of low‐impact interventions.
Sharpening of Hierarchical Visual Feature Representations of Blurred Images
Open the record for dataset details and reuse information.
Supraglacial features of debris covered glaciers in the Himalaya from Landsat-8 spectral umixing and Pleiades
<p>This dataset contains the spectral unmixing output files for the debris covered glacier surfaces based on Landsat-8 OLI imagery and Pleiades imagery of 2015. Files are provided for two domains, the Khumbu reference region of Nepal and the greater Himalaya region (76.3 to 92.6° W and 26.3 to 34.2° N), which covers covering most area from Himachal/Jammu and Kashmir border to Bhutan Himalaya. </p> <ul> <li>Landsat surface reflectance : Himalaya_L8_6S_surface_reflectance_scenes_2015 .zip <ul> <li>Contains surface reflectance images of Landsat-8 OLI scenes mostly from 2015 (two images are from 2014 and 2016 due to clouds in 2015) </li> <li>Collection 1 Level 1 (L1TP)</li> <li>Atmospherically and topographically corrected using the ARCSI routine, supplied in .kea format. These can be converted to GeoTifs using the GDAL command.</li> <li>Naming structure: LS8_yyyymmdd_latYYlongXXXX_rRRpPPP_vmsk_topshad_rad_srefdem_stdsref.kea</li> <li>Projection is UTM (zones depending on the image), from the original Landsat L1TP files</li> <li>The file naming convention, which is a standard output from ARCSI routine, include the image date ("yyyy" = year, mm = "month", "dd" = day), latitude ("YY") and longitude ("XXXX") of the image center, path/row ("PPP" = path, "RRR" = row), and the output products generated by ARCSI ("rad" = radiation, "topshad" = topographic shadows, "srefdem" indicates the use of elevation data, "stdsref" = standardized surface reflectance)</li> </ul> </li> <li>Fractional maps for the Khumbu: LS8_20150930_r41p140_frac_files. zip <ul> <li>Raster format (GeoTiffs) </li> <li>Non-normalized fractional water, light and dark debris and vegetation maps for the Khumbu reference image (Sept 30, 2015, path 140 row 40)</li> <li>Output from the linear mixing model routine used to produce binary maps of surfaces with values ranging from 0 to 1 (0% to 100% pixel coverage)</li> </ul> </li> <li>Binary surface maps for the Himalaya: Himalaya_L8_raw_binary_surface_maps.zip <ul> <li>Vector format (ArcGIS shapefiles)</li> <li>Raw, unprocessed binary maps of ponds, vegetation debris, ice and clouds over the debris covered glacier tongues in the Himalaya around the year 2015 (binary files) </li> <li>Derived from tresholding the fractional maps using a variable threshold (see publication)</li> <li>Maps in this pre-release version have not been manually corrected for misclassified areas due to confusion of classes, and the ice and cloud classes are not highly accurate</li> <li>These are not the final coverages of these surfaces over the domain and should not be used as such</li> <li>The supraglacial pond maps will undergo manual corrections and the datasets will be updated on this page</li> </ul> </li> <li>Dataset for analysis, glacier-by-glacier: Himalaya_SDC_LS_for_analysis_gt1km2_with_frac_and_debris_attributes.txt <ul> <li>original data from the SupraGlacial Debris Cover dataset (Sherler et al 2018)</li> <li>updated with the preliminary fractional cover of each surface (in %) on a glacier-by-glacier basis</li> <li>contains only debris covered tongues >1 km2 </li> <li>debris covered attributes were calculated from the ALOS Global Digital Surface Model (AW3D30 DEM) for each debris covered tongue <ul> <li>DC_area_km2 = recalculated debris covered area</li> <li>DCmin = minimum debris cover elevation (meters)</li> <li>DCmax = maximum debris cover elevation (meters)</li> <li>DCrange = altitudinal range (meters)</li> <li>DCmed = median elevation (meters)</li> <li>SLmean = mean slope (degrees)</li> <li>SLrange = slope range (degrees)</li> <li>SLmin = min slope (degrees)</li> <li>SLmax = max slope (degrees)</li> </ul> </li> </ul> </li> </ul>
City features collection
<div> <div># City features collection</div> <br> <div>A collection of features for ~700 European cities, for the reference year 2018.</div> <br> <div>## Features</div> <br> <div>The features are divided in three main thematic areas: land, climate and socioeconomic characteristics. Find more information about the features in the codebook `cities_features_collection_codebook.csv`.</div> <div>Codelists for categorical features are in the same folder `codelist_<feature>.csv`.</div> <br> <div>## Cities</div> <br> <div>City selection (and outline polygon) is taken from the Eurostat Urban Atlas. More information [here](https://ec.europa.eu/eurostat/web/gisco/geodata/reference-data/administrative-units-statistical-units/urban-audit). The original list of cities with geometries can be downloaded at these links:</div> <br> <div>- EPSG:4326 (WGS84) <https://gisco-services.ec.europa.eu/distribution/v2/urau/geojson/URAU_RG_01M_2018_4326_CITIES.geojson></div> <div>- EPSG:3035 <https://gisco-services.ec.europa.eu/distribution/v2/urau/geojson/URAU_RG_01M_2018_3035_CITIES.geojson></div> <br> <div>Note: the dataset `city_features_collection.geojson` only contains the city outline in CRS EPSG:4326.</div> <br><br> <div>## Example usage</div> <br> <div>Clustering analysis of European cities: check out this interactive demo notebook: `notebooks\demo\cities_clustering_interactive_demo.ipynb`.</div> </div> <p> </p>
Long term response of arctic tussock tundra to thermal erosion features: A modeling analysis. Undisturbed tussock tundra
The Multiple Element Limitation (MEL) model is used to simulate the recovery of Alaskan arctic tussock tundra to thermal erosion features (TEFs) caused by permafrost thaw and mass wasting. TEFs could be significant to regional carbon (C) and nutrient budgets because permafrost soils contain large stocks of soil organic matter (SOM) and TEFs are expected to become more frequent as climate warms. These simulations deal only with recovery following TEF stabilization and do not address initial losses of C and nutrients during TEF formation. To capture the variability among and within TEFs, we simulate a range of post-stabilization conditions by varying the initial size of SOM pools and nutrient supply rates. This file contains the results for 100 years of undisturbed tussock tundra. Data is presented for day 250 of each year.
Cover and frequency of biological soil crust community types, moss species, vascular plants, and abiotic land surface features, on gypsum & non-gypsum soils from the Chihuahuan and Mojave Deserts in 2023
This dataset contains raw and calculated percent cover and frequency data for biological soil crust (hereafter biocrust) functional groups, vascular plant functional groups, and abiotic land surface features on and off gypsum soils in the northern Chihuahuan and eastern Mojave Deserts. Abundance data were obtained from 20 study sites total, 10 located on soils derived from gypsum parent material and 10 located on soils derived from non-gypsum parent materials. Sites were grouped into 10 pairs, in which every gypsum site was partnered with a non-gypsum site located in the same region. Apart from soil type, partnered-site characteristics (topography, climate, elevation, slope, aspect, and presence of biocrusts) were held relatively constant. At each site, cover and frequency assessments were made using the line-point intercept method (LPI) and frequency quadrats (1.0 m^2), respectively. Biocrust functional groups included the following crusts: lichen, moss, incipient algal, light algal, dark algal, unknown photosynthetic crust, and vagrant cyanobacteria. Vascular plant categories included: perennial forbs, perennial graminoids, annual forbs, annual graminoids, subshrub, shrub, Yucca, and cacti. Abiotic land surface features included: woody litter, herbaceous litter, bare soil, rock, bedrock, and animal feces. Moss crusts identified within cover and frequency analyses were sampled, and classified to species level via microscopy. The resulting percent cover and frequency data was used to understand differences in biocrust and moss species abundance and diversity on and off gypsum soils; furthermore, how biocrust and moss species abundance was associated with the measured environmental variables. Soil physical and chemical data from this study can be accessed at knb-lter-jrn.210616002. This study and dataset are complete.
Fault Analysis Database with Features (FADbF)
<p>This repository is also available in GitHub: <a href="https://github.com/leandroensina/FADbF">https://github.com/leandroensina/FADbF</a></p><p>The FADbF dataset companions the paper entitled "Fault Distance Estimation for Transmission Lines with Dynamic Regressor Selection", published in <i>Neural Computing and Applications</i>, <strong>doi</strong>: <a href="https://doi.org/10.1007/s00521-023-09155-y">10.1007/s00521-023-09155-y</a>. More information about the dataset can be found in this reference.</p><p><strong>Associated Tasks</strong>: classification and regression</p><p><strong>Instances</strong>: 168,000</p><p><strong>Attributes</strong>: 128, including the two possible targets</p><p><strong>Additional Information</strong>: this database comprises several attributes extracted from time series of fault simulations of a transmission line with 500 kV, 414 km, and 60 Hz. In total, we extracted 21 features separately for each of the three phases for both voltage and current waveforms along two post-fault cycles from a single terminal, resulting in 126 attributes (21 * 3 * 2 = 126) in addition to the two possible targets, i.e., fault type (classification task) and fault location (regression task). If desired, the fault type can also be used as a feature for the fault location task.</p>
Catalogues of Semantic Artefacts - Maturity Dimensions and Features
<p>This dataset contains, in three different formats (XSLX, CSV, and PDF), a description of twelve maturity dimensions identified from the literature that can be used to measure the maturity of the semantic artefacts catalogues (SAC). For each dimension, a number from 2 to 6 features has been added, for a total of 43 features overall.</p>
PTC_Rhetorical_features
<p><a href="https://zenodo.org/api/records/14221532/draft/files/sample_sentences_table.csv/content" target="_blank" rel="noopener noreferrer">sample_sentences_table.csv</a> - 350 sentences selected from the PTC corpus for annotation by 3 human annotators</p> <p><a href="https://zenodo.org/api/records/14221532/draft/files/annotations_table.csv/content" target="_blank" rel="noopener noreferrer">annotations_table.csv</a> - annotations by 3 human annotators of the 350 sample sentences</p> <p><a href="https://zenodo.org/api/records/14221532/draft/files/ChatGPT_table.csv/content" target="_blank" rel="noopener noreferrer">ChatGPT_table.csv</a> - GPT4 annotations of the 350 sample sentences</p> <p><a href="https://zenodo.org/api/records/14221532/draft/files/PTC_table.csv/content" target="_blank" rel="noopener noreferrer">PTC_table.csv</a> - 20,000 sentences from the PTC corpus labeled with 18 propaganda techniques</p> <p><a href="https://zenodo.org/api/records/14221532/draft/files/PTC_annotations_table.csv/content" target="_blank" rel="noopener noreferrer">PTC_annotations_table.csv</a> - Fine-tuned GPT3.5 annotations of the 20,000 PTC sentences</p> <p><a href="https://zenodo.org/api/records/14221532/draft/files/PTC_interpretable_features_table.csv/content" target="_blank" rel="noopener noreferrer">PTC_interpretable_features_table.csv</a> - One-hot encoded features representing the GPT3.5 annotations for 20,000 PTC sentences. Also included are other features.</p> <p>For more information please see <a href="https://doi.org/10.1145/3589335.3651909" target="_blank" rel="noopener noreferrer">https://doi.org/10.1145/3589335.3651909</a></p> <p> </p>
Summary of the most important features for selected ABs
<p>These data summarizes the relevant findings and the identified limitations (in terms of "Category", "Technology", "Properties", "Limitation", and "Applicability to railway"), coming from the overview of different Alternative Bearers (ABs), carried out in deliverable D21 (AB4Rail project, www.ab4rail.eu).<br> The results have provided an overview of several technologies, each of them showing specific characteristics. The heterogeneous nature of different ABs allows to provide a plethora of available communication technologies to be potentially used by the Adaptable Communication System (ACS) for different railway scenarios. All the selected ABs provide the IP interconnection feature since they are Integrated within OSI reference model.<br> In this way, it collects the planned objectives of deliverable D2.1, expressed as a technological overview of selected ABs, as possible candidates coexisting with Traditional Bearers (TBs) for supporting railway applications.</p>
NAAMES_IFCB_diatoms_image_features
<p>This text file (NAAMES_IFCB_diatoms_image_features.txt) contains data on all Imaging FlowCytobot (IFCB) images identified as diatoms (either individual cells or chains) during the North Atlantic Aerosol and Marine Ecosystem Study (NAAMES). Data are from the western North Atlantic, 2015-2018; see Behrenfeld et al. (2019) <em>Frontiers in Marine Science </em>for expedition details. The images were collected using an IFCB deployed onboard the ship and connected to the flowing sea-water system with intake at ~5 m depth. The images were classified and the data here identified as diatoms using a convolutional neural network with 90% accuracy. See https://github.com/emmettFC/selected-projects/blob/master/plankton_vision/README.md for details of the neural network development and testing. Data are used in the manuscript "Plankton Imagery Data Inform Satellite-Based Estimates of Diatom Carbon", Geophysical Research Letters, 49, e2022GL098076. https://doi. org/10.1029/2022GL098076. Additional data from the NAAMES expedition are available at https://seabass.gsfc.nasa.gov/naames.</p>
CoAID dataset with multiple extracted features (both sparse and dense) and degraded by OCR
<p>This is the same datasets as:</p> <p>Guillaume Bernard. (2022). CoAID dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630405</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). CoAID dataset texts with OCR degradations (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630710</p>
Fibvid dataset with multiple extracted features (both sparse and dense)
<p>This is a publication of the FibVid dataset originaly dedicated to fake news detection. We changed here the purpose of this dataset in order to use it in the context of event tracking in press documents.</p> <p>Kim, Jisu, Jihwan Aum, SangEun Lee, Yeonju Jang, Eunil Park, et Daejin Choi. 2021. « FibVID: Comprehensive Fake News Diffusion Dataset during the COVID-19 Period ». <em>Telematics and Informatics</em> 64 (novembre): 101688. <a href="https://doi.org/10.1016/j.tele.2021.101688">https://doi.org/10.1016/j.tele.2021.101688</a>.</p> <p>In this dataset, we provide multiple features extracted from the text itself. <strong>Please note the text is missing from the dataset published in the CSV format for copyright reasons. You can download the original datasets and manually add the missing texts from the original publications.</strong></p> <p>Features are extracted using:</p> <p>- A corpus of reference articles in multiple languages languages for TF-IDF weighting. (<em>features_news</em>) [1]</p> <p>- A corpus of tweets reporting news for TF-IDF weighting. (<em>features_tweets)</em> [1]</p> <p>- A S-BERT model [2] that uses <em>distiluse-base-multilingual-cased-v1 </em>(called <em>features_use</em>) [3]</p> <p>- A S-BERT model [2] that uses <em>paraphrase-multilingual-mpnet-base-v2 </em>(called <em>features_mpnet</em>) [4]</p> <p><strong>References:</strong></p> <p>[1]: Guillaume Bernard. (2022). Resources to compute TF-IDF weightings on press articles and tweets (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6610406</p> <p>[2]: Reimers, Nils, et Iryna Gurevych. 2019. « Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks ». In <em>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</em>, 3982‑92. Hong Kong, China: Association for Computational Linguistics. <a href="https://doi.org/10.18653/v1/D19-1410">https://doi.org/10.18653/v1/D19-1410</a>.</p> <p>[3]: https://huggingface.co/sentence-transformers/distiluse-base-multilingual-cased-v1</p> <p>[4]: https://huggingface.co/sentence-transformers/paraphrase-multilingual-mpnet-base-v2</p>
Event Registry dataset with multiple extracted features (both sparse and dense)
<p>This is a republication of the Event Registry dataset originaly published by:</p> <p>Rupnik, Jan, Andrej Muhic, Gregor Leban, Primoz Skraba, Blaz Fortuna, et Marko Grobelnik. 2016. « News Across Languages - Cross-Lingual Document Similarity and Event Tracking ». <em>Journal of Artificial Intelligence Research</em> 55 (janvier): 283‑316. <a href="https://doi.org/10.1613/jair.4780">https://doi.org/10.1613/jair.4780</a>.</p> <p>And reorganised for document tracking by:</p> <p>Miranda, Sebastião, Artūrs Znotiņš, Shay B. Cohen, et Guntis Barzdins. 2018. « Multilingual Clustering of Streaming News ». In <em>2018 Conference on Empirical Methods in Natural Language Processing</em>, 4535‑44. Brussels, Belgium: Association for Computational Linguistics. <a href="https://www.aclweb.org/anthology/D18-1483/">https://www.aclweb.org/anthology/D18-1483/</a>.</p> <p>In this dataset, we provide multiple features extracted from the text itself. <strong>Please note the text is missing from the dataset published in the CSV format for copyright reasons. You can download the original datasets and manually add the missing texts from the original publications.</strong></p> <p>Features are extracted using:</p> <p>- A corpus of reference articles in multiple languages languages for TF-IDF weighting. (<em>features_news</em>) [1]</p> <p>- A corpus of tweets reporting news for TF-IDF weighting. (<em>features_tweets)</em> [1]</p> <p>- A S-BERT model [2] that uses <em>distiluse-base-multilingual-cased-v1 </em>(called <em>features_use</em>) [3]</p> <p>- A S-BERT model [2] that uses <em>paraphrase-multilingual-mpnet-base-v2 </em>(called <em>features_mpnet</em>) [4]</p> <p><strong>References:</strong></p> <p>[1]: Guillaume Bernard. (2022). Resources to compute TF-IDF weightings on press articles and tweets (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6610406</p> <p>[2]: Reimers, Nils, et Iryna Gurevych. 2019. « Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks ». In <em>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</em>, 3982‑92. Hong Kong, China: Association for Computational Linguistics. <a href="https://doi.org/10.18653/v1/D19-1410">https://doi.org/10.18653/v1/D19-1410</a>.</p> <p>[3]: https://huggingface.co/sentence-transformers/distiluse-base-multilingual-cased-v1</p> <p>[4]: https://huggingface.co/sentence-transformers/paraphrase-multilingual-mpnet-base-v2</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.