Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
753
datasets available to search
ShareScore release 0.9.0
Dataset results
753 results for “metrics”
Paired Vegetation and Soil Burn Severity Metrics and Associated Climate, Weather, Topographical, and Land Cover Attributes
<p>This dataset pairs differenced Normalized Burn Ratio (dNBR) and soil burn severity (SBS) for 254 large (>400 ha in size) fires across the western US. Dataset also includes climate, weather, topography, physical and chemical soil characteristics, and land cover attributes of each burned pixel at the time of fire. This effort provided a table of 16.3 million burned pixels and their associated characteristics including dNBR, SBS, and 94 biological and physical covariates. After removing correlated features, the final data includes 18 fire covariates namely: dNBR, elevation, slope, aspect, land cover type, wind speed, energy release component, vapor pressure deficit, annual precipitation, and annual average daily max temperature, as well as the clay, sand and silt content of the soil and volumetric fraction of coarse fragments and soil organic carbon content. We also included spatial coherence metrices for dNBR, including DVAR, SHADE and SAVG. This data is provided as CSV files in Xtrain, Xvalidation, Xtest, as well as Ytrain, Yvalidation, and Ytest; in which X files (model input) provide all features except for SBS and Y files (model output) include SBS.</p><p>We also provided this data for an additional 16 large fires across the western US ("Extra Test" folder, including Dataset – X file – and Label – Y file).</p><p>Finally, the trained XGBoost model to translate dNBR to SBS using the associated features is also provided in this folder.</p>
Trypanosoma Epitope Dataset: Valid Epitopes and Randomly Generated Peptides with Biochemical Metrics and AI-Generated Scores
<p>This dataset contains information about valid linear B-Cell epitopes from the Trypanosoma genus, as well as randomly generated peptides. It includes biochemical metrics generated by the EpiBuilder-1.0 tool and scores generated by the BepiPred-3.0 software. The data was originally collected from the IEDB and UniProtKB platforms and has been processed and enhanced with these informations for researchers interested in understanding the molecular interactions between Trypanosoma protozoans and the immune system.</p>
Forest-related landscape metrics in LandKlif project
<p>We calculated forest-related landscape metrics which have influence on insect diversity. Based on the detailed Landklif map (dataset 11560 at LandKlif database, https://www.landklif.biozentrum.uni-wuerzburg.de), we classify coniferous forest, decideous forest, mixed forest, small wood, and transitional woodland-shrub as forest features. Sub land use class and origical classification was kept as well. This dataset includes the area percentage (landscape composition) of these classes as well as edge length between forest features and non-forest features, in a scale of 100, 200, 500, 1000, 1500 meter radius around the study plots, as well as in TK 25 quadrant scale. TK 25 quadrant is common name in Germany for the topographical map unit at a scale of 1:25000 designated by four-digit numbers, which has a long history (from 1875) and is been used as unit for geographical survey and biodiversity mapping (http://maps.snsb.info/TK25/).</p> <p>Detailed Landklif map was created by combining 3 different land cover maps to create a detailed land cover map for 6 km buffer area around landklif study plots. We used ATKIS 2019 land cover as basis, added details from Invekos 2019 and Corine 2018. We categorized the land cover into 6 classes, further subcategorized them into sub land use classes. The original classification from different sources are kept. In case of overlapping, the priority goes (from high to low): natural > forest > grassland > arable > urban > water. In case of overlapping between data source: transitional woodland-shrub from Corine > Invekos > ATKIS. Areas outside of Bayern are filled with only Corine data. The coordinate system of the shapefile is ETRS89 / UTM zone 32N (EPSG:25832). This dataset is not open access due to its sensitivity but can be reached (https://www.landklif.biozentrum.uni-wuerzburg.de/Download/ShowXml.aspx?DatasetId=11560) and requested via the LandKlif database.</p> <p>LandKlif is funded by the Bavarian State Ministry of Science and the Arts within the Bavarian Climate Research Network (bayklif). Within the five year funding period of bayklif, five interdisciplinary senior research associations and five junior research groups are be financed with a total sum of 18 million Euro. LandKliF, as one of the five interdisciplinary senior research associations, addresses the effects of climate change on biodiversity and ecosystem services in semi-natural, agricultural and urban landscapes.</p>
Datasets for testing the robustness of LiDAR vegetation metrics to varying point densities
<p><span>The calculation of vegetation metrics from LiDAR point clouds might be affected by the available point density of a dataset. Testing how the same LiDAR vegetation metrics differ with different point densities can therefore inform about their robustness for upscaling metrics to other areas or other LiDAR point clouds. The datasets made available here were generated to test the robustness of LiDAR vegetation metrics to varying point densities. A total of 25 LiDAR vegetation metrics representing different aspects of vegetation height, vegetation cover and structural complexity were tested (see metric definition in Kissling et al. 2023, <a href="https://doi.org/10.1016/j.dib.2022.108798">https://doi.org/10.1016/j.dib.2022.108798</a>). The metric calculation was similar to the metric calculation in the Laserchicken software (Meijer et al. 2020, <span><a href="https://doi.org/10.1016/j.softx.2020.100626">https://doi.org/10.1016/j.softx.2020.100626</a>) and the Laserfarm workflow (Kissling et al. 2022, https://doi.org/10.1016/j.ecoinf.2022.101836). The Dutch AHN4 dataset from the years 2020–2022 with a point density of 20–30 points/m<sup>2</sup> was used. A number of plots (i.e., squared polygons around centre points) were randomly placed across the Netherlands within Dutch Natura 2000 sites (using shapefiles from the European Environmental Agency). Different Dutch Natura 2000 sites were distinguished based on their dominant habitat type (dunes, grassland, marsh, shrubland, and woodland). About 100 plots were randomly placed in each habitat type. The AHN4 point cloud of each plot was clipped and then randomly downsampled to 1, 2, 5, 10, 15, 20 points per square meter, respectively. This was done for six different spatial resolutions (1, 2, 5, 10, 20 and 30 meter). The clipped points were then used to calculate the 25 LiDAR vegetation metrics for the original point density and for the six down-sampled point densities.</span></span></p>
Data from: BioEncoder: a metric learning toolkit for comparative organismal biology
<p><strong>BioEncoder: a metric learning toolkit for comparative organismal biology</strong></p> <p><strong>Abstract </strong>- In the realm of biological image analysis, deep learning (DL) has become a core toolkit, e.g., for segmentation and classification. However, conventional DL methods are challenged by large biodiversity datasets characterized by unbalanced classes and hard-to-distinguish phenotypic differences between them. Here we present BioEncoder, a user-friendly toolkit for metric learning, which overcomes these challenges by focussing on learning relationships between individual data points rather than on the separability of classes. BioEncoder is released as a Python package, created for ease of use and flexibility across diverse datasets. It features taxon-agnostic data loaders, custom augmentation options, and simple hyperparameter adjustments through text-based configuration files. The toolkit's significance lies in its potential to unlock new research avenues in biological image analysis while democratizing access to advanced deep metric learning techniques. BioEncoder focuses on the urgent need for toolkits bridging the gap between complex DL pipelines and practical applications in biological research.</p> <p><strong>Dataset </strong>- This data repository includes two things: a snapshot of the BioEncoder package (BioEncoder-main.zip, version 1.0.0, downloaded from https://github.com/agporto/BioEncoder on 2024-07-19 at 17:20), and the damselfly dataset used for the case study presented in the paper (bioencoder_data.zip). The dataset archive also encompasses the configuration files and the final model checkpoints from the case study, as well as a script to reproduce the results and figures presented in the paper.</p> <p><strong>How to use - </strong>Get started by consulting the <a href="https://github.com/agporto/BioEncoder?tab=readme-ov-file#quickstart">GithHub repository</a> for information on how to install BioEncoder, then download the <a href="../records/10909614/files/BioEncoder-data.zip?download=1&preview=1">data archive</a> and run the script. Some parts of the script can be executed using the model checkpoints, for orther parts the training rountine needs to be run. </p>
Device Performance Metrics as Function of Absorption Onset
<p>Updated database and overview plot based on previously published version in J. Mater. Chem. A, 2017, 5, 11401 (Unger et al.)</p>
Financial Metrics Dataset of US companies
<p>Financial Metrics Dataset of US companies obtained from 10-K filings (in XBRL format) from SEC. The dataset contains financial metric answers for 9,263 US companies and in total 800,714 metric answers related to 28 financial metrics.</p>
Phenology metric layers and their classification layers for the NDVI approximated phenological cycle of Donana from 01/12/2015 to 31/11/2016.
<p>Analysis of changes in the phenological cycle of different plant species provide important information that may be used to assess the impact of seasonal and inter-annual climate variations on terrestrial vegetation. Phenex software has been used for estimating phenology related layers for Donana marshes relying on NDVI time series covering one year period from 01/12/2015 to 31/11/2016.</p> <p>“Phenology_metrics_layer_Dec2015_Nov2016.tif” includes the following layers: (i) green up day, (ii) senescence day, (iii) day of max NDVI value, and (iv) total number of NDVI peaks. These layers are also provided separately with the names: “Greenup_day_Dec2015_Nov2016.tif”, “Senescence_day_Dec2015_Nov2016.tif”, “Max_day_Dec2015_Nov2016.tif”, “Number_of_peaks_Dec2015_Nov2016.tif”.</p> <p>Classification layers based on these layers have been also generated. In particular, "ISODATA_classification_all_input_layers_Dec2015_Nov2016.tif" layer contains the classes generated when providing all phenology related layers as input to the ISODATA algorithm, while "ISODATA_classification_three_input_layers_Dec2015_Nov2016.tif" layer contains contains the classes generated when providing three penology related layers (i.e. greenup day, day of max NDVI value, senescence day layers) as input to the ISODATA algorithm.</p> <p>The above files are accompanied by INSPIRE metadata XML files. Detailed information can be found in the “Readme.pdf” included in the zip containing the dataset.</p> <p> </p>
Fed4Fire/CDN-X-ALL network metrics dataset for time series analysis in Media content delivery for 4G/5G networks
<p>The following dataset was generated at VICOMTECH (https://www.vicomtech.org) under project/experiment CDN-X-ALL: "CDN edge-cloud computing for efficient cache and reliable streaming aCROSS Aggregated unicast-multicast LinkS".</p> <p>Project funded by Fed4FIRE+ OC5 (<a href="https://www.fed4fire.eu/">https://www.fed4fire.eu</a>) under grant 732638.</p> <p>The dataset provides network metrics captures across several days employing a GStreamer-based MPEG-DASH player running on an UE connected to a LTE network.</p> <p>Nitos LTE/OpenAirInterface (OAI) testbed (<a href="https://nitlab.inf.uth.gr/NITlab/nitos/lte">https://nitlab.inf.uth.gr/NITlab/nitos/lte</a>) was used to deploy the LTE network.</p> <p><strong>CDN-like server/DASH Dataset -> Internet -> EPC/OAI -> eNodeB/OAI -> UE/DASH player</strong></p> <p>The player downloads MPEG-DASH video files provided by Distributed DASH dataset (<a href="https://dash.itec.aau.at/distributed-dash-datset/">https://dash.itec.aau.at/distributed-dash-datset/</a>), a dataset for CDN-like experiments, and captures the following data:</p> <ol> <li>Date: date when the data is collected</li> <li>Player: type of the player (in this case it is always "GStreamer")</li> <li>Num: identifier of the player</li> <li>URLVid: URL of the MPD file</li> <li>Latency: latency experienced by the player</li> <li>BW: bandwidth experienced by the player</li> <li>Quality: chosen DASH video representation</li> </ol> <p>During the experiments, other players run in order to generate realistic media streaming traffic at the CDN-like servers. These players start playing by following Poisson or Pareto distribution.</p> <p>The dataset was used to train Machine Learning Time Series predictor in order to forecast network capabilities and can be used for further experimentation concerning time series analysis.</p>
Dataset of Kantorovich-Rubinstein-Wasserstein Polytopes of Metric Spaces on up to 6 Points
<p>We present a complete list of all combinatorial types of generic Kantorovich-Rubinstein-Wasserstein (KRW) polytopes associated with metric spaces on up to 6 points that are generic in the sense of Gordon and Petrov, see [1]. These polytopes and their properties are described in detail in [2].</p> <p>The catalog of KRW polytopes was computed using certain regular triangulations of the full root polytope, see Section 4 in [2]. These regular triangulations were enumerated up to symmetry by Jörg Rambau using the new <em>topcom</em> package described in [3].</p> <p>The provided data comes in three parts.</p> <ul> <li>The files ending in ".result" contain the original <em>topcom</em> output including the specific regular triangulations of the root polytope.</li> <li>There are <em>julia</em> files that contain these triangulations ("triangulations_x.jl"), one triangulation per line.</li> <li>There is an <em>OSCAR</em> script ("read_triangulations.jl") that reads these triangulations and produces sample metrics associated with each of these triangulations. </li> </ul> <h3>References:</h3> <p>[1] J. Gordon and F. Petrov: Combinatorics of the Lipschitz polytope, 2017, Arnold Math. J. <em>3.2.</em></p> <p>[2] E. Delucchi, L. Kühne, and L. Mühlherr: <em>Combinatorial invariants of finite metric spaces and the Wasserstein arrangement</em>, 2024, in preparation.</p> <p>[3] J. Rambau: <em>Symmetric lexicographic subset reverse search for the enumeration of circuits, cocircuits, and triangulations up to symmetry, </em>2023, <a href="https://www.wm.uni-bayreuth.de/de/team/rambau_joerg/TOPCOM/SymLexSubsetRS-2.pdf" target="_blank" rel="noopener">preprint</a>.</p>
Patch metrics and landscape patterns of forest disturbances at the beginning of the 20th Century
<h1>Summary:</h1> <p>The database consists of a compressed .CSV file containing structural information of forest disturbance patches identified between 2002 and 2014 using the Global Forest Change Tree Cover Loss Year dataset version 1.6 (Hansen et al, 2013) available at https://earthenginepartners.appspot.com/science-2013-global-forest/download_v1.6.html. Each row in the database represents a patch (249,149,911 in total). The columns (15) represent the structural metrics calculated for each patch, as well as the landscape patterns identified using kmeans cluster analysis. </p> <p>The methods used for building this database are published in the paper: Acil, N., Sadler, J.P., Senf, C. <em>et al.</em> Landscape patterns in stand-replacing disturbances across the world’s forests. <em>Nat Sustain</em> <strong>8</strong>, 86–98 (2025). <a href="https://doi.org/10.1038/s41893-024-01450-3">https://doi.org/10.1038/s41893-024-01450-3</a></p> <p>Aggregated global maps of the patch metrics can be visualised in <a href="https://ee-treemort-disturbances-nacil.projects.earthengine.app/view/patchmetrics2002-2014">Google Earth Engine</a> and accessed in the asset "http://projects/ee-treemort-disturbances-nacil/assets/PatchMetrics_Means_nonLU_2002-2014/". </p> <p>Some of the scripts associated with this project are hosted in <a href="https://github.com/N-Acil/GlobalForestDisturbances_PatchMetrics">GitHub</a> and <a href="https://code.earthengine.google.com/?accept_repo=users/NXA807/%20GlobalForestDisturbances_PatchMetrics">Google Earth Engine</a>.</p> <p>Additional scripts and data will be made available upon request.</p> <p> </p> <p> </p> <h1>Database structure: </h1> <h2>Patch metrics</h2> <h3>Occurrence: </h3> <p>Patch form and year were retrieved from the Global Forest Change tree cover loss year dataset version 1.6 (Hansen et al, 2013).</p> <table> <tbody> <tr> <td><strong>Column name</strong></td> <td><strong>Description</strong></td> <td><strong>Unit</strong></td> <td><strong>Format</strong></td> <td><strong>Valid values</strong></td> </tr> <tr> <td><strong>PID</strong></td> <td>Patch unique identifier in the format Tile_Year_PatchNumber (e.g. 01U_02_00000001).</td> <td> </td> <td>Characters</td> <td> </td> </tr> <tr> <td><strong>X_INT_deg</strong></td> <td>Longitude of the patch's internal centroid</td> <td>Degrees</td> <td>Float</td> <td>[-180-180]</td> </tr> <tr> <td><strong>Y_INT_deg</strong></td> <td>Latitude of the patch's internal centroid</td> <td>Degrees</td> <td>Float</td> <td>[-90-90]</td> </tr> <tr> <td><strong>YEAR_maj</strong></td> <td>Year of patch majority occurrence. </td> <td> </td> <td>Integer</td> <td>[2-14]</td> </tr> <tr> <td><strong>YEAR_n</strong></td> <td>Number of years over which the patch exhibited continuous growth.</td> <td> </td> <td>Integer</td> <td>>0</td> </tr> </tbody> </table> <h3>Metrics: </h3> <p>These patch and landscape metrics were calculated from the patch delineated.</p> <table> <tbody> <tr> <td><strong>Column name</strong></td> <td><strong>Description</strong></td> <td><strong>Unit</strong></td> <td><strong>Format</strong></td> <td><strong>Valid values</strong></td> </tr> <tr> <td><strong>AREA_G_ha</strong></td> <td>Patch geodesic area</td> <td>Hectares</td> <td>Float</td> <td>>0</td> </tr> <tr> <td><strong>PERIM_G_m</strong></td> <td>Patch geodesic perimeter</td> <td>Meters</td> <td>Float</td> <td>>0</td> </tr> <tr> <td><strong>PARA</strong></td> <td>Perimeter-area ratio</td> <td> </td> <td>Float</td> <td>>0</td> </tr> <tr> <td><strong>SHAPE</strong></td> <td>Shape index</td> <td> </td> <td>Float</td> <td>>=1</td> </tr> <tr> <td><strong>ELONG</strong></td> <td>Elongation index</td> <td> </td> <td>Float</td> <td>[0-1[</td> </tr> <tr> <td><strong>FRAC</strong></td> <td>Fractal dimension index</td> <td> </td> <td>Float</td> <td>[1-2]</td> </tr> <tr> <td><strong>NN5000_T0_n</strong></td> <td>Number of patches assigned the same year within 5 km radius.</td> <td> </td> <td>Integer</td> <td>>0</td> </tr> <tr> <td><strong>NN5000_AREA_T0_perc</strong><strong><br></strong></td> <td>Percent of the total area disturbed over the period 2001-2018 within 5 km radius from the focal patch centroid.</td> <td>%</td> <td>Float</td> <td>[0-100]</td> </tr> </tbody> </table> <h3>Clusters:</h3> <p>Cluster identification was performed using AREA_G_ha, YEAR_n, SHAPE, ELONG, NN5000_T0_n and NN5000_AREA_T0_perc. </p> <table> <tbody> <tr> <td><strong>Column name</strong></td> <td><strong>Description</strong></td> <td><strong>Unit</strong></td> <td><strong>Format</strong></td> <td><strong>Valid values</strong></td> </tr> <tr> <td><strong>CLUSTER_CODE</strong></td> <td>Code assigned to each cluster</td> <td> </td> <td>Integer</td> <td>[1-4]</td> </tr> <tr> <td><strong>CLUSTER_LABEL</strong></td> <td>Name given to the cluster identified. </td> <td> </td> <td>Character</td> <td> <ul> <li>Small-isolated</li> <li>Clustered</li> <li>Complex</li> <li>Large-multiyear</li> </ul> </td> </tr> </tbody> </table> <p> </p>
Country-wide data products for the ecosystem structure metrics derived from ALS data across the Netherlands (AHN3)
<p>This data repository contains country-wide data products for the ecosystem structure metrics generated from Airborne Laser Scanning (ALS) data across the Netherlands (AHN3). Twenty-five ecosystem structure metrics (at 10-meter resolution, GeoTIFF format) were derived from AHN3 dataset (<a href="https://downloads.pdok.nl/ahn3-downloadpage/">https://downloads.pdok.nl/ahn3-downloadpage/</a>) using <a href="https://laserfarm.readthedocs.io/en/latest/">Laserfarm</a> workflow (<a href="../record/5636773">https://zenodo.org/record/5636773</a>). Laserfarm is a free and open-source workflow that enables efficient, scalable, and distributed processing of multi-terabyte LiDAR point clouds from national and regional ALS surveys into LiDAR metrics of ecosystem structure. All code of Laserfarm is hosted and freely available on GitHub (<a href="https://github.com/eEcoLiDAR/Laserfarm">https://github.com/eEcoLiDAR/Laserfarm</a>). The Jupyter Notebooks for the processing of the AHN3 dataset are available on GitHub (<a href="https://github.com/eEcoLiDAR/AHN/tree/main/AHN3">https://github.com/eEcoLiDAR/AHN/tree/main/AHN3</a>).</p> <p>The twenty-five LiDAR metrics are related to three key dimensions of ecosystem structure (ecosystem height, ecosystem cover, and ecosystem structural complexity), and a layer of point density and a layer of building/road/water mask are also provided. Each GeoTIFF layer represents one LiDAR metric at 10 m resolution covering the whole Netherlands (file name as "ahn3_10m_feature_name.tiff").</p> <p>An overview of all the listed metrics (maps) is also provided in the PDF version (AHN3.pdf).</p> <p>A detailed description of the dataset is available from the following data publication:<br>Kissling, W. D., Y. Shi, Z. Koma, C. Meijer, O. Ku, F. Nattino, A. C. Seijmonsbergen, and M. W. Grootes. 2022. Country-wide data of ecosystem structure from the third Dutch airborne laser scanning survey. Data in Brief: 108798.<br><a href="https://eur04.safelinks.protection.outlook.com/?url=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.dib.2022.108798&data=05%7C01%7Cy.shi%40uva.nl%7C177a19a4359a422b0ef808dad9d30ef8%7Ca0f1cacd618c4403b94576fb3d6874e5%7C0%7C0%7C638061797757145956%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&sdata=2R7NSGli4Mw6Pp5FAIyOzBu4USPZXigng46EFVT4X68%3D&reserved=0">https://doi.org/10.1016/j.dib.2022.108798</a></p> <p>A detailed description of all the metrics can be found in the README file (README.docx). </p> <p>A .zip file is also provided containing all the data for the validation of the AHN3 data products (AHN3_validation.zip). </p>
Mapping a novel metric for Flash Flood Recovery using Interpretable Machine Learning
<p>This data is supplementary to the paper titled "Mapping a novel metric for Flash Flood Recovery using Interpretable Machine Learning". The file contains the main results.<br><br>For any queries, please visit <a href="https://hydrosense.iitd.ac.in" target="_blank" rel="noopener">Hydrosense Lab (IIT Delhi)</a>.</p>
Data supporting "Burn Period: A use-inspired metric to track wildfire risk across the southwest U.S."
<p>Comma delimited data file of derived daily meteorological metrics from hourly, gap filled and quality controlled Remote Automated Weather Station (RAWS) data for Arizona and New Mexico (southwest U.S.) provided by the Climate, Ecosystems, and Fire Applications (CEFA) program at the Desert Research Institute (Brown, 2022, unpublished data). Data file contains daily average dewpoint temperature, air temperature, maximum Hot-Dry-Windy Index, maximum Fosberg Fire Weather Index, maximum vapor pressure deficit, and total number of hours/day with relative humidity below 20% for 124 RAWS from 2000-2022.</p>
Metrics for two-sample tests: results on JetNet dataset
<p>The repository includes version 1.0 (v1.0) of the code and results corresponding to the GitHub repository <a href="https://github.com/TwoSampleTests/JetNetMetrics" target="_blank" rel="noopener">JetNetMetrics</a>.</p> <p>Publishing information and arXiv identifier will be added after publication of the main manuscript related to the data.</p>
Metrics for two-sample tests: results on Mixture of Gaussians and Correlated Gaussians models
<p>The repository includes version 1.0 (v1.0) of the code and results corresponding to the GitHub repository <a href="https://github.com/TwoSampleTests/GenerativeModelsMetrics">GenerativeModelsMetrics</a>.</p> <p>Publishing information and arXiv identifier will be added after publication of the main manuscript related to the data.</p>
R scripts for analyzing LiDAR data to assess forest canopy structure and perform Principal Component Analysis (PCA) on derived metrics
<p>This repository contains R scripts for analyzing LiDAR data to assess forest canopy structure and perform Principal Component Analysis (PCA) on spectral and LiDAR-derived metrics. The scripts cover LiDAR data processing, canopy height model (CHM) generation, calculation of forest canopy metrics, and PCA analysis.</p>
Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis
<p>This repository contain datasets and results for the paper:</p> <p><strong>Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis</strong></p> <p> </p> <p><strong>Github repository for the code: </strong></p> <p><a href="https://github.com/siebeniris/QuantifyingLanguageConfusion/tree/main">Quantifying Language Confusion GitHub repo</a></p> <p> </p> <p><strong>DATA</strong> include the following datasets:</p> <p>i) raw language graphs and</p> <p>ii) the calculated language similarities from the language graphs,</p> <p>iii) <strong>MTEI</strong>: the files from the <a href="https://github.com/siebeniris/vec2text_exp/tree/aaai">experimental results of multilingual inversion attacks</a>, and calculated language confusion entropy from the data;</p> <p>iv) <strong>LCB</strong>: the files from the <a href="https://github.com/for-ai/language-confusion?tab=Apache-2.0-1-ov-file#readme">language confusion benchmark</a> and calculated language confusion entropy from the data </p> <p> </p> <p><strong>Results</strong> include aggregated results for further analysis:</p> <p>i) <strong>inversion_language_confusion</strong>: results from MTEI</p> <p>ii) <strong>prompting_language_confusion</strong>: results from LCB</p> <p> </p> <p> </p>
Metrics As Scores Dataset: Price, Weight, and Other Properties of Over 1,200 Ideal-Cut and Best-Clarity Diamonds
<p>This dataset is a subset of the original diamonds dataset with more than 54,000 diamonds. It was reduced to only contain diamonds of the best cut (ideal) and clarity (IF). The group is now given by the colors from J (worst) to D (best). This dataset comes from the R-package ggplot2 (Wickham 2016). For each color, we can examine the following attributes (<strong>features</strong>) of each diamond:</p> <ul> <li><em>Carat</em>: Weight of the diamond</li> <li><em>Depth</em>: Total depth percentage</li> <li><em>Price</em>: Price in US dollars [discrete]</li> <li><em>Table</em>: Width of top of diamond relative to widest point</li> <li><em>X</em>: Length in mm</li> <li><em>Y</em>: Width in mm</li> <li><em>Z</em>: Depth in mm</li> </ul> <p>It has a total of 7 Colors (<strong>groups</strong>): <em>D</em>, <em>E</em>, <em>F</em>, <em>G</em>, <em>H</em>, <em>I</em>, and <em>J</em>. The best color is <em>D</em> and the worst color is <em>J</em>. This dataset was created to analyze whether there are differences between the colors.</p>
Metrics As Scores Dataset: Elisa Spectrophotometer Positive Samples
<p>The ELISA dataset contains data from a spectrophotometer that determined the optical density of positive control samples, from five different lots, across five different runs. This dataset was introduced by (Schmid 1991).</p> <p>This dataset has the following <strong>Features</strong>:</p> <ul> <li><em>Lot1</em>: The first lot</li> <li><em>Lot2</em>: The second lot</li> <li><em>Lot3</em>: The third lot</li> <li><em>Lot4</em>: The fourth lot</li> <li><em>Lot5</em>: The fifth lot</li> </ul> <p>It has a total of 5 <strong>Groups</strong>: <em>Run1</em>, <em>Run2</em>, <em>Run3</em>, <em>Run4</em>, and <em>Run5</em>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.