Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,974
datasets available to search
ShareScore release 0.9.0
Dataset results
7,974 results for “V1”
BehaviouralTraitExpressionUnderClimaticForcing-V1.0
<p>R code used to generate statistical results, figures and models for williams et al. - <span>Species from regions of rapid climate transition show functionally important intra- and interspecific differences in trait expression</span>. Please see readme file for explanation of which scripts correspond to which results and figures.</p>
augMENTOR: Simulated Student Learning Profiles and their Engagement Metrics in TryHackMe Platform_V1
<p>The dataset provides simulated insights into student engagement and performance within the THM platform. It outlines mathematical representations of student learning profiles, detailing behaviors ranging from high achievers to inconsistent performers. Additionally, the dataset includes key performance indicators, offering metrics like room completion, points earned, and time spent to gauge student progress and interaction within the platform's modules.</p><p>Here are definitions of the learning profiles, along with mathematical representations of their behaviors:</p><ul><li>High Achiever: These are students who consistently perform well across all modules. Their performance can be described as a normal distribution centered at a high mean value. Their performance P in a given module can be modelled as: P = N(90, 5) where N is the normal distribution function, 90 is the mean, and 5 is the standard deviation.</li><li>Average Performer: These are students who typically perform at the average level across all modules. Their performance can be described as a normal distribution centered at a medium mean value: P = N(70, 10), where 70 is the mean, and 10 is the standard deviation.</li><li>Late Bloomer: These are students whose performance improves as they progress through the modules. Their performance can be modelled as: P = N(50 + i*10, 10), where i is the module index and shows an increasing trend.</li><li>Specialized Talent: These are students who have average performance in most modules but excel in a particular module (e.g., module5). Their performance can be described as: P = N(90, 5) if the module is module 5, else P = N(70, 10).</li><li>Inconsistent Performer: These are students whose performance varies significantly across modules. Their performance can be described as a normal distribution with a high standard deviation: P = N(70, 30), where 70 is the mean, and 30 is the high standard deviation, reflecting inconsistency.</li></ul><p>Note that the actual performances are bounded between 0 and 100 using the function max(0, min(100, performance)) to ensure valid percentages.</p><p>In these formulas, the <i>np.random.normal</i> function is used to simulate the variability in student performance around the mean values. The first argument to this function is the mean, and the second argument is the standard deviation, reflecting the level of variability around the mean. The function returns a number drawn from the normal distribution described by these parameters. Note that the proposed method is experimental and has not been validated. </p><p> </p><p>List of Key Performance Indicators (KPIs) for Student Engagement and Progress within the Platform:</p><ul><li>Room Name: This represents the unique identifier or name of a specific room (or module). Think of each room as a separate module or lesson within an educational platform. For example, Room1, Room2, etc.</li><li>Total rooms completed: Indicates the cumulative number of rooms that a student has fully completed. Completion is typically determined by meeting certain criteria, like answering all questions or achieving a certain score.</li><li>Rooms registered in: Represents the number of rooms a student has registered or enrolled in. This could be different from the total number of rooms they've completed.</li><li>Ratio of Questions completed per room: This gives an insight into a student's progress in a particular room. For instance, a ratio of 7/10 suggests the student has completed 7 out of 10 available questions in that room.</li><li>Room Completed (yes no): Indicates whether a student has fully completed a specific room or not. This could be determined by the percentage of material covered, questions answered, or a certain score achieved.</li><li>Room Last deploy (count of days): Refers to the number of days since the last update or deployment was made to that room. It can give an idea about the effort of the student.</li><li>Points in room used for the leaderboard (range 0-560): Each room assigns points based on student performance, and these points contribute to leaderboards. The range suggests that a student can earn anywhere from 0 to 560 points in a particular room.</li><li>Last answered question in a room (27th Jan 2023): This indicates the date when a student last answered a question in a specific room. It can provide insights into a student's recent activity and engagement.</li><li>Total points in all rooms (range 0-560): The cumulative score a student has achieved across all rooms.</li><li>Path Percentage completed (range 0-100): Indicates the percentage of the overall learning path that the student has completed. A path could consist of multiple modules or rooms.</li><li>Module Percentage completed (range 0-100): Represents how much of a specific module (which could have multiple lessons or topics) a student has completed.</li><li>Room Percentage completed (range 0-100): Shows the percentage of a specific room that has been completed by a student.</li><li>Time Spent on the platform (seconds): This provides an aggregate of the total time a student has spent on the entire educational platform.</li><li>Time spent on each room (seconds): Represents the amount of time a student has dedicated to a specific room. This can give insights into which rooms or modules are the most time-consuming or engaging for students.</li></ul>
5GENESIS-BERLIN-2021-BITMOVIN-V1.0
<p>Exported data from Bitmovin Analytics: a comprehensive video analytics system facilitating QoE assessment.</p>
5GENESIS-BERLIN-2021-NWI-V1.0
<p>Information about clients´ network connection in terms of general connection type, such as Wi-Fi or cellular.</p>
5GENESIS-BERLIN-2021-POSITION-V1.0
<p>Per-client position trace of longitute, latitude, and altitude, using available geolocation techniques, primarily GPS.Per-client position trace of longitute, latitude, and altitude, using available geolocation techniques, primarily GPS.</p>
VTEA-v1-ImageDataSets
<p>This combines the image datasets used for the manuscript describing an approach for the integrated tissue cytometry analysis of mesoscale confocal imaging datasets.</p> <p>Data use by figure:</p> <table> <tbody> <tr> <td><strong>Figure</strong></td> <td><strong>Image filename</strong></td> <td><strong>Image/data DOI</strong></td> </tr> <tr> <td>1</td> <td>NA</td> <td>NA</td> </tr> <tr> <td>2</td> <td>Kidney_Cortex_Human.tif</td> <td>10.5281/zenodo.5816199</td> </tr> <tr> <td>S2</td> <td>Kidney_Cortex_Human.tif</td> <td>10.5281/zenodo.5816199</td> </tr> <tr> <td>3</td> <td>Kidney_Cortex_Human_2.tif</td> <td>10.5281/zenodo.5842098</td> </tr> <tr> <td>4</td> <td>Human_Kidney_Cortex_Mesoscale.tif</td> <td>10.5281/zenodo.5842108</td> </tr> <tr> <td>5</td> <td>Human_Kidney_Cortex_Mesoscale.tif</td> <td>10.5281/zenodo.5842108</td> </tr> <tr> <td>6</td> <td>Kidney_Cortex_Human_Spectral_1.tif</td> <td>10.5281/zenodo.5842207</td> </tr> <tr> <td>7</td> <td>Kidney_Cortex_Human_CODEX.tif</td> <td>10.5281/zenodo.5826144</td> </tr> <tr> <td>S3</td> <td>Kidney_Cortex_Human_CODEX.tif</td> <td>10.5281/zenodo.5826144</td> </tr> </tbody> </table> <p>For channel information related to Kidney_Cortex_Human_CODEX.tif, please see 10.5281/zenodo.5826144.</p>
WP2_SIM-TST_DSB-SSBchirpPerformance_V1
<p>WP2_SIM_DSB-SSBperformance_V1.xslx</p> <p>WP2_TST_VCSEL20GHz-S-parametersBTB_V1.txt</p> <p>WP2_TST_VCSEL20GHz-S-parameters10km_V1.txt</p> <p>WP2_TST_VCSEL20GHz-S-parameters24km_V1.txt</p> <p>WP2_TST_VCSEL20GHz-S-parameters34km_V1.txt</p> <p>The results of the simulations performed to assess the performance of a directly-modulated source in different chirp conditions are reported in the Excel file. A standard dual sideband system and a single sideband condition are considered.</p> <p>The S-parameters of a short-cavity high-bandwidth VCSEL measured at various lengths (BTB, 10 km, 24 km, 34 km) are reported in the .txt files. The measured S-parameters are useful for calculating the transfer function of the SM fiber due to the interplay between the source chirp and the fiber chromatic dispersion, needed to obtain the chirp parameters <em>alfa</em> and <em>k</em> of the source.</p>
Matpower Ill-conditioned systems v1
<p>According to the definition of [1], an ill-conditioned power system can be described as this system which, despite its solution does exists, it is not reachable using the standard Newton-Raphson technique and a flat start. They are ill-conditioned systems building by combining some available cases from Matpower's database (see file MATPOWER_Ill-conditioned_systems.txt for information). These systems have been used in some papers for validating novel Power-Flow techniques.</p> <p>[1] F. Milano, "Continuous Newton's Method for Power Flow Analysis," in <em>IEEE Transactions on Power Systems</em>, vol. 24, no. 1, pp. 50-57, Feb. 2009.</p>
Data physicalization papers analysis (2019) V1
<p>Dataset containing a "report" of the reading of articles listed on <a href="http://dataphys.org/wiki/Bibliography">http://dataphys.org/wiki/Bibliography</a> with classification by human sense used in prototype/proposal.</p>
SMARTDEST DATASET WP2 V1
<p>SMARTDEST (https://smartdest.eu) is an EU-funded H2020 research project under the Grant Agreement no. 870753. It brings together 11 universities and 1 innovation centre from seven European and Mediterranean countries. It aims to develop innovative solutions in the face of the conflicts and externalities that are produced by tourism-related mobilities in cities, by informing the design of alternative policy options for more socially inclusive places in the age of mobilities.</p> <p>The SMARTDEST DATASET WP2 V1 includes data and indicators elaborated from different public sources (EUROSTAT, LFS, EU SILC) and some private sources that have been used to provide evidence on the territorial impacts of tourism mobilities and on exclusionary trends across the European space, which are collected in a technical report (SMARTDEST deliverable 2.2).</p> <p> </p> <p> </p>
Práctica web scraping motorflash_v1
<p>Esta práctica se ha realizado en el contexto de la asignatura de 'tipología y ciclo de vida de los datos' del Master en Ciencia de Datos de la Universitat Oberta de Catalunya. En ella, se aplican técnicas de web scraping utilizando el lenguaje de programación Python y la librería scrapy. Los datos se han extraído de la página web de anuncios de coches de segunda mano '<a href="https://www.motorflash.com/">https://www.motorflash.com/</a>', de la que se obtienen datos generales y características del vehículo anunciado. </p> <p>Para conocer más a fondo el proceso de extracción puede visitar el repositorio del proyecto <a href="https://github.com/CarlosRea/MotorflashScraper">https://github.com/CarlosRea/MotorflashScraper</a> </p>
Extended Data on China's 30-m Annual Cropland Dataset for 1990–2023 (CACD-v1)
<h2>CACD 2022 and 2023 are now available!</h2> <p>The 30-m Annual Cropland Dataset of China (CACD) provides long-term, high-resolution maps of cropland extent across the country and has been widely applied in diverse studies. To meet the growing needs of the research community, we have extended the dataset to include the years 2022 and 2023, reprocessing all spatial tiles on the Google Earth Engine platform. In this updated version, minor methodological adjustments were introduced to further improve classification accuracy. For details on the mapping procedures and performance, please refer to the attached document.</p> <p> </p> <p>Data description</p> <p>*Data format: GeoTIFF (.tif)</p> <p>*Pixel size: 30 m (∼ 0.00027°)</p> <p>*Projection: EPSG: 4326 (WGS84)</p> <p>*Values: 1 denotes cropland and 0 denotes non-cropland</p> <p> </p> <p>Reference: Ying Tu, Shengbiao Wu, Bin Chen, Qihao Weng, Yuqi Bai, Jun Yang, Le Yu, and Bing Xu*. A 30 m annual cropland dataset of China from 1986 to 2021. <em>Earth System Science Data</em> 16 (2024): 2297–2316. https://doi.org/10.5194/essd-16-2297-2024</p>
XRECO 3D Buildings and Monuments v1
<p>The dataset consists of 201 textured 3D models created with photogrammetry of monuments and buildings mainly across Europe. The 3D models depict various buildings and monuments mainly across Europe. They were cleaned manually by removing all background information from the scene and keeping only the main building. The data are annotated into 12 building classes including: castle, cathedral, church, city hall, factory, hotel, house, mosque, office, palace, school, villa.</p>
Daily 1-km gap-free PM2.5 grids in China, v1 (2000–2020)
<p>A Long-term Gap-free High-resolution Air Pollutants concentration dataset (abbreviated as LGHAP) is of great significance for environmental management and earth system science analysis. In the current release of LGHAP aerosol dataset (LGHAP.v1), we provide a 21-year-long (2000–2020) gap free PM2.5 concentration product with daily 1-km resolution covering the land area of China. The dataset was generated from the daily gap free AOD (https://doi.org/10.5281/zenodo.5652257) that was derived through an integration of a set of data tensors of AOD and other related datasets such as air pollutants concentration and atmospheric visibility acquired from diversified sensors or platforms via a machine learned regression model. The dataset was provided in the NetCDF format, while data in each individual year were archived in a zip file. Python, Matlab, R, and IDL codes were also provided to help users read and visualize the LGHAP data.</p>
Daily 1-km gap-free PM10 grids in China, v1 (2000–2020)
<p>A Long-term Gap-free High-resolution Air Pollutants concentration dataset (abbreviated as LGHAP) is of great significance for environmental management and earth system science analysis. In the current release of LGHAP aerosol dataset (LGHAP.v1), we provide a 21-year-long (2000–2020) gap free PM10 concentration product with daily 1-km resolution covering the land area of China. The dataset was generated from the daily gap free AOD (https://doi.org/10.5281/zenodo.5652257) that was derived through an integration of a set of data tensors of AOD and other related datasets such as air pollutants concentration and atmospheric visibility acquired from diversified sensors or platforms via a machine learned regression model. The dataset was provided in the NetCDF format, while data in each individual year were archived in a zip file. Python, Matlab, R, and IDL codes were also provided to help users read and visualize the LGHAP data.</p>
Daily 1-km gap-free AOD grids in China, v1 (2000–2020)
<p>A Long-term Gap-free High-resolution Air Pollutants concentration dataset (abbreviated as LGHAP) is of great significance for environmental management and earth system science analysis. In the current release of LGHAP aerosol dataset (LGHAP.v1), we provide a 21-year-long (2000–2020) gap free AOD product with daily 1-km resolution covering the land area of China. The dataset was generated via a seamless integration of the tensor flow based multimodal data fusion with ensemble learning based knowledge transfer in statistical data mining. The proposed method transformed a set of data tensors of AOD and other related datasets such as air pollutants concentration and atmospheric visibility that were acquired from diversified sensors or platforms via integrative efforts of spatial pattern recognition for high dimensional gridded data analysis toward data fusion and multiresolution image analysis. The daily gap free AOD was provided in the NetCDF format, while data in each individual year were archived in a zip file. Python, Matlab, R, and IDL codes were also provided to help users read and visualize the LGHAP data.</p>
Annual mean 1-km gap-free AOD, PM2.5, and PM10 grids in China, v1 (2000–2020)
<p>A Long-term Gap-free High-resolution Air Pollutants concentration dataset (abbreviated as LGHAP) is of great significance for environmental management and earth system science analysis. In the current release of LGHAP aerosol dataset (LGHAP.v1), we provide 21-year-long (2000–2020) gap free annual mean AOD, PM2.5 and PM10 concentration data with a 1-km resolution covering the land area of China. The dataset was generated from the daily gap free AOD (https://doi.org/10.5281/zenodo.5652257) that was derived through an integration of a set of data tensors of AOD and other related datasets such as air pollutants concentration and atmospheric visibility acquired from diversified sensors or platforms via a machine learned regression model. The dataset was provided in the NetCDF format, while data in each individual year were archived in a zip file. Python, Matlab, R, and IDL codes were also provided to help users read and visualize the LGHAP data.</p>
Historical Film Shot Dataset V1 (HistShotDS V1)
<p><em>Paper title: </em></p> <p><strong>HistShot: A Shot Type Dataset based on Historical Documentation during WWII</strong></p> <p><em>Conference title: </em></p> <p>International Conference on Pattern Recognition Applications and Methods (ICPRAM 2022)</p> <p> </p> <p><em>Description:</em></p> <p>Automated shot type classification plays a significant role in film preservation and indexing of film datasets. In this paper a historical shot type dataset (HistShot) is presented, where the frames have been extracted from original historical documentary films. A center frame of each shot has been chosen for the dataset and is annotated according to the following shot types: Close-Up (CU), Medium-Shot (MS), Long-Shot (LS), Extreme-Long-Shot (ELS), Intertitle (I), and Not Available/None (NA). The validity to choose the center frame is shown by a user study. Additionally, standard CNN-based methods (ResNet50, VGG16) have been applied to provide a baseline for the HistShot dataset.</p> <p> </p> <p><em>References: </em></p> <p>Github Repository: <a href="https://github.com/dahe-cvl/ICPRAM2022_histshotV1">https://github.com/dahe-cvl/ICPRAM2022_histshotV1</a></p> <p>VHH-MMSI: <a href="https://vhh-mmsi.eu/">https://vhh-mmsi.eu/</a></p> <p>VHH-project Page: <a href="https://www.vhh-project.eu/">https://www.vhh-project.eu/</a></p> <p> </p> <p> </p> <p> </p> <p> </p>
MOE_Golgi_Analyzer_v1
<p><strong>Description of the folder content :</strong></p> <p><br> 1) The macro in .ijm format.<br> Suited for analysis of 3-channel confocal fluorescence microscopy images of mammalian cells (~200*200µm). <br> Requires ImageJ v1.4 with Bio-render plugin.<br> Images should be as .nd2 format but it can easily be changed, simply search & replace all occurences of ".nd2" with your format in the macro code.<br> Images should be organized with every replicate of a same test-condition in a unique folder. The macro will analyze the whole folder at once and will create a folder in it to save results.</p> <p><br> 2) A folder named "example_data", it contains 3 representative images that can be used to test the macro. <br> It also contains a results folder with representative data obtained via the analysis of these representative images with the macro (see Description of the macro for description of the results obtained)</p> <p>____________________________</p> <p><strong>Description of the macro :</strong></p> <p>input : 3-channel image with </p> <p>C1 = nucleus labeling (e.g. DAPI, Hoechst, etc.) <br> C2 = signal of interest, the one you want to measure in whole cells & in the region of interest <br> C3 = region of interest (ROI) (e.g. an antibody directed against a particular organelle, in our case Golgi apparatus) <br> <br> this macro will : <br> count the cells according to C1 (user input of threshold values for C1) <br> create ROI(s) according to C3 (user input of threshold values, or manual setting of each image for C3) <br> measure signal of C2 (mean min max grey values, integrated density, area) in whole cells (user input of threshold values for C2) measure signal of C2 in ROI(s) <br> save results as a .csv file <br> <br> it will also create several .png images for each analyzed one : <br> C1+nucleusROI (to assess correct cell counting) <br> C3+ROIC3 (to assess correct creation of ROI(s) from C3 signal) <br> C2 (glow LUT) + ROIC3 (to assess correct thresholding of C2 signal) <br> C2+ROIC3 <br> merge C1+C2+C3</p>
MIDAS2 protocol example custom genome collection dataset v1
<p>Example input database of custom genome collection for MIDAS2 protocols.</p> <p>Two genomes from two species (<em>Staphylococcus epidermidis </em>and <em>Streptococcus mutans </em>)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.