Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
635
datasets available to search
ShareScore release 0.9.0
Dataset results
635 results for “attribution”
Dataset for the manuscript "Attribute Recognition: A New Method for Grouping Planetary Images by Visual Characteristics, Using the Example of Mn-Rich Rocks in the Floor of Gale Crater, Mars."
<p>This dataset supports the manuscript "Attribute Recognition: A New Method for Grouping Planetary Images by Visual Characteristics, Using the Example of Mn-Rich Rocks in the Floor of Gale Crater, Mars." The dataset is contained in a single CSV file with 201 data rows (one row per NASA Curiosity rover ChemCam instrument target used in the study). The columns in this dataset include the martian solar day (sol) on which each target was imaged by ChemCam; the standoff distance from ChemCam to each target (in meters); binary columns (values are either 1 or 0, indicating presence or absence, respectively) for each of the 17 visual attributes we documented for each target image; the corresponding greyscale ChemCam RMI mosaic file location (on the Planetary Data System); and columns indicating which group each target was sorted into under each classification algorithm discussed in the text (P_{SG}: simple graph method; P_{AP}: automatic partitioning method; P_{\lambda=1.6}: community detection method with \lambda=1.6). To obtain the binary strings used for the classification algorithms, the 17 visual attribute columns can be concatenated. </p> <p>Also included is a collection of HTML files that enables easy viewing of the RMI mosaics in each cluster, using the Planetary Data System links. To use it, download the <code>.zip</code> file, unzip it, and open the <code>index.html</code> file in the browser of your choice (likely will work to simply double-click <code>index.html</code>)</p>
CAMELS-IND: hydrometeorological time series and catchment attributes for 472 catchments in Peninsular India
<p>We introduce <strong>CAMELS-IND</strong> (<em><strong>C</strong>atchment <strong>A</strong>ttributes and <strong>ME</strong>teorology for <strong>L</strong>arge-sample <strong>S</strong>tudies – <strong>India</strong></em>), a dataset containing hydrometeorological time series, and catchment attributes for 472 catchments in Peninsular India, of which 242 catchments have observed streamflow data available for over 30% of the period between 1980 to 2020. This dataset aims to foster large-sample hydrological studies within India and encourage the inclusion of Indian catchments in global hydrological research.</p> <p>The data set covers <strong>41 years</strong> of data between <em><strong>1st January 1980</strong></em> and <em><strong>31st December 2020</strong></em> for each catchments: daily time series of available streamflow observations, meteorological data such as precipitation, air temperature, solar radiation, relative humidity, wind speed, potential and actual evapotranspiration, and soil moisture. Additionally, CAMELS-IND includes regionally trained LSTM-model predicted streamflow for all 472 catchments. The static catchment attributes includes location and topography, climate, hydrological signatures, land-use, land cover, soil, geology, and anthropogenic influences.</p> <p>The corresponding manuscript is published in the Journal "Earth System Science Data" (ESSD).<br>Mangukiya, N. K., Kumar, K. B., Dey, P., Sharma, S., Bejagam, V., Mujumdar, P. P., and Sharma, A.: CAMELS-IND: hydrometeorological time series and catchment attributes for 228 catchments in Peninsular India, Earth Syst. Sci. Data, 17, 461–491, <a href="https://doi.org/10.5194/essd-17-461-2025" target="_blank" rel="noopener">https://doi.org/10.5194/essd-17-461-2025</a>, 2025.</p> <p>The data description file (<strong><em>CAMELS_IND_Data_Description.pdf</em></strong>) contains a comprehensive list of all time series and attribute variables covered by the dataset and references to the original data sources.</p> <p> </p> <h3><strong>### CAMELS-IND attributes/forcings history</strong></h3> <p>--------------------------------------<br><strong>Version 2.2: March 2025</strong><br>--------------------------------------<br><strong>Major changes/additions:</strong><br>- Updated streamflow observations: CAMELS-IND now includes 242 catchments with streamflow observations for more than 30% of the period between 1980 and 2020.<br>- "<em>CAMELS_IND_Catchments_Streamflow_Sufficient.zip</em>" contains a subset of 242 catchments with observed streamflow data available for more than 30% of the duration between 1980 and 2020.</p> <p><strong>Minor changes:</strong><br>- Correction to the forcing column headings for 'evap_canopy (kg/m²/s)' and 'evap_surface (kg/m²/s)': the units have been updated to (mm/day).<br>- Corrections made to the gauge_id mapping for basin codes 12 and 15.</p> <p> </p> <p>--------------------------------------<br><strong>Version 2.1: October 2024</strong><br>--------------------------------------<br>Data described in revised ESSD paper - <br><strong>Major changes/additions:</strong><br>- The dataset name "<em>CAMELS-INDIA</em>" has been changed back to "<em>CAMELS-IND</em>" to align with the naming convention of other CAMELS datasets.<br>- A Python script file “<em>filter_catchment.py</em>” is added to filter out the subset of the dataset based on flow data availability.<br>- <strong>"<em>CAMELS_IND_Catchments_Streamflow_Sufficient.zip</em>" contains a subset of 228 catchments with observed streamflow data available for more than 30% of the duration between 1980 and 2020.</strong></p> <p> </p> <p>-----------------------------------<br><strong>Version 2: August 2024</strong><br>-----------------------------------</p> <p><strong>Major changes/additions:</strong><br>- The dataset name "<em>CAMELS-IND</em>" has beed changed to "<em>CAMELS-INDIA</em>".<br>- All forcing time series have been extended from 01/01/1980 to 31/12/2020.<br>- Two new forcings, "<em>pet_gleam</em>" and "<em>aet_gleam</em>", have been added.<br>- Several attributes have been added, including gauge elevation, mean drainage path slopes, precipitation uniformity, asynchronicity, gini coefficient, base flow index, stream elasticity, slope of FDC, water table depth, and anthropogenic influence.<br>- Available observed streamflow time series has been added for all 472 catchments for the period 01/01/1980 to 31/12/2020.<br>- Regionally trained LSTM model-predicted streamflow has been added for all 472 catchments for the period 01/01/1980 to 31/12/2020.</p> <p><strong>Minor changes:</strong><br>- Attribute files have been renamed to "<em>camels_India_XXXX</em>"</p> <p>A short descriptions of all attributes and time series is provided in "<em>camels_India_data_description.pdf"</em>.</p> <p><strong>The following attributes are included in CAMELS-INDIA v2:</strong><br>07 attributes : camels_India_name<br>16 attributes : camels_India_topo (topography and location)<br>42 attributes : camels_India_clim (climate indices)<br>73 attributes : camels_India_hydro (hydrological signatures)<br>13 attributes : camels_India_land (land cover characteristics)<br>28 attributes : camels_India_soil (soil characteristics)<br>07 attributes : camels_India_geol (geological characteristics)<br>25 attributes : camels_India_anth (anthropogenic influences)<br>---------------<br><strong>Total:</strong> 211 attributes, 19 catchment mean forcings, and available observed and LSTM-based predicted streamflow time series.</p> <p> </p> <p>--------------------------------<br><strong>Version 1 : April 2024</strong><br>--------------------------------<br>A short descriptions of all attributes and forcings are described in "<em>camels_ind_attributes.xlsx</em>" and "<em>camels_ind_forcings.xlsx</em>"<br><em>Following attributes were included in CAMELS-IND 1.0:</em><br>06 attributes : camels_ind_name<br>14 attributes : camels_ind_topo (topography and location)<br>36 attributes : camels_ind_clim (climate indices)<br>64 attributes : camels_ind_hydro (hydrological signatures)<br>13 attributes : camels_ind_land (land cover characteristics)<br>27 attributes : camels_ind_soil (soil characteristics)<br>07 attributes : camels_ind_geol (geological characteristics)<br>13 attributes : camels_ind_anth (anthropogenic influences)<br>-------------<br><strong>Total:</strong> 180 attributes & 17 catchment mean forcings. </p> <p> </p> <p>-----------------------------------------------------<br><strong>### CONTRIBUTE TO CAMELS-IND</strong><br>-----------------------------------------------------</p> <p>If you are working with a data set covering Indian catchments and would like to contribute catchment averages to <em>CAMELS-IND</em>, please get in touch.</p> <p>We are committed to identifying and correcting errors. If you encounter any unrealistic or suspicious values, please notify us as soon as possible. Thank you for your assistance.</p> <p><strong>Contacts:</strong><br>- Nikunj K. Mangukiya (<em>nikk.mangukiya@gmail.com</em>)<br>- Ashutosh Sharma (<em>ashutosh.sharma@hy.iitr.ac.in</em>)</p> <p> </p> <p>--------------------------------<br><strong>### Acknowledgments</strong><br>--------------------------------</p> <p>The authors gratefully acknowledge the Central Water Commission (CWC), the National Water Informatics Centre (NWIC), and the Ministry of Jal Shakti (MoJS) for providing the streamflow dataset through the online portal, India – Water Resources Information System (India-WRIS; <a href="https://indiawris.gov.in/wris/#/">https://indiawris.gov.in/wris/#/</a>). The authors also extend their gratitude to the India Meteorological Department (IMD), Ministry of Earth Sciences, Government of India, for providing the gridded rainfall and temperature datasets through their respective websites. Additionally, the authors gratefully acknowledge the National Centre for Medium Range Weather Forecasting (NCMRWF), Ministry of Earth Sciences, Government of India, for the Indian Monsoon Data Assimilation and Analysis (IMDAA) reanalysis. The IMDAA reanalysis was produced under the collaboration between UK Met Office, NCMRWF, and IMD, with financial support from the Ministry of Earth Sciences under the National Monsoon Mission programme. The authors utilized numerous publicly available datasets for compiling catchment attributes and meteorological forcing time series, duly acknowledging and citing them where applicable. The authors extend their gratitude to all the researchers and contributing authors of these open-source datasets.</p> <p> </p> <p>--------------------------------<br><strong>### Disclaimer</strong><br>--------------------------------</p> <p>The CAMELS-IND dataset provided on this webpage is openly accessible for academic and research purposes. While efforts have been made to ensure data accuracy, the authors do not take any responsibility for errors, omissions, or misuse of the data. Users must cite the following paper when utilizing the dataset and acknowledge that all interpretations and conclusions drawn from the data are their own. We encourage users to cite/acknowledge the original data sources wherever required based on the source data usage policies. The dataset is provided "as is" without any warranties, and users are advised to check for updates. It is strongly recommended that users exercise caution and verify the data before using it for any purpose. The authors assume no responsibility for any consequences arising from the use or misuse of this dataset.</p> <p><strong>How to cite:</strong> Mangukiya, N. K., Kumar, K. B., Dey, P., Sharma, S., Bejagam, V., Mujumdar, P. P., and Sharma, A.: CAMELS-IND: hydrometeorological time series and catchment attributes for 228 catchments in Peninsular India, Earth Syst. Sci. Data, 17, 461–491, <a href="https://doi.org/10.5194/essd-17-461-2025" target="_blank" rel="noopener">https://doi.org/10.5194/essd-17-461-2025</a>, 2025.</p>
Catchment attributes and hydro-meteorological time series for large-sample studies across hydrologic Switzerland (CAMELS-CH)
<p>CAMELS-CH (Catchment Attributes and MEteorology for large-sample Studies - Switzerland) is a large-sample hydro-meteorological data set for hydrological Switzerland in Central Europe that covers 331 basins within Switzerland and neighboring countries (Austria, France, Germany and Italy). CAMELS-CH comprises dynamic hydro-meteorological variables and static catchment attributes.</p> <p>The data set covers 40 years of data between 1st January 1981 and 31st December 2020 for each catchment: daily time series of stream flow and water levels, of meteorological data such as precipitation and air temperature and of daily snow water equivalent data. Additionally, CAMELS-CH encompasses annual time series of land cover change and glacier evolution per catchment. The static catchment attributes comprise the following categories: location and topography, climate, hydrology, soil, hydrogeology, geology, land use, human impact and glaciers.</p> <p>The corresponding manuscript is published at the journal "Earth System Science Data" (ESSD) and available <a href="https://essd.copernicus.org/articles/15/5755/2023/">here</a>. The code used to generate the dataset is available on <a href="https://github.com/camels-ch">Github</a>.</p> <p>The data description file below contains a comprehensive list of all time series and attribute variables covered by the dataset and references to the original data sources. Further, this repository contains the "Caravan extension CH" for the "Caravan - A global community dataset for large-sample hydrology" <a href="../records/7944025">Caravan dataset</a> (see the <a href="https://github.com/kratzert/Caravan/discussions/10">list of extensions</a>). This extension has the same format like other Caravan parts and is based on the same data sources. Note that some features like the annual glacier time series, etc. are therefore only available in the original CAMELS-CH dataset.</p> <p> </p> <h2>Updates:</h2> <p>- Update version 0.9: affects "Caravan_extension_CH" - In version 1.5 of the Caravan dataset, Penman-Monteith PET was added as an additional time series feature. Additional to the new time series feature, also all pet-related climate indices were recomputed using the new Penman-Monteith PET. For consistency, the old ERA5-Land potential_evaporation time series and climate indices were kept, but renamed for a better identification of the differences. </p> <p>- Update version 0.8: resolving projection issue for shapefiles in "Caravan_extension_CH" using EPSG:4326 (WGS84); updating readme file of "camels_ch" regarding the <a href="../communities/dischma/">Dischma</a> catchment</p> <p>- Update version 0.7: update corresponding to the revision of the manuscript at "Earth System Science Data" (ESSD)</p> <ul> <li>dataset file delimiters have been changed to commas from semicolons</li> <li>the "time_series" folder was renamed to "timeseries"</li> <li>in the simulation-based data, there was an error in the previous aggregation of precipitation and evapotranspiration. The corresponding time series, affected hydrologic signatures and climatic indices were corrected</li> <li>the order of simulation-based variables in the timeseries files was changed to resemble the order shown in the tables of the corresponding publication in ESSD</li> <li>blank values that were masked by "NA" are now consistently indicated by "NaN"</li> <li>the readme file has been extended</li> </ul> <p>- Update version 0.6: updating links to related material (all links and references are available in the preprint/manuscript) and abstract</p> <p>- Update version 0.5: adding the "camels_ch_data_description.pdf" file</p> <p>- Update version 0.4: update of several static attributes in "Caravan_extension_CH" following a general update in Caravan and all its extensions + adopting the geographic coordinate system to Caravan-standard EPSG:4326</p> <p>- Update version 0.3: renaming single files/entries in "Caravan_extension_CH" to start with "camelsch" as unique Caravan extension identifier</p> <p>- Update version 0.2: CH extension to <a href="../records/7944025">Caravan</a> added</p>
CAMELS-BR: Hydrometeorological time series and landscape attributes for 897 catchments in Brazil - link to files.
<blockquote> <h3><strong>Version 1.2 (March 2025): </strong>Now with longer time series, expanded stream gauge coverage, meteorological data from additional sources, soil moisture time series, and observed rainfall time series from 11,853 rain gauges.</h3> </blockquote> <p> </p> <p>This is the CAMELS-BR dataset (Catchment Attributes and MEteorology for Large-sample Studies – Brazil) accompanying the paper: Chagas, V. B. P., Chaffe, P. L. B., Addor, N., Fan, F. M., Fleischmann, A. S., Paiva, R. C. D., and Siqueira, V. A.: CAMELS-BR: hydrometeorological time series and landscape attributes for 897 catchments in Brazil, Earth Syst. Sci. Data, 12, 2075–2096, <a href="https://doi.org/10.5194/essd-12-2075-2020" target="_blank" rel="noopener">https://doi.org/10.5194/essd-12-2075-2020</a>, 2020.</p> <p>CAMELS-BR provides daily observed streamflow time series for 4,025 stream gauges, daily observed rainfall for 11,853 rain gauges, daily meteorological time series and 65 attributes for 897 catchments in Brazil.</p> <p>The daily hydrometeorological time series include (i) observed streamflow accompanied by quality control information, (ii) precipitation extracted from five products, (iii) actual evapotranspiration extracted from three products, (iv) potential evapotranspiration extracted from two products, (v) reference evapotranspiration extracted from one product, (vi) minimum, mean, and maximum temperature extracted from three products, and (vii) soil moisture extracted from two products.</p> <p>The 65 catchment attributes cover properties such as (i) topography, (ii) climate, (iii) hydrology, (iv) land cover, (v) geology, (vi) soil, and (vii) human intervention.</p> <p>The data follow the same standards as other CAMELS datasets such as for the United States (https://doi.org/10.5194/hess-21-5293-2017), Chile (https://doi.org/10.5194/hess-22-5817-2018), and Great Britain (https://doi.org/10.5194/essd-2020-49).</p> <p><strong>How to cite:</strong> Chagas, V. B. P., Chaffe, P. L. B., Addor, N., Fan, F. M., Fleischmann, A. S., Paiva, R. C. D., and Siqueira, V. A.: CAMELS-BR: hydrometeorological time series and landscape attributes for 897 catchments in Brazil, Earth Syst. Sci. Data, 12, 2075–2096, https://doi.org/10.5194/essd-12-2075-2020, 2020.</p> <p> </p> <h3><strong>Changes in CAMELS-BR version 1.2:</strong></h3> <p><strong>Major changes</strong></p> <ul> <li>Updated streamflow time series up to February 2025 (where available), as obtained from ANA's website on 27 February 2025 (ANA – Brazilian National Water and Sanitation Agency – http://www.snirh.gov.br/hidroweb/). Some historical records have changed slightly due to ANA's quality control procedures. For eight gauges (see the readme.txt file), data are merged from 2025 and 2019 records (i.e. from CAMELS-BR version 1.1).</li> <li>Increased the stream gauge coverage to 4025 stream gauges (including both quality-controlled and non-quality-controlled series), up from 3679 in version 1.1.</li> <li>Added daily observed rainfall time series for 11853 rain gauges (not catchment averages), as obtained from ANA's website on 27 February 2025 (ANA – Brazilian National Water and Sanitation Agency – http://www.snirh.gov.br/hidroweb/). Data include quality flags but are mostly not quality-controlled.</li> <li>Added a GeoPackage file with coordinates for 11853 rain gauges.</li> <li>Updated precipitation time series (catchment averages) up to October 2024 (where available). Now derived from: CHIRPS v2.0; CPC; ERA5-Land; MSWEP v2.8; and BR-DWGD v3.2.3 (when at least 95% of the catchment area lies within Brazil – 864 catchments).</li> <li>Updated actual evapotranspiration time series (catchment averages) up to October 2024 (where available). Now derived from: GLEAM v4.2a; ERA5-Land; and MGB-SA.</li> <li>Updated potential evapotranspiration time series (catchment averages) up to October 2024 (where available). Now derived from GLEAM v4.2a and ERA5-Land.</li> <li>Added reference evapotranspiration time series (catchment averages). Derived from BR-DWGD v3.2.3 (when at least 95% of the catchment area lies within Brazil).</li> <li>Updated daily maximum, mean, and minimum temperature time series (catchment averages) up to October 2024 (where available). Now derived from: CPC; ERA5-Land; and BR-DWGD v3.2.3 (when at least 95% of the catchment area lies within Brazil).</li> <li>Added daily soil moisture time series (catchment averages) up to December 2024 (where available). Computed from GLEAM v4.2a and ERA5-Land.</li> <li>Improved meteorological data processing. Catchment averages now account for pixel fraction coverage.</li> <li>Reformatted meteorological time series files. Files now includes data from different products, with columns renamed for clarity.</li> <li>Hydrological and climatic indices were not updated, despite the new streamflow and meteorological data.</li> </ul> <p><strong>Minor changes</strong></p> <ul> <li>Updated stream gauge coordinates based on ANA's website on 27 February 2025. Coordinates were updated for 73 gauges in the 897 selected catchments and for 298 gauges across all catchments.</li> <li>Streamflow time series now include quality flag values from 0 to 7 (see the readme.txt file), previously from 0 to 4 in CAMELS-BR version 1.1. Flags from 5 to 7 may be present only in the last few years of data.</li> <li>Streamflow time series files for the 897 selected gauges now include values in both millimeters per day and cubic meters per second.</li> <li>Removed streamflow time series with fewer than 180 days of measurement.</li> <li>Converted gauge and catchment spatial data from Shapefile (.shp) to GeoPackage (.gpkg).</li> <li>Catchment areas computed by GSIM (in files "camels_br_location.txt" and "location_gauges_streamflow.gpkg") flagged as "caution" for quality were set to "nan" due to low reliability.</li> <li>Updated catchment areas computed by ANA (in files "camels_br_location.txt" and "location_gauges_streamflow.gpkg") to reflect the newest ANA's data from 27 February 2025. Streamflow values in millimeters per day remain unchanged because unit conversions rely on GSIM areas.</li> <li>Set catchment areas with zero squared kilometers, as computed by ANA, to "nan".</li> <li>Removed CPC daily mean temperature time series (catchment averages) because they were a simple average of minimum and maximum temperatures. For daily mean temperatures, refer to ERA5-Land data (now included) as they are computed from hourly data.</li> </ul> <p> </p>
Database for the Geospatial Synthesis of Biogeochemical Attributions of Porphyrins to Oil Pollution in Marine Sediments of the Gulf of México
<p>This dataset is associated with the journal article "Geospatial Synthesis of Biogeochemical Attributions of Porphyrins to Oil Pollution in Marine Sediments of the Gulf of México" by Muñoz-Arriola and Macias-Zamora (2022). The porphyrin and biogeochemical data were obtained by and analyzed at the Universidad Autónoma de Baja California's Instituto de Investigaciones Oceanológicas. The samples were collected to identify the effects of natural and human-originated oil spills in the Campeche Sound, and these efforts are part of the oceanographic campaign Xaman-Ek.</p> <p> </p> <p>Associated references are:</p> <p>Munoz-Arriola, F. and V. Macias-Zamora (2022) <em>Geospatial Synthesis of Biogeochemical Attributions of Porphyrins to Oil Pollution in Marine Sediments of the Gulf of México</em>. Geosciences. https://doi.org/10.3390/ geosciences12020077.</p> <p>Macias-Zamora, J. V., J. A. Villaescusa-Celaya; A. Munoz-Barbosa; and G. Gold-Bouchot (1999). Trace metals in sediment cores from the Campeche shelf, Gulf of Mexico. Environmental Pollution. Vol 104:69-77.</p> <p> </p>
Simple attributes predict the value of plants as hosts to fungal and arthropod communities
Fungal and arthropod consumers constitute the vast majority of global terrestrial biodiversity. Yet, the link from richness and composition of producer (plant) communities to the richness of consumer communities is poorly understood. Fungal and arthropod species richness could be a simple function of producer species richness at a site. Alternatively, it could be a complex function of chemical and structural properties of the producer species making up communities. We used databases on plant-fungus and plant-arthropod trophic links to derive the richness of consumer biota per associated plant species (coined link score). We assessed how well link scores could be predicted by simple attributes of plant species. Next, we used a multi-taxon inventory of 130 sites, representing all major habitat types in a country (Denmark), to investigate whether link scores summed over plant species in communities (coined link sum) could outperform simple plant species richness as predictor of fungal and arthropod richness at the sites. We found plant species' link scores for both fungi and arthropods to be positively related to plant size, regional occupancy, nativeness and ectomycorrhizal status. Link-based indices generally improved the prediction of richness of fungal and arthropod communities. For fungal communities, both observed link sum (from databases) and predicted link sum (from plant attributes) had high predictive power, while plant richness alone had none. For arthropod communities, predictive performance varied between functional groups. For both fungi and arthropods, richness predictions were further improved by considering abiotic habitat conditions. Our results underline the importance of plants as niche space for the megadiverse groups of arthropods and fungi. The plant-attribute approach holds promise for predicting local and regional consumer richness in areas of the world lacking detailed plant-consumer databases.
FLOAT: Factorized Learning of Object Attributes for Improved Multi-object Multi-part Scene Parsing
<p>Pascal-Part-201 is the most comprehensive and challenging version of the Pascal-Part dataset for Multi-object multi-part scene parsing. The dataset is the part of the publication "FLOAT: Factorized Learning of Object Attributes for Improved Multi-object Multi-part Scene Parsing" published in CVPR 2022.</p>
Analysis of pangolin metagenomic datasets reveals significant contamination, raising concerns for pangolin CoV host attribution
<p>Supplementary Figures, Information and Data to accompany:</p> <p>Analysis of pangolin metagenomic datasets reveals significant contamination, raising concerns for pangolin CoV host attribution<br> <em>Adrian Jones, Daoyu Zhang, Yuri Deigin and Steven C. Quay</em></p> <p>https://arxiv.org/abs/2108.08163</p> <p> </p> <p> </p>
Two Dynamic Attributed Networks: Enron & Jazz LastFM
<p><strong>Description. </strong>This repository contains two dynamic and attributed social networks extracted from the well-known Enron email dataset, and from the LastFM online music platform. We used both networks in the following papers:</p> <ol> <li>G. K. Orman, V. Labatut, M. Plantevit, and J.-F. Boulicaut, “A Method for Characterizing Communities in Dynamic Attributed Complex Networks,” in <em>IEEE/ACM International Conference on Advances in Social Network Analysis and Mining (ASONAM)</em>, 2014, pp. 481–484. ⟨<a href="https://hal.archives-ouvertes.fr/hal-01011913">hal-01011913</a>⟩ DOI: <a href="http://doi.org/10.1109/ASONAM.2014.6921629">10.1109/ASONAM.2014.6921629</a></li> <li>G. K. Orman, V. Labatut, M. Plantevit, and J.-F. Boulicaut, “Interpreting communities based on the evolution of a dynamic attributed network,” <em>Social Network Analysis and Mining</em>, vol. 5, p. 20, 2015. ⟨<a href="https://hal.archives-ouvertes.fr/hal-01163778">hal-01163778</a>⟩ DOI: <a href="http://doi.org/10.1007/s13278-015-0262-4">10.1007/s13278-015-0262-4</a></li> </ol> <p><strong>Citation. </strong>If you use these data, please cite the paper [1].</p> <p><br><code>@InProceedings{Orman2014,</code><br><code> author = {Orman, Günce Keziban and Labatut, Vincent and Plantevit, Marc and Boulicaut, Jean-François},</code><br><code> title = {A Method for Characterizing Communities in Dynamic Attributed Complex Networks},</code><br><code> booktitle = {IEEE/ACM International Conference on Advances in Social Network Analysis and Mining},</code><br><code> year = {2014},</code><br><code> pages = {481-484},</code><br><code> address = {Beijing, CN},</code><br><code> publisher = {IEEE Publishing},</code><br><code> doi = {10.1109/ASONAM.2014.6921629},</code><br><code>}</code></p> <p>----------------------------------------</p> <p><strong>Enron dataset. </strong>Enron is a well-known dataset in network science and text mining. It has been widely studied in academia. In network science, several different static networks appear in the literature. However, up to now, no dynamic network has been published, even though the email conversations have timestamps.</p> <p>We processed the original dataset to extract a dynamic network. There are 158 nodes representing Enron employees between 1997 and 2002. All the addresses in the <em>From</em> and <em>To</em> fields of each email are considered, resulting in a network of 28,802 nodes representing a distinct email addresses. A time span of one month is chosen for the time slices, generating 46 time slices. Two nodes are connected if the corresponding persons emailed each other during the given time slice. We did not make any distinction between sender and receiver, and thus produced an undirected dynamic network. </p> <p>----------------------------------------</p> <p><strong>LastFM dataset. </strong>LastFM is a music website that allows its members to register and listen to music online. It is also a social network platform, because its members can declare friendship relationships. In LastFM, members can join a predefined group related to their music tastes, and participate in music-related events such as concerts. Using the LastFM API, One can retrieve the information of the artist and track a user has listened to, with the exact timestamp. Moreover, it is also possible to get some information regarding the music-related events the users joined, including the exact timestamps.</p> <p>We extracted a network by focusing on the members of the <em>Jazz</em> group, which is supposed to include users appreciating this type of music. We took advantage of the LastFM API to retrieve the members of this group and the existing friendship connection between them. In the end, our network contains 1,702 nodes representing the <em>Jazz</em> users. The friendship relationships between them is static, though, in the sense that the LastFM API does not give access to any temporal information regarding their beginning or end. So, we decided to take advantage of some additional information to get a dynamic structure. We put a link between two nodes if two conditions were simultaneously true: 1) both considered users listened to at least one common artist for a specific period of time, and 2) they are friends on the LastFM platform. For the mentioned period of time, we decided to use 3 months with 1 month overlap, after having analyzed the dynamics of the platform. In other words, we extracted a dynamic network in which each time slice represents three months of LastFM usage for our 1,702 users of interest. There are one month overlap between two consecutive time slices.</p>
Appendices of the work "On the perceived relevance of critical internal quality attributes when evolving software features"
<p><strong>Context:</strong> Several refactorings performed while evolving software features aim to improve internal quality attributes like cohesion and complexity. Studies shows that non-assisted refactorings might worsen, not improve, internal attributes. Current knowledge is scarce on how developers perceive the relevance of critical internal attributes while evolving features. Internal attributes are critical if their measurement assumes anomalous values. <strong>Objective:</strong> This qualitative study aims at revealing the developer's perception on the relevance of critical internal attributes when evolving features. We target six class-level critical attributes: low cohesion, high complexity, high coupling, large hierarchy depth, large hierarchy breadth, and large size. <strong>Method:</strong> We performed two industry case studies based on online focus group sessions. We asked developers to discuss how much (and why) critical attributes are relevant for adding or enhancing features. We assessed the relevance of critical attributes individually and relatively, reasons behind the relevance of each critical attribute, and interrelations of critical attributes. <strong>Results:</strong> Low cohesion and high complexity were perceived as very relevant because they often make evolving features hard while tracking failures and adding features. The other critical attributes were perceived as less relevant when reusing code or adopting design patterns, for instance. Examples of interrelations include large size leads to low cohesion and high complexity leads to high coupling. <strong>Conclusions:</strong> Our findings could be combined with previous results on how refactorings affect quality attributes to assist developers in applying refactorings that may have a practically relevant impact on critical attributes.</p>
MONACO: Modes of Narration and Attribution Corpus
<p><strong>MONACO: Modes of Narration and Attribution Corpus</strong></p> <p>This corpus is constructed by the project group Modes of Narration and Attribution (<a href="https://www.uni-goettingen.de/de/mona/626918.html">MONA</a>). We provide German literary texts annotated with three base phenomena: <strong>Generalising Interpretation</strong> (GI), <strong>Comment</strong>, and <strong>Non-fictional Speech</strong> (NfR), as well as <strong>Attribution</strong> on top of them.</p> <p><strong>DFG Schwerpunktprogramm SPP 2207 "Computational Literary Studies"</strong></p> <p>Online:</p> <ul> <li><a href="https://gepris.dfg.de/gepris/projekt/402743989">https://gepris.dfg.de/gepris/projekt/402743989</a></li> <li><a href="https://dfg-spp-cls.github.io/">https://dfg-spp-cls.github.io/</a></li> </ul> <p><strong>Teilprojekt: "Structuring Literature - Variants and Functions of Reflextive Passages in Narrative Fiction"</strong></p> <p>Online:</p> <ul> <li><a href="https://gepris.dfg.de/gepris/projekt/424264086">https://gepris.dfg.de/gepris/projekt/424264086</a></li> <li><a href="https://dfg-spp-cls.github.io/projects_en/2020/01/24/TP-Structuring_Literature/">https://dfg-spp-cls.github.io/projects_en/2020/01/24/TP-Structuring_Literature/</a></li> <li><a href="https://www.uni-goettingen.de/de/structuring+literature/626921.html">https://www.uni-goettingen.de/de/structuring+literature/626921.html</a></li> </ul>
Robust retrieval of forest canopy structural attributes using multi-platform airborne LiDAR
<p><strong>Data and R code to replicate the analyses presented in</strong>:<br>Zhang et al. (2024) Robust retrieval of forest canopy structural attributes using multi-platform airborne LiDAR. Remote Sensing in Ecology and Conservation, <a href="https://doi.org/10.1002/rse2.398">https://doi.org/10.1002/rse2.398</a></p> <p>If using these data and/or R code in your work please cite the original publication listed above, as well as this repository using the corresponding DOI.</p>
Predicting particle quality attributes of organic crystalline materials using Particle Informatics
<p>Dataset related to the publication: "Predicting particle quality attributes of organic crystalline materials using Particle Informatics" published in Powder Technology (<a title="Go to table of contents for this volume/issue" href="https://www.sciencedirect.com/journal/powder-technology/vol/443/suppl/C"><span>Volume 443</span></a>, 1 July 2024, 119927Volume 443, 1 July 2024, 119927). In this work, a novel quercetin solvate of dimethylformamide (QDMF) was studied. The crystal structure was solved using single crystal X-ray diffraction and analysed using synthon analysis and other particle informatics tools (<em>e.g.</em>, solvate analyser). The thermal behaviour and thermodynamic stability of QDMF were studied experimentally using Raman spectroscopy, ATR-FTIR spectroscopy, differential scanning calorimetry, and thermogravimetric analysis. A clear relationship between the two-step desolvation behaviour of QDMF and the type, strength, and directionality of the main bulk synthons characterizing the QDMF structure was observed. Additionally, the attachment energy model was used to predict the QDMF morphology, together with facet-specific topology and chemical nature of each of the dominant {001}, {110}, and {200} facets. The {200} facet was found to be significantly rougher than the other two; whereas, the {110} was characterized by a higher percentage of exposed DMF molecules compared to the other two facets. Specific scanning electron microscopy and contact angle measurements were used to experimentally detect differences among the three facets and validate the modelling results.</p>
Data to accompany "Automatic text clustering for audio attribute elicitation experiment responses", AES 143rd Convention, New York, NY, USA, 2017
<p>This work was supported by the EPSRC Programme Grant S3A: Future Spatial Audio for an Immersive Listener Experience at Home (EP/L000539/1) and the BBC as part of the BBC Audio Research Partnership. Details about the data underlying this work, along with the terms for data access, are available from http://dx.doi.org/10.15126/surreydata.00841589.</p> <p>If you use the data, please cite the following paper:</p> <p>J. Francombe, T. Brookes, and R. Mason, “Automatic text clustering for audio attribute elicitation experiment responses”, AES 143rd Convention, New York, NY, USA, 2017</p>
Fig. 22 in Taxonomic and stratigraphic update of the material historically attributed to Megalosaurus from Portugal
Fig. 22. Small dorsal vertebra of a juvenile indeterminate avetheropod (MG 4822) from the Kimmeridgian–Tithonian Sobral Formation, Paimogo (Lourinhã region, Portugal) in?right lateral (A1),?left lateral (A2),?posterior (A3), dorsal (?anterior to the right, A4), ventral (?anterior to the left, A5), and?anterior (A6) views.
Fig. 24 in Taxonomic and stratigraphic update of the material historically attributed to Megalosaurus from Portugal
Fig. 24. Sacral and caudal vertebrae attributed to indeterminate allosauroid theropods from Kimmeridgian Alcobaça Formation, Ourém region (Portugal). A. Sacral vertebra, MG 4824a. B. Posterior caudal vertebra, MG 4831a. Left lateral (A1, B2), right lateral (A2, B1), anterior (A3, B6), dorsal (anterior to the left, A4, B4), ventral (anterior to the right, A5, B5), and posterior (A6, B3) views.
Fig. 23 in Taxonomic and stratigraphic update of the material historically attributed to Megalosaurus from Portugal
Fig. 23. Tooth crown fragment (MG 15) attributed to an indeterminate allosauroid theropod from the Lower Cretaceous of Cabo Espichel (Setúbal region, Portugal), in labial or lingual (A1, A2), distal (A3), and mesial (A4) views, and basal cross-section (A5).
Fig. 21. Simplified strict consensus trees resulting from a in Taxonomic and stratigraphic update of the material historically attributed to Megalosaurus from Portugal
Fig. 21. Simplified strict consensus trees resulting from a cladistic analysis performed on a dentition-based data matrix and forcing the constrains defined by Hendrickx et al. (2020a) pruning a priori all morphotypes from the Lusitanian Basin but Morphotype 2 (A), simplified strict consensus tree resulting from the analysis pruning a priori all morphotypes from the Lusitanian Basin but MG 15 (B), and simplified strict consensus tree resulting from the analysis pruning a priori all morphotypes from the Lusitanian Basin but Morphotype 3 (C).
Fig. 19 in Taxonomic and stratigraphic update of the material historically attributed to Megalosaurus from Portugal
Fig. 19. Tooth crown (MNHN/UL.EPt.023) attributed to an indeterminate megalosauroid theropod from the upper Kimmeridgian–lowermost Tithonian Praia da Amoreira-Porto Novo Formation, Porto Dinheiro (Lourinhã region, Portugal) in labial (A3), distal (A4), mesial (A5), and lingual (A6) views; detail of the mesial and distal denticles in the apical part of the crown (A1), detail of the enamel ornamentation (A2), detail of the distal denticles in the central sector of the carina (A7).
Fig. 20 in Taxonomic and stratigraphic update of the material historically attributed to Megalosaurus from Portugal
Fig. 20. Tooth crowns attributed to the megalosauroid theropod cf. Torvosaurus gurneyi Hendrickx & Mateus, 2014 from different Kimmeridgian to Tithonian localities of Portugal. A. MNHN/UL.EPt.8628, Porto Dinheiro (Lourinhã region). B. MG 4813, Montoito (Lourinhã region). C. MG 4818, Praia de S. Bernardino (Peniche region). Labial (A1, B1, C1), lingual (A2, B2, C2), distal (A5, B4, C4), and mesial views (A6, B3, C3); detail of the distal denticles in the central sector of the carina (A3); basal cross-section (A4, B5, C6, C7); cross-section at mid-height of the crown (C5).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.