Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,221
datasets available to search
ShareScore release 0.9.0
Dataset results
1,221 results for “Aggregators”
Monthly aggregated Water Vapor MODIS MCD19A2 (1 km): Monthly time-series (2009-2011)
<p>This data is part of the <em>Monthly aggregated Water Vapor MODIS MCD19A2 (1 km)</em> dataset. Check the related identifiers section on the Zenodo side panel to access other parts of the dataset.</p><p><strong>General Description</strong></p><p>The monthly aggregated water vapor dataset is derived from <a href="https://ladsweb.modaps.eosdis.nasa.gov/missions-and-measurements/products/MCD19A2"><abbr title="MCD19A2 MODIS/Terra+Aqua daily product">MCD19A2 v061</abbr></a>. The Water Vapor data measures the column above ground retrieved from MODIS near-IR bands at 0.94μm. The dataset time spans from 2000 to 2022 and provides data that covers the entire globe. The dataset can be used in many applications like water cycle modeling, vegetation mapping, and soil mapping. This dataset includes:</p><ul><li><strong>Monthly time-series:</strong><br>Derived from <em>MCD19A2 v061</em>, this data provides a monthly aggregated mean and standard deviation of daily water vapor time-series data from 2000 to 2022. Only positive non-cloudy pixels were considered valid observations to derive the mean and the standard deviation. The remaining no-data values were filled using the <abbr title="Moving Window Median">TMWM</abbr> algorithm. This dataset also includes smoothed mean and standard deviation values using the Whittaker method. The quality assessment layers and the number of valid observations for each month can provide an indication of the reliability of the monthly mean and standard deviation values.</li><li><strong>Yearly time-series:</strong><br>Derived from <em>monthly time-series</em>, this data provides a yearly time-series aggregated statistics of the monthly time-series data.</li><li><strong>Long-term data (2000-2022):</strong><br>Derived from <em>monthly time-series</em>, this data provides long-term aggregated statistics for the whole series of monthly observations.</li></ul><p><strong>Data Details</strong></p><ul><li><strong>Time period:</strong> 2009–2011</li><li><strong>Type of data:</strong> Water vapor column above the ground (0.001cm)</li><li><strong>How the data was collected or derived:</strong> Derived from MCD19A2 v061 using <a href="https://earthengine.google.com">Google Earth Engine</a>. Cloudy pixels were removed and only positive values of water vapor were considered to compute the statistics. The time-series gap-filling and time-series smoothing were computed using the <a href="https://github.com/scikit-map/scikit-map">Scikit-map</a> Python package.</li><li><strong>Statistical methods used:</strong> Four statistics were derived: mean, standard deviation, smoothed mean, smoothed standard deviation.</li><li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Antarctica.</li><li><strong>Coordinate reference system:</strong> EPSG:4326</li><li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (-180.00000, -62.00081, 179.99994, 87.37000)</li><li><strong>Spatial resolution:</strong> 1/120 d.d. = 0.008333333 (1km)</li><li><strong>Image size:</strong> 43,200 x 17,924</li><li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li></ul><p><strong>Support</strong></p><p>If you discover a bug, artifact, or inconsistency, or if you have a question please use some of the following channels:</p><ul><li>Technical issues and questions about the code: <a href="https://gitlab.com/openlandmap/global-layers/-/issues">GitLab Issues</a></li><li>General questions and comments: <a href="https://disqus.com/home/forums/landgis">LandGIS Forum</a></li></ul><p><strong>Name convention</strong></p><p>To ensure consistency and ease of use across and within the projects, we follow the standard Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describes important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. The fields are:</p><ol><li>generic variable name: wv = Water vapor</li><li>variable procedure combination: mcd19a2v061.seasconv = MCD19A2 v061 with gap-filling algorithm</li><li>Position in the probability distribution / variable type: m = mean | sd = standard deviation | n = number of observations | qa = quality assessment</li><li>Spatial support: 1km</li><li>Depth reference: s = surface</li><li>Time reference begin time: 20090101 = 2009-01-01</li><li>Time reference end time: 20111231 = 2011-12-31</li><li>Bounding box: go = global (without Antarctica)</li><li>EPSG code: epsg.4326 = EPSG:4326</li><li>Version code: v20230619 = 2023-06-19 (creation date)</li></ol>
Monthly aggregated Water Vapor MODIS MCD19A2 (1 km): Monthly time-series (2006-2008)
<p>This data is part of the <em>Monthly aggregated Water Vapor MODIS MCD19A2 (1 km)</em> dataset. Check the related identifiers section on the Zenodo side panel to access other parts of the dataset.</p><p><strong>General Description</strong></p><p>The monthly aggregated water vapor dataset is derived from <a href="https://ladsweb.modaps.eosdis.nasa.gov/missions-and-measurements/products/MCD19A2"><abbr title="MCD19A2 MODIS/Terra+Aqua daily product">MCD19A2 v061</abbr></a>. The Water Vapor data measures the column above ground retrieved from MODIS near-IR bands at 0.94μm. The dataset time spans from 2000 to 2022 and provides data that covers the entire globe. The dataset can be used in many applications like water cycle modeling, vegetation mapping, and soil mapping. This dataset includes:</p><ul><li><strong>Monthly time-series:</strong><br>Derived from <em>MCD19A2 v061</em>, this data provides a monthly aggregated mean and standard deviation of daily water vapor time-series data from 2000 to 2022. Only positive non-cloudy pixels were considered valid observations to derive the mean and the standard deviation. The remaining no-data values were filled using the <abbr title="Moving Window Median">TMWM</abbr> algorithm. This dataset also includes smoothed mean and standard deviation values using the Whittaker method. The quality assessment layers and the number of valid observations for each month can provide an indication of the reliability of the monthly mean and standard deviation values.</li><li><strong>Yearly time-series:</strong><br>Derived from <em>monthly time-series</em>, this data provides a yearly time-series aggregated statistics of the monthly time-series data.</li><li><strong>Long-term data (2000-2022):</strong><br>Derived from <em>monthly time-series</em>, this data provides long-term aggregated statistics for the whole series of monthly observations.</li></ul><p><strong>Data Details</strong></p><ul><li><strong>Time period:</strong> 2006–2008</li><li><strong>Type of data:</strong> Water vapor column above the ground (0.001cm)</li><li><strong>How the data was collected or derived:</strong> Derived from MCD19A2 v061 using <a href="https://earthengine.google.com">Google Earth Engine</a>. Cloudy pixels were removed and only positive values of water vapor were considered to compute the statistics. The time-series gap-filling and time-series smoothing were computed using the <a href="https://github.com/scikit-map/scikit-map">Scikit-map</a> Python package.</li><li><strong>Statistical methods used:</strong> Four statistics were derived: mean, standard deviation, smoothed mean, smoothed standard deviation.</li><li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Antarctica.</li><li><strong>Coordinate reference system:</strong> EPSG:4326</li><li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (-180.00000, -62.00081, 179.99994, 87.37000)</li><li><strong>Spatial resolution:</strong> 1/120 d.d. = 0.008333333 (1km)</li><li><strong>Image size:</strong> 43,200 x 17,924</li><li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li></ul><p><strong>Support</strong></p><p>If you discover a bug, artifact, or inconsistency, or if you have a question please use some of the following channels:</p><ul><li>Technical issues and questions about the code: <a href="https://gitlab.com/openlandmap/global-layers/-/issues">GitLab Issues</a></li><li>General questions and comments: <a href="https://disqus.com/home/forums/landgis">LandGIS Forum</a></li></ul><p><strong>Name convention</strong></p><p>To ensure consistency and ease of use across and within the projects, we follow the standard Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describes important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. The fields are:</p><ol><li>generic variable name: wv = Water vapor</li><li>variable procedure combination: mcd19a2v061.seasconv = MCD19A2 v061 with gap-filling algorithm</li><li>Position in the probability distribution / variable type: m = mean | sd = standard deviation | n = number of observations | qa = quality assessment</li><li>Spatial support: 1km</li><li>Depth reference: s = surface</li><li>Time reference begin time: 20060101 = 2006-01-01</li><li>Time reference end time: 20081231 = 2008-12-31</li><li>Bounding box: go = global (without Antarctica)</li><li>EPSG code: epsg.4326 = EPSG:4326</li><li>Version code: v20230619 = 2023-06-19 (creation date)</li></ol>
NO2, O3, PM10 and PM2.5 concentrations - Daily geographical aggregates at NUTS3 level from CAMS European Air Quality Re-analyses.
<p>This dataset offers daily aggregated measurements of air pollutants – NO2, O3, PM10, and PM2.5 – across distinct NUTS3 regions in continetal Europe. The temporal coverage spans from January 1, 2013, to December 31, 2022, providing a comprehensive temporal context for analyzing long-term air quality dynamics.</p> <p>Each daily entry comprises key statistical descriptors, encompassing mean, maximum, minimum, and standard deviation values of pollutant concentrations specific to each NUTS3 area. Additionally, for O3, the dataset includes an eight-hour rolling mean daily maximum.</p> <p>Spatial reference is established via shapefiles (EPSG:4326) sourced from Eurostat's official repository (<a href="https://ec.europa.eu/eurostat/web/gisco/geodata/reference-data/administrative-units-statistical-units/nuts">https://ec.europa.eu/eurostat/web/gisco/geodata/reference-data/administrative-units-statistical-units/nuts</a>). These shapefiles link the air quality data to precise NUTS3 regions through unique identifiers.</p> <p>The concentration data spanning from 2018 to 2022 originate from the European Air Quality Reanalyses dataset of the Atmosphere Data Store (ADS), an initiative by the Copernicus Atmosphere Monitoring Service (CAMS). Accessible via <a href="https://ads.atmosphere.copernicus.eu/cdsapp#!/dataset/cams-europe-air-quality-reanalyses?tab=doc">https://ads.atmosphere.copernicus.eu/cdsapp#!/dataset/cams-europe-air-quality-reanalyses?tab=doc</a>, this dataset offers a robust foundation for assessing air quality. For the years 2013 to 2017, data were previously obtained from a former download platform for the same dataset. Important: in future all data will be migrated to the Atmosphere Data Store (ADS) platform.</p> <p>The native resolution of the CAMS data is 0.1° x 0.1° spatially and hourly temporally. To enhance spatial accuracy, the spatial resolution was virtually increased by a factor of 5 using bilinear interpolation, resulting in a refined grid. The daily mean concentrations were subsequently computed for this augmented grid.</p> <p>Aggregated statistics were derived for each NUTS3 polygon, employing all grid cells intersecting with the polygons. The computation was based on the proportion of cell area included within the respective polygons.</p> <p>This dataset constitutes a valuable resource for conducting ecologically designed epidemiological studies, as it facilitates the exploration of potential associations between air quality and health trends across broad geographical areas.</p>
With or without you: Gut microbiota does not predict aggregation behaviour in females of the European earwig
<p>Recent studies suggest that the gut microbiota could be one of the main driving forces behind the evolution of group living. However, these studies are mainly based on our knowledge of species in which non-social individuals are rare and abnormal, calling into question the adaptive value of the reported association between group-living and gut microbiota. In this study, we addressed this issue by testing this association in females of the European earwig, an insect showing frequent, naturaland wide inter-individual variation in the expression of group living. We video-tracked 320 field-sampled females to quantify their natural variation in aggregation and then tested whether the most and least gregarious females had different gut microbiota. We also compared the general activity, boldness, body size and body condition of these females and examined the association between each of these traits and the gut microbiota. Contrary to our predictions, we found no difference in gut microbiota between the most and least gregarious females, as well as no difference between these females in terms of general activity, boldness, body size and condition. We did show that the gut microbiota of females was overall linked to their body condition, even though it was also unrelated to the other measurements. Overall, these results demonstrate that a host's gut microbiota is not necessarily a major driver of aggregation behaviour in species with inter-individual variation in group living and call for future studies to investigate the determinants and role of gut microbiota in earwigs.</p>
Effects of ribbed mussel aggregation size on marsh invertebrate community structure and multiple eocsystem functions
Ribbed mussels (Geukensia demissa) were added in aggregations containing 0, 1, 3, 5, 10, 20, 40 or 80 mussels (N=3 replicates per aggregation size) in Spring of 2012 in a high marsh platform at the Airport Marsh on Sapelo Island, GA. In summer of 2013, we measured the response of invertebrate communities and six ecosystem functions. Specifically, we counted the number of Littoraria irrorata, Sesarma reticulatum burrows, Uca pugnax burrows (those > and <5mm in diameter were counted separately), and mud crab (Eurythium limosum and Panopeus herbstii) in 50cm x 50cm sampling frames. And, we measured aboveground cordgrass biomass, benthic algae biomass, invertebrate biomass, decomposition rate, infiltration rate, and soil accretion in the same size sampling frames in August of 2013.
Aggregate mesquite litter chemistry following soil-mixing and decomposition in a semi-arid grassland at the Jornada Basin LTER, 2010-2012
This dataset contains litter carbon content, nitrogen content, and associated chemistry data from a litter decomposition experiment at the Jornada Basin LTER in 2010 to 2012. To assess the role of soil-litter mixing (SLM) in aridland litter decomposition, litterbags were deployed in the Chihuahuan Desert and interrelationships between vegetation structure, SLM, and rates of decomposition were quantified. To assess the role of vegetation structure, litterbags were deployed in contrasting vegetation microsites, including grass, shrub, and bare ground microsites. This dataset contains litter chemistry data from the experiment including percent carbon, percent nitrogen, ash corrections, and the carbon to nitrogen ratio of litter in recovered bags. This study is complete.
Aggregate mesquite litter mass-loss following soil-mixing and decomposition in a semi-arid grassland at the Jornada Basin LTER, 2010-2012
This package contains litter mass loss data from a litter decomposition experiment at the Jornada Basin LTER. To assess the role of soil-litter mixing (SLM) in aridland litter decomposition, litterbags were deployed in the Chihuahuan Desert and interrelationships between vegetation structure, SLM, and rates of decomposition were quantified. To assess the role of vegetation structure, litterbags were deployed in contrasting vegetation microsites, including grass, shrub, and bare ground microsites. This dataset contains the mass-loss data (including ash-corrections) from the experiment. This study is complete.
SBC LTER: Santa Cruz Island: Aggregated mean abundance of black surfperch (Embiotoca jacksoni) and prey availability, 1994-2008
These data are the annual mean abundances of three age-classes of black surfperch and their prey at each of 11 sites at Santa Cruz Island, California. Reported values are a) the average number of adult black surfperch, young-of-year black surfperch, and one year-old black surfperch per 40m x 2m transect (for all transects), and b) total food availability (grams per 0.1 m squared) at each site. Food availability includes biomass of caprellid and gammarid amphipods and is calculated from the mean density within Gelidium spp. (for caprellids) and all other foliose or turfing algae (for gammarids). The data also include one year lags for each variable, i.e. the value of the variable at that site in the previous year.
Alpine ice sheet glacial cycle simulations aggregated variables
<p>These data contain time-integrated and otherwise time-reduced glacier model output variables.</p> <p><strong>Reference:</strong></p> <ul> <li>Seguinot, J., Ivy-Ochs, S., Jouvet, G., Huss, M., Funk, M., and Preusser, F.: Modelling last glacial cycle ice dynamics in the Alps, <em>The Cryosphere</em>, 12, 3265-3285, doi:<a href="https://doi.org/10.5194/tc-12-3265-2018">10.5194/tc-12-3265-2018</a>, 2018.</li> </ul> <p><strong>File names:</strong></p> <pre><code>alpcyc.{1km|2km}.{epic|grip|md01}.{cp|pp}.agg.nc</code></pre> <ul> <li>Horizontal resolution: <ul> <li><em>1km</em>: 1 km horizontal resolution</li> <li><em>2km</em>: 2 km horizontal resolution</li> </ul> </li> <li>Temperature forcing: <ul> <li><em>epic</em>: EPICA ice core temperature forcing</li> <li><em>grip</em>: GRIP ice core temperature forcing</li> <li><em>md01</em>: MD01-2444 core temperature forcing</li> </ul> </li> <li>Precipitation forcing: <ul> <li><em>cp</em>: constant precipitation</li> <li><em>pp</em>: palaeo-precipitation reduction</li> </ul> </li> </ul> <p><strong>Data format:</strong></p> <p>The data use compressed netCDF format. For quick inspection I recommend ncview. Conversion to GeoTIFF (and other GIS formats) can be achieved with e.g. GDAL::</p> <pre><code>gdal_translate NETCDF:filename.nc:variable filename.variable.tif</code></pre> <p>The list of variables (subdatasets) can be obtained from ncdump or gdalinfo. To convert all variables to separate files use:</p> <pre><code>gdalinfo $filename | grep NETCDF | cut -d '=' -f 2 | egrep -v '(lat|lon|time_bounds)' | while read sub do gdal_translate $sub ${filename%.nc}.${sub##*:}.tif done</code></pre> <p>Variable long names, units, PISM configuration parametres and additional information are contained within the netCDF metadata. Also see <a href="https://doi.org/10.5281/zenodo.1423175">continuous</a> variables.</p> <p><strong>Changelog:</strong></p> <ul> <li>Version 2: <ul> <li>Add age coordinate in kiloyears (ka) before present.</li> <li>Use ka units for covertime, deglacage and maxthkage.</li> </ul> </li> <li>Version 1: <ul> <li>Initial version</li> </ul> </li> </ul>
COVID-19 Mobility Data Aggregator
<p><strong>Description</strong></p> <p>This repository includes:<br> 1) Data scraper of Google, Apple and Waze Mobility data<br> 2) Preprocessed mobility reports in different formats<br> 3) Merged mobility reports in summary files</p> <p><strong>About data</strong></p> <p>About <a href="https://www.google.com/covid19/mobility/">Google COVID-19 Community Mobility Reports</a></p> <p>About <a href="https://www.apple.com/covid19/mobility">Apple COVID-19 Mobility Trends Reports</a></p> <p>About <a href="https://www.waze.com/covid19">Waze COVID-19 local driving trends</a></p> <p><strong>Description of data files</strong></p> <p><em><strong>Google reports (located in google_reports directory):</strong></em></p> <p>The raw report in ZIP format: Global_Mobility_Report.zip<br> Data for the worldwide (only 1st level of subregions): mobility_report_countries (CSV and Excel formats available)<br> Data for Brazil: mobility_report_brazil (CSV and Excel formats available)<br> Data for Europe: mobility_report_europe (CSV and Excel formats available)<br> Data for Asia + Africa: mobility_report_asia_africa (CSV and Excel formats available)<br> Data for North and South America + Oceania (Brazil and US excluded): mobility_report_america_oceania (CSV and Excel formats available)</p> <p><em><strong>Apple reports (located in apple_reports directory):</strong></em></p> <p>Raw report: applemobilitytrends.csv<br> Data for the worldwide: apple_mobility_report (Google Sheets, CSV and Excel formats available)<br> Data for the US: apple_mobility_report_US (CSV and Excel formats available)</p> <p><em><strong>Waze reports (located in waze_reports directory):</strong></em></p> <p>Raw CSV files: Waze_Country-Level_Data.csv, Waze_City-Level_Data.csv<br> Preprocessed report: waze_mobility (Google Sheets, CSV and Excel formats available)</p> <p><em><strong>Summary reports (located in summary_reports directory)</strong></em></p> <p>These are merged Apple and Google reports.</p> <p>Report by regions: summary_report_regions (CSV and Excel formats available)<br> Report by countries: summary_report_countries (Google Sheets, CSV and Excel formats available)<br> Report for the US: summary_report_US (CSV and Excel formats available)</p> <p><strong>License</strong></p> <p>See LICENSE.txt</p> <p><strong>Credits</strong></p> <p>If you use this dataset, please also cite the original data sources:</p> <p>1. Google LLC <em>"Google COVID-19 Community Mobility Reports"</em>. https://www.google.com/covid19/mobility/ Accessed: <date></p> <p>2. Apple Inc. "<em>Apple COVID-19 Mobility Trends Reports"</em>. https://www.apple.com/covid19/mobility Accessed: <date></p> <p>3. Waze Ltd "<em>Waze COVID-19 Impact Dashboard". </em>https://www.waze.com/covid19 Accessed: <date></p>
Supplementary material (aggregated data set): Egeler, G.-A. & Baur, P. (2020). Menüwahl in der Hochschulmensa: Fleisch oder Vegi? Ergebnisse eines 12-wöchigen Feldexperiments (NOVANIMAL Working Paper No. 5). ZHAW. https://doi.org/10.21256/zhaw-1405
<p><strong>Meal choice at two university canteens in a field experiment during 12 weeks: aggregated menu sales data</strong></p> <p>How do canteen visitors respond to a revised offer of meat-based and plant-based meals? Selected innovations were simultaneously implemented and tested in a trans­disciplinary field experiment in two university canteens over a 12-week period in the autumn semester 2017. Throughout this time, the meat dishes and ‘veg-meals’ (ovo-lacto-vegetarian and vegan meals) were randomly distributed among the three menu lines, the veg-meals were not marketed and advertised as such and the previous vegetarian menu line was abolished. Weeks where the usual number of meat dishes were on offer (the ‘base weeks’) alternated with weeks where the share of veg-meals was increased (the ‘intervention weeks’). <br> The field experiment did not have a negative impact on the number of meals sold or the turnover compared to the two previous years. Women choose meat dishes less often than men. This connection applies in the base weeks and intervention weeks, in all age groups, among both students and among staff. Remarkably, the share of (non-labelled) vegan dishes is comparable for women and men over all age groups, independent of university affiliation (student, staff). Authentic vegan dishes were particularly welcome. Veg-meals could also be sold on the more expensive menu line. There was a better correlation between meal choice, eating habits and attitudes (health, environment, animal welfare, social aspects) than expected. <br> One quarter of canteen visitors show ‘veg-oriented’ eating habits and three quarters thereof ‘meat-oriented’ eating habits. Only a minority of potential visitors eat regularly at the canteen, and those who do exhibit meat-oriented eating habits more often. We conclude, therefore, that the canteen’s usual menu offer is primarily aimed at visitors with meat-oriented eating habits at lunchtime. The most typical visitors to the canteen are male students who select meat dishes.<br> It has been shown, therefore, that the simultaneous changes in supply have worked. Veg-meals are preferred, particularly by women and those prone to flexitarian eating habits; however, also the canteen visitors with meat-oriented eating habits chose veg-meals during the intervention weeks. Catering in canteens has the great potential to expand the range of veg-meals at the expense of meat dishes, provided that the culinary quality is of a high enough standard and meals are not offered as vegetarian or vegan. The question arises as to whether canteens are not missing an economic opportunity if they only offer traditional meat dishes? Canteens are perfectly suited as real-world laboratories in which innovations for sustainable catering can be tried out. The field experiment in the two university canteens is a start; further experiments are needed.</p> <p><strong>The data set contains more than <em>26'000</em> aggregated menu sales. The analyses and results are summarized in the working paper No. 5 <a href="https://doi.org/10.21256/zhaw-1405">https://doi.org/10.21256/zhaw-1405</a></strong></p> <p>The corresponding scripts are: </p> <p>- <a href="http://doi.org/10.5281/zenodo.4034686">10.5281/zenodo.4034686</a></p> <p>- <a href="http://doi.org/10.5281/zenodo.4034698">10.5281/zenodo.4034698</a></p> <p>- <a href="http://doi.org/10.5281/zenodo.4244258">10.5281/zenodo.4244258</a> (newer Version)</p> <p>For more information visit the <a href="http://novanimal.ch">novanimal.ch</a> website.</p>
Open-data release of aggregated Australian school-level information. Edition 2016.1
<p>The file set is a freely downloadable aggregation of information about Australian schools. The individual files represent a series of tables which, when considered together, form a relational database. The records cover the years 2008-2014 and include information on approximately 9500 primary and secondary school main-campuses and around 500 subcampuses. The records all relate to school-level data; no data about individuals is included. All the information has previously been published and is publicly available but it has not previously been released as a documented, useful aggregation. The information includes:<br /> (a) the names of schools<br /> (b) staffing levels, including full-time and part-time teaching and non-teaching staff<br /> (c) student enrolments, including the number of boys and girls<br /> (d) school financial information, including Commonwealth government, state government, and private funding<br /> (e) test data, potentially for school years 3, 5, 7 and 9, relating to an Australian national testing programme know by the trademark 'NAPLAN'<br /> <br /> Documentation of this Edition 2016.1 is incomplete but the organization of the data should be readily understandable to most people. If you are a researcher, the simplest way to study the data is to make use of the SQLite3 database called 'school-data-2016-1.db'. If you are unsure how to use an SQLite database, ask a guru.<br /> <br /> The database was constructed directly from the other included files by running the following command at a command-line prompt:<br /> <em>sqlite3 school-data-2016-1.db < school-data-2016-1.sql</em><br /> Note that a few, non-consequential, errors will be reported if you run this command yourself. The reason for the errors is that the SQLite database is created by importing a series of '.csv' files. Each of the .csv files contains a header line with the names of the variable relevant to each column. The information is useful for many statistical packages but it is not what SQLite expects, so it complains about the header. Despite the complaint, the database will be created correctly.<br /> <br /> Briefly, the data are organized as follows.<br /> (a) The .csv files ('comma separated values') do not actually use a comma as the field delimiter. Instead, the vertical bar character '|' (ASCII Octal 174 Decimal 124 Hex 7C) is used. If you read the .csv files using Microsoft Excel, Open Office, or Libre Office, you will need to set the field-separator to be '|'. Check your software documentation to understand how to do this.<br /> (b) Each school-related record is indexed by an identifer called 'ageid'. The ageid uniquely identifies each school and consequently serves as the appropriate variable for JOIN-ing records in different data files. For example, the first school-related record after the header line in file 'students-headed-bar.csv' shows the ageid of the school as 40000. The relevant school name can be found by looking in the file 'ageidtoname-headed-bar.csv' to discover that the the ageid of 40000 corresponds to a school called 'Corpus Christi Catholic School'.<br /> (3) In addition to the variable 'ageid' each record is also identified by one or two 'year' variables. The most important purpose of a year identifier will be to indicate the year that is relevant to the record. For example, if one turn again to file 'students-headed-bar.csv', one sees that the first seven school-related records after the header line all relate to the school Corpus Christi Catholic School with ageid of 40000. The variable that identifies the important differences between these seven records is the variable 'studentyear'. 'studentyear' shows the year to which the student data refer. One can see, for example, that in 2008, there were a total of 410 students enrolled, of whom 185 were girls and 225 were boys (look at the variable names in the header line).<br /> (4) The variables relating to years are given different names in each of the different files ('studentsyear' in the file 'students-headed-bar.csv', 'financesummaryyear' in the file 'financesummary-headed-bar.csv'). Despite the different names, the year variables provide the second-level means for joining information acrosss files. For example, if you wanted to relate the enrolments at a school in each year to its financial state, you might wish to JOIN records using 'ageid' in the two files and, secondarily, matching 'studentsyear' with 'financialsummaryyear'.<br /> (5) The manipulation of the data is most readily done using the SQL language with the SQLite database but it can also be done in a variety of statistical packages.<br /> (6) It is our intention for Edition 2016-2 to create large 'flat' files suitable for use by non-researchers who want to view the data with spreadsheet software. The disadvantage of such 'flat' files is that they contain vast amounts of redundant information and might not display the data in the form that the user most wants it.<br /> (7) Geocoding of the schools is not available in this edition.<br /> (8) Some files, such as 'sector-headed-bar.csv' are not used in the creation of the database but are provided as a convenience for researchers who might wish to recode some of the data to remove redundancy.<br /> (9) A detailed example of a suitable SQLite query can be found in the file 'school-data-sqlite-example.sql'. The same query, used in the context of analyses done with the excellent, freely available R statistical package (http://www.r-project.org) can be seen in the file 'school-data-with-sqlite.R'.</p>
Growth of Merocyanine Dye J-Aggregate Nanosheets by Living Supramolecular Polymerization
<p>Data to report <a href="https://doi.org/10.1002/ange.202314667">https://doi.org/10.1002/ange.202314667</a>:</p><p>J-aggregates are highly desired dye aggregates but so far there has been no general concept how to accomplish the required slip-stacked packing arrangement for dipolar merocyanine (MC) dyes whose aggregation commonly affords one-dimensional aggregates composed of antiparallel, co-facially stacked MCs with H-type coupling. Herein we describe a strategy for MC J-aggregates based on our results for an amphiphilic MC dye bearing alkyl and oligo(ethylene glycol) side chains. In an aqueous solvent mixture, we observe the formation of two supramolecular polymorphs for this MC dye, a metastable off-pathway nanoparticle showing H-type coupling and a thermodynamically favored nanosheet showing J-type coupling. Detailed studies concerning the self-assembly mechanism by UV-Vis spectroscopy and the packing structure by atomic force microscopy and wide-angle X-ray scattering show how the packing arrangement of such amphiphilic MC dyes can afford slip-stacked two-dimensional nanosheets whose macrodipole is compensated by the formation of a bilayer structure. As an additional feature we demonstrate how the size of the nanosheets can be controlled by seeded living supramolecular polymerization.</p>
Collective Strong Coupling Modifies Aggregation and Solvation - Dataset
<p>Dataset to complement "Collective Strong Coupling Modifies Aggregation and Solvation" - includes output and cube files obtained using the <a href="https://etprogram.org/">eT program</a>, an open source electronic (and molecular-polaritonic) structure program.</p> <p>See the paper at <a title="DOI URL" href="https://doi.org/10.1021/acs.jpclett.3c03506">https://doi.org/10.1021/acs.jpclett.3c03506</a></p>
Polarized, color-selective and semi-transparent organic photodiode of aligned merocyanine H-aggregates
<p>Data to report <a href="https://doi.org/10.1039/D4TC00678J">https://doi.org/10.1039/D4TC00678J</a>:</p> <p><span><span>Highly anisotropic thin films of H-type coupled dipolar merocyanines </span></span><span><span>with large dichroic ratios of over 50 </span></span><span><span>were </span></span><span><span>deposited by solution shearing. These layers were incorporated into simultaneously color- and polarization-selective organic photodiodes. Using a transparent non-fullerene acceptor, polarization-sensitive planar-heterojunction devices with an average visible transmittance of 93% were obtained.</span></span></p> <p> </p>
Data from: Fluid flow and amyloid transport and aggregation in the Brain Interstitial Space
<p>This data accompanies the paper entitled <strong>Fluid flow and amyloid transport and aggregation in the Brain Interstitial Space.</strong></p> <p> </p> <p>The zip archive contains the results of Lattice Boltzmann Molecular Dynamics simulations of the systems investigated and presented in the manuscript. Computational Fluid Dynamics data are in VTK format. Molecular Dynamics trajectories are in XYZ format.</p>
Copernicus EMS fire activations delimitations (2012 - 2020) rasterised at 30m and aggregated per year and season
<p>This dataset was created as part of the <a href="https://opendatascience.eu/">Geo-harmonizer project</a>, with the scope of making open data easier to access. It contains all the fire activations (forest fire, wild fire, wildfire) mapped by the<a href="https://emergency.copernicus.eu/mapping/list-of-activations-rapid"> Copernicus Emergency Rapid Mapping Service</a> between 2012 and 2020. To obtain these GeoTIFFs, the vector data packages from CEMS were individually downloaded, rasterized and mosaicked per year and season, resampled at 30-m and reprojected to <a href="https://epsg.io/3035">EPSG 3035: ETRS89-extended / LAEA Europe</a>. If no CEMS fire activation was identified in a specific year and season, the raster was not created. The rasters are provided as COG files, type=16Int, nodata value is 255.</p> <p>To allow an easier and faster search through all 2012 - 2020 CEMS fire activations, we have prepared a point vector layer (geojson) containing one point for each fire activation area of interest with the following attributes attached: CEMS identification number <ems_id>, area of interest defined by CEMS <ems_aoi>, URL link to the CEMS activation <ems_link>, year of the event <year_start>, <year_end> , <season> and the name of the <geo_harmonizer_raster> where the 30m rasterised delimitations of the burned areas of the corresponding fire activation can be found. </p> <p>For any additional questions regarding the data please contact the author at codrina.ilie[at]terrasigna.com.</p> <p>The Copernicus Emergency Rapid Mapping Service data access policy is available <a href="https://emergency.copernicus.eu/mapping/sites/default/files/files/CopernicusEMS-Data_and_Dissemination_Policy.pdf">here</a>.</p>
Medieval manuscripts and their migrations: Using SPARQL to investigate the research potential of an aggregated Knowledge Graph
<p>This dataset contains the <strong>SPARQL queries</strong> presented and discussed in our article published in <em>Digital Medievalist</em> 2022 (as a PDF file), together with the <strong>results of those queries</strong> as CSV files. The query and step numbering follows that given in the article.</p> <p>The queries can be run against the SPARQL endpoint for the <strong>Mapping Manuscript Migrations</strong> project: <a href="https://ldf.fi/mmm/sparql">https://ldf.fi/mmm/sparql</a></p> <p>The full <strong>Mapping Manuscript Migrations dataset </strong>can also be downloaded from the Zenodo repository and installed in your own triple store: <a href="https://zenodo.org/record/4440464">https://zenodo.org/record/4440464</a></p> <p>When copying and pasting these SPARQL queries into a SPARQL client like <a href="https://yasgui.triply.cc/">YASGUI</a>, please check that the line numbering has been copied over correctly. Copying from a PDF file can sometimes break a single long line into multiple separate lines, which will cause a SPARQL validation error.</p> <p>The CSV files contain the results of the queries when run against the Mapping Manuscript Migrations SPARQL endpoint as of 17 December 2021. Please note that Query 2, Step 2, produces no results, so a CSV file has not been provided.</p> <p>The<strong> Mapping Manuscript Migrations portal </strong>can be found at <a href="https://mappingmanuscriptmigrations.org/en/">https://mappingmanuscriptmigrations.org/en/ </a></p> <p>SPARQL tutorials are included in the project's <strong>GitHub documentation</strong>: <a href="https://mapping-manuscript-migrations.github.io/">https://mapping-manuscript-migrations.github.io/</a></p>
Aggregated US Bank Dataset
<p>These two .csv files contain the US bank dataset for FETILDA, containing sections of 10-K reports submitted by US banks from 2006 to 2016. They are directly used by the Python scripts for training, validation, and testing. There are two files, one for Item 1A of the 10-K reports, and the other for Item 7/7A.</p>
Raw and aggregated data for the study introduced in the paper "The way we cite: common metadata used across disciplines for defining bibliographic references"
<p>These data have been gathered in the context of a study aiming to investigate citation practices for referencing different types of entities and, in particular, for understanding the most used metadata in bibliographic references. The data are stored in two documents in XLSX format:</p> <ul> <li>file "links-intext-pointers-and-cited-entity-types.xlsx" - it contains information about whether the in-text reference pointers of the various PDF articles of the corpus have specified hypertextual links from the in-text reference pointers to the denoted bibliographic reference, plus information about the types of all the entities cited by each article in the corpus;</li> <li>file "metadata-bibliographic-references.xlsm" - it contains information about the metadata used to identify the various descriptive elements of all the bibliographic references defined in the article of the corpus.</li> </ul> <p>The methodology used to gather all these data is described in:</p> <blockquote> <p>Santos, E. A. d., Peroni, S., Mucheroni, M. L.: Workflow for retrieving all the data of the analysis introduced in the article "Citing and referencing habits in Medicine and Social Sciences journals in 2019". (2020), <a href="https://doi.org/10.17504/protocols.io.bbifikbn">https://doi.org/10.17504/protocols.io.bbifikbn</a></p> </blockquote>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.