Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8,537
datasets available to search
ShareScore release 0.9.0
Dataset results
8,537 results for “Openness”
Alpha-2 Adrenoreceptor Antagonist Yohimbine Potentiates Consolidation of Conditioned Fear (Open Data and Open Materials)
<p><strong>Open Data and Open Materials of: Sperl, M. F. J., Panitz, C., Skoluda, N., Nater, U. M., Pizzagalli, D. A., Hermann, C., & Mueller, E. M. (2022). Alpha-2 adrenoreceptor antagonist yohimbine potentiates consolidation of conditioned fear. <em>International Journal of Neuropsychopharmacology</em>, 25(9), 759–773.</strong></p> <p><em>Background:</em> Hyperconsolidation of aversive associations and poor extinction learning have been hypothesized to be crucial in the acquisition of pathological fear. Previous animal and human research points to the potential role of the catecholaminergic system, particularly noradrenaline and dopamine, in acquiring emotional memories. Here, we investigated in a between-participants design with 3 groups whether the noradrenergic alpha-2 adrenoreceptor antagonist yohimbine and the dopaminergic D2-receptor antagonist sulpiride modulate long-term fear conditioning and extinction in humans.<br><em>Methods:</em> Fifty-five healthy male students were recruited. The final sample consisted of n = 51 participants who were explicitly aware of the contingencies between conditioned stimuli (CS) and unconditioned stimuli after fear acquisition. The participants were then randomly assigned to 1 of the 3 groups and received either yohimbine (10 mg, n = 17), sulpiride (200 mg, n = 16), or placebo (n = 18) between fear acquisition and extinction. Recall of conditioned (non-extinguished CS+ vs CS−) and extinguished fear (extinguished CS+ vs CS−) was assessed 1 day later, and a 64-channel electroencephalogram was recorded.<br><em>Results:</em> The yohimbine group showed increased salivary alpha-amylase activity, confirming a successful manipulation of central noradrenergic release. Elevated fear-conditioned bradycardia and larger differential amplitudes of the N170 and late positive potential components in the event-related brain potential indicated that yohimbine treatment (compared with a placebo and sulpiride) enhanced fear recall during day 2.<br><em>Conclusions:</em> These results suggest that yohimbine potentiates cardiac and central electrophysiological signatures of fear memory consolidation. They thereby elucidate the key role of noradrenaline in strengthening the consolidation of conditioned fear associations, which may be a key mechanism in the etiology of fear-related disorders.</p>
Open Surface Drifter Data - Pirita river
<p>This dataset contains the surface drifter tracks collected in Pirita River (Estonia) to test the open drifter presented in the publication https://doi.org/10.3390/s22249918</p>
Open Science Quest
<p>The Open Science Quest was an activity organised as part of a national Open Science event in Luxembourg at the university library and was displayed for two weeks from 12 to 23 November 2018. Its aim was for library users (mainly Bachelor and Master students) and the event’s attendees (early-career and senior researchers, librarians and research support staff) to explore and discover Open Science practices at their own pace.</p> <p><br> The activity was stand-alone and promoted independent learning – once set up no external help was needed apart from issuing the diploma and prize for completing the Quest. The aim was also to make the activity as informative and engaging as possible by requiring participants to use a mix of information gathering techniques – text, images and videos presented on the displays, websearch and online tools, (very simple) puzzle-solving. The Quest was created and displayed in such a way that allowed various types of individual learning goals – each display provided knowledge without requiring to do the Quest and the Quest itself could be completed by grasping a minimal of concepts, while allowing participants to get more in-depth knowledge of each subject if they wanted to.</p> <p>All resources and materials used to prepare and showcase the Quest can be found here.<strong> Please read the 'User Guide' and the README files</strong> included in each folder and sub-folder to get more detailed information. You may also be interested in the blogpost: <a href="https://www.openaire.eu/blogs/open-science-quest">https://www.openaire.eu/blogs/open-science-quest</a></p> <p>Share, reuse, adapt and organise your own Open Science Quest! #OpenScienceQuest</p>
Introductory Motus Prioritization Tool Data (Open first)
Addressing survival and movement of priority migratory avian species of concern along the Pacific Flyway is paramount for their conservation. Yet, the migratory life stage is understudied in many avian species. The Motus radiotelemetry receiver network is an established system for tracking survival and movement of avian species. This network is an international collaborative that successfully identifies stopover site duration, connected migratory routes, post-fledging dispersal and survival, and adult survival and fidelity on a landscape-scale; parameters that cannot be easily estimated using non-tagged birds. While the Motus network is highly connected in eastern North America, the western part of the continent is lagging in coverage and connectivity, limiting the ability to obtain sample sizes large enough to robustly model demographic parameters from tagged birds. Thus, the expansion of the Motus network is a high priority for Pacific Flyway State Agencies. To date, no method exists for determining priority locations for new Motus receiving stations. With collaborations from States and the Canadian Province of British Columbia, we used eBird citizen scientist data to prioritize strategic locations for new Motus receiving stations throughout the Pacific Flyway. We model priority species’ co-occupancy of varying abundance states (i.e., absent, present, abundant, abundant in multiple weeks) with spatially varying Landsat (red and near infrared), water, land cover types, and weather covariates while accounting for variable detection with temporally varying survey effort covariates. Using occupancy model predictions, we identify high-use areas of the Pacific Flyway for establishing new Motus receiving towers that have high probabilities of intercepting high presence and /or abundance of multiple species of interest in a series of predictive occupancy maps. This package contains required files to recreate the data analysis, print out maps based on predictions from the
Fungal litter mat cover in Cannopy Trimming Experiment (CTE) plots responses to canopy opening, hurricanes and drought
Fungi that bind leaf litter into mats and produce white-rot via degradation of lignin and other aromatic compounds influence forest nutrient cycling and soil fertility. Over three and a half years beginning in June 2014, 6 months before the second iteration of the Canopy Trimming Experiment (CTE), we measured quarterly the extent of white-rot litter mats formed by basidiomycete fungi in the Luquillo Mountains of Puerto Rico in response to disturbances – a simulated hurricane treatment executed by canopy trimming and debris addition in December 2014 (CTE0, a mid-year drought in 2015, and two hurricanes 10 days apart in September 2017. Percent fungal litter mat cover ranged from 0.4% after hurricanes Irma and Maria to a high of 53% in forest with undisturbed canopy prior to the 2017 hurricanes, with means mostly between 10 - 45% of fungal litter mat cover in undisturbed forest. Drought decreased litter mat cover in both treatments, except in one undisturbed plot dominated by a drought-resistant fungus, Marasmius crinis-equi. Percent fungal litter mat cover sharply declined after real hurricanes and the simulated hurricane treatment (CTE). We found that solar radiation had a significant treatment effect and was strongly negatively correlated with percent litter mat cover within each of the four climatic seasons. Solar radiation was also strongly negatively correlated with relative humidity, throughfall, rain and litter wetness. However, rainfall was negatively correlated with litter mat cover, possibly due to erosion or saturation during high rainfall events. Canopy opening reduced leaf litterfall rates but did not affect litter mat cover. The main negative effect on basidiomycete fungi that bind leaf litter into mats was lower litter moisture associated with increased solar radiation from canopy opening and high leaf fall during drought. Variation in drought tolerance among basidiomycete fungal litter mat formers provided some resilience to drought. \<para\> Support f
Seasonal sea ice indices including the timing of ice-edge advance and ice-edge retreat (in year day), the ice season duration (in days) and number of actual ice days (versus open water days) within the ice season, extracted for various PAL LTER sub-regions West of the Antarctic Peninsula and derived from passive microwave satellite data for 1979/80 to 2023/24 ice seasons.
Seasonal sea ice indices including the timing of ice-edge advance and ice-edge retreat (in year day), the ice season duration (in days) and number of actual ice days (versus open water days) within the ice season, extracted for various PAL LTER sub-regions West of the Antarctic Peninsula and derived from passive microwave satellite data for 1979/80 to 2023/24 ice seasons. The ice season duration is defined as the time elapsed between day of ice-edge advance and day of ice-edge retreat within a given sea ice year, which begins mid-February (mean minimum of summer sea ice extent for the Southern Ocean) and ends the following mid-February. See Stammerjohn et al (2008, JGR) for further details.
User-centered Usability Analysis of 41 Open Government Data Portals
<p>The data were collected during the user-centered analysis of usability of 41 open government data portals including EU27, applying a common methodology to them, considering aspects such as specification of open data set, feedback and requests, further broken down into 14 sub-criteria. Each aspect was assessed using a three-level Likert scale (fulfilled - 3, partially fulfilled - 2, and unfulfilled – 1), that belongs to the acceptability tasks. This dataset summarises a total of 1640 protocols obtained during the analysis of the selected portals carried out by 40 participants, who were selected on a voluntary basis. This is complemented with 4 summaries of these protocols, which include calculated average scores by category, aspect and country. These data allow comparative analysis of the national open data portals, help to find the key challenges that can negatively impact users’ experience, and identifies portals that can be considered as an example for the less successful open data portals.</p>
Open Database of Spatial Room Impulse Responses at Detmold University of Music
<p>This repository contains an open source database of Spatial Room Impulse Responses (SRIR) captured at three different performance spaces of the Detmold University of Music. It includes the following rooms: </p> <ul> <li>Detmold Konzerthaus (medium sized concert hall, ~600 seats).</li> <li>Brahmssaal (small music chamber room, ~100 seats).</li> <li>Detmold Sommertheater (theater, ~300 seats).</li> </ul> <p>The collection contains approximately 600 multichannel RIRs corresponding to several source and receiver configurations. For each room we include measurement positions on stage and at the audience area captured with both an artificial head and an open microphone array compatible with the Spatial Decomposition Method (SDM).</p> <p>The Detmold Konzerthaus holds a large scale Wave Field Synthesis system and a Room Acoustic Enhancement System. SRIRs of an ensemble of focused sources on stage and with conditions of increased artificial reverberation are also included.</p> <p>If you use this dataset for your research, please cite our work:</p> <p>Amengual Gari, S. V.; Sahin, B.; Eddy, D; Kob, M.: <strong>"Open Database of Spatial Room Impulse Responses at Detmold University of Music"</strong>, <em>149th Convention of the Audio Engineering Society, </em>2020.</p> <p> </p> <p>The database is organized in 3 sets:</p> <p><strong>- Set A: </strong></p> <p>Source: Single Source measurements.</p> <p>Receiver: Open Array and Dummy Head.</p> <p>Rooms: BS, DST, KH</p> <p>Special configurations: Artificial reverberation, music stand on stage</p> <p><strong>- Set B: </strong></p> <p>Source: Loudspeaker and WFS orchestra</p> <p>Receiver: Open Array.</p> <p>Rooms: KH</p> <p><strong>- Set C:</strong></p> <p>Source: Loudspeaker orchestra</p> <p>Receiver: Dummy Head and Omni8 array</p> <p>Rooms: KH</p> <p> </p> <p>Further details on the measurement procedure and acoustical analysis of the RIRs can be found in the following publications:</p> <p><strong>Set A</strong></p> <p>Amengual Gari, S. V., Investigations on the Influence of Acoustics on Live Music Performance using Virtual Acoustic Methods, Ph.D. thesis, 2017.</p> <p>Amengual Garí, S. V.; Kob, M: "Investigating the impact of a music stand on stage using spatial impulse responses". 142nd Convention of the Audio Engineering Society, Berlin, May 2017.</p> <p><strong>Set B</strong></p> <p>Amengual Garí, S. V.; Pätynen, J.; Lokki, T.: "Physical and perceptual comparison of real and focused sound sources in a concert hall". Journal of the Audio Engineering Society, vol. 64 (12), pp. 1014-1025, December 2016.</p> <p><strong>Set C</strong></p> <p>Sahin, B., ““Investigation of the Detmold Concert Hall auditorium acoustics by comparing preference ratings and objective measurements.”, M.Sc. Thesis, 2017.</p> <p>Sahin, B., Amengual, S. V., and Kob, M., “Investigating listeners’ preferences in Detmold Concert Hall by comparing sensory evaluation and objective measurements,” Proc. 43th DAGA, Kiel, 2017.<br> </p>
Cayla et al., 2020, Wellcome Open Research - Underlying data
<p>Underlying dataset for the identifications of low-complexity regions (LCRs) in the proteome of Trypanosoma brucei.</p> <p>- Supplement File 2.xlsx (Position of every InterPro domain and LCR identified. All genes are provided with indication on chromosome localisation, presence of transmembrane domains, signal peptides and the localisation of the encoded proteins, either predicted using DeepLoc<sup>37</sup> or observed (Tryptag<sup>38</sup>)).</p> <p>- Supplement File 3.xlsx (List of genes and Molecular Function gene ontology (GO) enrichment analysis of proteins with predicted LCRs in the N-terminal, central part or C-terminal or the different possible combinations.)</p> <p>- Supplement File 4.xlsx (Property analysis of sequences of every InterPro and LCRs identified.)</p> <p>- Supplement File 5.xlsx (List of genes and Molecular GO enrichment analysis of proteins presenting a Low (<8) or High (>9) polarity index level.)</p> <p>- Supplement File 6.xlsx (List and position of PTMs present on InterPro domains and LCRs. The different datasets from which the PTMs have been extracted can be found in the Zhang2020, Benz2019, Cayla2019, Urbaniak2013, Ooi2020, Fisk2012, Lott2012 and Moretti2017<sup>12,13,15,17–19,27,28</sup> columns. The sequence properties of the domains/LCRs on which these PTMs are located are also indicated. The list of modifications identified in Ooi <em>et al.</em> 2020<sup>28</sup> present on LCRs are indicated in the second sheet.)</p> <p>- Supplement File 7.xlsx (List and position of LCRs, signal peptides and their overlap.)</p>
Learning Dynamics of Electrophysiological Brain Signals During Human Fear Conditioning (Open Data and Open Materials)
<p><strong>Open Data and Open Materials of: Sperl, M. F. J., Wroblewski, A., Mueller, M., Straube, B., & Mueller, E. M. (2021). Learning Dynamics of Electrophysiological Brain Signals During Human Fear Conditioning. <em>NeuroImage</em>, <em>226</em>, 117569.</strong></p> <p>Electrophysiological studies in rodents allow recording neural activity during threats with high temporal and spatial precision. Although fMRI has helped translate insights about the anatomy of underlying brain circuits to humans, the temporal dynamics of neural fear processes remain opaque and require EEG. To date, studies on electrophysiological brain signals in humans have helped to elucidate underlying perceptual and attentional processes, but have widely ignored how fear memory traces <em>evolve</em> over time. The low signal-to-noise ratio of EEG demands aggregations across high numbers of trials, which will wash out transient neurobiological processes that are induced by learning and prone to habituation. Here, our goal was to unravel the plasticity and temporal emergence of EEG responses during fear conditioning. To this end, we developed a new sequential-set fear conditioning paradigm that comprises three successive acquisition and extinction phases, each with a novel CS+/CS- set. Each set consists of two different neutral faces on different background colors which serve as CS+ and CS-, respectively. Thereby, this design provides sufficient trials for EEG analyses while tripling the relative amount of trials that tap into more transient neurobiological processes. Consistent with prior studies on ERP components, data-driven topographic EEG analyses revealed that ERP amplitudes were potentiated during time periods from 33–60 ms, 108–200 ms, and 468–820 ms indicating that fear conditioning prioritizes early sensory processing in the brain, but also facilitates neural responding during later attentional and evaluative stages. Importantly, averaging across the three CS+/CS- sets allowed us to probe the temporal evolution of neural processes: Responses during each of the three time windows gradually increased from early to late fear conditioning, while long-latency (460–730 ms) electrocortical responses diminished throughout fear extinction. Our novel paradigm demonstrates how short-, mid-, and long-latency EEG responses change during fear conditioning and extinction, findings that enlighten the learning curve of neurophysiological responses to threat in humans.</p>
A Lagrangian study of the contribution of the Canary coastal upwelling to the nitrogen budget of the open North Atlantic
<p>The attached datasets constitute the particle trajectory data produced in the experiment for Hailegeorgis et al..</p> <p>The "traj_upwell_1d_70m_1d-variables.nc" contains variables that describe different aspects of each upwelled particle (mostly regarding a particle's release or its initial or final conditions).</p> <p>The rest of the files with the format "traj_upwell_1d_70m_XXX-traj.nc" describe an attribute XXX (location or nutrient concentration) along the trajectory of upwelled particles tracked as part of the experiment.</p> <p>With ARIANE, particles are released and tracked in a ROMS simulation of the Canary coastal upwelling region. Out of the ~10M particles, the trajectories of the ~353K (~3.6%) that upwell are included. The variable "index_in_full_exp" in file "traj_upwell_1d_70m_1d-variables.nc" shows the index of each of these upwelling particles in the larger pool of released particles. For each upwelled particle, out of the 720-day trajectories starting from its release into the coast, the values from its upwelling step to its exit from the experiment are included, with the values outside this range being filled with a generic value (1.e20). An upwelled particle exits the experiment when it leaves the regional ROMS simulation altogether or when it leaves the coast and returns to the coast to re-upwell (more details in the paper).</p> <p>Be mindful of the different values of time. In "traj_upwell_1d_70m_1d-variables.nc", the variable "release_time" tells each particle's release time, in days since onset of the ROMS simulation, while variable "coast_exit_time" tells each particle's day of exiting coast, in days since its release. In each particle's trajectory (in traj_lon, traj_lat, etc), the first and last steps with valid values are the same as the days of its upwelling and its exit, respectively, since its release.</p> <p>The files contain the name and description of each variable. Along with the details in the publication, the descriptions here should be enough to fully interpret the information and replicate our analysis.</p>
The Open Aurignacian Project. Volume 2: Grotta di Castelcivita in southern Italy
<h2><strong>Overview</strong></h2> <p>The repository contains an extensive dataset (n = 538) comprising 3D meshes representing various classes of lithic artifacts such as cores, blades, bladelets, flakes, and retouched tools. These artifacts originate from the Protoaurignacian (<em>rsa'</em>) and Early Aurignacian (<em>gic</em>, <em>ars</em>) layers of Grotta di Castelcivita (40.49563600N, 015.20922177E) in southern Italy (Gambassini, 1997). The layers date back to approximately 41,000 to 39,800 years ago (Douka<em> et al.</em>, 2014). A new technological assessment of the <em>rsa’</em>–<em>ars </em>sequence has been conducted utilizing the models included in this repository (Falcucci et al., 2024). Grotta di Castelcivita holds significant importance for the study of Early Upper Paleolithic cultural dynamics due to its substantial archaeological content and the presence of the Campanian Ignimbrite geochronological marker, which seals the archaeological sequence of the site (Giaccio<em> et al.</em>, 2008).</p> <p>The 3D scanning of artifacts was performed using the first models of the Artec Space Spider and Artec Micro scanners from Artec Inc., Luxembourg. The scanning process adhered to best practices for lithic digitization (Göldner <em>et al.</em>, 2022), ensuring accurate capture of artifact details. 3D scanning with the Artec Spider follows the third version of the <em>Styrostone </em>protocol outlined by Göldner <em>et al.</em> (2023). For detailed information, please refer to Part 8 (Artec scanning of larger artifacts) of the protocol: <a href="dx.doi.org/10.17504/protocols.io.4r3l24d9qg1y/v3" rel="noopener">dx.doi.org/10.17504/protocols.io.4r3l24d9qg1y/v3</a>. 3D scanning with the Artec Micro follows the <em>Microstone </em>protocol by Falcucci (2022): <a href="dx.doi.org/10.17504/protocols.io.81wgb6781lpk/v1" rel="noopener">dx.doi.org/10.17504/protocols.io.81wgb6781lpk/v1</a>. The use of the Artec Micro was particularly valuable for digitizing extremely small lithics, such as retouched bladelets with lengths around 1 cm.</p> <p>The creation of this open-access repository is intended to encourage archaeologists to participate in collaborative initiatives, thereby contributing to the advancement of research in the field of lithic technology and facilitating broader access to the prehistoric record. This initiative aligns with the promotion of Open Science practices in archaeological sciences, as advocated by Marwick<em> et al.</em> (2017). This dataset is part of the <a href="https://www.armandofalcucci.com/project/open_aurignacian/">Open Aurignacian Project</a>.</p> <h2>Author contact</h2> <p>Dr. Armando Falcucci</p> <p>armando.falcucci@uni-tuebingen.de; falcucciarmando@gmail.com</p> <h2><strong>Description of the dataset</strong></h2> <p>This repository includes the following components:</p> <ol> <li><code>CTC_3D_Meshes.zip</code>:<strong> </strong>Compressed folder containing 3D models in PLY format for the lithic artifacts.</li> <li><code>Readme_Castelcivita_3D.txt</code>: This README file provides detailed information about the 3D models and metadata associated with this repository. It includes descriptions of the dataset's structure, the scanning and postprocessing protocols, and detailed metadata variables for the lithic artifacts, including scanning technology, resolution, and file formats. The file serves as a comprehensive guide to understanding the dataset and how to properly use and cite the data for research purposes.</li> <li><code>Castelcivita_3D_metadata.csv</code>:<strong> </strong>CSV file containing information, characteristics, and metadata of the lithic artifacts.</li> </ol> <p> </p> <p>The <code>Castelcivita_3D_metadata.csv</code> file includes the following metadata attributes:</p> <ul> <li><strong>ID:</strong> Each artifact has been assigned a unique identifier in the format "CTC" followed by a sequential number, allowing for cross-referencing with techno-typological data presented in related publications.</li> <li><strong>Site:</strong> The archaeological site where the lithic was excavated.</li> <li><strong>Layer: </strong>The stratigraphic origin of the lithic.</li> <li><strong>Raw_material:</strong> Categorization by the type of raw material (e.g., Chert, Radiolarite).</li> <li><strong>Class:</strong> Broad artifact sorting (e.g., Blank, Core, Core-Tool, Tool), following common classifications in lithic analysis. Cores are pieces of any size that lack a dorsal/ventral surface but have two or more blade/bladelet/flake scars. Tools are pieces of any size that exhibit retouch along the margins. Core-tools are pieces that have produced bladelets but can also be classified as tools (e.g., carinated endscrapers and burin cores) following a typological classification. Blanks are flaked pieces with both a dorsal and ventral face.</li> <li><strong>Blank: </strong>Classification of the blank into flake, blade, and bladelet categories. A blade is defined as a flaked blank whose length is at least twice its width, regardless of shape. Bladelets are defined as blades whose maximum width is less than 12 mm.</li> <li><strong>Technology: </strong>Technological classification of the blanks into categories such as initialization, maintenance, optimal, semi-cortical, and others, following Falcucci <em>et al. </em>(2020) and Falcucci <em>et al. </em>(2024).</li> <li><strong>Core_classification: </strong>Technological categories for cores and core-tools (e.g., Carinated, Multi-platform, Narrow-sided, Semicircumferential) following Falcucci & Peresani (2018).</li> <li><strong>Cortex: </strong>Percentage of cortex coverage (0%, 1–33%, 33–66%, 66–99%, 100%), estimated visually.</li> <li><strong>Preservation: </strong>Breakage classification for blanks (e.g., Complete, Distal, Mesial, Proximal, Undetermined). For cores and most core-tools, preservation is marked as "Other".</li> <li><strong>Volume:</strong> The volume of the artifact in cubic millimeters.</li> <li><strong>Surface: </strong>The surface area of the artifact in square millimeters.</li> <li><strong>Length: </strong>Maximum length in millimeters based on technological orientation, recorded with a digital caliper.</li> <li><strong>Width:</strong> Maximum width in millimeters based on technological orientation, recorded with a digital caliper.</li> <li><strong>Thickness:</strong> Maximum thickness in millimeters based on technological orientation, recorded with a digital caliper.</li> <li><strong>File_list: </strong>The list of files in the dataset that correspond to this specific ID.</li> <li><strong>Model_unit:</strong> The unit of measurement used for the 3D model. When viewing the artifact in a 3D viewer that supports real-world units, this is the unit you enter into your program to ensure proper scaling. Note that this is not related to the object's resolution; it's simply the value needed for accurate scaling when importing the model into your 3D program.</li> <li><strong>#_of_polygons:</strong> The number of polygons in the 3D model of the artifact.</li> <li><strong>Avg_edge_length(mm)/Resolution: </strong>The average distance between points on the model, serving as an effective measure of the model's resolution.</li> <li><strong>Resolution_score:</strong> A qualitative value assigned to each model, reflecting its resolution. Based on the entire set of scans from the Open Aurignacian Project, it classifies artifacts into four categories (i.e., ultra-detailed, detailed, moderate detail, low detail) based on their average edge length, providing an assessment of the model's resolution relative to others in the project.</li> <li><strong>Scanner: </strong>The specific model of the scanner used to capture the 3D data of the lithic artifact.</li> <li><strong>Scan_software:</strong> The version of the software used in conjunction with the scanner to capture the 3D data of the artifact.</li> <li><strong>Postprocessing_software:</strong> The version of the software used to execute postprocessing algorithms and generate the final 3D mesh of the artifact.</li> <li><strong>Coating: </strong>Yes/No entry speifying if coating was used for any scan.</li> </ul> <h2><strong>Research and Usage Notes</strong></h2> <p>Users are encouraged to consult the <a href="https://github.com/ArmandoFalcucci/Castelcivita-Aur-Techno">GitHub</a> and <a href="https://doi.org/10.5281/zenodo.10639552">Zenodo</a> repositories associated with the main publication on the Aurignacian sequence at Grotta di Castelcivita for further techno-typological data and analytical resources. This dataset is intended to foster open collaboration and reproducibility in lithic analysis, aligning with best practices in archaeological research.</p> <h2><strong>Licensing and Citation</strong></h2> <p>Please cite this repository and related publications when using this dataset in your research. Licensing details and citation formats are provided in the repository documentation.</p> <h2><strong>References</strong></h2> <p>Douka K., Higham T., Wood R.<em> et al.</em> (2014) On the chronology of the Uluzzian. <em>Journal of Human Evolution</em>, 68: 1-13. doi:10.1016/j.jhevol.2013.12.007</p> <p>Falcucci A. (2022) MicroStone: Exploring the capabilities of the Artec Micro in scanning stone tools. <em>protocols.io</em>. doi:<a href="https://dx.doi.org/10.17504/protocols.io.81wgb6781lpk/v1">https://dx.doi.org/10.17504/protocols.io.81wgb6781lpk/v1</a></p> <p>Falcucci A. & Peresani M. (2018) Protoaurignacian Core Reduction Procedures: Blade and Bladelet Technologies at Fumane Cave. Lithic Technology 43: 125-140. doi:10.1080/01977261.2018.1439681</p> <p>Falcucci A., Conard N.J. & Peresani M. (2020) Breaking through the Aquitaine frame: A re-evaluation on the significance of regional variants during the Aurignacian as seen from a key record in southern Europe. Journal of Anthropological Sciences, 98: 99-140. doi:https://doi.org/10.4436/JASS.98021</p> <p>Falcucci A., Arrighi S., Spagnolo V., Rossini M., Higgins O.A., Muttillo B., Martini I., Crezzini J., Boschin F., Ronchitelli A. & Moroni A. (2024) A pre-Campanian Ignimbrite techno-cultural shift in the Aurignacian sequence of Grotta di Castelcivita, southern Italy. Scientific Reports, 14: 12783. doi:10.1038/s41598-024-59896-6</p> <p>Gambassini P. (1997) <em>Il Paleolitico di Castelcivita: Culture e Ambiente</em>. Electa, Naples</p> <p>Giaccio B., Isaia R., Fedele F.G.<em> et al.</em> (2008) The Campanian Ignimbrite and Codola tephra layers: Two temporal/stratigraphic markers for the Early Upper Palaeolithic in southern Italy and eastern Europe. <em>Journal of Volcanology and Geothermal Research</em>, 177: 208-226. doi:<a href="https://doi.org/10.1016/j.jvolgeores.2007.10.007">https://doi.org/10.1016/j.jvolgeores.2007.10.007</a></p> <p>Göldner D., Karakostis F.A. & Falcucci A. (2022) Practical and technical aspects for the 3D scanning of lithic artefacts using micro-computed tomography techniques and laser light scanners for subsequent geometric morphometric analysis. Introducing the StyroStone protocol. PLoS One, 17: e0267163. doi:10.1371/journal.pone.0267163</p> <p>Göldner D., Karakostis F.A. & Falcucci A. (2023) <em>StyroStone</em>: A protocol for scanning and extracting three-dimensional meshes of stone artefacts using Micro-CT scanners V.3. protocols.io. <a href="dx.doi.org/10.17504/protocols.io.4r3l24d9qg1y/v3">dx.doi.org/10.17504/protocols.io.4r3l24d9qg1y/v3</a></p> <p>Marwick B., d’Alpoim Guedes J., Barton C.M.<em> et al.</em> (2017) Open science in archaeology. <em>SAA Archaeological Record</em>, 17: 8-14. doi:10.17605/OSF.IO/3D6XX</p>
The Open Aurignacian Project. Volume 1: Grotta di Fumane in northeastern Italy
<h2><strong>Overview</strong></h2> <p>This repository contains a large dataset (n = 948) of 3D meshes of different classes of lithic artifacts (blade and bladelet cores, blades, bladelets, flakes, and retouched tools) from the Aurignacian (A2, A1, D6, D3+D6, D3l, D3d base, D3d, D3b alpha, D3b, and D1c) and Gravettian (D1d, D1e, and D1f) units at Fumane Cave in northeastern Italy (see Bartolomei et al., 1992). The Upper Paleolithic sequence spans from about 41 to 33 ky cal BP (Higham et al., 2009) and several studies have focused on the lithic technology (Bertola et al., 2013; Broglio et al., 2005; Falcucci et al., 2017; Falcucci, 2018; Falcucci & Peresani, 2018; Falcucci et al., 2018; Falcucci et al., 2020). The importance of the site for understanding the earliest phases of the Upper Paleolithic in Mediterranean Europe is well acknowledged (Conard & Bolus, 2015). Recently, all complete blades and bladelets from the best-preserved area of the cave (i.e., the external area of the excavation) were 3D-scanned using a protocol that relies on both Micro-CT and Artec Spider scanners (Göldner et al., 2022). Our main goal was to conduct a geometric morphometric assessment of the laminar products and test hypotheses related to stone tool production and, more broadly, past human behavior (Falcucci et al., 2022; Falcucci & Peresani, 2022). Furthermore, all core types have been scanned throughout the years of research at the site with an Artec Spider (Falcucci<em> et al.</em>, 2024a; Lombao<em> et al.</em>, 2023).</p> <p>The 3D scanning of artifacts was performed using the first model of the Artec Space Spider and a micro-CT scanner. The scanning process adhered to best practices for lithic digitization (Göldner <em>et al.</em>, 2022), ensuring accurate capture of artifact details. 3D scanning and postprocessing for both micro-CT and Artec Spider follow the third version of the <em>Styrostone </em>protocol outlined by Göldner <em>et al.</em> (2023): <a href="dx.doi.org/10.17504/protocols.io.4r3l24d9qg1y/v3" rel="noopener">dx.doi.org/10.17504/protocols.io.4r3l24d9qg1y/v3</a>.</p> <p>The creation of this open-access repository is intended to encourage archaeologists to participate in collaborative initiatives, thereby contributing to the advancement of research in the field of lithic technology and facilitating broader access to the prehistoric record. This initiative aligns with the promotion of Open Science practices in archaeological sciences, as advocated by Marwick<em> et al.</em> (2017). This dataset is part of the <a href="https://www.armandofalcucci.com/project/open_aurignacian/">Open Aurignacian Project</a>.</p> <h2>Author contact</h2> <p>Dr. Armando Falcucci</p> <p>armando.falcucci@uni-tuebingen.de; falcucciarmando@gmail.com</p> <h2><strong>Description of the dataset</strong></h2> <p>This repository includes the following components:</p> <ol> <li><code>RF_3D_Meshes.zip</code>:<strong> </strong>Compressed folder containing 3D models in PLY format for the lithic artifacts.</li> <li><code>Readme_Fumane_3D.txt</code>: This README file provides detailed information about the 3D models and metadata associated with this repository. It includes descriptions of the dataset's structure, the scanning and postprocessing protocols, and detailed metadata variables for the lithic artifacts, including scanning technology, resolution, and file formats. The file serves as a comprehensive guide to understanding the dataset and how to properly use and cite the data for research purposes.</li> <li><code>Fumane_3D_metadata.csv</code>:<strong> </strong>CSV file containing information, characteristics, and metadata of the lithic artifacts.</li> </ol> <p>Each artifact has been assigned a unique identifier in the format "RF.b" (for blanks and tools) and "RF.c" (for cores) followed by a sequential number, allowing for cross-referencing with the techno-typological data presented in related publications.</p> <p>The <code>Fumane_3D_metadata.csv</code> file includes the following metadata attributes:</p> <ul> <li><strong>ID:</strong> Each artifact has been assigned a unique identifier in the format "RF.b" (for blanks and tools) and "RF.c" (for cores) followed by a sequential number, allowing for cross-referencing with the techno-typological data presented in related publications.</li> <li><strong>Site:</strong> The archaeological site where the lithic was excavated.</li> <li><strong>Layer: </strong>The stratigraphic origin of the lithic.</li> <li><strong>Raw_material:</strong> Categorization by the type of raw material (e.g., Maiolica, Scaglia Variegata, Scaglia Rossa).</li> <li><strong>Class:</strong> Broad artifact sorting (e.g., Blank, Core, Core-Tool, Tool), following common classifications in lithic analysis. Cores are pieces of any size that lack a dorsal/ventral surface but have two or more blade/bladelet/flake scars. Tools are pieces of any size that exhibit retouch along the margins. Core-tools are pieces that have produced bladelets but can also be classified as tools (e.g., carinated endscrapers and burin cores) following a typological classification. Blanks are flaked pieces with both a dorsal and ventral face.</li> <li><strong>Blank: </strong>Classification of the blank into flake, blade, and bladelet categories. A blade is defined as a flaked blank whose length is at least twice its width, regardless of shape. Bladelets are defined as blades whose maximum width is less than 12 mm.</li> <li><strong>Technology: </strong>Technological classification of the blanks into categories such as initialization, maintenance, optimal, semi-cortical, and others, following Falcucci <em>et al. </em>(2020) and Falcucci <em>et al. </em>(2024b).</li> <li><strong>Core_classification: </strong>Technological categories for cores and core-tools (e.g., Carinated, Multi-platform, Narrow-sided, Semicircumferential) following Falcucci & Peresani (2018).</li> <li><strong>Cortex: </strong>Percentage of cortex coverage (0%, 1–33%, 33–66%, 66–99%, 100%), estimated visually.</li> <li><strong>Preservation: </strong>Breakage classification for blanks (e.g., Complete, Distal, Mesial, Proximal, Undetermined). For cores and most core-tools, preservation is marked as "Other".</li> <li><strong>Volume:</strong> The volume of the artifact in cubic millimeters.</li> <li><strong>Surface: </strong>The surface area of the artifact in square millimeters.</li> <li><strong>Length: </strong>Maximum length in millimeters based on technological orientation, recorded with a digital caliper.</li> <li><strong>Width:</strong> Maximum width in millimeters based on technological orientation, recorded with a digital caliper.</li> <li><strong>Thickness:</strong> Maximum thickness in millimeters based on technological orientation, recorded with a digital caliper.</li> <li><strong>File_list: </strong>The list of files in the dataset that correspond to this specific ID.</li> <li><strong>Model_unit:</strong> The unit of measurement used for the 3D model. When viewing the artifact in a 3D viewer that supports real-world units, this is the unit you enter into your program to ensure proper scaling. Note that this is not related to the object's resolution; it's simply the value needed for accurate scaling when importing the model into your 3D program.</li> <li><strong>#_of_polygons:</strong> The number of polygons in the 3D model of the artifact.</li> <li><strong>Avg_edge_length(mm)/Resolution: </strong>The average distance between points on the model, serving as an effective measure of the model's resolution.</li> <li><strong>Resolution_score:</strong> A qualitative value assigned to each model, reflecting its resolution. Based on the entire set of scans from the Open Aurignacian Project, it classifies artifacts into four categories (i.e., ultra-detailed, detailed, moderate detail, low detail) based on their average edge length, providing an assessment of the model's resolution relative to others in the project.</li> <li><strong>Scanner: </strong>The specific model of the scanner used to capture the 3D data of the lithic artifact.</li> <li><strong>Scan_software:</strong> The version of the software used in conjunction with the scanner to capture the 3D data of the artifact.</li> <li><strong>Postprocessing_software:</strong> The version of the software used to execute postprocessing algorithms and generate the final 3D mesh of the artifact.</li> <li><strong>Coating: </strong>Yes/No entry speifying if coating was used for any scan.</li> </ul> <h2><strong>What's new in this release (Version 3.0.1)</strong></h2> <p>In this new version, we have reworked all 3D models of cores and core-tools to enhance their overall quality and improve analysis. This was accomplished using Artec Studio Professional software by adjusting the settings for Global Registration and, in particular, Sharp Fusion (i.e., using 0.1 instead of 0.3 in 3D Resolution, mm) . These changes mainly affect models with IDs starting with "RF.c". This change was applied only to the PLY files, while the WRL files were not included in this release. The WRL files can be downloaded from previous versions of this repository.</p> <h2><strong>Research and Usage Notes</strong></h2> <p>Users are encouraged to consult the <a href="https://github.com/ArmandoFalcucci/Refitting-The-Context">GitHub</a> and <a href="https://zenodo.org/doi/10.5281/zenodo.10965413">Zenodo</a> repositories associated with the main publication on the Aurignacian sequence at Grotta di Fumane for further techno-typological data and analytical resources. This dataset is intended to foster open collaboration and reproducibility in lithic analysis, aligning with best practices in archaeological research.</p> <h2><strong>Licensing and Citation</strong></h2> <p>Please ensure that this dataset is properly cited in any research or publication that utilizes it. Detailed licensing and citation information is provided within the dataset documentation.</p> <h2><strong>References</strong></h2> <p>Bartolomei G., Broglio A., Cassoli P. et al. (1992) La Grotte de Fumane. Un site aurignacien au pied des Alpes. Preistoria Alpina, 28: 131-179</p> <p>Bertola S., Broglio A., Cristiani E. et al. (2013) La diffusione del primo Aurignaziano a sud dell'arco alpino. Preistoria Alpina, 47: 17-30</p> <p>Broglio A., Bertola S., De Stefani M. et al. (2005) La production lamellaire et les armatures lamellaires de l’Aurignacien ancien de la grotte de Fumane (Monts Lessini, Vénétie). In F. Le Brun-Ricalens (ed.): Productions lamellaires attribuées à l’Aurignacien, pp. 415-436. MNHA, Luxembourg.</p> <p>Conard N.J. & Bolus M. (2015) Chronicling modern human’s arrival in Europe. Science. doi:10.1126/science.aab0234</p> <p>Falcucci A., Conard N.J. & Peresani M. (2017) A critical assessment of the Protoaurignacian lithic technology at Fumane Cave and its implications for the definition of the earliest Aurignacian. PLoS One, 12: e0189241. doi:10.1371/journal.pone.0189241</p> <p>Falcucci A. & Peresani M. (2018) Protoaurignacian Core Reduction Procedures: Blade and Bladelet Technologies at Fumane Cave. Lithic Technology 43: 125-140. doi:10.1080/01977261.2018.1439681</p> <p>Falcucci A. (2018) Towards a renewed definition of the Protoaurignacian. Mitteilungen der Gesellschaft für Urgeschichte, 27: 87-130</p> <p>Falcucci A., Peresani M., Roussel M. et al. (2018) What’s the point? Retouched bladelet variability in the Protoaurignacian. Results from Fumane, Isturitz, and Les Cottés. Archaeol. Anthropol. Sci., 10: 539-554. doi:10.1007/s12520-016-0365-5</p> <p>Falcucci A., Conard N.J. & Peresani M. (2020) Breaking through the Aquitaine frame: A re-evaluation on the significance of regional variants during the Aurignacian as seen from a key record in southern Europe. J. Anthropol. Sci., 98: 99-140. doi:10.4436/JASS.98021</p> <p>Falcucci A., Karakostis F.A., Göldner D. et al. (2022) Bringing shape into focus: Assessing differences between blades and bladelets and their technological significance in 3D form. Journal of Archaeological Science: Reports, 43: 103490. doi:https://doi.org/10.1016/j.jasrep.2022.103490</p> <p>Falcucci A. & Peresani M. (2022) The contribution of integrated 3D model analysis to Protoaurignacian stone tool design. PLoS One, 17: e0268539. doi:10.1371/journal.pone.0268539</p> <p>Falcucci A., Giusti D., Zangrossi F., De Lorenzi M., Ceregatti L. & Peresani M. (2024a) Refitting the Context: A Reconsideration of Cultural Change among Early Homo sapiens at Fumane Cave through Blade Break Connections, Spatial Taphonomy, and Lithic Technology. Journal of Paleolithic Archaeology, 8: 2. doi:10.1007/s41982-024-00203-0</p> <p>Falcucci A., Arrighi S., Spagnolo V., Rossini M., Higgins O.A., Muttillo B., Martini I., Crezzini J., Boschin F., Ronchitelli A. & Moroni A. (2024b) A pre-Campanian Ignimbrite techno-cultural shift in the Aurignacian sequence of Grotta di Castelcivita, southern Italy. Scientific Reports, 14: 12783. doi:10.1038/s41598-024-59896-6</p> <p>Göldner D., Karakostis F.A. & Falcucci A. (2022) Practical and technical aspects for the 3D scanning of lithic artefacts using micro-computed tomography techniques and laser light scanners for subsequent geometric morphometric analysis. Introducing the StyroStone protocol. PLoS One, 17: e0267163. doi:10.1371/journal.pone.0267163</p> <p>Göldner D., Karakostis F.A. & Falcucci A. (2023) <em>StyroStone</em>: A protocol for scanning and extracting three-dimensional meshes of stone artefacts using Micro-CT scanners V.3. protocols.io. <a href="dx.doi.org/10.17504/protocols.io.4r3l24d9qg1y/v3">dx.doi.org/10.17504/protocols.io.4r3l24d9qg1y/v3</a></p> <p>Lombao D., Falcucci A., Moos E. & Peresani M. (2023) Unravelling technological behaviors through core reduction intensity. The case of the early Protoaurignacian assemblage from Fumane Cave. Journal of Archaeological Science, 160: 105889. doi:https://doi.org/10.1016/j.jas.2023.105889</p>
Global Naturalized Alien Flora (GloNAF). Open access data to support research on understanding global plant invasions.
<p>This dataset is a snapshot of the Global Naturalized Alien Flora (GloNAF) database, version 2.02. GloNAF is a continuously updated, curated compilation of alien naturalized vascular plant inventories for geographic regions from around the world. The dataset has 16,429 unique taxa reported as naturalized or invasive and covers 1,343 regions (including 427 islands) from 336 data sources. For each region, the status (invasive, naturalized) is provided as listed in the original source. We provide the scientific names included with the original data source, and the matching accepted name or synonym of the taxon as given in the World Checklist of Vascular Plants (WCVP) Version 12. In addition, we provide an ESRI shapefile of polygons for each region. We also provide several variables that can be used to filter the data according to quality and completeness of alien taxon lists, which vary among the combinations of regions and data sources.</p> <p>The 'glonaf_flora2.csv' file lists the IDs ('taxon_wcvp_id') of all naturalized taxa contained in GloNAF and the regions they occur in. The 'glonaf_taxon_wcvp.csv' lists the original taxon names provided in the source data along with the corresponding accepted taxon name from the WCVP (version 12) for all alien taxa in GloNAF, regardless of their naturalization status. To link taxon names with naturalization records, join the 'id' column of the 'glonaf_taxon_wcvp.csv' file to the 'taxon_wcvp_id' column in 'glonaf_flora2.csv' . Additional information regarding the original source of the data ('glonaf_reference.csv'), specific attributes of the taxon lists ('glonaf_list.csv') and the region ('glonaf_region.csv') can also be joined similarly to 'glonaf_flora2.csv '. </p> <p> </p>
AN OPEN-SOURCE, THREE-DIMENSIONAL GROWTH MODEL OF THE MANDIBLE
<p>This repository contains all geometrical data and metadata belonging to the paper AN OPEN-SOURCE, THREE-DIMENSIONAL GROWTH MODEL OF THE MANDIBLE by the MAGIC Amsterdam research consortium. The following contents are uploaded:</p><p><strong>shapeVectors_original.csv</strong> | shape vectors of the original data<br><strong>shapeVectors_rescaled.csv</strong> | shape vectors of the rescaled data<br>678 x 62589 matrices where the rows are samples and the columns are shape vectors. The shape vectors are formatted<i> [x1, x2, x3, ..., y1, y2, y3, ..., z1, z2, z3, ...].</i></p><p><strong>PCA_coeff_original.csv</strong> | principal component coefficients of the original data<br><strong>PCA_coeff_rescaled.csv</strong> | principal component coefficients of the rescaled data<br>62589 x 677 matrices where each row of these matrices is a variable (x-, y-, or z-coordinate of a vertex) and each column is a principal component.</p><p><strong>PCA_score_original.csv</strong> | principal component scores of the original data<br><strong>PCA_score_rescaled.csv</strong> | principal component scores of the rescaled data<br>678 x 677 matrices where rows correspond to samples and columns correspond to principal components.</p><p><strong>PCA_latent_original.csv</strong> | principal component variances of the original data<br><strong>PCA_latent_rescaled.csv</strong> | principal component variances of the rescaled data<br>677 x 1 vectors where each element is an eigenvalue of a principal component.</p><p><strong>PCA_mu_original.csv</strong> | mean of the original data<br><strong>PCA_mu_rescaled.csv</strong> | mean of the rescaled data<br>1 x 62589 vectors that represent the average shape vector. All (centered) data can be reconstructed as follows: <i>shapeVectors = PCA_score * PCA_coeff' + PCA_mu.</i></p><p><strong>PCA_standardDeviations_original.csv</strong> | standard deviations of each sample for each principal component of the original data.<br><strong>PCA_standardDeviations_rescaled.csv</strong> | standard deviations of each sample for each principal component of the rescaled data.<br>677 x 678 matrices where the rows are principal components and the columns are samples. The standard deviations were calculated as follows: <i>PCA_standardDeviations = PCA_score' ./ sqrt(PCA_latent).</i></p><p><strong>metadata.csv</strong> | This matrix contains the age in years (first column) and biological sex (second column, 1 = male and 2 = female) for all samples (rows).</p><p><strong>connectivityList.csv</strong> | This matrix defines the mesh of the 3D model of the mandible. The vector in each row represents which vertices define a triangle. Indexing starts at 0, so for use in e.g. Matlab, add 1 to all elements.</p>
BASE-9 binarity and stellar masses from Gaia DR3, 2MASS, and Pan-STARRS data for six open clusters: NGC 2168, NGC 7789, NGC 6819, NGC 2682, NGC 188, NGC 6791
<h2>Data sets as described in "Goodbye to Chi-by-Eye: A Bayesian Analysis of Photometric Binaries in Six Open Clusters", Childs et al. 2023 <a href="https://ui.adsabs.harvard.edu/abs/2023arXiv230816282C/abstract">https://ui.adsabs.harvard.edu/abs/2023arXiv230816282C/abstract</a></h2>
Open-source traffic and CO2 emission dataset for commercial aviation
<p>This record is a global open-source passenger air traffic dataset primarily dedicated to the research community. <br>It gives a seating capacity available on each origin-destination route for a given year, 2019, and the associated aircraft and airline when this information is available. </p> <p>Context on the original work is given in the related articles (<a href="https://doi.org/10.59490/joas.2024.7365">https://doi.org/10.59490/joas.2024.7365,</a> <a href="https://doi.org/10.59490/joas.2023.7201">https://doi.org/10.59490/joas.2023.7201)</a> and on the associated GitHub page (<a href="https://github.com/AeroMAPS/AeroSCOPE/">https://github.com/AeroMAPS/AeroSCOPE/</a>).<br>A simple data exploration interface will be available at <a href="www.aeromaps.eu/aeroscope">www.aeromaps.eu/aeroscope.</a><br>The dataset was created by aggregating various available open-source databases with limited geographical coverage. It was then completed using a route database created by parsing Wikipedia and Wikidata, on which the traffic volume was estimated using a machine learning algorithm (XGBoost) trained using traffic and socio-economical data.<br> </p> <h4><br><strong>1- DISCLAIMER</strong></h4> <p><br>The dataset was gathered to allow highly aggregated analyses of the air traffic, at the continental or country levels. At the route level, the accuracy is limited as mentioned in the associated article and improper usage could lead to erroneous analyses. </p> <p>Although all sources used are open to everyone, the Eurocontrol database is only freely available to academic researchers. It is used in this dataset in a very aggregated way and under several levels of abstraction. As a result, it is not distributed in its original format as specified in the contract of use.</p> <p>As a general rule, we decline any responsibility for any use that is contrary to the terms and conditions of the various sources that are used. In case of commercial use of the database, please contact us in advance.</p> <h4><br><strong>2- DESCRIPTION</strong></h4> <p>Each data entry represents an (Origin-Destination-Operator-Aircraft type) tuple.</p> <p><em>Please </em>refer<em> to </em>the<em> support article for more details (see above).</em></p> <p>The dataset contains the following columns:</p> <ul> <li>"First column" : index</li> <li><strong>airline_iata : </strong>IATA code of the operator in nominal cases. An ICAO -> IATA code conversion was performed for some sources, and the ICAO code was kept if no match was found.</li> <li><strong>acft_icao : </strong>ICAO code of the aircraft type</li> <li><strong>acft_class : </strong>Aircraft class identifier, own classification. <ul> <li>WB: Wide Body</li> <li>NB: Narrow Body</li> <li>RJ: Regional Jet</li> <li>PJ: Private Jet</li> <li>TP: Turbo Propeller</li> <li>PP: Piston Propeller</li> <li>HE: Helicopter</li> <li>OTHER</li> </ul> </li> <li><strong>seymour_proxy: </strong>Aircraft code for Seymour Surrogate (https://doi.org/10.1016/j.trd.2020.102528), own classification to derive proxy aircraft when nominal aircraft type unavailable in the aircraft performance model.</li> <li><strong>source: </strong>Original data source for the record, before compilation and enrichment. <ul> <li>ANAC: Brasilian Civil Aviation Authorities</li> <li>AUS Stats: Australian Civil Aviation Authorities</li> <li>BTS: US Bureau of Transportation Statistics T100</li> <li>Estimation: Own model, estimation on Wikipedia-parsed route database</li> <li>Eurocontrol: Aggregation and enrichment of R&D database</li> <li>OpenSky</li> <li>World Bank</li> </ul> </li> <li><strong>seats: </strong>Number of seats available for the data entry, AFTER airport residual scaling</li> <li><strong>n_flights: </strong>Number of flights of the data entry, when available</li> <li><strong>iata_departure</strong>, <strong>iata_arrival : </strong>IATA code of the origin and destination airports. Some BTS inhouse identifiers could remain but it is marginal.</li> <li><strong>departure_lon</strong><em>, </em><strong>departure_lat</strong><em>, </em><strong>arrival_lon</strong><em>, </em><strong>arrival_lat : </strong>Origin and destination coordinates, could be NaN if the IATA identifier is erroneous</li> <li><strong>departure_country, arrival_country</strong>: Origin and destination country ISO2 code. <strong>WARNING: </strong>disable NA (Namibia) as default NaN at import</li> <li><strong>departure_continent, arrival_continent: </strong>Origin and destination continent code. <strong>WARNING: </strong>disable NA (North America) as default NaN at import</li> <li><strong>seats_no_est_scaling: </strong>Number of seats available for the data entry, BEFORE airport residual scaling</li> <li><strong>distance_km: </strong>Flight distance (km)</li> <li><strong>ask: </strong>Available Seat Kilometres</li> <li><strong>rpk: </strong>Revenue Passenger Kilometres (simple calculation from ASK using IATA average load factor)</li> <li><strong>fuel_burn_seymour: </strong>Fuel burn <em>per flight</em> (kg) when seymour proxy available</li> <li><strong>fuel_burn: </strong>Total fuel burn of the data entry (kg)</li> <li><strong>co2: </strong>Total CO2 emissions of the data entry (kg)</li> <li><strong>domestic: </strong>Domestic/international boolean (Domestic=1, International=0)</li> </ul> <p> </p> <h4><strong>3- Citation</strong></h4> <p>Please cite the support paper instead of the dataset itself. </p> <blockquote> <p>Salgas, A., Sun, J., Delbecq, S., Planès, T., & Lafforgue, G. (2024). Compilation and Applications of an Open-Source Dataset on Global Air Traffic Flows and Carbon Emissions. <em>Journal of Open Aviation Science</em>. <a href="https://doi.org/10.59490/joas.2024.7365">https://doi.org/10.59490/joas.2023.7201</a></p> </blockquote>
Open Research Skills Workshops - Open access publishing Workshop
<p><strong>This is the first workshop on Open Access Publishing in a series of workshops about Open Research Skills.</strong></p><p>This workshop covers:</p><p>Introduction to open access publishing</p><ul><li>Types of open access publishing</li><li>Examples of open access publishing journals and platforms</li><li>Benefits of open access publishing</li><li>Types of outputs that can be published</li></ul><p>Demonstration </p><ul><li>Demonstrating open publishing </li><li>Showing how a reproducible article is published and all the different outputs that are linked to it and how to do this</li></ul><p>Exercise</p><ul><li>Discuss and explore open publishing giving examples of different articles that show how open publishing works. We will pick those that show data and code deposited in repositories and also that use of protocol.io for publishing open methods</li></ul><p><strong>List of training workshops in Open Research Skills:</strong></p><ul><li><strong>24th February 2023 - Open access publishing</strong></li><li>24th March 2023 - Using repositories</li><li>21st April 2023 - GitHub basics</li><li>28th April 2023 - GitHub collaborative workflows</li><li>26th May 2023 - Standard vocabularies and ontologies</li><li>30th June 2023 - FAIR data</li></ul><p><strong>Project overview:</strong></p><p>Our project aims to upskill participants in open research skills to increase the quality and reusability of phytolith research and related disciplines such as archaeology, palaeosciences and plant sciences. We will run six hands-on training workshops on open access publishing and research outputs, using repositories, ontologies and standard vocabularies, implementation of FAIR Guidelines for phytolith research, and two workshops on Github basic and advanced skills. The materials from all workshops will be archived as self-study courses on our website (<a href="https://open-phytoliths.netlify.app/">https://open-phytoliths.netlify.app/</a>). We will also provide translation during workshops and training materials into multiple languages.</p>
ALL-READY Questionnare on potential drivers and barriers to the adoption of innovation management, open science, and Intellectual Property Rights (IPR) among the members of the Pilot Network
<p><strong>Background & Summary</strong>: </p><p>The ALL-READY project unites a diverse consortium of Research Infrastructures (RI) and Living Labs, instrumental in developing new methodologies and technologies in agroecology. The project focuses on effective management of innovation, adherence to open science principles, and strategic application of Intellectual Property Rights (IPR). Task 6.4 of the project, which concentrates on Innovation and IPR Management, seeks to understand the dynamics influencing the adoption of these practices among its members. Recognizing the need for end-to-end data management, the project emphasizes standardized data collection and management while adhering to FAIR principles.</p><p><strong>Methods</strong>: </p><p>The questionnaire was developed by LifeWatch ERIC to capture data reflecting current practices and perceptions in agroecology. It included 26 questions divided into four sections, focusing on existing practices, potential drivers, and barriers in innovation management, open science, and IPR. The survey was disseminated via an online platform to the ALLREADY Pilot Network, ensuring a representative sample from diverse organizations. The data collection process was closely monitored, and the responses were analyzed using a mixed-methods approach to extract meaningful insights.</p><p><strong>Data Records of the ALLREADY Project Questionnaire</strong>: </p><p>The dataset, collected through an online survey platform, underwent a meticulous process of data preparation, download, formatting, and anonymization. It consists of one text file containing metadata (Readme.txt) and a single CSV file encompassing all questionnaire responses. The dataset provides a comprehensive view of innovation management, open science adoption, and IPR handling within the agroecology sector, particularly among the network of RIs and Living Labs involved in the project.</p><p><strong>Technical Validation of the ALLREADY Project Questionnaire</strong>: </p><p>Several critical steps were taken to ensure the accuracy, reliability, and overall quality of the data collected. This included development and testing of the questionnaire, rigorous monitoring of the data collection process, and thorough checks for data quality and completeness. The representativeness of the sample was analyzed specifically with respect to the Pilot Network rather than the broader population involved in agroecology. Strategies were employed to counter survey fatigue and maintain respondent engagement.</p><p><strong>Usage Notes for the ALLREADY Project Questionnaire</strong>: </p><p>The dataset's proper usage is vital for ensuring the validity and reproducibility of research. Researchers are advised to consider the nature of the data, the representativeness of the dataset, and its generalizability. The dataset allows for comprehensive analysis and integration of different sections, and analysts have the flexibility to handle open and write-in responses according to their research needs. Additional information to facilitate analysis is provided in a separate documentation file.</p><p> </p>
MarFERReT: an open-source, version-controlled reference library of marine microbial eukaryote functional genes
<p>Metatranscriptomics generates large volumes of sequence data about transcribed genes in natural environments. Taxonomic annotation of these datasets depends on availability of curated reference sequences. For marine microbial eukaryotes, current reference libraries are limited by gaps in sequenced organism diversity and barriers to updating libraries with new sequence data, resulting in taxonomic annotation of only about half of eukaryotic environmental transcripts. Here, we introduce version 1.0 of the Marine Functional EukaRyotic Reference Taxa (MarFERReT), an updated marine microbial eukaryotic sequence library with a version-controlled framework designed for taxonomic annotation of eukaryotic metatranscriptomes. We gathered 902 marine eukaryote genomes and transcriptomes from multiple sources and assessed these candidate entries for sequence quality and cross-contamination issues, selecting 800 validated entries for inclusion in the library. MarFERReT v1 contains reference sequences from 800 marine eukaryotic genomes and transcriptomes, covering 453 species- and strain-level taxa, totaling nearly 28 million protein sequences with associated NCBI and PR2 Taxonomy identifiers and Pfam functional annotations. An accompanying MarFERReT project repository hosts containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT.<br><br>MarFERReT is linked to a code repository hosting containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT here: <a href="https://github.com/armbrustlab/marferret">https://github.com/armbrustlab/marferret</a></p> <p>The raw source data for the 902 candidate entries considered for MarFERReT v1.1.1, including the 800 accepted entries, are available for download from their respective online locations. The source URL for each of the entries is listed here in MarFERReT.v1.1.1.entry_curation.csv, and detailed instructions and code for downloading the raw sequence data from source are available in the MarFERReT code repository (<a href="https://github.com/armbrustlab/marferret/blob/main/docs/process_clean_marmicrodb.log.sh">link</a>). </p> <p>This repository release contains MarFERReT database files from the v1.1.1 MarFERReT release using the following MarFERReT library build scripts: <strong>assemble_marferret.sh</strong>, <strong>pfam_annotate.sh</strong>, and <strong>build_diamond_db.sh</strong><br><br>The following MarFERReT data products are available in this repository:</p> <p><strong>MarFERReT.v1.1.1.metadata.csv</strong><br>This CSV file contains descriptors of each of the 902 database entries, including data source, taxonomy, and sequence descriptors. Data fields are as follows:</p> <ol> <li><strong>entry_id</strong>: Unique MarFERReT sequence entry identifier.</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y/N). The Y/N values can be adjusted to customize the final build output according to user-specific needs.</li> <li><strong>marferret_name</strong>: A human and machine friendly string derived from the NCBI Taxonomy organism name; maintaining strain-level designation wherever possible.</li> <li><strong>tax_id</strong>: The NCBI Taxonomy ID (taxID).</li> <li><strong>pr2_accession</strong>: Best-matching PR2 accession ID associated with entry</li> <li><strong>pr2_rank</strong>: The lowest shared rank between the entry and the pr2_accession</li> <li><strong>pr2_taxonomy</strong>: PR2 Taxonomy classification scheme of the pr2_accession</li> <li><strong>data_type</strong>: Type of sequence data; transcriptome shotgun assemblies (TSA), gene models from assembled genomes (genome), and single-cell amplified genomes (SAG) or transcriptomes (SAT).</li> <li><strong>data_source</strong>: Online location of sequence data; the Zenodo data repository (<a href="../">Zenodo</a>), the datadryad.org repository (<a href="http://datadryad.org/">datadryad.org</a>), MMETSP re-assemblies on Zenodo (MMETSP)17, NCBI GenBank (<a href="https://www.ncbi.nlm.nih.gov/genbank/">NCBI</a>), JGI Phycocosm (<a href="https://phycocosm.jgi.doe.gov/phycocosm/home">JGI-Phycocosm</a>), the TARA Oceans portal on Genoscope (<a href="http://www.genoscope.cns.fr/tara/">TARA</a>), or entries from the Roscoff Culture Collection through the METdb database repository (<a href="https://metdb.sb-roscoff.fr/metdb/">METdb</a>).</li> <li><strong>source_link</strong>: URL where the original sequence data and/or metadata was collected.</li> <li><strong>pub_year</strong>: Year of data release or publication of linked reference.</li> <li><strong>ref_link</strong>: Pubmed URL directs to the published reference for entry, if available.</li> <li><strong>ref_doi</strong>: DOI of entry data from source, if available.</li> <li><strong>source_filename</strong>: Name of the original sequence file name from the data source.</li> <li><strong>seq_type</strong>: Entry sequence data retrieved in nucleotide (nt) or amino acid (aa) alphabets.</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file.</li> <li><strong>source_name:</strong> Full organism name from entry source</li> <li><strong>original_taxID</strong>: Original NCBI taxID from entry data source metadata, if available</li> <li><strong>alias:</strong> Additional identifiers for the entry, if available</li> </ol> <p><br><strong>MarFERReT.v1.1.1.curation.csv</strong><br>This CSV file contains curation and quality-control information on the 902 candidate entries considered for incorporation into MarFERReT v1, including curated NCBI Taxonomy IDs and entry validation statistics. Data fields are as follows:</p> <ol> <li><strong>entry_id:</strong> Unique MarFERReT sequence entry identifier</li> <li><strong>marferret_name: </strong>Organism name in human and machine friendly format, including additional NCBI taxonomy strain identifiers if available.</li> <li><strong>tax_id</strong>: Verified NCBI taxID used in MarFERReT</li> <li><strong>taxID_status</strong>: Status of the final NCBI taxID (Assigned, Updated, or Unchanged)</li> <li><strong>taxID_notes</strong>: Notes on the original_taxID</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file</li> <li><strong>n_pfams</strong>: Number of Pfam domains identified in protein sequences</li> <li><strong>qc_flag</strong>: Early validation quality control flags for the following: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.</li> <li><strong>flag_Lasek</strong>: Flag notes from Lasek-Nesselquist and Johnson (2019); contains the flag 'FLAG_LASEK' indicating ciliate samples reported as contaminated in this study.</li> <li><strong>VV_contam_pct</strong>: Estimated contamination reported for MMETSP entries in Van Vlierberghe et al., (2021).</li> <li><strong>flag_VanVlierberghe: </strong>Flag for a high level of estimated contamination, from 'flag_VanVlierberghe' values over 50%: FLAG_VV.</li> <li><strong>rp63_npfams</strong>: Number of ribosomal protein Pfam domains out of 63 total.</li> <li><strong>rp63_contam_pct</strong>: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, as described in the Technical Validation section from analysis of 63 Pfam ribosomal protein domains.</li> <li><strong>flag_rp63</strong>: Flag for a high level of estimated contamination, from 'rp63_contam_pct' values over 50%: FLAG_RP63.</li> <li><strong>flag_sum: </strong>Count of the number of flag columns (`qc_flag`, `flag_Lasek`, `flag_VanVlierberghe`, and `flag_rp63`). All entries with one or more flag are nominally rejected ('accepted' = N); entries without any flags are validated and accepted ('accepted' = Y).</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y or N).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins.faa.gz</strong><br>This Gzip-compressed FASTA file contains the 27,951,013 final translated and clustered protein sequences for all 800 accepted MarFERReT entries. The sequence defline contains the unique identifier for the sequence and its reference (mftX, where 'X' is a ten-digit integer value). </p> <p> </p> <p><strong>MarFERReT.v1.1.1.taxonomies.tab.gz</strong><br>This Gzip-compressed tab-separated file is formatted for interoperability with the DIAMOND protein alignment tool commonly used for downstream analyses and contains some columns without any data. Each row contains an entry for one of the MarFERReT protein sequences in MarFERReT.v1.proteins.faa.gz. Note that 'accession.version' and 'taxid' are populated columns while 'accession' and 'gi' have NA values; the latter columns are required for back-compatibility as input for the DIAMOND alignment software and LCA analysis. </p> <p>The columns in this file contain the following information:</p> <ol> <li><strong>accession</strong>: (NA)</li> <li><strong>accession.version</strong>: The unique MarFERReT sequence identifier ('mftX').</li> <li><strong>taxid</strong>: The NCBI Taxonomy ID associated with this reference sequence.</li> <li><strong>gi</strong>: (NA).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins_info.tab.gz</strong><br>This Gzip-compressed tab-separated file contains a row for each final MarFERReT protein sequence with the following columns:</p> <ol> <li><strong>aa_id</strong>: the unique identifier for each MarFERReT protein sequence.</li> <li><strong>entry_id</strong>: The unique numeric identifier for each MarFERReT entry.</li> <li><strong>source_defline</strong>: The original, unformatted sequence identifier</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.best_pfam_annotations.csv.gz<br></strong>This Gzip-compressed CSV file contains the best-scoring Pfam annotation for intra-species clustered protein sequences from the 800 validated MarFERReT entries; derived from the hmmsearch annotations against Pfam 34.0 functional domains. This file contains the following fields:</p> <ol> <li><strong>aa_id</strong>: The unique MarFERReT protein sequence ID ('mftX').</li> <li><strong>pfam_name</strong>: The shorthand Pfam protein family name.</li> <li><strong>pfam_id</strong>: The Pfam identifier.</li> <li><strong>pfam_eval</strong>: hmm profile match e-value score</li> <li><strong>pfam_score:</strong> hmm profile match bitscore</li> </ol> <p><br><strong>MarFERReT.v1.1.1.dmnd</strong><br>This binary file is the indexed database of the MarFERReT protein library with embedded NCBI taxonomic information generated by the DIAMOND makedb tool using the build_diamond_db.sh script from the MarFERReT /scripts/ library. This can be used as the reference DIAMOND database for annotating environment sequences from eukaryotic metatranscriptomes. <br><br></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.