Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
48
datasets available to search
ShareScore release 0.9.0
Dataset results
48 results for “collaborative projects”
Mangrove Coast Collaborative Project, Post-hurricane Irma mangrove forest structure data, Rookery Bay NERR, February 2022 - March 2023
The dataset describes the structure, composition, and condition of mangrove forests in Rookery Bay National Estuarine Research Reserve (NERR) assessed Feb 2022 - Mar 2023, approximately 5 years post Hurricane Irma (2017) and concurrent with Hurricane Ian (Sep 2022). The dataset includes information on each stem greater than or equal to 1 cm DBH (diameter at breast height) rooted within 69 100 m2 circular plots. Information collected includes site ID, location, species, DBH, status (live or dead), damage associated with the hurricane, presence/absence of regrowth, presence/absence of adventitious roots, whether or not stem is part of a multi-stemmed individual, and the canopy conditions (whether the tree has a canopy or is only sprouting at the base of the tree or trunk). This dataset is associated with the Mangrove Coast Collaborative project (2020 - 2024).
Mangrove Coast Collaborative Project, Post-hurricane Irma mangrove forest understory data, Rookery Bay NERR, February 2022 - March 2023
This dataset describes the structure and composition of the mangrove forest understory in Rookery Bay National Estuarine Research Reserve (NERR) assessed Feb 2022 - Mar 2023, approximately 5 years post Hurricane Irma (2017) and concurrent with Hurricane Ian (2022). The dataset includes information on the number of seedlings and saplings of each species as sampled in four 1 m2 quadrats in each 100 m2 structural sampling plot. The dataset also includes counts of the five tallest pneumatophores occurring in each quadrat. The height of pneumatophores was used as a proxy for the average maximum high water level in a plot. This dataset is associated with the Mangrove Coast Collaborative project (2020 - 2024).
Mangrove Coast Collaborative Project, Post-hurricane Maria mangrove forest understory data, Jobos Bay NERR, March 2022 - August 2022
This dataset describes the structure and composition of the mangrove forest understory in Jobos Bay National Estuarine Research Reserve (NERR) assessed approximately 5 years after disturbance from Hurricane Maria (2017). The dataset includes information on the number of seedlings and saplings of each species as sampled in four 1 m2 quadrats in each 100 m2 structural sampling plot. The dataset also includes counts of the five tallest pneumatophores occurring in each quadrat. The height of pneumatophores was used as a proxy for the average maximum high water level in a plot. This dataset is associated with the Mangrove Coast Collaborative project (2020 - 2024).
Mangrove Coast Collaborative Project, Post-hurricane Maria mangrove forest structure data, Jobos Bay NERR, March 2022 - August 2022
The dataset describes the structure, composition, and condition of mangrove forests in Jobos Bay National Estuarine Research Reserve (NERR) assessed approximately 5 years after disturbance from Hurricane Maria (2017). The dataset includes information on each stem greater than or equal to 1 cm DBH (diameter at breast height) rooted within 64 100 m2 circular plots. Information collected includes site ID, location, species, DBH, status (live or dead), damage associated with the hurricane, presence/absence of regrowth, presence/absence of adventitious roots, whether or not stem is part of a multi-stemmed individual, and the canopy conditions (whether the stem has a leafed canopy or is only sprouting at the base if live). This dataset is associated with the Mangrove Coast Collaborative project (2020 - 2024).
Mangrove Coast Collaborative Project, Post-hurricane Maria mangrove forest coarse woody debris data, Jobos Bay NERR, March 2022 - August 2022
This dataset describes the quality and size of coarse woody debris (downed woody debris > 7.5 cm in diameter) for each mangrove forest plot in Jobos Bay National Estuarine Research Reserve (NERR) assessed approximately 5 years after disturbance from Hurricane Maria (2017). Data was collected along three 20 m transects beginning at each plot center point and heading toward a randomly-selected azimuth. This dataset is associated with the Mangrove Coast Collaborative project (2020 - 2024).
Mangrove Coast Collaborative Project, Post-hurricane mangrove forest coarse woody debris data, Rookery Bay NERR, February 2022 - March 2023
This dataset describes the quality and size of coarse woody debris (downed woody debris > 7.5 cm in diameter) for each mangrove forest plot in Rookery Bay National Estuarine Research Reserve (NERR) assessed approximately 5 years after disturbance from Hurricane Irma (2017) and concurrent with Hurricane Ian (2022). Data was collected along three 20 m transects beginning at each plot center point and heading toward a randomly-selected azimuth. This dataset is associated with the Mangrove Coast Collaborative project (2020 - 2024).
Mangrove Coast Collaborative Project, Post-hurricane mangrove forest downed woody debris data, Rookery Bay NERR, February 2022 - March 2023
This dataset describes the quantity and size distribution of downed woody debris for each mangrove forest plot in Rookery Bay National Estuarine Research Reserve (NERR) assessed approximately 5 years after disturbance from Hurricane Irma (2017) and concurrent with Hurricane Ian (2022). Data was collected along three 20 m transects beginning at each plot center point and heading toward a randomly-selected azimuth. This dataset is associated with the Mangrove Coast Collaborative project (2020 - 2024).
Mangrove Coast Collaborative Project, Post-hurricane Maria mangrove forest downed woody debris data, Jobos Bay NERR, March 2022 - August 2022
This dataset describes the quantity and size distribution of downed woody debris for each mangrove forest plot in Jobos Bay National Estuarine Research Reserve (NERR) assessed approximately 5 years after disturbance from Hurricane Maria (2017). Data was collected along three 20 m transects beginning at each plot center point and heading toward a randomly-selected azimuth. This dataset is associated with the Mangrove Coast Collaborative project (2020 - 2024).
Mangrove Coast Collaborative Project, Hydrologic monitoring data in mangrove forests, Jobos Bay NERR, April 2024 - December 2024
The dataset describes the hydrologic conditions of the soil porewater (water level, conductivity, and temperature) at a depth of ~70 cm below ground in six mangrove forest locations in Jobos Bay National Estuarine Research Reserve (JBNERR) at 30-minute intervals between April 2024 to December 2024. Locations of minimal forest recovery following the effects of Hurricane Maria (September 2017) were identified and selected for hydrologic monitoring coincident with sites sampled for structural metrics in 2022. One reference site, defined as a site that was observed to be recovering following the hurricane, was selected in black mangrove forest. Two of the six sampling locations were selected to monitor effects of human encroachment on the western boundary of the reserve. These two sites were not coincident with structural sampling plots established in 2022. This dataset is associated with the MCC Catalyst Project entitled Limits of Resilience (2023-2025) funded by the National Estuarine Research Reserve System (NERRS) Science Collaborative.
Mangrove Coast Collaborative Project, Hydrologic monitoring data in mangrove forests, Rookery Bay NERR, April 2024 - December 2024
The dataset describes the hydrologic conditions of the soil porewater (water level, conductivity, and temperature) at a depth of ~70 cm below ground in six black mangrove forest locations in Rookery Bay National Esturarine Research Reserve (NERR) at 30-minute intervals between April 2024 to December 2024. Locations of minimal forest recovery following the effects of Hurricane Irma (September 2017) were identified and selected for hydrologic monitoring. The design consists of three sites in mainland/interior black mangroves and three sites on ocean-facing islands, all of which are located on the east side of Hurricane Irma eyewall. In each group, two of the sites selected were considered sites of minimal recovery whereas one site was selected as a reference (location of recovering mangroves). This dataset is associated with the MCC Catalyst Project entitled Limits of Resilience (2023-2025) funded by the National Estuarine Research Reserve System (NERRS) Science Collaborative.
Dataset - Survey results - Applying Model-based Requirements Engineering in Three Large European Collaborative Projects
<p>This dataset and its associated report contain the results of an online survey on using a model-based requirements engineering approach in three European projects. </p>
Educational transformation and network learning dataset – qualitative data from an international collaborative EU-project
<p>We are releasing our dataset of workshop outcomes acquired from the annual consortium conferences organized by the international “NextFood” consortium. The purpose of this project is to develop new ways of educating the future sustainability leaders of the agrifood sector, making sure that the professionals (farmers, advisers, businesses, students) have the right set of skills and competences needed to tackle the sustainability challenges we face ahead. Data gathering started from May 2018 yielding considerable amount of data on achievements, challenges and action plans related to educational transformation. This dataset will be updated by the time of project finalization. This work was funded by the European Union, through the Horizon 2020 project “NextFood”, Grant agreement No. 771738.</p>
The e-NDP project : collaborative digital edition of the Chapter registers of Notre-Dame of Paris (1326-1504). Ground-truth for handwriting text recognition (HTR) on late medieval manuscripts.
<p>The <a href="https://endp.hypotheses.org/">e-NDP project</a>, funded by the ANR, is led by the <a href="https://lamop.hypotheses.org/6870">LaMOP</a> (Julie Claustre and Darwin Smith).</p> <p>The project's partners are the Archives nationales, the Bibliothèque nationale de France (Department of Manuscripts, Bibliothèque de l'Arsenal), the École nationale des chartes and the Bibliothèque Mazarine.</p> <p>The e-NDP project aims at renewing our knowledge on <strong>Notre-Dame de Paris cathedral</strong> through the creation of a collaborative digital edition of the registers of its Chapter (1326-1504, <em>AN LL 105-128</em>), the community of 51 canons meeting three times a week on set days to take all administrative, financial and practical decisions pertaining to the cathedral, its estate and the society living in its cloister. This corpus has never been the object of a comprehensive study to understand the workings and history of this urban enclave and powerful community. The collaborative digital edition is based on a process of<strong> handwriting text recognition (HTR)</strong>, tested and supervised by scholars, researchers and engineers combining expertise in Medieval history, paleography, philology and digital humanities. The edition shall allow a better insight into the Chapter’s administration, into its economical and political power within Paris, and the relationships it maintained with other institutions in the city.</p> <p> </p> <p><strong>Section 1 : The e-NDP ground-truth dataset for Handwriting text recognition.</strong></p> <p>The full e-NDP corpus kept today in the French National Archives and was entirely digitized and described in its <a href="https://www.siv.archives-nationales.culture.gouv.fr/siv/rechercheconsultation/consultation/ir/consultationIR.action?formCaller=GENERALISTE&irId=FRAN_IR_059635">catalog</a> in 2022.</p> <p>The first major goal of the e-NDP projet is to propose a first automatic transcription of the 14k pages composing the 26 chapter registers. To achieve this goal representative samples from each one of the volumes were selected and transcribed in order to train a specialized HTR model able to propose a high quality automatic transcription. The collected ground-truth released on this repository currently has <strong>512 pages from the 26 registers</strong> of the cathedral chapter preserved in the National Archives (LL105 - LL128, <strong>1326-1504</strong>). The transcriptions were manually completed in <strong>two rounds</strong> by a group of 12 contributors, historians and paleographers, over the course of 2021-2022 using <a href="https://escriptorium.paris.inria.fr/">eScriptorium </a>as annotation environment. </p> <p> </p> <p><strong>Ground-truth features :</strong></p> <p><br> <em>Number of hands </em>: according to our estimates no fewer than 18 main hands were involved in the writing of the registers during the medieval period. </p> <p><em>Language</em> : More than 98% of the content of the registers was written in Latin, the rest in French. The exact percentage is hard to estimate because the vernacular language is often used in formulae, notes and comments. It is rare to find entire pages or blocks written in French. </p> <p><em>Script family</em> : The registers were written using a Cursive script (ca. late XIIIe - XVIe).</p> <p><em>Documental typology</em> : The volumes containing the chapter conclusions were conceived to serve as memorial records, but above all as documents for regular use and consultation in the daily practice of administration and management. In diplomatics the notion of "documentary manuscripts" is used to describe this kind of sources also by opposition to books and litterary or normative manuscripts.</p> <table align="center"> <caption><strong>Ground truth statistics</strong></caption> <tbody> <tr> <th>Text units</th> <th>Count</th> </tr> <tr> <td>Pages</td> <td>512</td> </tr> <tr> <td>Annotated regions (see section 2)</td> <td>2448</td> </tr> <tr> <td>Lines of text</td> <td>34231</td> </tr> <tr> <td>Tokens</td> <td>205083</td> </tr> <tr> <td>Characters</td> <td>3320407</td> </tr> </tbody> </table> <p> </p> <p><strong>Rules of transcription :</strong></p> <ul> <li>The abbreviations have been resolved, both those by suspension (<code>facimꝰ</code> ---> <code>facimus</code>) and by contraction (<code>dñi</code> --> <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --> <code>et</code> ; <code>ꝓ</code> --> <code>pro</code>) have been resolved. </li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the notary are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the text: <code>.</code> and <code>/</code> have been transcribed, but the transcription has not been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> <li>More specific transcription rules can be found into the file <code>transcription_guidelines.pdf</code></li> </ul> <p> </p> <p><strong>Section 2. e-NDP Layout Segmentation.</strong></p> <p>Layout segmentation is a compulsory step before HTR recognition in order to distinguish sections and regions inside a document. This process intend to separate interdependant page zones to produce a recognition in a section-sequence order and not in a line-sequence order which mix textual and peri-textual content.</p> <p>The regions of 364 pages (see <code>GT-layout_list</code>) of the e-NDP corpus were annotated using a 5 sections vocabulary (see <code>endp_layout_regions</code>) in order to describe the page distribution in all the 26 volumes :</p> <ol> <li><em>Block</em> : All the central text blocks, that normally corresponds to the main content called "conclusions" in registers.</li> <li><em>Liste</em> : List of names of the canons who were present during the meeting. Normally located before the <em>conclusions</em>.</li> <li><em>Entrée</em> : Marginal notes or entries to inform about the content of <em>conclusions</em>.</li> <li><em>Date</em> : Paragraph contending the date. Normally at the head of a <em>conclusion</em>, but separate of the main body.</li> <li><em>Numérotation</em> : Page numbers in roman or arabic. Usually appear in the top corners of the pages.</li> </ol> <table align="center"> <caption><strong>Layout GT statistics</strong></caption> <tbody> <tr> <th>Region</th> <th>Count</th> </tr> <tr> <td>block</td> <td>833</td> </tr> <tr> <td>liste</td> <td>431</td> </tr> <tr> <td>date</td> <td>448</td> </tr> <tr> <td>entrée</td> <td>205</td> </tr> <tr> <td>numérotation</td> <td>531</td> </tr> </tbody> </table> <p> </p> <p><strong>Section 3. The e-NDP HTR modeling.</strong></p> <p>The e-NDP project has progressively trained several HTR models adapted to work on late medieval cursive in order to accelerate the production of ground truth. Currently the best model delivers an average <strong>CER (Character error ratio) of 9.7%</strong> in handwriting recognition on the 26 registers (see <code>endp_learning_curve</code>) and can serve as generalist model for other manuscripts of the same period and similar script family. These models and their training implementation details can be found in the project's github <a href="https://github.com/chartes/e-NDP_HTR">repository</a>. </p> <p>Additionally, the automatic HTR transcriptions of the 26 registers (14k pages, 4.5M tokens) enriched with lexical and semantical information has been the subject of a first <a href="https://nosketch-engine.lamop.fr/#dashboard?corpname=endp">online publication</a> using the NoSketch engine that allows advanced data mining based on the combination of data, metadata and NLP features. </p> <p> </p> <p><strong>Section 4. Dataset content.</strong></p> <p>This zip dataset contains :</p> <p>- <code>HTR_ground_truth</code> : Two folders containing the jpg / jpeg images and their curated transcriptions in PAGE XML format.</p> <p>- <code>images_docs</code> : 4 files illustrating the different phases of the project (list of GT for layout segmentation, layout ontologie, transcription guideline and HTR evaluation curves)</p>
Dataset - Survey results - Applying Model-based Requirements Engineering in AIDOaRt Collaborative Project
<p>This dataset and its associated report contain the results of an online survey on using a model-based requirements engineering approach in AIDOaRT project in 2022.</p>
Collaborative Innovation Project funding launch - Dr Samantha Kanza (University of Reading, University of Southampton)
<p>This video is the eighth talk from our two day Future Blood Testing: Challenges & Opportunities Event that took place on the 13/09/2022.</p> <p>Collaborative Innovation Project funding launch - Dr Samantha Kanza (University of Reading)</p> <p>Bio: Dr Samantha Kanza is a Senior Enterprise Fellow at the University of Southampton. She completed her MEng in Computer Science at the University of Southampton and then worked for BAE Systems Applied Intelligence for a year before returning to do an iPhD in Web Science (in Computer Science and Chemistry), which focused on Semantic Tagging of Scientific Documents and Electronic Lab Notebooks. She was awarded her PhD in April 2018. Samantha works in the interdisciplinary research area of applying computer science techniques to the scientific domain, specifically through the use of semantic web technologies and artificial intelligence. Her research includes looking at electronic lab notebooks and smart laboratories, to improve the digitization and knowledge management of the scientific record using semantic web technologies; and using IoT devices in the laboratory. She has also worked on a number of interdisciplinary Semantic Web projects in different domains, including agriculture, chemistry and the social sciences.</p> <p>Further details on this event can be found at: https://futurebloodtesting.org/event/13-14-09-2022/ </p> <p>This video is an output from the Future Blood Testing Network which is funded by EPSRC under Grant Number EP/W000652/1</p> <p>YouTube Link: https://youtu.be/PDWZkZBzfqw</p>
Experiences Applying Lean R&D in Industry-Academia Collaboration Projects
<p>Supplementary materials of the paper Experiences Applying Lean R&D in Industry-Academia Collaboration Projects</p>
2x6 - a collaborative project in generative literature and translation
<p>5<sup>th</sup> Project Presentation</p>
globalbioticinteractions/AEC-DBCNet: Collaborative databasing of North American bee collections within a global informatics network project archive
<p>Data in this archive are from the <em>Collaborative databasing of North American bee collections within a global informatics network project</em>. Data was originally captured using Arthropod Easy Capture software developed at the American Museum of Natural History (AMNH), New York. Project lead investigators are John Ascher (Principal Investigator) and Jerome Rozen (Co-Principal Investigator) at the AMNH, and Douglas Yanega (Principal Investigator), University of California Riverside.</p> <p><strong>Please use this citation for this archive: </strong>John Ascher, Digital Bee Collections Network data archive from the C<em>ollaborative databasing of North American bee collections within a global informatics network project</em>. Version: 08 Mar 2016. https://doi.org/10.5281/zenodo.1436853</p> <p>This project was supported by the National Science Foundation grant <a href="https://nsf.gov/awardsearch/showAward?AWD_ID=0956388">DBI 0956388</a> and <a href="https://nsf.gov/awardsearch/showAward?AWD_ID=0956340">DBI 0956340</a></p> <p><strong>ABSTRACT</strong> Natural history collections contain millions of bee specimens documenting the geographic ranges, temporal occurrence patterns, and floral associations of the 20,000 described bee species. This project will digitize and consolidate specimen records from 10 bee collections across the United States. The investigators will make or verify species identifications, capture full label data, georeference and error-check localities, and upload this information to publicly accessible databases. Web-based tools will be used to capture data across collections efficiently, validate bee and plant names through automated comparison with taxonomic authority files, and synthesize data on species pages with images, digitized literature records, and other information about bees and their host plants. Data will be uploaded to the Global Biodiversity Information Facility and to Discover Life (www.discoverlife.org), a website that features customizable global maps for all global bee species and dynamic identification keys for North American species. To obtain information needed to conserve and manage pollinators, the investigators will work with ecologists to model geographic and temporal trends in bee populations in relation to environmental variables. Bees are the most important pollinators of the approximately 1/3 of crops that require animal pollination. Recent declines in honey bee populations highlight the need to understand better the roles of native bees in agricultural and natural systems. This project will help predict risks to bees and their pollination services from climate change, habitat loss, and other factors. The outreach program Bee Hunt (www.discoverlife.org/bee) will educate the public, including students in underserved communities, about bee diversity and the importance of pollination services. Using digital photography and rigorous research protocols, Bee Hunt will empower people at biological field stations, nature centers, parks, schools, and other sites to collect high-quality data to augment information from specimen records.</p>
STRIDE Project D5: Overcoming Barriers to Freight and Logistics Firm Collaboration with Urban Planning
<p>This dataset contains a list of Reddit posts and the corresponding threads for those posts resulting from targeted searches of four delivery-related subreddits conducted in Fall 2021. We used this dataset to understand driver practices in and views on delivering in urban areas and the challenges they face.</p> <p>It was downloaded using Reddit's API through the RedditExtractoR package for the R programming language. Re-use of this data is subject to Reddit API terms.</p>
Dataset and scripts for "Sentiment Analysis over Collaborative Relationships in Open Source Software Projects"
<p>Dataset and scripts for "Sentiment Analysis over Collaborative Relationships in Open Source Software Projects".</p> <p>README is included in the files</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.