Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
12,906
datasets available to search
ShareScore release 0.9.0
Dataset results
12,906 results for “programming”
Programming Problems Submitted for Evaluation of LLMs GPT-3.5 and Gemini Pro 1.0
<p>Problems extracted from platforms LeetCode and BeeCrowd for evaluation of LLMs GPT3.5 and Gemini Pro 1.0.</p> <p>The data from the plataforms has the following columns and values:</p> <p> </p> <table> <tbody> <tr> <td><strong>LeetCode Data</strong></td> <td> </td> <td><strong>BeeCrowd Data</strong></td> <td> </td> </tr> <tr> <td><strong>Column</strong></td> <td><strong>Doc</strong></td> <td><strong>Column</strong></td> <td><strong>Doc</strong></td> </tr> <tr> <td>problem_level</td> <td>easy | medium | hard</td> <td>problem_level</td> <td><1...10></td> </tr> <tr> <td>problem_link</td> <td><LeetCode link for the problem></td> <td>problem_link</td> <td><BeeCrowd link for the problem></td> </tr> <tr> <td>prompt</td> <td><text submitted to LLM></td> <td>prompt</td> <td><text submitted to LLM></td> </tr> <tr> <td>response_code</td> <td><code provided by the LLM></td> <td>response_code</td> <td><code provided by the LLM></td> </tr> <tr> <td>response_evaluation</td> <td>True | False</td> <td>response_evaluation</td> <td>True | False</td> </tr> <tr> <td>execution_time_ms</td> <td><time></td> <td>execution_time_ms</td> <td><time></td> </tr> <tr> <td>memory_usage_mb</td> <td><memory></td> <td>error_generated</td> <td>Wrong Answer | Time Limit Exceeded | Memory Limit Exceeded....</td> </tr> <tr> <td>error_generated</td> <td>Wrong Answer | Time Limit Exceeded | Memory Limit Exceeded....</td> <td>attempts_number</td> <td>1 | 2 | 3</td> </tr> <tr> <td>attempts_number</td> <td>1 | 2 | 3</td> <td>author</td> <td><author's name></td> </tr> <tr> <td>contains_image</td> <td>True | False</td> <td>source</td> <td><origin institution></td> </tr> <tr> <td>related_topic_1</td> <td><topic></td> <td>origin_country</td> <td><origin country></td> </tr> <tr> <td>related_topic_2</td> <td><topic></td> <td>contains_image</td> <td>True | False</td> </tr> <tr> <td>related_topic_3</td> <td><topic></td> <td>related_topic_1</td> <td><topic></td> </tr> <tr> <td>related_topic_4</td> <td><topic></td> <td> </td> <td> </td> </tr> <tr> <td>related_topic_5</td> <td><topic></td> <td> </td> <td> </td> </tr> </tbody> </table> <p> </p>
Film Programming in the USSR: A Case Study of Moscow Cinemas (1946–1955)
<p>This paper presents a database on film programming in Moscow cinemas between 1946 and 1955. It outlines the place of this research at the intersection of new cinema history and academic debates on film distribution in the field of Soviet history. The paper describes the data collection, the coding to present the data, and the structure of the database, which consists of the three datasets on Moscow film programming (1946–1955), Moscow cinemas (1946–1955), and the 1952 film calendar. Concluding remarks summarize the knowledge obtained from the database and introduce the Soviet case into the international context of digital data collections for historical cinema studies.</p> <p>The database includes Cyrillic; it might require additional text encoding. </p> <p>The updated version includes the unique IDs for cinemas to increase the usability of the data. </p>
A Surface-Induced Asymmetric Program Promotes Tissue Colonization by a Human Pathogen - Supplemental data Fig 4A
<p>Raw data used for Fig. 4A of the article "A Surface-Induced Asymmetric Program Promotes Tissue Colonization by a Human Pathogen" published in Cell Host & Microbe. This Western Blot dataset is composed of 4 images:</p> <ul> <li>Western Blot 1: whole cell lysate (raw image, and annotated image)</li> <li>Western Blot 2: purified pili (raw image, and annotated image)</li> </ul>
Database of infection control and surveillance program, 2011-2020
<p>A full anonymized data set was collected as a part of the ICU infection control and surveillance program; 01/01/2011-12/31/2020</p> <p>File "Zenodo_DB_v4<a href="https://zenodo.org/api/files/6d089d03-7a43-4513-b476-92f438837941/VAE_Data_Main_0821_1338.csv">.csv</a>" contains daily data (one row is one day) on infection surveillance ordered by date.</p> <p>File "<a href="https://zenodo.org/api/files/6d089d03-7a43-4513-b476-92f438837941/Data_Dictionary_MainDB.csv">Data_Dictionary_MainDB_2021.csv</a>" contains the description of all variables from the data set.</p> <p> </p>
Speculative Automated Refactoring of Imperative Deep Learning Programs to Graph Execution
<p>Efficiency is essential to support ever-growing datasets, especially for Deep Learning (DL) systems. DL frameworks have traditionally embraced deferred execution-style DL code---supporting symbolic, graph-based Deep Neural Network (DNN) computation. While scalable, such development is error-prone, non-intuitive, and difficult to debug. Consequently, more natural, imperative DL frameworks encouraging eager execution have emerged but at the expense of run-time performance. Though hybrid approaches aim for the "best of both worlds," using them effectively requires subtle considerations. Our key insight is that, while DL programs typically execute sequentially, hybridizing imperative DL code resembles parallelizing sequential code in traditional systems. Inspired by this, we present an automated refactoring approach that assists developers in determining which otherwise eagerly-executed imperative DL functions could be effectively and efficiently executed as graphs. The approach features novel static imperative tensor and side-effect analyses for Python. Due to its inherent dynamism, analyzing Python may be unsound; however, the conservative approach leverages a speculative (keyword-based) analysis for resolving difficult cases that informs developers of any assumptions made. The approach is: (i) implemented as a plug-in to the PyDev Eclipse IDE that integrates the WALA Ariadne analysis framework and (ii) evaluated on nineteen DL projects consisting of 132 KLOC. The results show that 326 of 766 candidate functions (42.56%) were refactorable, and an average relative speedup of 2.16x on performance tests was observed with negligible differences in model accuracy. The results indicate that the approach is useful in optimizing imperative DL code to its full potential.</p>
Invasive pneumococcal diseases in children and adults before and after introduction of the 10-valent pneumococcal conjugate vaccine into the Austrian national immunization program
<p>The dataset contains case-based data on invasive pneumococcal disease in Austria, 2009/01 to 2017/02, by year and month of diagnosis, serotype and clinical presentation. Cases are anonymised by using a random ID.</p>
The Thousand-Pulsar-Array program on MeerKAT -- IX. The time-averaged properties of the observed pulsar population: data set
<p>This archive contains pulsar data presented as part of the MNRAS paper: <em>"The Thousand-Pulsar-Array program on MeerKAT -- IX. The time-averaged properties of the observed pulsar population"</em>.</p> <p>Folded, time-averaged pulse profiles (4 Stokes parameters, 8 frequency channels, 1024 time bins across the period) of the 1271 pulsars listed in Table 1 of the MNRAS paper are included in the ar_files.zip. Ephemerides of these pulsars (as used in the MNRAS paper) are included in the eph_files.zip. The pulsar data are readable by the PSRCHIVE package, see e.g. van Straten et al., Astronomical Research and Technology 9, 237 (2012).</p> <p>Tables 1, 5, and 6 from the MNRAS paper are included in tables_files.zip as .csv files. The file column_descriptions.txt describes the quantities in columns of these tables.<br> </p>
Dataset to Model the Sustainability of a Primary School Digital Education Curricular Reform and Professional Development Program
<p>This dataset contains the quantitative teacher data used to analyse the sustainability of an in-service teacher training program for Digital Education that took place from September 2019 to March 2020 in the Canton Vaud in Switzerland. As such, the study follows up on the 350 teachers over a year after the end of their professional development program had ended in order to model the sustainability of the reform, understand to what extent sustainability had been reached, thus validating the curricular reform model and helping draw recommendations for researchers and practitioners involved in Digital Education curricular reforms. As such, approximately 290 teachers from grades 1-4 in primary school (ages 5-9) responded to two sustainability surveys using web-based questionnaire to provide information relating to their perception of the training sessions and adoption of the computer science activities.</p> <p>The study is accepted for publication in Education and Information Technologies. </p> <p>A README is included and provides additional information regarding :</p> <p>- the requirements for re-use. </p> <p>- the specific content of the 2 csv files</p>
Dataset for "Too Simple? Notions of Task Complexity used in Maintenance-based Studies of Programming Tools"
<p>This dataset contains the data to replicate the findings in the publication "Too Simple? Notions of Task Complexity used in Maintenance-based Studies of Programming Tools". The dataset includes the bibliographies and intermediate analysis tables for the analyzed literature. Further, it contains the experiment materials providing context for the task discussed in Section V of the publication.</p> <p>The artifact contains the following files:</p> <ul> <li>icpc-from-survey.bib: The bibliographic entries of all publications selected from the corpus of the paper "Confounding parameters on program comprehension: a literature survey" by Janet Siegmund and Jana Schumann (https://doi.org/10.1007/s10664-014-9318-8)</li> <li>icpc-manually-added.bib: The bibliographic entries of manually added publications.</li> <li>icpc-phase-2-and-3.[csv|xlsx]: The intermediate analysis tables for phases 2 and 3 as described in the paper. The table also contains columns transferred from the original dataset of the Siegmund and Schumann study (marked as "transferred")</li> <li>experiment-materials.zip: Includes the materials for the experiment used in Section V <ul> <li>introduction-english.[docx|rtf]: The experiment protocol we read to participants, including the introduction to the system architecture.</li> <li>JumpODrom.package.zip: The code of the used system.</li> <li>task-16-description.txt: The textual description of the task discussed in Section V.</li> <li>task-16.patch: The change applied to the system code prior to the task, which causes the faulty behavior outlined in the task description.</li> </ul> </li> </ul> <p> </p>
SCShores: time-series of shorelines from Spanish Sandy beaches from citizen-science monitoring program.
<p>This repository contains 5 years of sandy beaches shorelines deriverd from a citizen-science monitoring program in the Spanish coast. The methodology and the dataset are described in:</p> <p><em><strong>González-Villanueva, R., Soriano-González, J., Alejo, I., Criado-Sudau, F., Plomaritis, T., Fernàndez-Mora, À., Benavente, J., Del Río, L., Nombela, M. Á., and Sánchez-García, E.: SCShores: a comprehensive shoreline dataset of Spanish sandy beaches from a citizen-science monitoring programme, Earth System Science Data. V. 15, 4613-4629 , <a href="https://essd.copernicus.org/articles/15/4613/2023/essd-15-4613-2023.html">https://doi.org/10.5194/essd-15-4613-2023</a>, 2023. </strong></em></p> <p>The shoreline dataset is provided in 1 GEOJSON file: SCShores.geojson. This dataset covers five<strong> </strong>sandy beaches located on the Atlantic and Mediterranean coasts of Spain where CoastSnap stations were available, and it includes a total of 1721 shorelines. The coordinate system for the geospatial layer is WGS84.</p> <ul> <li><strong><em>SCShores.geojson</em></strong>: this layer contains the sandy shorelines . Each feature in this layer is a multipoint with the following attributtes: <ul> <li><strong>site</strong>: CoastSnap station name id, e.g. agrelo, samarador, cadiz, ….</li> <li><strong>date</strong>: date and time of the shoreline, yyyyy-mm-dd hh:mm:ss</li> <li><strong>timezone</strong>: Coordinated Universal Time, UTC</li> <li><strong>timestampQuality</strong>: quality flag indicating the confidence in the date-time indicated by the image provider, e.g. 1, 2</li> <li><strong>imageSource</strong>: source of the original image from which the shoreline has been derived, e.g. Instagram, Twitter, Facebook, Email, CoastSnapApp</li> <li><strong>elevation_m:</strong> same as Z coordinate, defined by the observed tide and the tidal offset, in meters, Tide+tide offset</li> <li><strong>verticalDatum</strong>: mean sea level in Alicante, which is considered the zero topographic reference in the Spanish territory, NMMA</li> <li><strong>geometry</strong>: type of geometry used in the file, MultiPoint</li> <li><strong>coordinates</strong>: Geographic WGS84 coordinates for each point in the geometry, longitude, latitude, Z</li> </ul> </li> </ul> <p> </p>
Interagency Ecological Program: Benthic invertebrate monitoring in the Sacramento-San Joaquin Bay-Delta, collected by the Environmental Monitoring Program, 1975-2024.
The Interagency Ecological Program’s (IEP) Environmental Monitoring Program (EMP) was initiated in compliance with the Water Right Decision D-1379 (now mandated by Water Right Decision D-1641) and has monitored benthic invertebrate macrofauna in the upper San Francisco Estuary (SFE) since 1975. The objectives of the EMP are to obtain consistent and accurate monthly data at established monitoring stations and to report this information for the purpose of management and conservation of the upper San Francisco Estuary. While the EMP also collects discrete and continuous water quality data, along with phytoplankton and zooplankton data, this dataset only includes the benthic invertebrate data collected by the EMP from 1975-2024. EMP monitors invertebrate communities in the benthos of the SFE by collecting dredge samples with a Ponar sampler. Sediment and particles smaller than 0.5mm are removed from the sample using a sieve table. The invertebrates present are preserved in formalin, identified to the lowest possible taxonomic level, and enumerated. Currently, samples are collected monthly at 10 sites across the range of salinities found in the SFE. Four replicate dredge samples are collected at each site. The frequency of sampling, number and identity of sampling sites, and number of replicate samples has changed through the 45+ years of monitoring effort, in response to changes in perceived need for data. Links to other EMP datasets can be found on the EMP website: https://emp-des.github.io/emp-reports/data-links.html, or can be found on EDI for searching for "Environmental Monitoring Program" and "San Francisco".
Sacramento trawl, Delta Juvenile Fish Monitoring Program, Genetic Determination of Population of Origin 2017-2021
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Chipps Island trawl, Delta Juvenile Fish Monitoring Program, Genetic Determination of Population of Origin 2017-2021
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Interagency Ecological Program: Zooplankton and water quality data in the San Francisco Estuary collected by the Summer Townet and Fall Midwater Trawl monitoring programs.
The Interagency Ecological Program’s (IEP) Summer Townet Survey (STN) and Fall Midwater Trawl (FMWT) are two long-term monitoring projects conducted by the California Department of Fish and Wildlife (CDFW) to monitor fish abundance and distribution trends in the San Francisco Estuary (SFE) since 1959 and 1967, respectively. Starting in 2005, zooplankton monitoring was added and paired with fish tows to investigate food availability for young fishes. Food limitation has been a long-term issue and a focus of the Pelagic Organism Decline (POD) studies that began in 2005. By 2011, STN routinely conducted zooplankton monitoring at 40 stations, and FMWT at 32 stations in the upper SFE from Carquinez Strait to the Sacramento Deep Water Ship Channel and into the South Delta. STN samples every other week from June to August and FMWT samples once monthly from September to December. Both projects collect mesozooplankton samples using a modified Clarke-Bumpus (CB) net to target copepods and cladocerans, and FMWT also samples macrozooplankton (i.e. mysids and amphipods) using a mysid net. Flowmeters are used to measure the volume sampled to determine zooplankton catch per unit effort. Environmental variables such as water temperature, turbidity, secchi, and electrical conductivity are collected with each zooplankton sample. Concurrent fish and zooplankton tows conducted by STN and FMWT have allowed for comparisons of fish diet to the available zooplankton prey at the time of collection.
Sacramento-San Joaquin Bay-Delta Continuous (15 Minute) water quality monitoring data collected by the Continuous Environmental Monitoring Program, DWR, 2005- ongoing.
The Continuous Environmental Monitoring Program (CEMP) plays an instrumental role in overseeing real-time water quality in the Sacramento-San Joaquin Delta (the Delta) and Suisun Bay. The program harnesses wireless telemetry to transmit crucial data to the California Data Exchange Center (CDEC), making high-resolution environmental data pertaining to the Delta and Suisun Bay publicly accessible. The extensive dataset captures information at 15-minute intervals from 15 monitoring stations, utilizing YSI 6600 and YSI EXO sondes to obtain standalone water quality measurements. This extensive dataset informs the operations of the California State Water Project, ensuring it adheres to mandated water quality standards set by Water Right Decision 1641. This data compilation incorporates all information since the transition to YSI multiparameter sondes in 2005. It is important to note that the commencement dates and subsequent upgrades vary between stations, leading to slight discrepancies in the dataset's date ranges. Since its inception in the mid-1980s, CEMP has progressively expanded its monitoring capabilities, consistently augmenting the number of monitoring locations and the array of water quality parameters assessed. Its commitment to utilizing the most advanced water quality monitoring technology reaffirms its position as an environmental monitoring leader in the Delta and Suisun Bay. Today, the program oversees 15 water quality stations that reliably capture data every 15 minutes, each day of the year, transmitting this data in real-time. The core tenents of CEMP: • to obtain consistent and accurate data in real-time at established monitoring stations • to provide data necessary to achieve compliance with salinity, flow, and dissolved oxygen standards • to perform data analyses for further understanding of estuarine ecology • to report information to other government agencies, as well as the public, for the purpose of management and conservation of the upper San F
Data from “A Mixed Method Approach to Understanding the Public Health Impact of a School-Based Citizen Science Program to Reduce Arsenic in Private Well Water”
Objectives We have approached the problem of low well water testing rates in Maine and New Hampshire communities by developing the All About Arsenic (AAA) project, which engages secondary school teachers and students as citizen scientists in collecting well water samples for analysis of arsenic and other toxic metals and supports their outreach efforts to their communities. Methods We assessed this project’s public health impact by analyzing student data relative to existing well water quality datasets in both states. In addition, we surveyed private well owners who contributed well water samples to the project to determine the actions taken to mitigate arsenic in well water. Data The data presented here are used in the analyses performed for the publication: "A Mixed Method Approach to Understanding the Public Health Impact of a School-Based Citizen Science Program to Reduce Arsenic in Private Well Water.” Additional data may be available at: The Anecdata Project Page: https://anecdata.org/projects/view/299 The project website: https://www.allaboutarsenic.org/
Interagency Ecological Program: Discrete water quality and phytoplankton data from the Sacramento River floodplain and Yolo Bypass tidal slough, collected by the Yolo Bypass Fish Monitoring Program, 1998 - 2022
The Yolo Bypass Fish Monitoring Program (YBFMP) operates a rotary screw trap and fyke trap and conducts biweekly beach seine and lower trophic surveys in addition to maintaining water quality instrumentation in the bypass. The YBFMP serves to fill information gaps regarding environmental conditions in the bypass that trigger migrations and enhanced survival and growth of native fishes, as well as provide data for IEP synthesis efforts. YBFMP staff also conduct analyses of YBFMP monitoring data to address pertinent management related questions as identified by IEP. The Yolo Bypass has been identified as a high restoration priority by the National Marine Fisheries Service and US Fish and Wildlife Service Biological Opinions for Delta Smelt, Winter and Spring-run Chinook salmon and by California EcoRestore. The YBFMP informs the restoration actions that are mandated or recommended in these plans and provides critical baseline data on the ecology of the bypass and how it interacts with the broader San Francisco Estuary. Program objectives include: Collecting baseline data on water quality, chlorophyll, lower trophic level biota, and fish in the Yolo Bypass to monitor spatial and temporal changes in trends and abundance; Analyzing and communicating Yolo Bypass data with stakeholders and the scientific and management communities to address pertinent management related questions; Providing technical expertise on Yolo Bypass aquatic ecology and monitoring and sampling methods. We collect discrete water quality data using a YSI ProDSS and sample phytoplankton, chlorophyll and nutrients as discrete water grabs taken biweekly (or weekly during Yolo Bypass inundation) along with lower trophic tows. Water is sampled at three sites along the Yolo Bypass and Sacramento River, then processed and analyzed by an internal DWR laboratory.
Interagency Ecological Program: Summer Townet Survey for Young Pelagic Fishes in the San Francisco Estuary
The Interagency Ecological Program’s (IEP) Summer Townet Survey (STN) is a long-term effort to monitor the annual recruitment success of young pelagic fishes in the upper San Francisco Estuary (California, United States). Conducted by the California Department of Fish and Wildlife (CDFW) since 1959, STN has sampled fixed locations from Eastern San Pablo Bay to Rio Vista on the Sacramento River, and to Stockton on the San Joaquin River; and a single station in the lower Napa River. The study area was expanded in 2011 to include the Sacramento Deep Water Ship Channel (SDWSC) and Cache Slough (CS). Currently, 40 stations are sampled as a “survey” every other week June through August for 6 surveys. A conical net, lashed to a fixed metal “D” frame, is pulled obliquely through the water column 2 to 3 times at each station. All fish and macro-invertebrates are identified and enumerated from each tow, with fork lengths (mm) of the first 50 of each fish species also recorded. Fish catch, length-frequency, and catch-per-unit-effort (CPUE) among stations is available with data visualization tools on the study website for all species (https://wildlife.ca.gov/Conservation/Delta/Townet-Survey). Data collected at historic 31 stations are used to calculate annual relative abundance indices for age-0 Striped Bass (Morone saxatilis) and Delta Smelt (Hypomesus transpacificus). The remaining 9 stations are sampled to expand our sampling range further upriver and increase our understanding of larval and juvenile fish abundance and distribution in the lower Napa River and the North Delta. A meso-zooplankton net targeting copepods and cladocerans is also used in parallel to assess fish food resources at each station and a subset of the fish collected are retained for diet analysis by CDFW researchers (beginning 2005, see STN and FMWT zooplankton data on EDI). The STN also measures habitat conditions via water temperature (°C), water clarity (Secchi disk depth in cm), Turbidity (NTU) and s
Interagency Ecological Program: Data for Synthesis for Ecological Impacts of Drought in the Upper San Francisco Estuary, 1975-2021
These data were collected for the special issue on Drought in the Sacramento-San Joaquin Delta published in San Francisco Estuary and Watershed Sciences, Volume 22, Issue 1, March of 2024. Multi-year droughts are important and impactful features of California’s Mediterranean climate and can fundamentally affect the water quality and the ecosystem response of the San Francisco Estuary and the Sacramento-San Joaquin Delta. This data set was assembled from data collected by long-term monitoring programs over the past 46 years (1975-2021) to evaluate how the estuary changed during multi-year droughts. Data include fish, zooplankton, jellyfish, clams, temperature, Secchi depth, and salinity as measured by fish and water quality surveys, as well as nutrients and chlorophyll measured by water quality surveys. These data are accompanied by hydrologic metrics including Delta Inflow, Delta Outflow, the position of the 2-PSU isohaline, residence time, and daily values of various continuous water quality stations located in the Delta.
Interagency Ecological Program San Francisco Estuary Larval Entrainment Study (LES)
Fish entrainment into California's State Water Project (SWP) and Central Valley Project (CVP) has been a source of mortality for native osmerid species, including Delta Smelt (Hypomesus transpacificus) and Longfin Smelt (Spirinchus thaleichthys). While entrained adult and juvenile fishes are collected, identified and enumerated, the facilities design does not allow for a quantifiable estimate of larval fishes being entrained. This knowledge gap makes it difficult for managers to make scientifically informed decisions regarding the endangered Longfin & Delta Smelt. The Larval Entrainment Study (LES) aims to address this gap by providing a high-resolution dataset and insights into the variable abundances of Smelt at risk of entrainment, the community composition surrounding them, and associated water quality data. Our study design utilizes a variety of gear types and maximizes replicates to improve the understanding of larval Smelt abundances within the southern reach of the Sacramento and San Joaquin River Delta, the zone of entrainment for the SWP and CVP. Sampling began in 2022 using 10-minute stepped oblique tows of a Smelt Larval Survey (SLS) sled mounted with a 500 μm net. Tows were conducted in West Canal, outside Clifton Court Forebay (CCF), where the risk of entrainment is highest. Sampling was conducted five days a week in 2022 and three days a week in the following years. . Beginning in mid-January 2023, a 940 μm mesh net with the same dimensions as the SLS net was added, with sampling alternating between the two. This provided an intermediate gear between SLS (500 μm) and 20mm (1600 μm) nets. In 2024 we expanded to sample a transect of sites within the Old and Middle River (OMR) corridor. This provides broader spatial resolution and improves our ability to detect Osmerids within the South Delta. Sampling alternated between a high priority day (with intensive sampling in the lower San Joaquin River, SLS stations 809 & 812) and a longer transect day that ran
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.