Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,505
datasets available to search
ShareScore release 0.9.0
Dataset results
3,505 results for “completeness”
Towards a more complete quantification of the global carbon cycle
<p>These are the data and IDL code required to create Table3 from the paper.</p> <p><strong>Abstract.</strong></p> <p>The main components of global carbon budget calculations are the emissions from burning fossil fuels, cement production, and net land-use change, partly balanced by ocean CO<sub>2</sub> uptake and CO<sub>2</sub> increase in the atmosphere. The difference between these terms is referred to as the residual sink, assumed to correspond to increasing carbon storage in the terrestrial biosphere through physiological plant responses to changing conditions (Δ<em>B</em><sub>phys</sub>). It is often used to constrain carbon exchange in global earth-system models. More broadly, it guides expectations of autonomous changes in global carbon stocks in response to climatic changes, including increasing CO<sub>2</sub>, that may add to, or subtract from, anthropogenic CO<sub>2</sub> emissions.</p> <p>However, a budget with only these terms omits some important additional fluxes that are important to correctly infer Δ<em>B</em><sub>phys</sub>. They are cement carbonation and fluxes into increasing pools of plastic, bitumen, harvested-wood products, and landfill deposition after disposal of these products, and carbon fluxes to the oceans via wind erosion and non-CO<sub>2</sub> fluxes of the intermediate break-down products of methane and other volatile organic compounds. While the global budget includes river transport of dissolved inorganic carbon, it omits river transport of dissolved and particulate organic carbon, and the deposition of carbon in inland water bodies.</p> <p>Each one of these terms is relatively small, but together they can constitute important additional fluxes that would significantly reduce the size of the inferred Δ<em>B</em><sub>phys</sub>. We estimate here that inclusion of these fluxes would reduce Δ<em>B</em><sub>phys</sub> from the currently reported 3.6 GtC yr<sup>–1 </sup>down to about 2.1 GtC yr<sup>–1</sup> (excluding losses from land-use change). The implicit reduction in the size of ΔB<sub>phys</sub> has important implications for the inferred magnitude of current-day biospheric net carbon uptake and the consequent potential of future biospheric feedbacks to amplify or negate net anthropogenic CO<sub>2</sub> emissions.</p>
Datasets for Non-Parametric Class Completeness Estimators for Collaborative Knowledge Graphs
<p><strong>Non-Parametric Class Completeness Estimators for Collaborative Knowledge Graphs</strong></p> <p>This are intermediary datasets used for the calculation of the Class Completeness Estimators on Wikidata. For more information see: https://github.com/eXascaleInfolab/cardinal/</p> <p><strong>edits_wikidatawiki-20181001-pages.csv</strong></p> <p>This is an extract from <em>wikidatawiki-20181001-pages-meta-history</em> (All pages with complete page edit history (.bz2)) found at <a href="https://dumps.wikimedia.org/wikidatawiki/">https://dumps.wikimedia.org/wikidatawiki/</a>.</p> <p>The extract was created by the following SQL query:</p> <pre> SELECT page_title, rev_comment, rev_user_text, rev_timestamp FROM revisions WHERE rev_comment LIKE '%[[Property:%]]%[[Q%' ORDER BY rev_id INTO OUTFILE 'edits_wikidatawiki-20181001-pages.csv'; </pre> <p> </p> <p><strong>wikidata-20180813-all.json.bz2.universe.noattr.gt.bz2</strong></p> <p>This is a graph-tool representation of the WikiData graph. Output of <a href="https://github.com/eXascaleInfolab/cardinal/blob/master/1_create_inmemory_graph.py">https://github.com/eXascaleInfolab/cardinal/blob/master/1_create_inmemory_graph.py</a>.</p> <p><strong>observations_wikidatawiki-20181001-pages.pickle</strong></p> <p>Extracted observations. Output of <a href="https://github.com/eXascaleInfolab/cardinal/blob/master/2_extract_observations.py">https://github.com/eXascaleInfolab/cardinal/blob/master/2_extract_observations.py</a>.</p> <p> </p> <p><strong>estimates_wikidatawiki-20181001-pages.pickle</strong></p> <p>Extracted estimates. Output of <a href="https://github.com/eXascaleInfolab/cardinal/blob/master/3_calculate_estimates.py">https://github.com/eXascaleInfolab/cardinal/blob/master/3_calculate_estimates.py</a></p> <p> </p> <p><strong>results_wikidatawiki-20181001-pages.pickle </strong></p> <p>Results. Output of <a href="https://github.com/eXascaleInfolab/cardinal/blob/master/4_draw_graphs.py">https://github.com/eXascaleInfolab/cardinal/blob/master/4_draw_graphs.py</a></p>
Query auto-completions for German politicians of the 18th Bundestag
<p><strong>bundestag.csv</strong> - UTF-8 encoded comma separated text file</p> <p>This dataset contains the members of the 18th German Bundestag in the constitution of late 2016.</p> <p><Name>: name of the politician</p> <p><Born>: birthday</p> <p><Party>: party membership of the politician</p> <p><Bundesland> state of the politician</p> <p><Gender> gender of the politician</p> <p><Age> age of the politician (as of 2017)</p> <p><Cluster 3> number of unique auto-completions assigned to topic: "location information"</p> <p><Cluster 2> number of unique auto-completions assigned to topic: "personal and emotional"</p> <p><Cluster 1> number of unique auto-completions assigned to topic: "politics and economics"</p> <p><Total> total number of unique auto-completions</p> <p> </p> <p><strong>terms.csv </strong>- UTF-8 encoded comma separated text file</p> <p>This dataset contains the unordered and pooled auto-completions for the German politicians from Bing search (http://api.bing.net/osjson.aspx), from Duck-Duck-Go (https://duckduckgo.com/ac/) and from Google search (http://clients1.google.de/complete/search). The data was crawled on (mostly) two times per day from 2017/02/03 to 2017/06/19. German language settings were used for Google and Bing, English language setting was used for Duck-Duck-Go. The API requests were sent with an IP address from Cologne, Germany. </p> <p><source>: google, bing or ddg</p> <p><queryterm>: the query term, matches the name of the politican in the file <bundestag.csv></p> <p><suggestterm>: the suggested query auto-completion</p>
Supplementary Materials associated with paper 'Complete linear mitochondrial genomes for Cephea cephea and Mastigias albipunctata (Scyphozoa: Rhizostomeae), with an analysis of phylogenetic relationships'
<p>This is a repository for coverage depth graphs and ML-phylogenetic trees that are associated with the paper 'Complete linear mitochondrial genomes for Cephea cephea and Mastigias albipunctata (Scyphozoa: Rhizostomeae), with an analysis of phylogenetic relationships' by Tan KC, Collins AG and Ames CL.</p>
The complete corpus of #COVID-19 Twitter dataset
<p><br> COVID-19 pandemic initiated over a year ago continues to spread around the globe and the ongoing research regarding COVID-19 is on a continues growth as well. The online discourse on social media regarding COVID-19 has been growing along with the timeline of the pandemic.</p> <p>Open data on Twitter have been released and offer the research community the opportunity for new findings and resolving this new threat. In this dataset, we open a corpus of Twitter's data from March 2020 till today, that is being updated every day based on the two most important hashtags regarding COVID-19. This dataset will offer the research community the opportunity to explore the social extensions of this pandemic including topic analysis, hate speech sentiment analysis, regarding either the opinion of the users on the pandemic, the comments on the public discourse, or the vaccination releases. The dataset has been collected by retrieving all the tweets that contain the hashtags: #coronavirus and #COVID19 including approximately 208M tweets for hashtags #coronavirus and 392M tweets for hashtag #COVID-19, resulting in a total of 600M tweets. </p>
Data set accompanying the research article "Complete representation of action space and value in all striatal pathways"
<p>GCaMP6s calcium imaging data set recorded from freely behaving mice performing open field and 2-choice decision-making tasks using miniscopes. Mice were implanted in the right dorsomedial striatum and three types of output neurons were genetically targeted using transgenic Cre-lines. The data set comprises single-cell spatial filters and calcium activity traces extracted using CaImAn (https://github.com/flatironinstitute/CaImAn) as well as behavioral event logs and tracking coordinates. For more details please refer to the article "Complete representation of action space and value in all striatal pathways" published by the data sets' authors. Analysis code can be found at https://doi.org/10.5281/zenodo.5034618.</p>
The IHA database of human geometries including torso, head and complete outer ears for acoustic research
<p>This is the first version of the IHA database, which is being created in the project HAPPAA C1 funded by the Deutsche Forschungsgemeinschaft (DFG) – Projektnummer 352015383 - SFB 1330 C1. (https://uol.de/en/sfb-1330-hearing-acoustics)</p> <p>The database includes a subsample of 10 human geometries comprising the torso, head and the entire outer ear including the ear canal and eardrum. The data are available in two different 3D object formats: ply binary file, stl binary file.</p>
Data from: Flattening the curve: approaching complete sampling for diverse beetle communities
<p><strong>DATA FROM:</strong></p> <p>Burner, R., J. Åstrom, T. Birkemoe, A. Sverdrup-Thygeson. 2021. Flattening the curve: approaching complete sampling for diverse beetle communities. <em>Insect Conservation and Diversity</em> <a href="https://doi.org/10.1111/icad.12540">https://doi.org/10.1111/icad.12540</a> </p> <p> </p> <p><strong>ACKNOWLEDGEMENTS</strong></p> <p>This research was funded by the Norwegian Environment Directorate as part of an ‘Agreement on monitoring hollow oaks and insects in hollow oaks’. The Norwegian University of Life Sciences (NMBU) workshop designed and produced the cross-pane flight intercept traps. Thanks to Sindre Ligaard for identifying the beetle species, and to Lindsay Burner, Ruben Roos, and Ross Wetherbee for assistance in the field. High-performance computing resources were provided by Frederick H. Sheldon and Louisiana State University (LSU HPC).</p> <p><strong>INFORMATION</strong></p> <p>This dataset contains all data necessary to reproduce the analysis in the resulting manuscript. Briefly, 110 insect traps were set for 3 months in a single forest stand in Ås, Norway in 2020. This dataset includes trap locations, number of individuals of each species captured in each trap, trap type, and forest covariates collected around the traps.</p> <p>For more detailed information see manuscript and README file.</p> <p>From abstract of manuscript:</p> <ol> <li>Insects are a hyper diverse and ecologically important group. Their high diversity, however, presents challenges in sampling methodology, because rare species are unreliably detected with low sampling effort. However, the relationship between effort and species detections, critical for effective monitoring and evaluation of population trends, is too seldom quantified.</li> <li>We sampled forest beetles for three months in a 4-ha stand of mixed deciduous forest in southeastern Norway using 110 flight intercept (four types) and Malaise traps, the highest trap density (29 traps/ha) that we have seen reported. We examined species accumulation curves to quantify the benefits of each additional trap, compared capture rates among several trap designs and trap emptying frequencies, and tested for spatial autocorrelation.</li> <li>In total we captured 566 beetle taxa (19,854 individuals) from 52 families, yet our species accumulation curve was only beginning to flatten. Trap types differed considerably in their effectiveness. Nevertheless, twenty of our most effective window traps detected 75% of all taxa in our dataset. We found no evidence of spatial correlation within the scale of the study (100 m radius), nor did trap-level forest covariates (5 m radius) explain much variation.</li> <li>This implies that low to moderate sampling effort dramatically underestimates species richness, but that a limited number of effective traps can nonetheless achieve relatively thorough sampling for some applications. Immediate trap surroundings and spacing appeared unimportant. But, insect ecologists should take particular care in selecting trap types and be cautious comparing studies that employed different trap types.</li> </ol> <p> </p>
btw17 query auto completion - query suggestions for German politicians and parties before the federal election 2017
<p>The dataset contains the query suggestions for 5 major German parties (terms: "afd", "csu", "dielinke", "fdp", "grüne", "spd") and ten popular politicians and party leaders (terms: "Alexander Gauland", "Alice Weidel", "Angela Merkel", "Cem Özdemir", "Christian Lindner", "Dietmar Bartsch", "Katrin Göring-Eckardt", "Martin Schulz", "Sahra Wagenknecht").</p> <p>The data was crawled on (mostly) two times per day from Tue Aug 04, 2017 to Tue Oct 31, 2017. The dataset contains 20001 suggestions from Bing search (http://api.bing.net/osjson.aspx), 11935 suggestions from Duck-Duck-Go (https://duckduckgo.com/ac/) and 33521 suggestions from Google search (http://clients1.google.de/complete/search). Note, that for some terms and dates no suggestions were returned by some of the APIs.</p> <p>German language settings were used for Google and Bing, English language setting was used for Duck-Duck-Go. The API requests were sent with an IP address from Cologne, Germany. </p> <p>The UTF-8 encoded comma separated text file contains the following columns:</p> <p><source>: google, bing or ddg</p> <p><queryterm>: the query term</p> <p><date>: the date and time of the API call formatted as ISO8601</p> <p><suggestterm>: the suggested query completion (the query term was removed from the suggestion)</p> <p><position>: the position of the query suggestion within the list returned by the API (ranges from 0 to 19)</p> <p> </p> <p> </p> <p><br> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
Versailles, Urban driving, complete sessions
<p>Scenario description: Car sharing and urban driving with cyclist and pedestrian detection</p> <p>Session description: Booking of a vehicle at the car sharing station, manual driving through the city until reaching the part where the AD mode can be switched on. Then there are two rounds: the first one without IoT (VRUs not connected) then with IoT (pedestrian with smartphone/watch and cyclist with connected bike). Then manual driving back to car sharing station.</p> <p>Datasets descriptions:</p> <p><strong>AUTOPILOT_Versailles_UrbanDriving_Vehicle: </strong>Data generated from the vehicle sensors</p> <p>Vehicle datasets generated by the vehicle sensors during urban driving at Versailles. This includes the data coming from the CAN bus and GPS. It includes following kind of datasets: Vehicle: general data (speed, battery), PositioningSystem: data from GPS, VehicleDynamics: data about dynamic (acceleration...), Accel: acceleration data, EnvironmentSensorsAbsolute: environment sensors in absolute coordinates</p> <p><strong>AUTOPILOT_Versailles_UrbanDriving_V2X: </strong>V2X messages during uraban driving sessions</p> <p>Data exchanged with other vehicles and pedestrian during urban driving. SortOfCam: messages sent or received from bicycles</p> <p><strong>AUTOPILOT_Versailles_UrbanDriving_IoT: </strong>Data extracted from IoT oneM2M platform</p> <p>This dataset refers to messages exchanged by urban driving and car sharing application with vehicle, across oneM2M platform.</p> <p>oneM2M: car sharing status data</p> <p><strong>AUTOPILOT_Versailles_UrbanDriving_CAM: </strong>CAM messages</p> <p>This dataset refers to messages captured inside the vehicle during car sharing.</p> <p> </p>
Raw EEG-EOG data used in the publication "Auditory Electrooculogram-based Communication System for ALS Patients in Transition from Locked-in to Complete Locked-in State"
<p>The dataset includes raw EEG and EOG recordings during BCI experiments for three patients: p11, p13, p15, and p16. The structure of the dataset is the following: patient/visit/day.</p> <p>The experiment is described in detail in the publication "Auditory Electrooculogram-based Communication System for ALS Patients in Transition from Locked-in to Complete Locked-in State". The correspondence between raw file and BCI session is reported in the attached pdf file "Supplementary Table S5 Session to Raw File Recordings Correspondence".</p> <p>The datasets include EEG and EOG channels. The data are raw (i.e. non filtered and non processed). Data have been acquired with a sampling rate of 500Hz using active electrodes and the amplifier V-Amp DC (Brain Products, Germany). EOG channels are labeled EOGU, EOGD, EOGR, EOGL namely for EOG up, down, right, left; the location in the 10-20 system are respectively SO1, IO1, LO1, LO2.</p> <p>The data are marked with triggers: for each session two markers indicate start and end of the session; for each trial markers indicate start of baseline, start of presentation of question, start of response time, start of feedback. Each trial was marked in a different way if it was a yes trial belonging to a training or feedback session, a no trial belonging to a training or feedback session, or a trial belonging to a speller session. The markers that have been used are the following:<br> <strong>start</strong> 9<br> <em> yes no speller</em><br> <strong>baseline</strong> 10 11 12<br> <strong>presentation</strong> 5 6 7<br> <strong>response</strong> 4 8 13<br> <strong>feedback</strong> 1 2 3</p> <p><strong>end</strong><strong> </strong> 15</p>
Figure 7 in Fifty years of devotion to spiders: a concise biography of Christo Deltshev, with a complete list of his publications and described taxa
Figure 7. Christo Deltshev at his first International Arachnological Congress in Brno, Czechoslovakia, August, 1971 (congress photo).
Figure 5 in Fifty years of devotion to spiders: a concise biography of Christo Deltshev, with a complete list of his publications and described taxa
Figure 5. Christo, Peter van Helsdingen and Konrad Thaler in Szombathely, Hungary, July, 2002 (photo Stoyan Lazarov).
Fig. 11 in The Earwig Collection (Dermaptera) of the Manchester Museum, UK, with a complete type catalogue
Fig. 11. The only specimen at MMUE mounted with wings unfurled, Purex remotus (Burr, 1899) from Panama, the Manchester Museum. Scale bar = 1 cm.
Figs 59–63 in The Earwig Collection (Dermaptera) of the Manchester Museum, UK, with a complete type catalogue
Figs 59–63. Other material, presumed manuscript names, in the collection of the Manchester Museum. 59. Nesogaster spatulus Brindle (det. 1983). 60. Metalabis subcarinata Brindle (det. 1987). 61. Carcinophora rossi Hincks (no date). 62. Tagalina flavolineata Brindle (det. 1980). 63. Brindle's handwritten label for the Tagalina flavolineata specimen (Fig. 62). Scale bars = 1 cm.
Fig. 6 in The Earwig Collection (Dermaptera) of the Manchester Museum, UK, with a complete type catalogue
Fig. 6. Walter Douglas Hincks (1906–1961), Keeper of Entomology at the Manchester Museum (1947– 1961), photographed in the late 1950s.
Fig. 5 in The Earwig Collection (Dermaptera) of the Manchester Museum, UK, with a complete type catalogue
Fig. 5. The collection's earliest dated specimen, Apachyus chartaceus (Haan, 1842), dated 1878, the Manchester Museum. Scale bar = 1 cm.
Figs 34–39 in The Earwig Collection (Dermaptera) of the Manchester Museum, UK, with a complete type catalogue
Figs 34–39. Holotypes (♂) of Pygidicranidae and Labiduridae in the collection of the Manchester Museum. 34. Esphalmenus mucronatus Hincks, 1959. 35. E. kuscheli Hincks, 1959. 36. E. argentinus Hincks, 1959. 37. E. dentatus Hincks, 1959. 38. Tagalina curta Brindle, 1975. 39. Gonolabidura nathani Brindle, 1965. Scale bars = 1 cm.
Fig. 10 in The Earwig Collection (Dermaptera) of the Manchester Museum, UK, with a complete type catalogue
Fig. 10. Alan Brindle (1915–2001), Keeper of Entomology at the Manchester Museum (1961–1982), photographed in the mid 1980s.
Figs 7–8 in The Earwig Collection (Dermaptera) of the Manchester Museum, UK, with a complete type catalogue
Figs 7–8. Two items from Hincks' archive retained at the Manchester Museum. 7. An illustration of Pyragra paraguayensis Borelli, 1904 (♂) by Hincks for his Dermaptera monograph (Hincks 1959: 195) (item 597). 8. A page from Hincks' 'Typomap of Africa' (item 15).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.