Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11
datasets available to search
ShareScore release 0.9.0
Dataset results
11 results for “2012-2018”
Fish Counts and Lengths in South Bay and Hog Island Bay, Virginia 2012-2018
To study how seagrass restoration affects coastal fish communities over time, we sampled fishes at each site once or twice per year with beach seines (7.6 m wide × 1.8 m tall; 1.5 m deep pocket with 6.4 mm mesh) hauled along 25-m transects in the summer (May or June) and autumn (September or October) from 2012 through 2018. Researchers ceased sampling at the 4 initially unvegetated sites in South Bay after 2015, when these sites were colonized by seagrass, although seining occurred once more at these sites during the autumn of 2017. During each sampling event, we counted, measured (total length), and identified fish to the lowest possible taxon in the field prior to release. All seine hauls occurred during the day and within 3 hours of low tide for logistical reasons (n = 204). Due to methodological changes, after 2018 surveys are recorded in a different dataset VCR22364 "Abundance and Size of Seagrass-Associated Fishes in the Virginia Coastal Lagoons, 2019-xxxx" https://doi.org/10.6073/pasta/400c84b859e81e9a1e5212bccb37b759.
Time series of high-frequency sensors measuring water temperature and dissolved oxygen at discrete depths in Falling Creek Reservoir, Virginia, USA in 2012-2018
We measured water temperature and dissolved oxygen at multiple depths in Falling Creek Reservoir (Vinton, Virginia, USA) with high-frequency (10 to 15-minute) sensors for different durations during 2012 to 2018. Falling Creek Reservoir is owned and managed by the Western Virginia Water Authority as a primary drinking water source for Roanoke, Virginia. All measurements were collected at discrete depths at the deepest site of the reservoir adjacent to the dam. The sensors consisted of: 1) InsiteIG dissolved oxygen and water temperature sensors (Model 20 dissolved oxygen sensor) at both 1 m (November 2015 - December 2018) and 8 m (September 2012 - December 2018) and 2) HOBO (HOBO Pendant Temperature/Light 64K Data Logger) water temperature loggers deployed at 1, 2, 3, 4, 5, 6, 7, 8, and 9.3 m depths (September 2015 - January 2018).
Murphy Dome study site soil temperature measurements, 2012-2018
Soil temperature was monitored from August 2012 to July 2013 using two Thermochron iButtons (Maxim Integrated Products, San Jose, CA) installed 10 cm below the SOL surface in each plot. A difference in SOL depth between the stand types resulted in the sensors being in the SOL in black spruce and in the mineral soil in birch. Data is presented as mean daily values (measurement logged every 4 hours) for each study plot.
Historikertage auf Twitter (2012-2018). Datenreport und Datenset
<p>This data report contains the annotated figures, statistics and visualisations of the project "Die twitternde Zunft. Historikertage auf Twitter (2012-2018)" by Mareike König and Paul Ramisch. In addition, the methodological approach to corpus creation, data cleaning, coding, network and text analysis as well as the legal and ethical considerations of the project are described.</p> <p>The datasheets contain the dehydrated and annotated tweet ids that were used for our study. With the Twitter API this can be used to hydrate and restore the whole corpus, apart from deleted tweets. There are two versions of the CSV file, one with clean id values, the other where the id values are prepended with an “x”. This prevents certain tools from using scientific notation for the ids and breaking them, with the R library rtweet function read_twitter_csv() this is automatically resolved on import.</p> <p>The files contain the following data:</p> <ul> <li> status_id: The Twitter status id of the tweet</li> <li> corpus_user_id: A corpus specific id for each user within the corpus (not the Twitter user id)</li> <li> hauptkategorie_1: Primary category</li> <li> hauptkategorie_2: Primary category 2</li> <li> Gender: Gender of the user</li> <li> Nebenkategorie: Secondary category</li> </ul> <p>Furthermore, the following boolean variables describe what sub corpus each tweet is in, the main corpus per year that contains of both data sources (TAGS and API) and the yearly sub corpora divided by their data source (TAGS: orig_, API: api_):</p> <p>You can find the code on R on GitHub: <a href="https://github.com/dhiparis/historikertag-twitter">https://github.com/dhiparis/historikertag-twitter</a>.</p>
Monthly word embeddings for Twitter random sample (English, 2012-2018)
<p>This dataset contains monthly word embeddings created from the tweets available via the statuses/sample endpoint of the Twitter Streaming API from 2012 to 2018. Full details of the creation of the dataset are given in <a href="https://www.aclweb.org/anthology/D19-1007/">Room to Glo: A Systematic Comparison of Semantic Change Detection Approaches with Word Embeddings</a>. </p> <p>The md5sum of the gzipped tarball file is a76888ffec8cc7aebba09d365ca55ace .</p>
Murphy Dome: annual litter inputs from 2012-2018
This dataset contains annual litter inputs collected in 2012-2018 from the Murphy Dome study site.
Murphy Dome: Hourly temperature of air and soil, PAR, relative humidity and soil moisture in three pairs of adjacent black spruce and birch stands 2012-2018
This dataset contains weather station data (air and soil temperature, relative humidity, moisture, and PAR) from 2012 to 2019. The data was collected in three blocks (A, B, and C) of adjacent black spruce and paper birch stands at Murphy Dome (access via Cache Creek road).
Wbbyyr: FastText language models for Mandarin Chinese, trained on 14m Sina Weibo posts for each year in 2012-2018 (Fold 1 of 10)
<p>Wbbyyr: FastText language models for Mandarin Chinese, trained on 14,440,000 Sina Weibo posts for each year in 2012-2018.</p> <p>The 14,440,000 posts from each year are split into 10 folds. Due to Zenodo size limit, this dataset contains only the first fold from each year.</p> <p>Each model is trained for 20 iterations. Each vector is 300 dimensions long.</p>
Quebec Ministry of Agriculture, Fisheries and Food from 2012-2018 web archive collection derivatives
<p>Web archive derivatives of the Quebec Ministry of Agriculture, Fisheries and Food from 2012-2018 collection from the <a href="https://www.banq.qc.ca/accueil/">Bibliothèque et Archives nationales du Québec</a>. The derivatives were created with the <a href="https://github.com/archivesunleashed/aut/">Archives Unleashed Toolkit</a>. Merci beaucoup BAnQ!</p> <p>These derivatives are in the <a href="https://parquet.apache.org/">Apache Parquet format</a>, which is a <a href="http://en.wikipedia.org/wiki/Column-oriented_DBMS">columnar storage</a> format. These derivatives are generally small enough to work with on your local machine, and can be easily converted to Pandas DataFrames. See <a href="https://github.com/archivesunleashed/notebooks/blob/master/parquet_pandas_example.ipynb">this</a> notebook for examples.</p> <p><strong>Domains</strong></p> <pre><code class="language-java">.webpages().groupBy(ExtractDomainDF($"url").alias("url")).count().sort($"count".desc)</code></pre> <p>Produces a DataFrame with the following columns:</p> <ul> <li>domain</li> <li>count</li> </ul> <p><strong>Web Pages</strong></p> <pre><code class="language-java">.webpages().select($"crawl_date", $"url", $"mime_type_web_server", $"mime_type_tika", RemoveHTMLDF(RemoveHTTPHeaderDF(($"content"))).alias("content"))</code></pre> <p>Produces a DataFrame with the following columns:</p> <ul> <li>crawl_date</li> <li>url</li> <li>mime_type_web_server</li> <li>mime_type_tika</li> <li>content</li> </ul> <p><strong>Web Graph</strong></p> <pre><code class="language-java">.webgraph()</code></pre> <p>Produces a DataFrame with the following columns:</p> <ul> <li>crawl_date</li> <li>src</li> <li>dest</li> <li>anchor</li> </ul> <p><strong>Image Links</strong></p> <pre><code class="language-java">.imageLinks()</code></pre> <p>Produces a DataFrame with the following columns:</p> <ul> <li>src</li> <li>image_url</li> </ul> <p><a href="https://github.com/archivesunleashed/aut-docs/blob/master/current/binary-analysis.md#binary-analysis"><strong>Binary Analysis</strong></a></p> <ul> <li>Audio</li> <li>Images</li> <li>PDFs</li> <li>Presentation program files</li> <li>Spreadsheets</li> <li>Text files</li> <li>Videos</li> <li>Word processor files</li> </ul>
Gilbert Islands, Kiribati Relative Abundance of Benthic Taxa (2012-2018)
<p>This dataset contains the relative abundance of key benthic taxa in Abaiang and Tarawa Atolls, Gilbert Islands, Republic of Kiribati. We collected data every two years from 2012 through 2018 and identified coral and algae to the genus level (and sometimes the species if identification was possible and the species was common).</p>
Open Access Publications at the Université de Lorraine (France) 2012-2018
<p>This file contains three datasets regarding the publications by authors affiliated with Université de Lorraine in open access journals from 2012 to 2018 :</p> <ul> <li>first tab: data from a large bibliographic database (Web of Science) enriched with affiliation data and APC expenses</li> <li>second tab: data from the university's accounting software (SIFAC) enriched with affiliation data</li> <li>third tab: data from the OpenEdition Journals platform enriched with affiliation data</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.