Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

10

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

10 results for “media text”

Learn how ShareScore rates datasets ↗
zenodo40/100

Text-fig. 5. Reyispermum parvum gen. et sp. nov. seeds from the Early Cretaceous Vale de Água locality, Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings). a) Holotype (S174178; Vale de Agua sample 141) in lateral view showing shape and cell pattern; remains of mounting media (¤). b) Cut volume rendering of seed (cut at yz0553) showing the slightly raised tissue immediately adjacent to the lower edge of the hilum (arrow head) and palisade-shaped cells of exotesta. c) Apical view of seed showing hilar depression (hi), position of micropylar slit (mi) and the slightly raised raphal ridge (ra). d) Seed in lateral view showing raised tissue immediately adjacent to the lower edge of the hilum (arrow head) (S174495, Vale de Água sample 300). e) Cut volume rendering (cut at yz0500) of the seed in (5d) showing the raised tissue (arrow head) immediately adjacent to the lower edge of the hilum and sclerenchyma cells of exotesta. f) Detail of seed in (5d) showing micropylar slit (mi), hilum (hi) and raised tissue (arrow head) immediately adjacent to the lower edge of the hilum. g, h) Seed in lateral view (g) and view towards raphe (h) showing seed shape, the raised tissue below hilum (arrow head) and the raphal ridge (ra); note pointed micropylar area (S174179, Vale de Água sample 141). i) Seed surface of seed in (5d) showing the raised outlines of the undulate anticlinal walls of the exotestal cells. Scale bars = 250 µm (a–e, g, h); 125 µm (f, i). in Extinct Taxa Of Exotestal Seeds Close To Austrobaileyales And Nymphaeales From The Early Cretaceous Of Portugal

Text-fig. 5. Reyispermum parvum gen. et sp. nov. seeds from the Early Cretaceous Vale de Água locality, Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings). a) Holotype (S174178; Vale de Agua sample 141) in lateral view showing shape and cell pattern; remains of mounting media (¤). b) Cut volume rendering of seed (cut at yz0553) showing the slightly raised tissue immediately adjacent to the lower edge of the hilum (arrow head) and palisade-shaped cells of exotesta. c) Apical view of seed showing hilar depression (hi), position of micropylar slit (mi) and the slightly raised raphal ridge (ra). d) Seed in lateral view showing raised tissue immediately adjacent to the lower edge of the hilum (arrow head) (S174495, Vale de Água sample 300). e) Cut volume rendering (cut at yz0500) of the seed in (5d) showing the raised tissue (arrow head) immediately adjacent to the lower edge of the hilum and sclerenchyma cells of exotesta. f) Detail of seed in (5d) showing micropylar slit (mi), hilum (hi) and raised tissue (arrow head) immediately adjacent to the lower edge of the hilum. g, h) Seed in lateral view (g) and view towards raphe (h) showing seed shape, the raised tissue below hilum (arrow head) and the raphal ridge (ra); note pointed micropylar area (S174179, Vale de Água sample 141). i) Seed surface of seed in (5d) showing the raised outlines of the undulate anticlinal walls of the exotestal cells. Scale bars = 250 µm (a–e, g, h); 125 µm (f, i).

opencc-by-4.0Aug 2018View details →
zenodo40/100

Text-fig. 3. Pazlia hilaris gen. et sp. nov. (a–e) from the Early Cretaceous Famalicão locality (sample 025), Portugal (holotype, S175096) and Pazliopsis reyi gen. et sp. nov. (f–i) from the Early Cretaceous Torres Vedras locality, Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings a–f, i) and scanning electron microscopy (SEM, g, h). a, b) Seed in lateral (a) and oblique apical (b) views showing the truncate hilar-micropylar region; note prominent hilar scar (hi) and micropyle (mi) at the seed apex and the raphe (ra) seen as slightly raised ridge; remains of mounting media (¤). c) Cut volume rendering (cut at yz0647) showing course of raphe (ra), hilar scar (hi) and micropyle (mi); note the strongly radially elongated cells below the hilar scar. d) Seed in antiraphal view. e) Seed surface showing the raised undulate anticlinal walls of the exotestal cells. f) Seed enclosed in remains of thin-walled fruit (fr) (S174632, Torres Vedras sample 298). g) Holotype, seed enclosed in remains of fruit (fr); raphal view showing the faintly ribbed surface of the seed (S171534, Torres Vedras sample 043). h) Apical view of seed fragment showing hilar scar (hi), position of raphe (ra) and the ribbed seed surface (S136683, Torres Vedras sample 044). i) Seed surface showing the raised undulate anticlinal walls of the exotestal cells (S171534; Torres Vedras sample 043). Scale bars = 250 µm (a–d, f–h); 125 µm (e, i). in Extinct Taxa Of Exotestal Seeds Close To Austrobaileyales And Nymphaeales From The Early Cretaceous Of Portugal

Text-fig. 3. Pazlia hilaris gen. et sp. nov. (a–e) from the Early Cretaceous Famalicão locality (sample 025), Portugal (holotype, S175096) and Pazliopsis reyi gen. et sp. nov. (f–i) from the Early Cretaceous Torres Vedras locality, Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings a–f, i) and scanning electron microscopy (SEM, g, h). a, b) Seed in lateral (a) and oblique apical (b) views showing the truncate hilar-micropylar region; note prominent hilar scar (hi) and micropyle (mi) at the seed apex and the raphe (ra) seen as slightly raised ridge; remains of mounting media (¤). c) Cut volume rendering (cut at yz0647) showing course of raphe (ra), hilar scar (hi) and micropyle (mi); note the strongly radially elongated cells below the hilar scar. d) Seed in antiraphal view. e) Seed surface showing the raised undulate anticlinal walls of the exotestal cells. f) Seed enclosed in remains of thin-walled fruit (fr) (S174632, Torres Vedras sample 298). g) Holotype, seed enclosed in remains of fruit (fr); raphal view showing the faintly ribbed surface of the seed (S171534, Torres Vedras sample 043). h) Apical view of seed fragment showing hilar scar (hi), position of raphe (ra) and the ribbed seed surface (S136683, Torres Vedras sample 044). i) Seed surface showing the raised undulate anticlinal walls of the exotestal cells (S171534; Torres Vedras sample 043). Scale bars = 250 µm (a–d, f–h); 125 µm (e, i).

opencc-by-4.0Aug 2018View details →
zenodo40/100

Text-fig. 1. Gastonispermum portugallicum gen. et sp. nov. seeds from the Early Cretaceous Famalicão locality (sample 025), Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings). Note remains of mounting media on several seeds (¤). a) Seed in oblique view showing seed shape, the slightly raised raphal ridge and the position of hilum (hi) and micropyle (mi) on the raphal side of the seed (S170218). b, c) Seeds in lateral view (b, S170234; c, S175095). d–f) Holotype (S174820); seed in lateral view (d) and cut volume rendering (e, f) through the median plane of the seed showing palisade-shaped sclerenchyma cells of exotesta and remains of embryo (emb) and surrounding nutritive tissue (e, cut between yz0440-0530; f, cut between slices yz440-480). g) Hilum (hi) and micropyle (mi) of seed in (1a) showing the Y-shaped micropylar slit in the outer integument. h) Cut volume rendering through the median plane of the seed (cut at yz0492) showing seed coat mainly composed of palisade-shaped cells of the exotesta (S174435). i) Seed surface showing the raised outlines of the undulate anticlinal walls of the exotestal cells (S175045). Scale bars = 500 µm (a–e); 250 µm (g); 125 µm (f, i). in Extinct Taxa Of Exotestal Seeds Close To Austrobaileyales And Nymphaeales From The Early Cretaceous Of Portugal

Text-fig. 1. Gastonispermum portugallicum gen. et sp. nov. seeds from the Early Cretaceous Famalicão locality (sample 025), Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings). Note remains of mounting media on several seeds (¤). a) Seed in oblique view showing seed shape, the slightly raised raphal ridge and the position of hilum (hi) and micropyle (mi) on the raphal side of the seed (S170218). b, c) Seeds in lateral view (b, S170234; c, S175095). d–f) Holotype (S174820); seed in lateral view (d) and cut volume rendering (e, f) through the median plane of the seed showing palisade-shaped sclerenchyma cells of exotesta and remains of embryo (emb) and surrounding nutritive tissue (e, cut between yz0440-0530; f, cut between slices yz440-480). g) Hilum (hi) and micropyle (mi) of seed in (1a) showing the Y-shaped micropylar slit in the outer integument. h) Cut volume rendering through the median plane of the seed (cut at yz0492) showing seed coat mainly composed of palisade-shaped cells of the exotesta (S174435). i) Seed surface showing the raised outlines of the undulate anticlinal walls of the exotestal cells (S175045). Scale bars = 500 µm (a–e); 250 µm (g); 125 µm (f, i).

opencc-by-4.0Aug 2018View details →
zenodo40/100

Text-fig. 11. Silutanispermum kvacekiorum gen. et sp. nov. seeds from the Early Cretaceous Famalicão locality (sample 025), Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings). a, b) Holotype (S170238), seed in oblique (a) and lateral (b) view showing large triangular hilar scar (hi) and transverse micropylar slit (mi). c) Oblique view of seed showing slightly raised raphal area (S174352); remains of mounting media (¤). d) Details of holotype showing triangular hilum (hi) and transverse micropylar slit in the exotesta (mi). e) Cut volume rendering of holotype (cut at yz1170) through the median plane showing hilum (hi) and micropylar slit (mi) lined by radially expanded exotestal cells. Scale bars = 500 µm (a–c); 250 µm (d, e). in Extinct Taxa Of Exotestal Seeds Close To Austrobaileyales And Nymphaeales From The Early Cretaceous Of Portugal

Text-fig. 11. Silutanispermum kvacekiorum gen. et sp. nov. seeds from the Early Cretaceous Famalicão locality (sample 025), Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings). a, b) Holotype (S170238), seed in oblique (a) and lateral (b) view showing large triangular hilar scar (hi) and transverse micropylar slit (mi). c) Oblique view of seed showing slightly raised raphal area (S174352); remains of mounting media (¤). d) Details of holotype showing triangular hilum (hi) and transverse micropylar slit in the exotesta (mi). e) Cut volume rendering of holotype (cut at yz1170) through the median plane showing hilum (hi) and micropylar slit (mi) lined by radially expanded exotestal cells. Scale bars = 500 µm (a–c); 250 µm (d, e).

opencc-by-4.0Aug 2018View details →
zenodo28/100

MediaText: a media industry-based dataset for scene text detetcion

<h1>Media-Text</h1> <p>Media-Text dataset comprising images of banners, posters, covers and another images characterised for media industry.</p> <p>Full paper is available here: <a title="Media-Text: a Media Industry-Based Dataset for Scene Text Detection" href="https://www.researchgate.net/publication/385351709_Media-Text_a_Media_Industry-Based_Dataset_for_Scene_Text_Detection" target="_blank" rel="noopener">Media-Text: a Media Industry-Based Dataset for Scene Text Detection</a></p> <h3><strong>DATASET DESCRIPTION</strong></h3> <ul> <li>400 images</li> <li>7 744 annotated text instances</li> <li>973 annotations have been marked as illegible for the task of text recognition</li> <li>659 texts have been markes as do not care (###) for scene text detection.</li> <li>Images are represented by 193 unique resolutions.</li> </ul> <p>Annotation Format - Each image has corresponding&nbsp; gt_*.txt file, which contains annotations in bounding box format (defined by 4 courners), transcription, and bool flag which determines that text is illegible for OCR. Proposed format is similar to ICDAR15 annotations.</p> <p>x1, x2, ..., x4, y4, transcription, OCR Flag&nbsp;</p> <p><strong>Example:<br></strong><br>37,68,198,49,214,181,52,200,LADIES,False</p> <p><strong>ACKNOWLEDGMENT</strong></p> <p>This work was supported by the Silesian University of Technology (SUT) through the subsidy for maintaining and developing research potential grant in 2024 for young researchers, No. 2/070/BKM24/0058, and by the Ministry of Science and Higher Education "Implementation Doctorate" No. DWD/5/0511/2021.</p> <p>Thanks to the graphic department of media-press group for the preparation and possibility of sharing graphics thematically related to the prepared dataset.</p> <p>&nbsp;</p> <p><strong>LICENSE</strong></p> <p>Annotations created by authors are licesned under CC-BY-4.0 license.Images from the Open-Image-V7 dataset and are licensed according to their source information. Source information is defined in a file metadata.csv file that defines all the metadata of each file (File name corresponds to the ImageID column).</p> <p><strong>Images whose name corresponds to the media_press pattern are provided for academic use.</strong></p> <div> <div><strong>CITING THE RELATED WORKS</strong></div> <div>&nbsp;</div> </div> <div> <div>Please cite the related works in your publications if it helps your research:</div> <br> <div>```</div> <div>@inproceedings{inproceedings,</div> <div>author = {Kalisz, Seweryn and Marczyk, Michał and Polanska, Joanna},</div> <div>booktitle = {Modelling and simulation 2024. The 2024 European Simulation and Modelling Conference}</div> <div>editor = {Manuel Gra&ntilde;a; J. David Nu&ntilde;ez-Gonzalez}</div> <div>year = {2024},</div> <div>month = {10},</div> <div>pages = {138-144},</div> <div>publisher = {EUROSIS-ETI},</div> <div>title = {Media-Text: a Media Industry-Based Dataset for Scene Text Detection}</div> <div>}</div> <div>```</div> </div>

opencc-by-4.0Jul 2024View details →
zenodo28/100

The texts related to Postgraduate stress from mainstream social media in China.

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo24/100

Here comes everything - managing media, text, audio and electronic versioning

<p>Recording of the presentation &quot;Here comes everything - managing media, text, audio and electronic versioning&quot;</p> <p>Case studies in managing multimedia at Disney and Pearson including cataloguing for interoperability, media management for physical and digital, electronic versioning, intellectual property and rights management</p>

opencc-ncJun 2009View details →
zenodo20/100

PyTAIL Benchmark of Active Learning on Social Media Text Classification

<p>PyTAIL Benchmark of Active Learning on Social Media Text Classification</p><p>Read our paper for details: https://arxiv.org/abs/2211.13786</p><ul><li>ArXiv: https://arxiv.org/abs/2211.13786</li><li>Dataset: https://doi.org/10.5281/zenodo.7236430</li><li>Code: https://github.com/socialmediaie/pytail</li><li>Video: https://www.youtube.com/watch?v=AwDu64gN8t4&nbsp;</li></ul>

restrictedcc-by-4.0Oct 2022View details →
zenodo20/100

THE USE OF SOCIO-POLITICAL TERMS IN THE NEWS GENRE OF MEDIA TEXT

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo20/100

Datasets from `Discovering and analysing lexical variation in social media text'

<p>This repository contains the datasets that were&nbsp;used in the following three papers,&nbsp;which are also included within P. Shoemark&#39;s PhD dissertation `Discovering and analysing lexical variation in social media text&#39;:</p> <ul> <li>P. Shoemark, D. Sur, L. Shrimpton, I. Murray, and S. Goldwater.&nbsp;<a href="https://www.aclweb.org/anthology/E17-1116/"><em>Aye or naw, whit dae ye hink? Scottish independence and linguistic identity on social media.</em></a>&nbsp;15th Conference of the European Chapter of the Association for Computational Linguistics (EACL). 2017.&nbsp;</li> <li>P. Shoemark, J. Kirby, and S. Goldwater.&nbsp;<a href="https://www.aclweb.org/anthology/W17-4908/"><em>Topic and audience effects on distinctively Scottish vocabulary usage in Twitter data.</em>&nbsp;</a>Workshop on Stylistic Variation at EMNLP. 2017.</li> <li>P. Shoemark, J. Kirby, and S. Goldwater.&nbsp;<a href="https://www.aclweb.org/anthology/W18-6101/"><em>Inducing a lexicon of sociolinguistic variables from code-mixed text.</em>&nbsp;</a>Workshop on Noisy User Generated Text at EMNLP. 2018.&nbsp;</li> </ul> <p>Datasets&nbsp;consist&nbsp;of&nbsp;tab-separated-values files, in which rows correspond&nbsp;to tweets, with columns for user ID, tweet ID, and timestamp.&nbsp;</p> <p>The text of the tweets (and&nbsp;associated metadata) can be re-downloaded (in batches of&nbsp;100 per request) using Twitter&#39;s <a href="https://developer.twitter.com/en/docs/tweets/post-and-engage/api-reference/get-statuses-lookup">GET Statuses/Lookup</a> API endpoint&nbsp;<em>(NB: Tweets which have been deleted or made private since the original datasets were collected can <strong>not </strong>be re-downloaded, so it may not be possible to reconstruct the original datasets in their entirety).&nbsp;</em></p> <p>&nbsp;</p> <p>Most of these datasets were originally drawn from <a href="https://developer.twitter.com/en/docs/tweets/sample-realtime/api-reference/get-statuses-sample">the Sample endpoint of Twitter&rsquo;s Streaming API (</a>a.k.a. the &lsquo;Spritzer&rsquo;), which provides a random 1% sample of all public tweets in near real-time:</p> <p>&nbsp;</p> <p><strong>GU Dataset:&nbsp;<a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Geotagged-UK.zip?versionId=a9cc222e-1d07-4026-9b55-cf2189b58191">Geotagged-UK.zip</a></strong></p> <p><em>Tweets from Sept 2013 - Sept 2014 which are&nbsp;geotagged to locations within the UK.</em></p> <p>The file <strong>GU_pre-filtering.tsv </strong>contains the IDs for all&nbsp;tweets from the &lsquo;Spritzer&rsquo; stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are&nbsp;not retweets or quotes, and&nbsp;are geotagged to locations within the UK.<strong> </strong><em><strong>&bull;&nbsp; Tweets: </strong>1,768,334<strong>&nbsp;&nbsp;&bull;&nbsp;&nbsp;Unique Users: </strong>455,075 &nbsp;<strong>&bull;&nbsp;</strong></em></p> <p>The file&nbsp;<strong>GU.tsv&nbsp;</strong>contains the IDs for tweets in the <strong>final</strong> GU dataset that was used for the analyses in our&nbsp;<a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, after applying additional pre-processing heuristics&nbsp;to&nbsp;filter out tweets by&nbsp;bots and spammers.&nbsp;<em><strong>&bull; &nbsp;Tweets: </strong>1,654,204<strong>&nbsp; &bull;&nbsp;&nbsp;Unique Users: </strong>446,510&nbsp;&nbsp;<strong>&bull;&nbsp;</strong><sub>(the number of users in the GU dataset&nbsp;was slightly over-counted when reported in the paper; this is the actual number)</sub></em></p> <p>&nbsp;</p> <p><strong>GS Dataset: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Geotagged-Scotland.zip?versionId=146a0589-ddc0-4c42-86f6-7da79686ae56">Geotagged-Scotland.zip&nbsp;</a></strong></p> <p><em>The subset of Tweets in the GU dataset which are geo-tagged to locations within Scotland, specifically.</em></p> <p>The file <strong>GS_pre-filtering.tsv </strong>contains the IDs for all&nbsp;tweets from the &lsquo;Spritzer&rsquo; stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by&nbsp;<a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets or quotes, and&nbsp;are geotagged to locations within Scotland.&nbsp;<em><strong>&bull; &nbsp;Tweets: </strong>178,401<strong>&nbsp; &bull;&nbsp;&nbsp;Unique Users: </strong>41,685&nbsp;&nbsp;<strong>&bull;&nbsp;</strong></em></p> <p>The file <strong>GS.tsv&nbsp;</strong>contains the IDs for tweets in the <strong>final</strong> GS dataset that was used for the analyses in our&nbsp;<a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, after applying additional pre-processing heuristics&nbsp;to&nbsp;filter out tweets by&nbsp;bots and spammers.&nbsp;<em><strong>&bull; &nbsp;Tweets: </strong>166,992<strong>&nbsp; &bull;&nbsp;&nbsp;Unique Users: </strong>40,837&nbsp;&nbsp;<strong>&bull;&nbsp;</strong><sub>(the number of users in the GS&nbsp;dataset&nbsp;was slightly over-counted when reported in the paper; this is the actual number)</sub></em></p> <p>&nbsp;</p> <p><strong>IT Dataset &amp;&nbsp;Controls: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Indyref-Tweets.zip?versionId=41042a3e-f247-4493-9abe-29fb9ae81ee1">Indyref-Tweets.zip</a></strong></p> <p><em>Tweets from Sept 2013 - Sept 2014 which contain hashtags relating to the 2014 Scottish Independence Referendum (plus &#39;control&#39; tweets&nbsp;which are by the same users but do not contain referendum-related hashtags)</em></p> <p>The file <strong>IT_</strong><strong>pre-filtering.tsv </strong>contains the IDs for all&nbsp;tweets from the &lsquo;Spritzer&rsquo; stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by&nbsp;<a href="https://github.com/saffsd/langid.py">langid.py</a>,&nbsp;are not retweets or quotes, and&nbsp;contain at least one of 47 hashtags we identified as relating to the 2014 Scottish Independence Referendum (see paper for hashtag list).&nbsp;<em><strong>&bull; &nbsp;Tweets: </strong>77,708<strong>&nbsp; &bull;&nbsp;&nbsp;Unique Users: </strong>26,019 &nbsp;<strong>&bull;&nbsp;</strong></em></p> <p>The file <strong>IT</strong><strong>.tsv&nbsp;</strong>contains the IDs for tweets in the <strong>final</strong>&nbsp;IT&nbsp;dataset that was used for the analyses in our&nbsp;<a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, after applying additional pre-processing heuristics&nbsp;to&nbsp;filter out tweets by&nbsp;bots and spammers, and tweets which do not contain hashtags that we judged to&nbsp;<em>unambiguously</em>&nbsp;relate to the referendum.&nbsp;<em><strong>&bull; &nbsp;Tweets: </strong>59,664 &nbsp;<strong>&bull; &nbsp;Unique Users: </strong>18,589&nbsp;&nbsp;<strong>&bull;&nbsp;</strong></em></p> <p>The file&nbsp;<strong>IT_controls_pre-filtering.tsv&nbsp;</strong>contains the IDs for all&nbsp;tweets from the &lsquo;Spritzer&rsquo; stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>,&nbsp;are not retweets or quotes, and&nbsp;do <em><strong>not</strong></em> contain any of the hashtags we identified as relating to the 2014 Scottish Independence Referendum, but&nbsp;are authored by a user who has <em><strong>also</strong> </em>authored a tweet in&nbsp;<strong>IT_</strong><strong>pre-filtering.tsv</strong>.&nbsp;<em><strong>&bull; &nbsp;Tweets: </strong>1,354,701&nbsp;&nbsp;<strong>&bull; &nbsp;Unique Users: </strong>26,019&nbsp;&nbsp;<strong>&bull;&nbsp;</strong></em></p> <p>The file <strong>IT_controls</strong><strong>.tsv&nbsp;</strong>contains the IDs for tweets in the <strong>final</strong>&nbsp;set of Control tweets that was used for the analyses in our&nbsp;<a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, i.e. tweets which do not contain referendum-related hashtags but are authored by users who also have also authored tweets in&nbsp;<strong>IT</strong><strong>.tsv</strong>.&nbsp;<em><strong>&bull; &nbsp;Tweets: </strong>881,679&nbsp;<strong>&bull; &nbsp;Unique Users: </strong>18,589&nbsp;&nbsp;<strong>&bull;&nbsp;</strong></em></p> <p>&nbsp;</p> <p><strong>SG-Users&rsquo; and IH-Users&#39; Autumn 2014 Timeline Datasets: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Autumn-2014_Timelines.zip?versionId=c456b530-361e-4b88-9006-5bebd4a43c92">Autumn-2014_Timelines.zip</a></strong></p> <p><em>Complete tweet&nbsp;histories&nbsp;from Aug-Oct 2014&nbsp;for users from the GS and IT datasets.</em></p> <p>The file <strong>SG-Users_Autumn_2014_timelines_pre-filtering.tsv&nbsp;</strong>contains the IDs for tweets which were posted in August, September, or October 2014 by users from the GS&nbsp;dataset, i.e. users we know to have used Scottish geotags. This dataset is not restricted to tweets which appear in the &lsquo;Spritzer&rsquo; sample; instead&nbsp;it consists of complete User Timelines for the months concerned, retrieved using&nbsp;the <a href="https://developer.twitter.com/en/docs/tweets/timelines/api-reference/get-statuses-user_timeline">statuses/user timeline</a> endpoint of Twitter&rsquo;s REST API in March 2017. Because there are limits on the number of tweets that can be retrieved using this endpoint, we were not able to retrieve complete Autumn 2014 tweet histories for <em>all</em> of the users in the GS&nbsp;dataset.&nbsp;<em><strong>&bull; &nbsp;Tweets: </strong>3,014,029&nbsp;</em>&nbsp;<em><strong>&bull; &nbsp;Unique Users: </strong>18,274&nbsp;&nbsp;<strong>&bull;&nbsp;</strong></em></p> <p>The file&nbsp;<strong>SG-Users_Autumn_2014_timelines.tsv&nbsp;</strong>contains the IDs for tweets in the <strong>final</strong> SG-Users dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/W17-4908/">StyleVar 2017</a>&nbsp;paper. This dataset consists only of tweets which contain at least one&nbsp;instance&nbsp;of one of 50 lexical variables that were the focus of the study, and has also undergone various other filtering steps; see the paper for full details.&nbsp;<em><strong>&bull; &nbsp;Tweets: </strong>1,112,931</em>&nbsp;&nbsp;<em><strong>&bull; &nbsp;Unique Users: </strong>10,103&nbsp;&nbsp;<strong>&bull;&nbsp;</strong></em></p> <p>The file <strong>IH-Users_Autumn_2014_timelines_pre-filtering.tsv&nbsp;</strong>contains the IDs for tweets which were posted in August, September, or October 2014 by users from the IT dataset, i.e. users we know to have used Indyref-related hashtags.&nbsp;This dataset was collected in the same manner as&nbsp;SG-Users_Autumn_2014_timelines_pre-filtering.tsv;&nbsp;however, due to an error in this process, <strong>the IDs of most of the tweets&nbsp;in this dataset were not recorded</strong>. For such tweets the tweet ID column instead contains a placeholder tweet ID of the form _&lt;user_ID&gt;_&lt;month&gt;_&lt;integer&gt;, where the integer denotes&nbsp;the tweet&#39;s position in the reverse-chronological list of tweets that were retrieved for that user from that month (e.g.&nbsp;_147527441_09_286 &nbsp;is the placeholder tweet ID we assigned to the 286th September tweet we retrieved from the user whose Account ID is&nbsp;147527441). Unfortunately, therefore, the tweets in this file whose &#39;IDs&#39; begin with an underscore cannot be straightforwardly re-downloaded using Twitter&#39;s free&nbsp;<a href="https://developer.twitter.com/en/docs/tweets/post-and-engage/api-reference/get-statuses-lookup">GET Statuses/Lookup</a> API endpoint; but since their user IDs and timestamps are intact, it would still be possible to retrieve them&nbsp;using the (paid-for) <a href="http:// https://developer.twitter.com/en/docs/tutorials/choosing-historical-api">Historical APIs</a>.&nbsp;<em><strong>&bull; &nbsp;Tweets: </strong>6,997,858&nbsp;&nbsp;<strong>&bull; &nbsp;Tweets whose IDs were recorded:&nbsp;</strong>288,394</em><em><strong>&nbsp;</strong>&nbsp;<strong>&bull; &nbsp;Unique Users: </strong>14,645</em><em>&nbsp;<strong>&bull;&nbsp;</strong></em></p> <p>The file <strong>IH-Users_Autumn_2014_timelines.tsv&nbsp;</strong>contains the IDs for tweets in the <strong>final</strong>&nbsp;IH-Users dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/W17-4908/">StyleVar 2017</a>&nbsp;paper. Like the SG-Users dataset, this dataset consists only of tweets which contain at least one&nbsp;instance&nbsp;of one of 50 lexical variables that were the focus of the study, and has also undergone various other filtering steps; see the paper for full details. As with the pre-filtered version, most of the tweet IDs are unfortunately missing in this dataset.&nbsp;<em><strong>&bull; &nbsp;Tweets: </strong>2,165,320 &nbsp;<strong>&bull; &nbsp;Tweets whose IDs were recorded:&nbsp;</strong>115,366</em><em><strong>&nbsp;</strong>&nbsp;<strong>&bull; &nbsp;Unique Users: </strong>10,784&nbsp;<strong>&bull;&nbsp;</strong></em></p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>US Geotags: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Geotagged-USA.zip">Geotagged-USA.zip</a></strong></p> <p><em>Tweets from June 2013 - July 2016&nbsp;which are&nbsp;geotagged to locations within the USA.</em></p> <p>The file&nbsp;<strong>GUSA.tsv</strong> contains all tweets from the &lsquo;Spritzer&rsquo; sample which were posted between June 30th 2013 to July 1st 2016, are classified as English by&nbsp;<a href="https://github.com/saffsd/langid.py">langid.py</a>,&nbsp;are not retweets, do not&nbsp;contain urls or embedded media, are not by users with more than 1000 friends or followers, and&nbsp;are geotagged to locations within the USA. This dataset (along with the GU&nbsp;Dataset) was used in our <a href="http://www.aclweb.org/anthology/W18-6101/">WNUT 2018</a> paper. &nbsp;&nbsp;<em><strong>&bull; &nbsp;Tweets: </strong></em>&nbsp;<em>8,375,573&nbsp;</em>&nbsp;<em><strong>&bull; &nbsp;Unique Users: </strong>1</em>,<em>826,260</em><em>&nbsp;&nbsp;<strong>&bull;&nbsp;</strong></em></p> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Nov 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record