Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
10
datasets available to search
ShareScore release 0.9.0
Dataset results
10 results for “media text”
Text-fig. 5. Reyispermum parvum gen. et sp. nov. seeds from the Early Cretaceous Vale de Água locality, Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings). a) Holotype (S174178; Vale de Agua sample 141) in lateral view showing shape and cell pattern; remains of mounting media (¤). b) Cut volume rendering of seed (cut at yz0553) showing the slightly raised tissue immediately adjacent to the lower edge of the hilum (arrow head) and palisade-shaped cells of exotesta. c) Apical view of seed showing hilar depression (hi), position of micropylar slit (mi) and the slightly raised raphal ridge (ra). d) Seed in lateral view showing raised tissue immediately adjacent to the lower edge of the hilum (arrow head) (S174495, Vale de Água sample 300). e) Cut volume rendering (cut at yz0500) of the seed in (5d) showing the raised tissue (arrow head) immediately adjacent to the lower edge of the hilum and sclerenchyma cells of exotesta. f) Detail of seed in (5d) showing micropylar slit (mi), hilum (hi) and raised tissue (arrow head) immediately adjacent to the lower edge of the hilum. g, h) Seed in lateral view (g) and view towards raphe (h) showing seed shape, the raised tissue below hilum (arrow head) and the raphal ridge (ra); note pointed micropylar area (S174179, Vale de Água sample 141). i) Seed surface of seed in (5d) showing the raised outlines of the undulate anticlinal walls of the exotestal cells. Scale bars = 250 µm (a–e, g, h); 125 µm (f, i). in Extinct Taxa Of Exotestal Seeds Close To Austrobaileyales And Nymphaeales From The Early Cretaceous Of Portugal
Text-fig. 5. Reyispermum parvum gen. et sp. nov. seeds from the Early Cretaceous Vale de Água locality, Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings). a) Holotype (S174178; Vale de Agua sample 141) in lateral view showing shape and cell pattern; remains of mounting media (¤). b) Cut volume rendering of seed (cut at yz0553) showing the slightly raised tissue immediately adjacent to the lower edge of the hilum (arrow head) and palisade-shaped cells of exotesta. c) Apical view of seed showing hilar depression (hi), position of micropylar slit (mi) and the slightly raised raphal ridge (ra). d) Seed in lateral view showing raised tissue immediately adjacent to the lower edge of the hilum (arrow head) (S174495, Vale de Água sample 300). e) Cut volume rendering (cut at yz0500) of the seed in (5d) showing the raised tissue (arrow head) immediately adjacent to the lower edge of the hilum and sclerenchyma cells of exotesta. f) Detail of seed in (5d) showing micropylar slit (mi), hilum (hi) and raised tissue (arrow head) immediately adjacent to the lower edge of the hilum. g, h) Seed in lateral view (g) and view towards raphe (h) showing seed shape, the raised tissue below hilum (arrow head) and the raphal ridge (ra); note pointed micropylar area (S174179, Vale de Água sample 141). i) Seed surface of seed in (5d) showing the raised outlines of the undulate anticlinal walls of the exotestal cells. Scale bars = 250 µm (a–e, g, h); 125 µm (f, i).
Text-fig. 3. Pazlia hilaris gen. et sp. nov. (a–e) from the Early Cretaceous Famalicão locality (sample 025), Portugal (holotype, S175096) and Pazliopsis reyi gen. et sp. nov. (f–i) from the Early Cretaceous Torres Vedras locality, Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings a–f, i) and scanning electron microscopy (SEM, g, h). a, b) Seed in lateral (a) and oblique apical (b) views showing the truncate hilar-micropylar region; note prominent hilar scar (hi) and micropyle (mi) at the seed apex and the raphe (ra) seen as slightly raised ridge; remains of mounting media (¤). c) Cut volume rendering (cut at yz0647) showing course of raphe (ra), hilar scar (hi) and micropyle (mi); note the strongly radially elongated cells below the hilar scar. d) Seed in antiraphal view. e) Seed surface showing the raised undulate anticlinal walls of the exotestal cells. f) Seed enclosed in remains of thin-walled fruit (fr) (S174632, Torres Vedras sample 298). g) Holotype, seed enclosed in remains of fruit (fr); raphal view showing the faintly ribbed surface of the seed (S171534, Torres Vedras sample 043). h) Apical view of seed fragment showing hilar scar (hi), position of raphe (ra) and the ribbed seed surface (S136683, Torres Vedras sample 044). i) Seed surface showing the raised undulate anticlinal walls of the exotestal cells (S171534; Torres Vedras sample 043). Scale bars = 250 µm (a–d, f–h); 125 µm (e, i). in Extinct Taxa Of Exotestal Seeds Close To Austrobaileyales And Nymphaeales From The Early Cretaceous Of Portugal
Text-fig. 3. Pazlia hilaris gen. et sp. nov. (a–e) from the Early Cretaceous Famalicão locality (sample 025), Portugal (holotype, S175096) and Pazliopsis reyi gen. et sp. nov. (f–i) from the Early Cretaceous Torres Vedras locality, Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings a–f, i) and scanning electron microscopy (SEM, g, h). a, b) Seed in lateral (a) and oblique apical (b) views showing the truncate hilar-micropylar region; note prominent hilar scar (hi) and micropyle (mi) at the seed apex and the raphe (ra) seen as slightly raised ridge; remains of mounting media (¤). c) Cut volume rendering (cut at yz0647) showing course of raphe (ra), hilar scar (hi) and micropyle (mi); note the strongly radially elongated cells below the hilar scar. d) Seed in antiraphal view. e) Seed surface showing the raised undulate anticlinal walls of the exotestal cells. f) Seed enclosed in remains of thin-walled fruit (fr) (S174632, Torres Vedras sample 298). g) Holotype, seed enclosed in remains of fruit (fr); raphal view showing the faintly ribbed surface of the seed (S171534, Torres Vedras sample 043). h) Apical view of seed fragment showing hilar scar (hi), position of raphe (ra) and the ribbed seed surface (S136683, Torres Vedras sample 044). i) Seed surface showing the raised undulate anticlinal walls of the exotestal cells (S171534; Torres Vedras sample 043). Scale bars = 250 µm (a–d, f–h); 125 µm (e, i).
Text-fig. 1. Gastonispermum portugallicum gen. et sp. nov. seeds from the Early Cretaceous Famalicão locality (sample 025), Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings). Note remains of mounting media on several seeds (¤). a) Seed in oblique view showing seed shape, the slightly raised raphal ridge and the position of hilum (hi) and micropyle (mi) on the raphal side of the seed (S170218). b, c) Seeds in lateral view (b, S170234; c, S175095). d–f) Holotype (S174820); seed in lateral view (d) and cut volume rendering (e, f) through the median plane of the seed showing palisade-shaped sclerenchyma cells of exotesta and remains of embryo (emb) and surrounding nutritive tissue (e, cut between yz0440-0530; f, cut between slices yz440-480). g) Hilum (hi) and micropyle (mi) of seed in (1a) showing the Y-shaped micropylar slit in the outer integument. h) Cut volume rendering through the median plane of the seed (cut at yz0492) showing seed coat mainly composed of palisade-shaped cells of the exotesta (S174435). i) Seed surface showing the raised outlines of the undulate anticlinal walls of the exotestal cells (S175045). Scale bars = 500 µm (a–e); 250 µm (g); 125 µm (f, i). in Extinct Taxa Of Exotestal Seeds Close To Austrobaileyales And Nymphaeales From The Early Cretaceous Of Portugal
Text-fig. 1. Gastonispermum portugallicum gen. et sp. nov. seeds from the Early Cretaceous Famalicão locality (sample 025), Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings). Note remains of mounting media on several seeds (¤). a) Seed in oblique view showing seed shape, the slightly raised raphal ridge and the position of hilum (hi) and micropyle (mi) on the raphal side of the seed (S170218). b, c) Seeds in lateral view (b, S170234; c, S175095). d–f) Holotype (S174820); seed in lateral view (d) and cut volume rendering (e, f) through the median plane of the seed showing palisade-shaped sclerenchyma cells of exotesta and remains of embryo (emb) and surrounding nutritive tissue (e, cut between yz0440-0530; f, cut between slices yz440-480). g) Hilum (hi) and micropyle (mi) of seed in (1a) showing the Y-shaped micropylar slit in the outer integument. h) Cut volume rendering through the median plane of the seed (cut at yz0492) showing seed coat mainly composed of palisade-shaped cells of the exotesta (S174435). i) Seed surface showing the raised outlines of the undulate anticlinal walls of the exotestal cells (S175045). Scale bars = 500 µm (a–e); 250 µm (g); 125 µm (f, i).
Text-fig. 11. Silutanispermum kvacekiorum gen. et sp. nov. seeds from the Early Cretaceous Famalicão locality (sample 025), Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings). a, b) Holotype (S170238), seed in oblique (a) and lateral (b) view showing large triangular hilar scar (hi) and transverse micropylar slit (mi). c) Oblique view of seed showing slightly raised raphal area (S174352); remains of mounting media (¤). d) Details of holotype showing triangular hilum (hi) and transverse micropylar slit in the exotesta (mi). e) Cut volume rendering of holotype (cut at yz1170) through the median plane showing hilum (hi) and micropylar slit (mi) lined by radially expanded exotestal cells. Scale bars = 500 µm (a–c); 250 µm (d, e). in Extinct Taxa Of Exotestal Seeds Close To Austrobaileyales And Nymphaeales From The Early Cretaceous Of Portugal
Text-fig. 11. Silutanispermum kvacekiorum gen. et sp. nov. seeds from the Early Cretaceous Famalicão locality (sample 025), Portugal; Synchrotron radiation X-ray tomographic microscopy (SRXTM, volume renderings). a, b) Holotype (S170238), seed in oblique (a) and lateral (b) view showing large triangular hilar scar (hi) and transverse micropylar slit (mi). c) Oblique view of seed showing slightly raised raphal area (S174352); remains of mounting media (¤). d) Details of holotype showing triangular hilum (hi) and transverse micropylar slit in the exotesta (mi). e) Cut volume rendering of holotype (cut at yz1170) through the median plane showing hilum (hi) and micropylar slit (mi) lined by radially expanded exotestal cells. Scale bars = 500 µm (a–c); 250 µm (d, e).
MediaText: a media industry-based dataset for scene text detetcion
<h1>Media-Text</h1> <p>Media-Text dataset comprising images of banners, posters, covers and another images characterised for media industry.</p> <p>Full paper is available here: <a title="Media-Text: a Media Industry-Based Dataset for Scene Text Detection" href="https://www.researchgate.net/publication/385351709_Media-Text_a_Media_Industry-Based_Dataset_for_Scene_Text_Detection" target="_blank" rel="noopener">Media-Text: a Media Industry-Based Dataset for Scene Text Detection</a></p> <h3><strong>DATASET DESCRIPTION</strong></h3> <ul> <li>400 images</li> <li>7 744 annotated text instances</li> <li>973 annotations have been marked as illegible for the task of text recognition</li> <li>659 texts have been markes as do not care (###) for scene text detection.</li> <li>Images are represented by 193 unique resolutions.</li> </ul> <p>Annotation Format - Each image has corresponding gt_*.txt file, which contains annotations in bounding box format (defined by 4 courners), transcription, and bool flag which determines that text is illegible for OCR. Proposed format is similar to ICDAR15 annotations.</p> <p>x1, x2, ..., x4, y4, transcription, OCR Flag </p> <p><strong>Example:<br></strong><br>37,68,198,49,214,181,52,200,LADIES,False</p> <p><strong>ACKNOWLEDGMENT</strong></p> <p>This work was supported by the Silesian University of Technology (SUT) through the subsidy for maintaining and developing research potential grant in 2024 for young researchers, No. 2/070/BKM24/0058, and by the Ministry of Science and Higher Education "Implementation Doctorate" No. DWD/5/0511/2021.</p> <p>Thanks to the graphic department of media-press group for the preparation and possibility of sharing graphics thematically related to the prepared dataset.</p> <p> </p> <p><strong>LICENSE</strong></p> <p>Annotations created by authors are licesned under CC-BY-4.0 license.Images from the Open-Image-V7 dataset and are licensed according to their source information. Source information is defined in a file metadata.csv file that defines all the metadata of each file (File name corresponds to the ImageID column).</p> <p><strong>Images whose name corresponds to the media_press pattern are provided for academic use.</strong></p> <div> <div><strong>CITING THE RELATED WORKS</strong></div> <div> </div> </div> <div> <div>Please cite the related works in your publications if it helps your research:</div> <br> <div>```</div> <div>@inproceedings{inproceedings,</div> <div>author = {Kalisz, Seweryn and Marczyk, Michał and Polanska, Joanna},</div> <div>booktitle = {Modelling and simulation 2024. The 2024 European Simulation and Modelling Conference}</div> <div>editor = {Manuel Graña; J. David Nuñez-Gonzalez}</div> <div>year = {2024},</div> <div>month = {10},</div> <div>pages = {138-144},</div> <div>publisher = {EUROSIS-ETI},</div> <div>title = {Media-Text: a Media Industry-Based Dataset for Scene Text Detection}</div> <div>}</div> <div>```</div> </div>
The texts related to Postgraduate stress from mainstream social media in China.
Open the record for dataset details and reuse information.
Here comes everything - managing media, text, audio and electronic versioning
<p>Recording of the presentation "Here comes everything - managing media, text, audio and electronic versioning"</p> <p>Case studies in managing multimedia at Disney and Pearson including cataloguing for interoperability, media management for physical and digital, electronic versioning, intellectual property and rights management</p>
PyTAIL Benchmark of Active Learning on Social Media Text Classification
<p>PyTAIL Benchmark of Active Learning on Social Media Text Classification</p><p>Read our paper for details: https://arxiv.org/abs/2211.13786</p><ul><li>ArXiv: https://arxiv.org/abs/2211.13786</li><li>Dataset: https://doi.org/10.5281/zenodo.7236430</li><li>Code: https://github.com/socialmediaie/pytail</li><li>Video: https://www.youtube.com/watch?v=AwDu64gN8t4 </li></ul>
THE USE OF SOCIO-POLITICAL TERMS IN THE NEWS GENRE OF MEDIA TEXT
Open the record for dataset details and reuse information.
Datasets from `Discovering and analysing lexical variation in social media text'
<p>This repository contains the datasets that were used in the following three papers, which are also included within P. Shoemark's PhD dissertation `Discovering and analysing lexical variation in social media text':</p> <ul> <li>P. Shoemark, D. Sur, L. Shrimpton, I. Murray, and S. Goldwater. <a href="https://www.aclweb.org/anthology/E17-1116/"><em>Aye or naw, whit dae ye hink? Scottish independence and linguistic identity on social media.</em></a> 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL). 2017. </li> <li>P. Shoemark, J. Kirby, and S. Goldwater. <a href="https://www.aclweb.org/anthology/W17-4908/"><em>Topic and audience effects on distinctively Scottish vocabulary usage in Twitter data.</em> </a>Workshop on Stylistic Variation at EMNLP. 2017.</li> <li>P. Shoemark, J. Kirby, and S. Goldwater. <a href="https://www.aclweb.org/anthology/W18-6101/"><em>Inducing a lexicon of sociolinguistic variables from code-mixed text.</em> </a>Workshop on Noisy User Generated Text at EMNLP. 2018. </li> </ul> <p>Datasets consist of tab-separated-values files, in which rows correspond to tweets, with columns for user ID, tweet ID, and timestamp. </p> <p>The text of the tweets (and associated metadata) can be re-downloaded (in batches of 100 per request) using Twitter's <a href="https://developer.twitter.com/en/docs/tweets/post-and-engage/api-reference/get-statuses-lookup">GET Statuses/Lookup</a> API endpoint <em>(NB: Tweets which have been deleted or made private since the original datasets were collected can <strong>not </strong>be re-downloaded, so it may not be possible to reconstruct the original datasets in their entirety). </em></p> <p> </p> <p>Most of these datasets were originally drawn from <a href="https://developer.twitter.com/en/docs/tweets/sample-realtime/api-reference/get-statuses-sample">the Sample endpoint of Twitter’s Streaming API (</a>a.k.a. the ‘Spritzer’), which provides a random 1% sample of all public tweets in near real-time:</p> <p> </p> <p><strong>GU Dataset: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Geotagged-UK.zip?versionId=a9cc222e-1d07-4026-9b55-cf2189b58191">Geotagged-UK.zip</a></strong></p> <p><em>Tweets from Sept 2013 - Sept 2014 which are geotagged to locations within the UK.</em></p> <p>The file <strong>GU_pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets or quotes, and are geotagged to locations within the UK.<strong> </strong><em><strong>• Tweets: </strong>1,768,334<strong> • Unique Users: </strong>455,075 <strong>• </strong></em></p> <p>The file <strong>GU.tsv </strong>contains the IDs for tweets in the <strong>final</strong> GU dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, after applying additional pre-processing heuristics to filter out tweets by bots and spammers. <em><strong>• Tweets: </strong>1,654,204<strong> • Unique Users: </strong>446,510 <strong>• </strong><sub>(the number of users in the GU dataset was slightly over-counted when reported in the paper; this is the actual number)</sub></em></p> <p> </p> <p><strong>GS Dataset: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Geotagged-Scotland.zip?versionId=146a0589-ddc0-4c42-86f6-7da79686ae56">Geotagged-Scotland.zip </a></strong></p> <p><em>The subset of Tweets in the GU dataset which are geo-tagged to locations within Scotland, specifically.</em></p> <p>The file <strong>GS_pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets or quotes, and are geotagged to locations within Scotland. <em><strong>• Tweets: </strong>178,401<strong> • Unique Users: </strong>41,685 <strong>• </strong></em></p> <p>The file <strong>GS.tsv </strong>contains the IDs for tweets in the <strong>final</strong> GS dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, after applying additional pre-processing heuristics to filter out tweets by bots and spammers. <em><strong>• Tweets: </strong>166,992<strong> • Unique Users: </strong>40,837 <strong>• </strong><sub>(the number of users in the GS dataset was slightly over-counted when reported in the paper; this is the actual number)</sub></em></p> <p> </p> <p><strong>IT Dataset & Controls: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Indyref-Tweets.zip?versionId=41042a3e-f247-4493-9abe-29fb9ae81ee1">Indyref-Tweets.zip</a></strong></p> <p><em>Tweets from Sept 2013 - Sept 2014 which contain hashtags relating to the 2014 Scottish Independence Referendum (plus 'control' tweets which are by the same users but do not contain referendum-related hashtags)</em></p> <p>The file <strong>IT_</strong><strong>pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets or quotes, and contain at least one of 47 hashtags we identified as relating to the 2014 Scottish Independence Referendum (see paper for hashtag list). <em><strong>• Tweets: </strong>77,708<strong> • Unique Users: </strong>26,019 <strong>• </strong></em></p> <p>The file <strong>IT</strong><strong>.tsv </strong>contains the IDs for tweets in the <strong>final</strong> IT dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, after applying additional pre-processing heuristics to filter out tweets by bots and spammers, and tweets which do not contain hashtags that we judged to <em>unambiguously</em> relate to the referendum. <em><strong>• Tweets: </strong>59,664 <strong>• Unique Users: </strong>18,589 <strong>• </strong></em></p> <p>The file <strong>IT_controls_pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets or quotes, and do <em><strong>not</strong></em> contain any of the hashtags we identified as relating to the 2014 Scottish Independence Referendum, but are authored by a user who has <em><strong>also</strong> </em>authored a tweet in <strong>IT_</strong><strong>pre-filtering.tsv</strong>. <em><strong>• Tweets: </strong>1,354,701 <strong>• Unique Users: </strong>26,019 <strong>• </strong></em></p> <p>The file <strong>IT_controls</strong><strong>.tsv </strong>contains the IDs for tweets in the <strong>final</strong> set of Control tweets that was used for the analyses in our <a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, i.e. tweets which do not contain referendum-related hashtags but are authored by users who also have also authored tweets in <strong>IT</strong><strong>.tsv</strong>. <em><strong>• Tweets: </strong>881,679 <strong>• Unique Users: </strong>18,589 <strong>• </strong></em></p> <p> </p> <p><strong>SG-Users’ and IH-Users' Autumn 2014 Timeline Datasets: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Autumn-2014_Timelines.zip?versionId=c456b530-361e-4b88-9006-5bebd4a43c92">Autumn-2014_Timelines.zip</a></strong></p> <p><em>Complete tweet histories from Aug-Oct 2014 for users from the GS and IT datasets.</em></p> <p>The file <strong>SG-Users_Autumn_2014_timelines_pre-filtering.tsv </strong>contains the IDs for tweets which were posted in August, September, or October 2014 by users from the GS dataset, i.e. users we know to have used Scottish geotags. This dataset is not restricted to tweets which appear in the ‘Spritzer’ sample; instead it consists of complete User Timelines for the months concerned, retrieved using the <a href="https://developer.twitter.com/en/docs/tweets/timelines/api-reference/get-statuses-user_timeline">statuses/user timeline</a> endpoint of Twitter’s REST API in March 2017. Because there are limits on the number of tweets that can be retrieved using this endpoint, we were not able to retrieve complete Autumn 2014 tweet histories for <em>all</em> of the users in the GS dataset. <em><strong>• Tweets: </strong>3,014,029 </em> <em><strong>• Unique Users: </strong>18,274 <strong>• </strong></em></p> <p>The file <strong>SG-Users_Autumn_2014_timelines.tsv </strong>contains the IDs for tweets in the <strong>final</strong> SG-Users dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/W17-4908/">StyleVar 2017</a> paper. This dataset consists only of tweets which contain at least one instance of one of 50 lexical variables that were the focus of the study, and has also undergone various other filtering steps; see the paper for full details. <em><strong>• Tweets: </strong>1,112,931</em> <em><strong>• Unique Users: </strong>10,103 <strong>• </strong></em></p> <p>The file <strong>IH-Users_Autumn_2014_timelines_pre-filtering.tsv </strong>contains the IDs for tweets which were posted in August, September, or October 2014 by users from the IT dataset, i.e. users we know to have used Indyref-related hashtags. This dataset was collected in the same manner as SG-Users_Autumn_2014_timelines_pre-filtering.tsv; however, due to an error in this process, <strong>the IDs of most of the tweets in this dataset were not recorded</strong>. For such tweets the tweet ID column instead contains a placeholder tweet ID of the form _<user_ID>_<month>_<integer>, where the integer denotes the tweet's position in the reverse-chronological list of tweets that were retrieved for that user from that month (e.g. _147527441_09_286 is the placeholder tweet ID we assigned to the 286th September tweet we retrieved from the user whose Account ID is 147527441). Unfortunately, therefore, the tweets in this file whose 'IDs' begin with an underscore cannot be straightforwardly re-downloaded using Twitter's free <a href="https://developer.twitter.com/en/docs/tweets/post-and-engage/api-reference/get-statuses-lookup">GET Statuses/Lookup</a> API endpoint; but since their user IDs and timestamps are intact, it would still be possible to retrieve them using the (paid-for) <a href="http:// https://developer.twitter.com/en/docs/tutorials/choosing-historical-api">Historical APIs</a>. <em><strong>• Tweets: </strong>6,997,858 <strong>• Tweets whose IDs were recorded: </strong>288,394</em><em><strong> </strong> <strong>• Unique Users: </strong>14,645</em><em> <strong>• </strong></em></p> <p>The file <strong>IH-Users_Autumn_2014_timelines.tsv </strong>contains the IDs for tweets in the <strong>final</strong> IH-Users dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/W17-4908/">StyleVar 2017</a> paper. Like the SG-Users dataset, this dataset consists only of tweets which contain at least one instance of one of 50 lexical variables that were the focus of the study, and has also undergone various other filtering steps; see the paper for full details. As with the pre-filtered version, most of the tweet IDs are unfortunately missing in this dataset. <em><strong>• Tweets: </strong>2,165,320 <strong>• Tweets whose IDs were recorded: </strong>115,366</em><em><strong> </strong> <strong>• Unique Users: </strong>10,784 <strong>• </strong></em></p> <p> </p> <p> </p> <p><strong>US Geotags: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Geotagged-USA.zip">Geotagged-USA.zip</a></strong></p> <p><em>Tweets from June 2013 - July 2016 which are geotagged to locations within the USA.</em></p> <p>The file <strong>GUSA.tsv</strong> contains all tweets from the ‘Spritzer’ sample which were posted between June 30th 2013 to July 1st 2016, are classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets, do not contain urls or embedded media, are not by users with more than 1000 friends or followers, and are geotagged to locations within the USA. This dataset (along with the GU Dataset) was used in our <a href="http://www.aclweb.org/anthology/W18-6101/">WNUT 2018</a> paper. <em><strong>• Tweets: </strong></em> <em>8,375,573 </em> <em><strong>• Unique Users: </strong>1</em>,<em>826,260</em><em> <strong>• </strong></em></p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.