RedMed: Extending drug lexicons for social media applications
<p>Data associated with the RedMed project.</p> <p>Details for the process behind the data creation can be found in the associated paper:</p> <p><strong>Lavertu, A. & Altman, R. B. </strong>"RedMed: Extending drug lexicons for social media applications"<br> Journal of Biomedical Informatics, (2019)</p> <p><a href="https://doi.org/10.1016/j.jbi.2019.103307">https://doi.org/10.1016/j.jbi.2019.103307</a></p> <p><strong>RedMed embedding model:</strong></p> <p>Word vectors trained on comments from health related subreddits and optimized for drug synonym retrieval.</p> <p>The Redmed model was train using only social media data from Reddit and achieves comparable performance on the UMNSRS and MayoSRS similarity tasks. Vectors are 64 dimensional.</p> <p><strong><strong>redmed_model_vectors.tsv.gz - </strong></strong>Tab-separated word vectors (token\tdim1\tdim2\t...dim64)</p> <p><strong><strong>redmed_model.bin - </strong></strong>Binary word2vec file saved using gensim, can be loaded into python gensim</p> <p>Other Files:</p> <p>supp_file_1_sidebar_subreddits.txt - List of health-related subreddits based on "r/Health" and "r/Drugs" sidebars<br> supp_file_2_enrichment_based_subreddits.txt - List of health-related subreddits based on amount of health-related content<br> supp_file_3_custom_stopword_list.txt - List of stopwords based on counts derived from Reddit comments</p>
ShareScore
44/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 8
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 0