Skip to main content
zenodoopen

RedMed: Extending drug lexicons for social media applications

<p>Data associated with the RedMed project.</p> <p>Details for the process behind the data creation can be found in the associated paper:</p> <p><strong>Lavertu, A. &amp; Altman, R. B. </strong>&quot;RedMed: Extending drug lexicons for social media applications&quot;<br> Journal of Biomedical Informatics, (2019)</p> <p><a href="https://doi.org/10.1016/j.jbi.2019.103307">https://doi.org/10.1016/j.jbi.2019.103307</a></p> <p><strong>RedMed embedding model:</strong></p> <p>Word vectors trained on comments from health related subreddits and optimized for drug synonym retrieval.</p> <p>The Redmed model was train using only social media data from Reddit and achieves comparable performance on the UMNSRS and MayoSRS similarity tasks. Vectors are 64 dimensional.</p> <p><strong><strong>redmed_model_vectors.tsv.gz - </strong></strong>Tab-separated word vectors (token\tdim1\tdim2\t...dim64)</p> <p><strong><strong>redmed_model.bin - </strong></strong>Binary word2vec file saved using gensim, can be loaded into python gensim</p> <p>Other Files:</p> <p>supp_file_1_sidebar_subreddits.txt - List of health-related subreddits based on &quot;r/Health&quot; and &quot;r/Drugs&quot; sidebars<br> supp_file_2_enrichment_based_subreddits.txt - List of health-related subreddits based on amount of health-related content<br> supp_file_3_custom_stopword_list.txt - List of stopwords based on counts derived from Reddit comments</p>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
8
Access
20
Reuse readiness
8
Engagement
0

Topics