Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
296
datasets available to search
ShareScore release 0.9.0
Dataset results
296 results for “language models”
Automated Programming Exercise Generation in the Era of Large Language Models
<p>Lecturers are increasingly attempting to use large language models (LLMs) to simplify and make the creation of exercises for students more efficient. Efforts are also being made to automate the exercise creation process in software engineering (SE) education. This study explores the use of advanced LLMs, including GPT-4 and LaMDA, for automated programming exercise creation in higher education and compares the results with related work using GPT-3.5-turbo. Utilizing applications such as ChatGPT, Bing AI Chat, and Google Bard, we identify LLMs capable of initiating different exercise designs. However, manual refinement is crucial for accuracy. Common error patterns across LLMs highlight challenges in complex programming concepts, while specific strengths in various topics showcase model distinctions. This research underscores LLMs' value in exercise generation, emphasizing the critical role of human supervision in refining these processes. Our concise insights cater to educators, practitioners, and other researchers seeking to enhance SE education through LLM applications.</p>
Specification-Driven Code Translation By Large Language Models: How Far Are We?
<p>The artifacts and dataset for "Specification-Driven Code Translation By Large Language Models: How Far Are We?"</p>
Replication Package of "Exploiting Vision-Language Models in GUI Reuse"
<p>This replication package is for the paper entitled "Exploiting Vision-Language Models in GUI Reuse". The authors remain anonymous for double-blind review purposes. The package contains six files. If the paper is accepted, then the authors will move the replication package to a public repository hosted by an institution.</p>
USPTO-LLM: A Large Language Model-Assisted Information-enriched Chemical Reaction Dataset
<p>USPTO-LLM is an <strong>information-enriched chemical reaction dataset</strong> that provides more side information (reaction conditions and reaction steps division) for developing new reaction prediction and retrosynthesis methods and inspires new problems, such as reaction condition prediction. It comprises over <strong>247K chemical reactions</strong> extracted from the patent documents of USPTO (United States Patent and Trademark Office), encompassing abundant information on reaction conditions. </p> <p>We employ large language models to expedite the data collection procedures automatically with a reliable quality control process. The extracted chemical reactions are organized as <strong>heterogeneous directed graphs</strong>, allowing us to formulate a series of prediction tasks, such as reaction prediction, retrosynthesis, and reaction condition prediction, in a unified graph-filling framework.</p>
Paper information in the topic of large language models
<p>This dataset supports the findings in the preprint 'Academic collaboration on large language model studies increases overall but varies across disciplines.' The study aims to explore the application of large language models (LLMs) in scientific disciplines and their implications for interdisciplinary collaboration.</p> <p>To build LLM paper group, we start with a broad search using general terms related to LLMs and popular models based on the MMLU benchmark spanning from October 2018 to September 2024. We apply this search to the title and abstract to avoid excessive noise in the dataset and then undergo a series of filtering steps<br>to enhance relevance and remove duplicates. The resulting dataset contains 59,293 papers.</p> <p>In addition to the paper group in the topic of LLMs, we establish two control groups. The first control group focuses on machine learning (ML) papers. We select ML as a control because it is a well-established field from which LLM emerged as a subfield. To construct this group, we collect a random sampling of 70,945 papers containing the phrase ''machine learning'' in either their title or abstract. To provide an even broader perspective beyond AI-related fields, we create a second control group consisting of a random sample of 73,110 papers from all other research categories---specifically, papers that belong neither to the ML nor LLM categories. </p> <p>The three files below contain the cleaned samples collected from OpenAlex, which are derived from the original files. </p> <ul> <li>LLM: llm-cleaned-samples.csv</li> <li>ML: ml-cleaned-samples.csv</li> <li>Non-LLM/ML: non-llm-cleaned-samples.csv</li> </ul> <p>The three zip files below contain author affiliation information (including departmental discipline) extracted by GPT-4o-mini to support the departmental analysis in the paper:</p> <ul> <li>LLM: llm-author-affiliations.zip</li> <li>ML: ml-author-affiliations.zip</li> <li>Non-LLM/ML: non-llm-author-affiliations.zip</li> </ul> <p>The three files below contain the paper information used to support all the analysis in our paper:</p> <ul> <li>LLM: llm-information-entropy.csv</li> <li>ML: ml-information-entropy.csv</li> <li>Non-LLM/ML: non-llm-information-entropy.csv</li> </ul> <p>If you have any additional questions, please feel free to contact <a rel="noreferrer">lingyaol@umich.edu or lydinh@usf.edu.</a></p>
Distinguishing GUI Component States for Blind Users using Large Language Models
<p><strong># Data Code Repository</strong></p><p> </p><p>This repository contains open-source data code that provides utilities for the paper named "Here comes trouble! Distinguishing GUI Component States for Blind Users using Large Language Models". The code is designed to facilitate data-related tasks and promote reproducibility in research and data analysis projects.</p><p> </p><p><strong>## Features</strong></p><p> </p><p>- Attribute identification and extraction: Including real-time recognition and extraction of GUI components in the view type, resource-id, color, action of four attributes</p><p>- Components State Distinction: Provides the prompt needed for large language models, covering their specific design schemes and chain of thought reasoning processes as well as contextual learning content.</p><p>- Implementation: Offers specific methods to realize the process, including the setting of relevant parameters and the use of functions.</p><p> </p><p><strong>## Installation</strong></p><p> </p><p>To use the data code, you can down or clone the required code.</p><p>Notably, before using the code, make sure the necessary environment configuration is done.</p><p> </p><p><strong>## Dependencies</strong></p><p>The data code has the following dependencies:</p><p> </p><p>Python (version 3.6 or higher)</p><p>NumPy</p><p>Pandas</p><p>Seaborn</p><p>Scikit-learn</p><p>Openai</p><p>Android Studio (version 4.0)</p><p> </p><p>Install the required dependencies using pip:</p><p>pip install numpy..</p><p> </p><p><strong>##License</strong></p><p>This data code is distributed under the MIT License. See LICENSE for more information.</p><p> </p><p><strong>##Copyright</strong></p><p>All copyright of the tool is owned by the author of the paper.</p>
BioVAE: a pre-trained latent variable language model for biomedical text mining
<p>We release BioVAE, the first large-scale pre-trained latent variable language model for the biomedical domain, which uses the OPTIMUS framework to train on large volumes of biomedical text.</p> <p>This version contains the pre-trained models for text mining tasks such as named entity recognition or relation extraction, and text generation task.</p> <p>Explanation of each file: (lt32: latent_size = 32, beta05: beta=0.5)</p> <ul> <li>pm-full-lt32-beta00</li> <li>pm-full-lt32-beta05</li> <li>pm-full-lt768-beta00</li> <li>pm-full-lt768-beta05</li> <li>pm-full-generation</li> </ul>
Protein language model embeddings and predictions for the fly proteome (FlyBase)
<p>Residue and sequence embeddings of the fly (drosophila melanogaster) proteome (FlyBase for organism drosophila melanogaster, downloaded on 2022.03.01) computed using bio_embeddings (bioembeddings.com) using the ProtT5 embedder at full precision (https://www.biorxiv.org/content/10.1101/2020.07.12.199554v3). To open the embeddings file, please see <a href="https://github.com/sacdallago/bio_embeddings/blob/develop/notebooks/open_embedding_file.ipynb">this notebook</a>. The embeddings will be indexed by numbers according to the mapping file (mapping_file.csv) in this dataset. All following results will share the same mapping (for instance, to access the variation prediction results, by accessing index "0", you will query results for the sequence "FBpp0304622").</p> <p>Additionally:</p> <p>- Sequence-level predictions of subcellular localization in 10 classes using LA (https://www.biorxiv.org/content/10.1101/2021.04.25.441334v1)</p> <p>- Residue-level three state secondary structure prediction (alpha, sheet or other) using models reported in the ProtTrans paper (https://www.biorxiv.org/content/10.1101/2020.07.12.199554v3)</p> <p>- Residue-level prediction of conservation (in 9 states) and of variation effect (from 0 [no-effect] to 1 [effect]) using VESPAl (https://doi.org/10.1007/s00439-021-02411-y)</p> <p> </p> <p>Files included:</p> <p>- dmel-all-translation-r6.44.fasta --> FASTA-formatted sequences of drosophila melanogaster from FlyBase</p> <p>- mapping_file.csv --> A CSV file mapping the identifiers used in the following files (from 0 to 30737) to the identifiers in the FlyBase fasta file (dmel-all-translation-r6.44.fasta).</p> <p>- DSSP3_fly_ProtT5Sec.fasta --> Secondary structure predictions in three states for each residue of each protein in dmel-all-translation-r6.44.fasta. "H" stands for Helix; "E" stands for Sheet; "C" stands for Other.</p> <p>- subcell_fly_LA_ProtT5.csv --> Subcellular location (10 states) and memrane-boundness (2 states) for each protein in dmel-all-translation-r6.44.fasta</p> <p>- embeddings_file.h5 --> per-residue embeddings of sequences in dmel-all-translation-r6.44.fasta. Each dataset in the .h5 file represents a protein sequence and contains a matrix of length Lx1024, with L being the length of the protein sequence. Datasets are indexed using integers. The original sequence identifier (from the FASTA header) can be accessed through the "original_id" attribute. See https://docs.bioembeddings.com/v0.2.0/notebooks/open_embedding_file.html for information on how to open the file.</p> <p>- reduced_embeddings_file.h5 --> per-sequence embeddings of sequences in dmel-all-translation-r6.44.fasta (obtained by mean-pooling the residue-embeddings along the length dimension of the protein sequence). Each dataset in the .h5 file represents a protein sequence and contains a vector of size 1024 (meaning, each sequence has the same dimension).</p> <p>- conspred_probs.h5 --> per-sequence conservation probability (softmax) prediction of sequences in dmel-all-translation-r6.44.fasta in 9 classes. Each dataset in the .h5 file represents a protein sequence and contains a matrix of length 9xL, with L being the length of the protein sequence, and 9 being the predicted conservation class (index 0 = very variable; index 8 = very conserved)</p> <p>- vespal_SAVeffect_fly.zip --> zipped .h5 file of per-sequence variation predictions of sequences in dmel-all-translation-r6.44.fasta on a scale from 0 (neutral) to 1 (effect). -1 indicates WT substitution. Each dataset in the .h5 file represents a protein sequence and contains a matrix of length 20xL, with L being the length of the protein sequence, and 20 being the predicted variation score for each residue substitution (AAs in the following order: "<strong>ALGVSREDTIPKFQNYMHWC</strong>" . Meaning that index 0 = substitution of the residue to "A", index = 1 substitution to residue "L", aso.)</p>
Custom language model checkpoints used in "Testing the limits of natural language models for predicting human language judgments"
<p>Checkpoint files for an RNN, LSTM, BILSTM and n-gram models used the paper "Testing the limits of natural language models for predicting human language judgments"</p>
Trained Models from "General Cross-Architecture Distillation of Pretrained Language Models into Matrix Embeddings"
<p>Trained models from the paper:</p> <p>Lukas Galke, Isabell Cuber, Christoph Meyer, Henrik Ferdinand Noelscher, Angelina Sonderecker, and Ansgar Scherp: <strong>General Cross-Architecture Distillation of Pretrained Language Models into Matrix Embeddings</strong>, in: <em>International Joint Conference on Neural Networks (IJCNN), </em>2022.</p> <ul> <li>File seq2mat_hybrid_bidirectional_sbertlike-100p-bsz512 holds the model from pretraining</li> <li>File ws2020_transformer_final_models holds the fine-tuned models for each task of the GLUE benchmark</li> </ul>
Extract from the PEAPL framework : Modelling "Writing sentences" competency (French as the schooling language)
<p>Competences, skills and knowledges that make up "writing sentences" competency, based on the linguistic praxeological organization of French as the schooling language (https://doi.org/10.5281/zenodo.4001381). This is an extract of the general framework, some of the visible objects are linked to other objects in other main competences. Orange links show how pedagogic ressources (game levels) are linked to framework objects.</p> <p>This framework is used for the PEAPL (peapl.eu) project to link activities in the GamesHub platform (for example activities using "L'Orthodyssée des Gram" grammar online game, available at https://www.lafamillegram.ch/#)</p> <p> </p> <p> </p> <p> </p>
Prediction and Visualization of Human Transmembrane Proteins using AlphaFold and Protein Language Models
<p><strong>Description:</strong> <strong>TMvis</strong> ("TMvis496.tar.gz") is a dataset containing 496 3D-structures of predicted human transmembrane proteins (TMP) and their predicted membrane embedding. The method TMbed [1], based on the protein language model ProtT5 [2] predicted 4.967 TMP for the human proteome (20,375 proteins, UniProt [3] version April 2022; excluding TITIN_HUMAN due to length). For these proteins, we obtained AlphaFold [4] structures from AlphaFoldDB [5] with an average per-residue confidence score (pLDDT) of more than 90%. This resulted in the 496 proteins of TMvis, as can be found in "TMvis496.fasta". The membrane embedding was predicted using the methods ANVIL [6], PPM3 [7], and per-residue TMbed predictions. As the three methods are based on different approaches, we decided to publish results for all. The figure “TMvis_project_overview.png” provides a graphical overview for each step described above.</p> <p><strong>TMvis Folder Structure:</strong> TMvis is separated into “alpha” containing predicted alpha-helical TMPs, and “beta” containing predicted beta-barrel TMPs. Within these folders, each protein is assigned one folder, identifiable by the respective unique UniProt ID. Each protein folder consists of:<br> - “UniprotID.fasta” with UniProt ID, sequence, TMbed per-residue prediction<br> - “AF-UniprotID-F1-model_v2.pdb” with the AlphaFold structure<br> - “AF-UniprotID-F1-model_v2.cif” with the AlphaFold structure<br> - “AF-UniprotID-F1-model_v2_ANVIL.pdb” with predicted ANVIL membrane embedding<br> - “AF-UniprotID-F1-model_v2_ppm.pdb” predicted PPM3 membrane embedding</p> <p>TMvis <br> | <br> ├── alpha <br> │ │ <br> │ ├── A0A087X1C5 <br> │ │ ├── A0A087X1C5.fasta <br> │ │ ├── AF-A0A087X1C5-F1-model_v2.pdb <br> │ │ ├── AF-A0A087X1C5-F1-model_v2.cif <br> │ │ ├── AF-A0A087X1C5-F1-model_v2_ANVIL.pdb <br> │ │ └── AF-A0A087X1C5-F1-model_v2_ppm.PDB <br> │ └── ... <br> └── beta <br> └── P45880</p> <p><strong>TMvis visualization:</strong> The 3D-visualization of every protein in the dataset TMvis can be easily accessed using the Jupyter Notebook “TMvis.ipynb”. It contains detailed descriptions the different membrane prediction tools ANVIL, PPM3, and TMbed as well as the respective code. Additionally, it allows to visualize the per-residue confidence scores (pLDDT) of AlphaFold.</p> <p>——————————————————————————————————————————————————————————————————————————</p> <p><strong>References:</strong></p> <p>[1] TMbed - TMbed Bernhofer, Michael, and Burkhard Rost. 2022. “TMbed – Transmembrane Proteins Predicted through Language Model Embeddings.” bioRxiv.</p> <p>[2] ProtT5 - A. Elnaggar et al., "ProtTrans: Towards Cracking the Language of Lifes Code Through Self-Supervised Deep Learning and High Performance Computing," in IEEE Transactions on Pattern Analysis and Machine Intelligence, doi: 10.1109/TPAMI.2021.3095381.</p> <p>[3] UniProt - UniProt Consortium (2021). UniProt: the universal protein knowledgebase in 2021. Nucleic acids research, 49(D1), D480–D489.</p> <p>[4] AlphaFold - AlphaFold Jumper, John, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, et al. 2021. “Highly Accurate Protein Structure Prediction with AlphaFold.” Nature 596 (7873): 583–89.</p> <p>[5] Alphafold DB - Varadi, Mihaly, Stephen Anyango, Mandar Deshpande, Sreenath Nair, Cindy Natassia, Galabina Yordanova, David Yuan, et al. 2022. “AlphaFold Protein Structure Database: Massively Expanding the Structural Coverage of Protein-Sequence Space with High-Accuracy Models.” Nucleic Acids Research 50 (D1): D439–44.</p> <p>[6] ANVIL - ANVIL Postic, Guillaume, Yassine Ghouzam, Vincent Guiraud, and Jean-Christophe Gelly. 2016. “Membrane Positioning for High- and Low-Resolution Protein Structures through a Binary Classification Approach.” Protein Engineering, Design & Selection: PEDS 29 (3): 87–91.</p> <p>[7] PPM3 - PPM3 Lomize, Mikhail A., Irina D. Pogozheva, Hyeon Joo, Henry I. Mosberg, and Andrei L. Lomize. 2012. “OPM Database and PPM Web Server: Resources for Positioning of Proteins in Membranes.” Nucleic Acids Research 40 (Database issue): D370–76.</p> <p>——————————————————————————————————————————————————————————————————————————</p> <p><strong>License:</strong></p> <p>This work is licensed under a Creative Commons Attribution 4.0 International License (CC-BY 4.0).</p> <p> </p>
Exploiting Pretrained Biochemical Language Models for Targeted Drug Design
<p>This repository contains materials for the paper,<em> Exploiting Pretrained Biochemical Language Models for Targeted Drug Design, </em>which<em> </em>has been accepted for publication in <em>Bioinformatics</em> Published by Oxford University Press.</p> <p><em>data.zip</em> contains vocabulary files for the pretrained models, additional information regarding proteins (PFAM family, protein similarity) and interactions filtered from <a href="https://www.bindingdb.org/bind/index.jsp">BindingDB</a> which are further split into train, validation and test sets and used to train target specific molecule generation models. </p> <p><em>models.zip </em>includes files for the models trained in this study. </p> <p><em>predictions.zip </em>comprises the compounds generated with the targeted models and the result of their evaluation with respect to benchmarking metrics. </p> <p><em>docking.zip </em>contains <em>targets/ </em>including PDB files of the test proteins selected for docking evaluation, <em>ligands/ </em>including SDF files for molecules generated with the targeted models and two decoding strategies (i.e. beam search and sampling) and <em>complex/ </em>including docking outputs. </p> <p> </p> <p> </p>
CNN for Modeling Sanskrit Originated Bengali and Hindi Language Dataset
<p>Though recent works have focused on modeling high resource languages, the area is still unexplored for low resource languages like Bengali and Hindi. We propose an end-to-end trainable memory efficient CNN architecture named CoCNN to handle specific characteristics such as high inflection, morphological richness, flexible word order and phonetical spelling errors of Bengali and Hindi. In particular, we introduce two learnable convolutional sub-models at word and at sentence level that are end-to-end trainable. We show that state-of-the-art (SOTA) Transformer models including pretrained BERT do not necessarily yield the best performance for Bengali and Hindi. CoCNN outperforms pretrained BERT with 16X less parameters and achieves much better performance than SOTA LSTMs on multiple real-world datasets. This is the first study on the effectiveness of different architectures from Convolution, Recurrent, and Transformer neural net paradigm for modeling Bengali and Hindi.</p>
Code and Dataset for "Examining Zero-Shot Vulnerability Repair with Large Language Models"
<p><strong>Code and Dataset for "Examining Zero-Shot Vulnerability Repair with Large Language Models"</strong></p> <p>The following Zenodo contains the resources associated with the S&P accepted paper ‘Examining Zero Shot Vulnerability Repair with Large Language Models’, https://arxiv.org/abs/2112.02125</p> <p>In this resource, you can find the following.</p> <p> - 'important_results' directory:<br> This directory is for containing the final raw results as generated by the framework, including a global CSV of all generations and an HTML file containing all of the diffs generated for the 'high-confidence' real-world scenarios.<br> - final_results.csv<br> - This contains the final results of all generated software patches.<br> - Note the nomenclature differences with the manuscript tables. These are explained in the README in the framework.<br> - Original vs LLM-Generated Vulnerability Fixes.html<br> - This contains all of the diffs for the real-world patches versus the canonical developer-provided patches.</p> <p> - 'framework' directory:<br> This directory contains the complete archive of the code framework and all results at the time of the paper’s submission. It contains every language model prompt, suggestions, assembled repair patch and analysis data. It contains every script used for generation and analysis. It is a large archive, and within it contains an included README describing how to understand and use it.<br> - For convenience, we include a copy of the README external to the zipped archive.</p> <p> - 'resources' directory: <br> This directory contains the resources used by our large associated tools, including:<br> - 'gpt2-csrc' subdirectory:<br> Everything to do with the gpt2-csrc model, including the trained files, training scripts, and training data.</p> <p> - 'ExtractFix' subdirectory:<br> A docker image containing all ExtractFix scenarios, even those we did not use. Provided for interest (not required for usage).</p>
Data for 'VespaG: Expert-guided protein language models enable accurate and blazingly fast fitness prediction'
<div>Datasets used for development of VespaG and VespaG predictions generated with <a href="https://github.com/JSchlensok/VespaG">https://github.com/JSchlensok/VespaG</a>. </div> <div> </div> <div>Uploads contain:</div> <div> <ol> <li><strong>Performance</strong> summaries for ProteinGym [1]:<br>- Spearman and Pearson correlation for VespaG: <em>proteingym_performance_vespag.csv </em>(columns: 'DMS_id', 'Spearman', 'Pearson')<br>- Spearman correlation for evaluated methods VespaG, GEMME [2], VESPA [3], TranceptEVE [4], AlphaMissense [5], PoET [6]: <em>proteingym_spearman_allmethods.csv </em>(columns: 'DMS_id', 'Trancept EVE-L', 'VESPA', 'VespaG', 'GEMME', 'AlphaMissense', 'PoET', 'UniProt_ID', 'coarse_selection_type' (function), 'taxon')</li> <li><strong>Fasta</strong> files with sequences for all train sets (<em>vespag_fasta_training_datasets.zip</em> with seq_all9k.fasta, seq_human5k.fasta, seq_droso4k.fasta, seq_ecoli2k.fasta, seq_virus1k.fasta) and test set (<em>proteingym_217.fasta</em>)</li> <li><strong>VespaG</strong> <strong>Predictions</strong> for test set: <em>vespag_proteingym_rawpreds_by_training_dataset.zip</em> with raw_preds_ecoli.csv, raw_preds_human.csv, raw_preds_virus.csv, raw_preds_all.csv, raw_preds_droso.csv (columns: 'DMS_id', 'mutation', 'DMS_score', 'VespaG'). Predictions are based on different training data, the final model VespaG was trained on a subset of the human proteome and <strong>raw VespaG predictions</strong> <strong>for</strong> <strong>the</strong> <strong>ProteinGym benchmark are in raw_preds_human.csv </strong>(used to calculate the performances above).</li> <li><strong>GEMME predictions</strong> for train sets: <em>vespag_proteingym_rawpreds_by_training_dataset.zip </em>with folders 'human', 'droso', 'ecoli', 'virus', 'all' for respective fasta file (each containing GEMME mutational landscape output files named '<em>ID' + '</em>_normPred_evolCombi.txt')</li> <li><strong>ESM-2</strong> <strong>embeddings</strong> [7] for test set (<em>proteingym_217_esm2.h5</em>)</li> </ol> </div> <div>For details on VespaG see:</div> <div> <div> <div>VespaG: Expert-guided protein Language Models enable accurate and blazingly fast fitness prediction</div> </div> <div>Celine Marquet, Julius Schlensok, Marina Abakarova, Burkhard Rost, Elodie Laine</div> <div>bioRxiv 2024.04.24.590982; doi: https://doi.org/10.1101/2024.04.24.590982</div> <div> </div> <div>For more information on data usage and generation please see <a href="https://github.com/JSchlensok/VespaG">https://github.com/JSchlensok/VespaG</a>.</div> <div> </div> <div>Abstract:</div> <div>Exhaustive experimental annotation of the effect of all known protein variants remains daunting and expensive, stressing the need for scalable effect predictions. We introduce VespaG, a blazingly fast single amino acid variant effect predictor, leveraging embeddings of protein Language Models as input to a minimal deep learning model. To overcome the sparsity of experimental training data, we created a dataset of 39 million single amino acid variants from the human proteome applying the multiple sequence alignment-based effect predictor GEMME as a pseudo standard-of-truth. Assessed against the ProteinGym Substitution Benchmark (217 multiplex assays of variant effect with 2.5 million variants), VespaG achieved a mean Spearman correlation of 0.48 +/- 0.01, matching state-of-the-art methods such as GEMME, TranceptEVE, PoET, AlphaMissense, and VESPA. VespaG reached its top-level performance several orders of magnitude faster, predicting all mutational landscapes of the human proteome in 30 minutes on a consumer laptop (12-core CPU, 16 GB RAM).</div> <div> </div> <div>[1] Notin, Pascal, et al. "ProteinGym: large-scale benchmarks for protein fitness prediction and design." <em>Advances in Neural Information Processing Systems</em> 36 (2024).<br>[2] Laine, Elodie, Yasaman Karami, and Alessandra Carbone. "GEMME: a simple and fast global epistatic model predicting mutational effects." <em>Molecular biology and evolution</em> 36.11 (2019): 2604-2619.</div> <div>[3] Marquet, Céline, et al. "Embeddings from protein language models predict conservation and variant effects." <em>Human genetics</em> 141.10 (2022): 1629-1647.</div> <div>[4] Notin, Pascal, et al. "TranceptEVE: Combining family-specific and family-agnostic models of protein sequences for improved fitness prediction." <em>bioRxiv</em> (2022): 2022-12.</div> <div>[5] Cheng, Jun, et al. "Accurate proteome-wide missense variant effect prediction with AlphaMissense." <em>Science</em> 381.6664 (2023): eadg7492.</div> <div>[6] Truong Jr, Timothy, and Tristan Bepler. "PoET: A generative model of protein families as sequences-of-sequences." <em>Advances in Neural Information Processing Systems</em> 36 (2024).</div> <div>[7] Lin, Zeming, et al. "Evolutionary-scale prediction of atomic-level protein structure with a language model." <em>Science</em>379.6637 (2023): 1123-1130.</div> </div>
Data and code from: Learning a deep language model for microbiomes: The power of large scale unlabeled microbiome data
<p>We use open source human gut microbiome data to learn a microbial "language" model by adapting techniques from Natural Language Processing (NLP). Our microbial "language" model is trained in a self-supervised fashion (i.e., without additional external labels) to capture the interactions among different microbial species and the common compositional patterns in microbial communities. The learned model produces contextualized taxa representations that allow a single bacteria species to be represented differently according to the specific microbial environment it appears in. The model further provides a sample representation by collectively interpreting different bacteria species in the sample and their interactions as a whole. We show that, compared to baseline representations, our sample representation consistently leads to improved performance for multiple prediction tasks including predicting Irritable Bowel Disease (IBD) and diet patterns. Coupled with a simple ensemble strategy, it produces a highly robust IBD prediction model that generalizes well to microbiome data independently collected from different populations with substantial distribution shift.</p> <p>We visualize the contextualized taxa representations and find that they exhibit meaningful phylum-level structure, despite never exposing the model to such a signal. Finally, we apply an interpretation method to highlight bacterial species that are particularly influential in driving our model's predictions for IBD.</p>
CausalBench A Comprehensive Benchmark for Evaluating Causal Reasoning Capabilities of Large Language Models
<p>CausalBench is a comprehensive benchmark dataset designed to evaluate the causal reasoning capabilities of large language models. The primary uses of this dataset include, but are not limited to:</p> <p>- Testing the performance of large language models on causal reasoning tasks</p> <p>- Serving as a benchmark dataset for causal reasoning research</p> <p>- Improving and developing new causal reasoning algorithms and models</p>
Transformers Model Zoos and Soups: A Population of Language and Vision Models
<p>Model Zoos submitted to the NeurIPS 2024 Dataset & Benchmark track: "<em>Transformer Model Zoos and Soups: A Population of Language and Vision Models</em>"</p> <p>We generate two model zoos, one for computer vision built on the ViT-S architecture, and one for language modeling based on the BERT architecture. For each, we train several backbone models with varying hyperparameters, and further fine-tune them using multiple hyperparameter combinations. We further annotate every model with performance metrics. These include test accuracy and F1-score, as well as the generalization gap. For the vision models, we also include the robust accuracy after a FGSM attack.</p>
Lost in Translation? Not for Large Language Models: Automated Divergent Thinking Scoring Performance Translates to Non-English Contexts (Datasets)
<p>Datasets for: Zielińska, A., Organisciak, P., Dumas, D., & Karwowski, M. (2023). Lost in translation? Not for large language models: Automated divergent thinking scoring performance translates to non-English contexts. <em>Thinking Skills and Creativity, 50</em>, 101414. https://doi.org/10.1016/j.tsc.2023.101414</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.