Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13,446
datasets available to search
ShareScore release 0.9.0
Dataset results
13,446 results for “targeted”
GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data
<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>
Resources for Mitigating Chemotherapy Side Effects through Targeted Gamma-Ray Delivery and CNNs
<p>This repository includes datasets and code used in the study "Mitigating Chemotherapy Side Effects through Targeted Gamma-Ray Delivery and Convolutional Neural Networks." The resources comprise:<br>- Binding Affinity Data: Used for simulations.<br>- Brain Tumor MRI and Chest CT Scan Datasets: Used for model training.<br>- Lightweight Deep CNN: Code for building and testing models.</p>
Data associated with the article 'Intervention factors associated with efficacy, when targeting oral language comprehension of children with or at risk for (Developmental) Language Disorder: A meta-analysis'
<p>The efficacy of oral language comprehension interventions varies, but the reasons for this variation have received little attention. A meta-analysis was conducted to examine intervention factors associated with the efficacy (as expressed with effect sizes) of oral language comprehension interventions in children under the age of 18 with or at risk for (Developmental) Language Disorder, (D)LD.</p> <p>The meta-analysis article together with this additional material comprise the content needed for a thorough understanding and replication of the results.</p> <p>This dataset is based on two systematic scoping reviews on oral language comprehension interventions (Tarvainen et al., 2020, 2021). Further information from the sourced articles was extracted for this study titled ‘Intervention factors associated with efficacy, when targeting oral language comprehension of children with or at risk for (Developmental) Language Disorder: A meta-analysis’. </p> <p>In the future, we hope that this data is used with a growing body of oral language comprehension interventions to conduct further and more detailed examinations of intervention factors associated with efficacy.</p> <p>References:</p> <p>Tarvainen, S., Launonen, K., & Stolt, S. (2021). Oral language comprehension interventions in school-age children and adolescents with developmental language disorder: A systematic scoping review. <em>Autism & Developmental Language Impairments</em>, <em>6</em>, 1–24. https://doi.org/10.1177/23969415211010423</p> <p>Tarvainen, S., Stolt, S., & Launonen, K. (2020). Oral language comprehension interventions in 1–8-year-old children with language disorders or difficulties: A systematic scoping review. <em>Autism & Developmental Language Impairments</em>, <em>5</em>, 1–24. https://doi.org/10.1177/2396941520946</p> <p> </p>
Cyberhate that targets people who are plus-size in the news: The role of bystanders in mitigating social pathologies (CYBERPLUS)
<p>The dataset was created for the project "Cyberhate that targets people who are plus-size in the news: The role of bystanders in mitigating social pathologies (CYBERPLUS)". The data was collected between July 12 and July 26, 2024, from 1,030 young Czech people aged 16-25. The survey asked young people about their sociodemographic information, attitudes toward and perceptions of entitativity of three groups (overweight people, underweight people, people with physical disabilities), group identification, bystander appraisals and behavioural intentions, hate speech perception, and internet use. It included an experimental part in which the participants were exposed as bystanders to social media news posts about overweight people and comments under the posts. The dataset is accompanied by a data dictionary and a technical report.</p>
BALTRAD_VPTS - Vertical profiles of biological targets derived from European weather radars
<p><em>BALTRAD_VPTS - Vertical profiles of biological targets derived from European weather radars</em> is a vertical profile time series dataset published by the <a href="https://www.inbo.be/en">Research Institute for Nature and Forest (INBO)</a>. It contains animal movement data derived from 151 European weather radars in 18 countries, with varying coverage from 2012 to 2023. These data were created by processing weather radar data - provided by the Operational Programme for the Exchange of Weather Radar Information (<a href="https://www.eumetnet.eu/activities/observations-programme/current-activities/opera/">OPERA</a>) - with methods optimized for extracting bird targets. The resulting data are vertical profile time series (VPTS), containing the density, speed and direction of biological targets within a weather radar (<code>radar</code>) volume, grouped into altitude bins (<code>height</code>) and measured over time (<code>datetime</code>). The data are also available in the <a href="https://aloftdata.eu/browse/?prefix=baltrad/">Aloft bucket</a>.</p> <div> <div>See Desmet et al. (2025, <a href="https://doi.org/10.1038/s41597-025-04641-5">https://doi.org/10.1038/s41597-025-04641-5</a>) for a more detailed description of this dataset.</div> </div> <h2>Files</h2> <p>VPTS data in this deposit are organized per country (.tgz file), radar (directory), year (directory) and month (.csv.gz file). Fields in the data follow the <a href="https://aloftdata.eu/vpts-csv/">VPTS CSV</a> format and are described in <code>vpts-csv-table-schema.json</code>. An overview of what data are available is provided in <code>coverage.csv</code>. Radar metadata can be found at <a href="https://aloftdata.eu/radars/">https://aloftdata.eu/radars/</a>.</p> <ul> <li><strong>coverage.csv</strong>: coverage of the VPTS data, representing the number of unique hours, heights, source files and records for each radar and date combination.</li> <li><strong>vpts-csv-table-schema.json</strong>: technical description of the fields in the VPTS data.</li> <li><strong>be.tgz</strong>: VPTS data from 2 radars in Belgium.</li> <li><strong>ch.tgz</strong>: VPTS data from 5 radars in Switzerland.</li> <li><strong>cz.tgz</strong>: VPTS data from 2 radars in Czechia.</li> <li><strong>de.tgz</strong>: VPTS data from 20 radars in Germany.</li> <li><strong>dk.tgz</strong>: VPTS data from 5 radars in Denmark.</li> <li><strong>ee.tgz</strong>: VPTS data from 2 radars in Estonia.</li> <li><strong>es.tgz</strong>: VPTS data from 15 radars in Spain.</li> <li><strong>fi.tgz</strong>: VPTS data from 13 radars in Finland.</li> <li><strong>fr.tgz</strong>: VPTS data from 26 radars in France.</li> <li><strong>hr.tgz</strong>: VPTS data from 7 radars in Croatia.</li> <li><strong>il.tgz</strong>: VPTS data from 1 radar in Israel.</li> <li><strong>nl.tgz</strong>: VPTS data from 3 radars in the Netherlands.</li> <li><strong>no.tgz</strong>: VPTS data from 11 radars in Norway.</li> <li><strong>pl.tgz</strong>: VPTS data from 8 radars in Poland.</li> <li><strong>pt.tgz</strong>: VPTS data from 3 radars in Portugal.</li> <li><strong>se.tgz</strong>: VPTS data from 22 radars in Sweden.</li> <li><strong>si.tgz</strong>: VPTS data from 2 radars in Slovenia.</li> <li><strong>sk.tgz</strong>: VPTS data from 4 radars in Slovakia.</li> </ul> <h2>Acknowledgements</h2> <p>This dataset was processed using infrastructure provided by the University of Amsterdam, SURF Cooperative, Ghent University and the Research Institute for Nature and Forest (INBO). It was mainly supported by the <a href="https://globam.science/">GloBAM project</a>, funded through the 2017-18 Belmont Forum and BiodivERsA joint call for research proposals under the BiodivScen ERA-Net COFUND programme.</p>
UVA_VPTS - Vertical profiles of biological targets derived from weather radars in Belgium, Germany and the Netherlands
<p><em>UVA_VPTS - Vertical profiles of biological targets derived from weather radars in Belgium, Germany and the Netherlands</em> is a vertical profile time series dataset published by the <a href="https://www.inbo.be/en">Research Institute for Nature and Forest (INBO)</a>. It contains animal movement data derived from 24 weather radars in Belgium, Germany and the Netherlands, with varying coverage from 2008 to 2023. These data were created by processing weather radar data - provided by the Royal Meteorological Institute of Belgium (<a href="https://www.meteo.be/">RMI</a>), German Meteorological Service (<a href="https://www.dwd.de/">DWD</a>) and Royal Netherlands Meteorological Institute (<a href="https://www.knmi.nl/">KMNI</a>) - with methods optimized for extracting bird targets. The resulting data are vertical profile time series (VPTS), containing the density, speed and direction of biological targets within a weather radar (<code>radar</code>) volume, grouped into altitude bins (<code>height</code>) and measured over time (<code>datetime</code>). The data are also available in the <a href="https://aloftdata.eu/browse/?prefix=uva/">Aloft bucket</a>.</p> <p>See Desmet et al. (2025, <a href="https://doi.org/10.1038/s41597-025-04641-5">https://doi.org/10.1038/s41597-025-04641-5</a>) for a more detailed description of this dataset.</p> <h2>Files</h2> <p>VPTS data in this deposit are organized per country (.tgz file), radar (directory), year (directory) and month (.csv.gz file). Fields in the data follow the <a href="https://aloftdata.eu/vpts-csv/">VPTS CSV</a> format and are described in <code>vpts-csv-table-schema.json</code>. An overview of what data are available is provided in <code>coverage.csv</code>. Radar metadata can be found at <a href="https://aloftdata.eu/radars/">https://aloftdata.eu/radars/</a>.</p> <ul> <li><strong>coverage.csv</strong>: coverage of the VPTS data, representing the number of unique hours, heights, source files and records for each radar and date combination.</li> <li><strong>vpts-csv-table-schema.json</strong>: technical description of the fields in the VPTS data.</li> <li><strong>be.tgz</strong>: VPTS data from 3 radars in Belgium.</li> <li><strong>de.gz</strong>: VPTS data from 18 radars in Germany.</li> <li><strong>nl.gz</strong>: VPTS data from 3 radars in the Netherlands.</li> </ul> <h2>Acknowledgements</h2> <p>This dataset was processed using infrastructure provided by the University of Amsterdam, SURF Cooperative, Ghent University and the Research Institute for Nature and Forest (INBO). It was mainly supported by the <a href="https://globam.science/">GloBAM project</a>, funded through the 2017-18 Belmont Forum and BiodivERsA joint call for research proposals under the BiodivScen ERA-Net COFUND programme.</p>
Supplementary material for Targeted gene knock-in reduces variation between transformants in the mushroom-forming fungus Schizophyllum commune
<p>Supplementary data for "Targeted gene knock-in reduces variation between transformants in the mushroom-forming fungus <em>Schizophyllum commune</em>"</p> <p>Dataset consists of fluorescent images of <em>S. commune</em> strains with an ectopic or targeted integration of <em>dTomato </em>under the control of the <em>tubulin </em>promoter and <em>hom2 </em>terminator and the obtained fluorescent intensity of each strain. For thesholding the mean intensity of all pixels above 14 (range 0, 255) was calculated.</p> <p>Files are names according to strain (E1-E12 for ectopic integrations and TI1-TI6 for targeted integrations and WT for wildtype) and replicate.</p>
Targeted Re-sequencing Identifies Candidate Fusiform Rust Resistance Genes in Loblolly Pine
<p>A fasta file containing the subset of the v2.01 Pita genome in addition to the novel NLR genes that were targeted by hybridization probes. </p> <p>A bed file describing the intervals targeted by the hybridization probes.</p> <p>Trinity assemblies of the 30 RNAseq libraries along with predictions by transdecoder of CDS and peptide sequences from those trinity assemblies. </p>
KiSSim: Predicting off-targets from structural similarities in the kinome
<p><strong>KiSSim: Predicting off-targets from structural similarities in the kinome</strong></p> <p><strong>Project description.</strong></p> <p>KiSSim (Kinase Structural Similarity) is a novel fingerprint designed specifically for kinase pockets, allowing for similarity studies across the structurally covered kinome. The kinase fingerprint is based on the <a href="https://klifs.net/">KLIFS</a> pocket alignment, which defines 85 pocket residues for all kinase structures. This enables a residue-by-residue comparison without a computationally expensive alignment step.</p> <p>The pocket fingerprint encodes each pocket residue’s spatial and physicochemical properties. The spatial properties describe the residue’s position in relation to the kinase pocket center and important kinase subpockets, i.e. the hinge region, the DFG region, and the front pocket. The physicochemical properties encompass for each residue its size and pharmacophoric features, solvent exposure, and side chain orientation.</p> <p>Some datasets are not part of the `kissim_app` GitHub repository due to their size but can be downloaded from here to the respective kissim_app folders.</p> <p><strong>Data.</strong></p> <ul> <li>`20210902_KLIFS_HUMAN.tar.gz` --- save in `kissim_app/data/external/structures`</li> <li>`complete_SiteAlign.txt.gz` --- save in `kissim_app/data/external/sitealign`</li> </ul> <p><strong>Results.</strong></p> <ul> <li>`results.tar.bz2`--- save as `kissim_app/results`</li> </ul> <p>These are the KiSSim results: fingerprints, feature/fingerprint distances, kinase matrices, and kinase trees for structures in all (`all`), DFG-in (`dfg_in`), and DFG-out (`dfg_out`) conformation. In the case of the DFG-in conformation, we also have KiSSim runs with fingerprint subsets based on only residues that interact with certain ligands in KLIFS IFPs: Erlotinib (`dfg_in_IRE`), Imatinib (`dfg_in_STI`), Bosutinib (`dfg_in_DB8`), and Dopamapimod (`dfg_in_B96`). The folder contains README with a detailed file list.</p> <p><strong>Usage.</strong></p> <p>This dataset can be used to run the notebooks available on <a href="https://github.com/volkamerlab/kissim_app">https://github.com/volkamerlab/kissim_app</a>.</p> <ol> <li>Clone the kissim_app repository.</li> <li>Download the files provided here.</li> <li>If applicable, extract the archive content to the folders as indicated above and run the notebooks.</li> </ol> <pre><code class="language-bash">cd /path/to/your/download tar -xvf results.tar.bz2 -C /path/to/kissim_app/ tar -xvf 20210902_KLIFS_HUMAN.tar.bz2 -C /path/to/kissim_app/data/external/structures/ # In case you want the raw SiteAlign data mv complete_SiteAlign.txt.gz /path/to/kissim_app/data/external/sitealign</code></pre> <p><strong>Citation.</strong></p> <p>These datasets are part of the KiSSim publication: TBA</p>
Cysteine dependence of Lactobacillus iners is a potential therapeutic target for vaginal microbiota modulation
<p>Compressed directories containing code and data files sufficient to reproduce analysis from Bloom et al paper on <em>Lactobacillus iners</em> (<em>Nature Microbiology</em>). An earlier, non-peer-reviewed manuscript version containing largely the same analysis was posted as a pre-print in <em>bioRxiv</em> at (https://doi.org/10.1101/2021.06.12.448098). Three compressed directory for analyses of:</p> <ol> <li>Vaginal <em>Lactobacillus </em>genome catalog characterization and gene content analysis.</li> <li>Analysis of relationship between cervicovaginal microbiota composition and cysteine concentrations in vaginal fluid from a South African cohort</li> <li>Analysis of results of <em>in vitro </em>mixed culture competition assays including: <ol> <li>Pairwise competition between <em>L. iners</em> and <em>Lactobacillus crispatus</em> in <em>Lactobacillus</em> MRS broth containing L-cysteine +/- S-methyl-L-cysteine (SMC)</li> <li>Defined bacterial-vaginosis (BV)-like communities including <em>L. iners</em>, <em>L. crispatus</em>, and BV-associated species <em>Gardnerella vaginalis</em>, <em>Prevotella bivia</em>, and <em>Atopobium (Fannyhessea) vaginae</em> cultured in S-broth with or without SMC and/or metronidazole.</li> </ol> </li> </ol>
Extended data for Manuscript: Identification of potential biological targets of oxindole scaffolds via in silico repositioning strategies
<p>This is the Extended Data for the manuscript "<strong>Identification of potential biological targets of oxindole scaffolds via <em>in silico</em> repositioning strategies" </strong>submitted to F1000 Research.</p> <p>Extended Data include a list of all the accession codes as mentioned in the text, the results of 2D fingerprint-based similarity analyses and ligand-protein complexes predicted by rigid docking and Induced Fit Docking calculations.</p>
Dipeptidyl peptidase 11 (PgDPP11); A Target Enabling Package
<p><em>Porphyromonas</em> gingivalis (<em>P. gingivalis</em>) is the main causative agent of Periodontitis, the most widespread inflammatory condition world-wide. Recently this organism has been implicated in several systemic conditions, such as Alzheimer’s disease and type 2 diabetes. <em>P. gingivalis</em> does not ferment carbohydrates, instead it uses proteases to generate energy and carbon source. Dipeptidyl peptidase 11 plays a central role in the energy metabolism of this bacterium and has been proposed as an attractive drug target. This TEP provide early tools to develop inhibitors of PgDPP11, including purification protocols of recombinant proteins, a crystal structure of the protein in complex with a dipeptide, crystallisation conditions suitable for crystallography-based fragment screening, an inhibition assay and fragment hits in the active site and an allosteric site. These molecules provide a promising starting point for the development of more specific and potent PgDPP11 inhibitors.</p>
Data for "Measurement of temperature induced X-ray tube transmission target displacements for dimensional computed tomography"
<p>Raw data used to create figures for the paper "Measurement of temperature induced X-ray tube transmission target displacements for dimensional computed tomography" <a href="https://doi.org/10.1016/j.precisioneng.2021.06.002">https://doi.org/10.1016/j.precisioneng.2021.06.002</a></p> <p>Data is available in tab delimited format (.txt) and in Excel (.xls).</p> <p> </p>
Testing of AgReFed FAIR data Minimum Thresholds and Stretch Targets
<p>This dataset is a testing of the FAIR thresholds for participation in The Australian Research Federation (AgReFed). The participants in the project assessed their data products before and after project works to improve the maturity of their datasets. The technology and information employed to progress the FAIR maturity of the data was recorded here.</p> <p>This data was used in the testing of the Minimum Thresholds and Stretch Targets developed by Box et al. (2019). Box, Paul, Levett, Kerry, Simons, Bruce, & Wong, Megan. (2019). Guidelines for the development of a Data Stewardship and Governance.</p>
Regional summary statistics for 1107 protein targets based on the Olink technology
<p>This data set contains regional summary statistics (±500kb around the protein coding gene) for a total of 1107 protein - gene combinations as measured by the Olink Proximity Extension Assay in the Fenland study (https://www.mrc-epid.cam.ac.uk/research/studies/fenland/) among 485 individuals. A detailed description of the genetic analysis can be found here https://www.nature.com/articles/s41467-021-27164-0. </p>
Open Satellite Video Single Target Tracking Datasets (OpenSatSTTD)
<p>We collect the latest open-source datasets for satellite video single target tracking (SatSTT) and launch the OpenSatSTTD project to promote the sharing of the latest research datasets in the SatSTT field. Satellite videos in the OpenSatSTTD project are collected from different sensors and platforms, and four targets (i.e., vehicles, trains, airplanes and vessels) are annotated by oriented bounding boxes. Users can obtain all satellite videos in the OpenSatSTTD project from links in the files.</p> <p>Source:</p> <p>Zheng, Ying., Zhu, Q., Luo, J., Li, Z., Lin, Z., Huang, X., and Zhang L.: Single Target Tracking in High-Resolution Satellite Videos: A Comprehensive Review (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6780820, 2022.</p> <p> </p>
Data files: Single-cell RNA profiling of Plasmodium vivax-infected hepatocytes reveals parasite- and host- specific transcriptomic signatures and therapeutic targets
<p>Scripts, preprocessed count matrices, and single-cell data objects generated in <strong>“Single-cell RNA profiling of <em>Plasmodium vivax</em><em>-</em>infected hepatocytes reveals parasite- and host- specific transcriptomic signatures and therapeutic targets” </strong></p>
Mars Target Encyclopedia - Labeled LPSC abstracts for four Mars missions
<p>This data set contains annotated text versions of 1635 two-page abstracts published at the Lunar and Planetary Science Conference from 1998 to 2020 of relevance to four Mars missions. The annotations were generated using named entity recognition and relation extraction provided by the MTE processing pipeline (available at https://github.com/wkiri/MTE), followed by manual review. Annotated entities include Element, Mineral, Property, and Target. Annotated relations include <strong>Contains</strong>(Target, Element | Mineral) and <strong>HasProperty</strong>(Target, Property). The extracted information (without full texts) is also available as a database (stored in .csv files) at https://pds-geosciences.wustl.edu/missions/mte/mte.htm . The complete annotated texts are provided here as a resource for further research and experimentation on information extraction methods. For more information about the Mars Target Encyclopedia and these annotations, please see:</p> <ul> <li>"<a href="https://www.hou.usra.edu/meetings/lpsc2022/pdf/1231.pdf">Targets from the Spirit Mars Exploration Rover in the Mars Target Encyclopedia</a>", Kiri L. Wagstaff, Raymond Francis, Matthew Golombek, Steven Lu, Ellen Riloff, Leslie Tamppari, Yuan Zhuang, and Thomas Stein.<br> <em>53rd Lunar and Planetary Science Conference</em>, Abstract #1231, March 2022.</li> <li>"<a href="https://www.hou.usra.edu/meetings/lpsc2021/pdf/1278.pdf">The Mars Target Encyclopedia Now Includes Mars Pathfinder and Mars Phoenix Targets</a>", Kiri L. Wagstaff, Raymond Francis, Matthew Golombek, Steven Lu, Ellen Riloff, Leslie Tamppari, and Thomas C. Stein.<br> <em>52nd Lunar and Planetary Science Conference</em>, Abstract #1278, March 2021.</li> </ul> <p>The original PDF abstracts are available at: </p> <ul> <li>For years prior to 2000: https://www.lpi.usra.edu/meetings/LPSC${two-digit-year}/pdf/${id}.pdf</li> <li>For year 2000: https://www.lpi.usra.edu/meetings/LPSC${four-digit-year}/pdf/${id}.pdf</li> <li>For years 2001-2017 (note lower-case lpsc): https://www.lpi.usra.edu/meetings/lpsc${four-digit-year}/pdf/${id}.pdf</li> <li>For years 2018-2020: https://www.hou.usra.edu/meetings/lpsc${four-digit-year}/pdf/${id}.pdf</li> </ul> <p>where ${id} is a four-digit abstract number, starting with 1001 (if available).</p> <p>The text files provided in this archive were extracted from the PDF files using the Apache Tika PDF parsing tool. They are named as ${four-digit-year}_${id}.txt. The text is provided here so that the annotations can be viewed in context. The text content remains copyright of the original abstract authors.</p> <p>The annotations (entities and relations) are provided in the format used by the brat annotation tool. They are named as ${four-digit-year}_${id}.ann. To view the annotations in a web-based graphical form, install the brat tool (http://brat.nlplab.org/). These annotations were generated using brat v1.3. The annotation files are also human-readable and can be parsed in to be used directly in code. If the .ann file is empty, then there are no relevant annotations for the associated text file.</p> <p><strong>Contents</strong>:</p> <ul> <li>mpf.zip: 591 abstracts relating to the Mars Pathfinder mission (1998-2020)</li> <li>mer-a.zip: 397 abstracts relating to the MER-A (Spirit) rover mission (2004-2020)</li> <li>mer-b.zip: 256 abstracts relating to the MER-B (Opportunity) rover mission (2005-2020)</li> <li>phx.zip: 391 abstracts relating to the Mars Phoenix Lander mission (2009-2020)</li> </ul> <p>Each directory contains a .txt and .ann file for each abstract. The .ann file is in brat standoff format (http://brat.nlplab.org/standoff.html). Additional .conf files are provided to generate color highlighting and keyboard shortcuts. These are used by the brat tool.</p> <p>Note: the same abstract may appear in more than one mission directory, if it discusses targets from more than one mission. It will have a different .ann file for each such appearance. Within each directory, a "Target" annotation is understood to refer to a target of the relevant mission.</p> <p><strong>Attribution</strong>:</p> <p>If you use this data set in your own work, please cite it as follows:</p> <p>Kiri L. Wagstaff, Raymond Francis, Matthew Golombek, Leslie Tamppari, and Steven Lu. (2022). Mars Target Encyclopedia - Labeled LPSC abstracts for four Mars missions (1.0.0.0) [Data set]. Zenodo. DOI: 10.5281/zenodo.7066107</p>
Mars Target Encyclopedia - LPSC abstracts labeled data set
<p>This data set contains annotated text versions of 2-page abstracts published at the Lunar and Planetary Science Conference in 2015 and 2016.</p> <p>The original PDF abstracts are available at:</p> <ul> <li>https://www.hou.usra.edu/meetings/lpsc2015/programAbstracts/view/</li> <li>https://www.hou.usra.edu/meetings/lpsc2016/programAbstracts/view/</li> </ul> <p>The text files in this archive were extracted using the Apache Tika PDF parsing tool. The text is provided here so that the annotations can be viewed. The text content remains copyright of the original abstract authors.</p> <p>The annotations (entities and relations) are provided in the format used by the brat annotation tool. To view the annotations in a web-based graphical form, install the brat tool (http://brat.nlplab.org/). These annotations were generated using brat v1.3. The annotation files are also human-readable and can be parsed in to be used directly in code.</p> <p><strong>Contents</strong>:</p> <ul> <li>lpsc15/: 62 abstracts</li> <li>lpsc16/: 55 abstracts</li> </ul> <p>Each directory contains a .txt and .ann file for each abstract. The .ann file is in brat standoff format (http://brat.nlplab.org/standoff.html).</p> <p>Additional .conf files are provided to generate color highlighting and keyboard shortcuts. These are used by the brat tool.</p> <p><strong>Attribution</strong>:</p> <p>If you use this data set in your own work, please cite this DOI:</p> <p>10.5281/zenodo.1048419</p> <p>Please also cite this paper, which provides additional details about the data set.</p> <p>Kiri L. Wagstaff, Raymond Francis, Thamme Gowda, You Lu, Ellen Riloff, Karanjeet Singh, and Nina Lanza. "Mars Target Encyclopedia: Rock and Soil Composition Extracted from the Literature." <em>Proceedings of the Thirtieth Annual Conference on Innovative Applications of Artificial Intelligence</em>, 2018.</p>
Ion Implantation Sensor and Process Target Data for Predicting Ion Beam Tuning in Semiconductor Manufacturing
<h2><strong>Dataset Description:</strong></h2> <p>This dataset is designed to predict ion beam tuning setup processes in semiconductor manufacturing, in terms of tuning success or failure, and tuning duration. It is split into <strong><code>X</code></strong> and <code><strong>y</strong></code> to allow for supervised learning approaches.</p> <ul> <li><code><strong>X</strong></code> represents the current equipment condition and the process targets of the currently processed and the upcoming lot, as defined within recipes.</li> <li><code><strong>y</strong></code> represents the ion beam tuning setup report, which informs about the tuning success ratio and tuning duration. These setups are necessary, when switching between recipes to prepare the equipment for processing the next lot. <strong><code>y</code></strong> contains three labels, enabling classification of (1) tuning success or fail, and (2) prolonged tuning, as well as (3) estimation of tuning duration as a regression task.</li> </ul> <p>About <strong><code>X</code></strong>:</p> <p>Each lot is processed with a specific recipe to achieve the process target. The tuning takes place before the first wafer of the to-be-tuned recipe is processed. Each row in <strong><code>X</code></strong> includes logistical information such as the equipment used for processing and parsed recipe / process target information for the current and upcoming lot. The majority of data consists out of aggregated metrics of equipment-internally tracked sensor traces, recording physical parameters such as gas flows, temperatures, voltages and currents. When analyzed in conjunction with the processed recipe, these sensors provide insights into the current equipment condition. </p> <p>About <code><strong>y</strong></code>:</p> <p>The <code>setup_result</code> column indicates the success or failure of tuning - with <code>setup_result=0</code> indicating tuning success, while <code>setup_result=1</code> signals tuning failure. If the first tuning attempt fails, there may be follow-up attempts, but these are not included in this dataset. The <code>duration</code> column represents the tuning duration in seconds, as used for regression analysis. The <code>duration_interval</code> column is a binary label for prolonged tunings, i.e. <code>duration_interval=1</code> for instances, which take more than 6 minutes to tune.</p> <p>For reproducibility of the corresponding paper's results:</p> <ol> <li>The dataset contains the same carefully curated subset of features.</li> <li>The train_test_split() has already been performed, thus we provide <code>x_train</code> and <code>x_valid</code> separately.</li> <li>To reduce the effect of outliers in the data, the sensor data has already been scaled, as derived from <code>x_train</code>.</li> </ol> <p>In summary, these datasets (<code><strong>X</strong></code>, <code><strong>y</strong></code>) provide comprehensive information for predicting ion beam tuning in semiconductor manufacturing, making it a valuable resource for researchers and practitioners in the field.</p> <h2><strong>Python Code for Reproducibility:</strong></h2> <p>Furthermore, we share a jupyter notebook <code>ionbeamtuning.ipynb</code> with Python code to train the best performing model on the provided data, as described in the paper. To execute the code, you may need to install any missing packages specified in the <code>requirements.txt</code>, as indicated within the notebook.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.