Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,129
datasets available to search
ShareScore release 0.9.0
Dataset results
2,129 results for “scores”
Real-bogus scores for active anomaly detection
<p>Data description for <a href="https://arxiv.org/abs/2409.10256">Semenikhin et al., 2024</a></p> <p>The dataset consists of the following files:</p> <p><strong>feature_snad4_r_100.dat</strong> contains light curve feature data for objects, where each object is represented by 54 feature values. These values are encoded as little-endian single-precision IEEE-754 floating-point numbers (32-bit floats). Feature names are listed in the plain text file <strong>feature_snad4_r_100.name</strong>, with one name per line.<br><strong>sid_snad4_r_100.dat</strong> contains ZTF DR object identifiers, encoded as little-endian 64-bit unsigned integers.</p> <p><strong>exp_feature_snad4_r_100.dat</strong> contains the same features as <strong>feature_snad4_r_100.dat</strong>, but with an additional column representing the real-bogus classifier prediction. Each object in this file corresponds to 55 features: the original 54 features plus 1 additional feature. Feature names for this file are provided in <strong>exp_feature_snad4_r_100.name</strong>.</p> <p>The files <strong>sid_snad4_r_100.dat</strong>, <strong>feature_snad4_r_100.dat</strong>, and <strong>exp_feature_snad4_r_100.dat</strong> share the same object order.</p> <p><br>Below is a sample Python script for accessing the data using NumPy:</p> <p><code>import numpy as np</code></p> <p><code># Load object IDs</code><br><code>oid = np.memmap('sid_snad4_r_100.dat', mode='c', dtype=np.uint64)</code></p> <p><code># Load features and reshape</code><br><code>feature = np.memmap('feature_snad4_r_100.dat', mode='c', dtype=np.float32).reshape(oid.shape[0], -1)</code></p> <p><code># Print dataset information</code><br><code>print(f'Number of objects: {len(oid)}')</code><br><code>print(f'Features shape: {feature.shape}')</code></p>
Data from: Coronary artery segmentation in non-contrast calcium scoring CT images using deep learning
<p><strong>Abstract</strong></p> <p>Precise segmentation of coronary arteries in non-contrast Computed Tomography (CT) scans plays an important role in the assessment of the coronary artery disease, where it is the key component for evaluating the Calcium Score (Agatston et al. 1990). In the paper by Bujny et al. (2024), a deep-learning approach for high-precision segmentation of coronary arteries in non-contrast CT was proposed along with a novel method for generating Ground Truth (GT) test data (<em>test-GT</em>) via manual registration of high-resolution coronary tree models obtained based on contrast CT with the non-contrast CT scans. In this dataset, we present the inferences of the neural network model together with the corresponding <em>test-GT</em> samples, based on 6 CT scans from the openly available OrCaScore dataset (Wolterink et al. 2016). The geometrical models included in the dataset can be used both for inspection of the proposed deep learning model and for testing of new non-contrast coronary vessel segmentation approaches, which is a unique opportunity since, to the best of our knowledge, manual generation of GT for non-contrast coronary artery segmentation was not addressed so far due to very challenging character of this particular segmentation task.</p> <p> </p> <p><strong>Methods</strong></p> <p><strong><em>Manual Generation of test-GT</em></strong></p> <p>The geometric models of coronary arteries used for the evaluation of the proposed neural network model were generated according to the manual mesh-to-image registration process as described by Bujny et al. (2024). In this approach, the high-resolution coronary artery masks obtained based on contrast CT scans are manually aligned with the corresponding non-contrast CT images using tools available in the open-source 3D computer graphics software, Blender (<a href="https://www.blender.org/">https://www.blender.org/</a>). To ease the manual alignment process, specialized add-ons for medical image processing such as Cardiac add-on for Blender of Graylight Imaging (<a href="https://graylight-imaging.com/3d-modelling/">https://graylight-imaging.com/3d-modelling/</a>) can be used, as well. The STL models in this dataset were manually generated by a medical expert with 4 years of experience.</p> <p><strong><em>Segmentation of Coronary Arteries using a Deep Learning Model</em></strong></p> <p>For each of the cases presented in this dataset, we run an inference of an nnU-Net (Isensee et al. 2021) model trained according to the process described in our paper (Bujny et al. 2024). Since we use a standard nnU-Net, which utilizes a sliding window approach for processing of the CT scan, the context information within a patch is limited, which can lead to some false-positive detections. To mitigate this problem, we additionally post-process the inferences by eliminating small vessel fragments of less than 50 [mm^3] volume and structures outside of pericardium, which we segment using another nnU-Net model, SegTHOR (Lambert et al. 2020). The resulting geometric models are stored using the STL format and presented as green masks in the HTML reports with an embedded viewer based on the K3D-jupyter library (<a href="https://k3d-jupyter.org/">https://k3d-jupyter.org/</a>).</p> <p> </p> <p><strong>Dataset organization</strong></p> <p>The root folder contains 6 folders whose names correspond to the CT scans from the OrCaScore dataset (Wolterink et al. 2016). In each of the folders, there are the following 4 files available:</p> <ul> <li><span>‘manualGT_rater1.stl’ – high-resolution STL model of coronary arteries obtained via manual alignment of the geometric model segmented in contrast CT with the corresponding non-contrast CT scan by the first rater.</span> A sample belonging to the <em>test-GT</em> set (Bujny et al. 2024).</li> <li>‘manualGT_rater2.stl’ – corresponding <em>test-GT</em> sample by the second rater.</li> <li>‘ML.stl’ – post-processed inference of the nnU-Net ML model in the STL format.</li> <li>‘report.html’ – interactive HTML report consisting of a manually-aligned <em>test-GT</em> sample (red mask), the ML segmentation based on the non-contrast CT scan (green mask), and selected slices of the non-contrast CT scan. The reports contain the relevant information related to the scanning device and present the main segmentation quality metrics for the ML model inference.</li> </ul>
New and modified scores in the phylogenetic matrix from Schoch (2013)
<p>Temnospondyl amphibians are a common component of non-marine Triassic assemblages, including in the Fremouw Formation (Lower to Middle Triassic) of Antarctica. Temnospondyls were among the first tetrapods to be collected from Antarctica, but their record from the lower Fremouw Formation has long been tenuous. One taxon, '<i>Austrobrachyops jenseni</i>,' is represented by a type specimen comprising only a partial pterygoid, which is now thought to belong to a dicynodont. A second taxon, '<i>Cryobatrachus kitchingi</i>,' is represented by a type specimen comprising a nearly complete skull, but the specimen is only exposed ventrally, and uncertainty over its ontogenetic maturity and some aspects of its anatomy has led it to be designated as a nomen dubium by previous workers. Here we redescribe the holotype of '<i>C. kitchingi</i>,' an undertaking that is augmented by tomographic analysis. Most of the original interpretations and reconstructions cannot be substantiated, and some are clearly erroneous. Although originally classified as a lydekkerinid, the purported lydekkerinid characteristics are shown to be unfounded or no longer diagnostic for the family. We instead identify numerous features shared with highly immature capitosaurs, a large-bodied clade documented in the upper Fremouw Formation of Antarctica and elsewhere in the Lower Triassic. Additionally, we describe a newly collected partial skull from the lower Fremouw Formation that represents a relatively mature, small-bodied individual and that we provisionally refer to Lydekkerinidae; this specimen represents the most confident identification of a lydekkerinid from Antarctica to date.</p>
AA-Score: a New Scoring Function Based on Amino Acid Specific Interaction for Molecular Docking
<p>The protein-ligand scoring function plays an important role in computer-aided drug discovery, which is heavily used in virtual screening and lead optimization. In this study, we developed a new empirical protein-ligand scoring function, which is a linear combination of empirical energy components, including hydrogen bond, van der Waals, electrostatic, hydrophobic, π-stacking, π-cation, and metal-ligand interaction. Different from previous empirical scoring functions, AA-Score uses several amino acid-specific empirical interaction components. We tested AA-Score on several test sets. The resulting performance shows AA-Score performs well on scoring, docking, and ranking compared with other widely used traditional scoring functions. Our results suggest that AA-Score gains substantial improvements from using detailed protein-ligand interaction components. Besides, we developed an easy-to-use tool to analyze protein-ligand interaction fingerprint and predict binding affinity using AA-Score.</p>
TocoDecoy: a new approach to design unbiased datasets for training and benchmarking machine-learning scoring functions
<p>This dataset file contains TocoDecoy datasets generated based on the targets and active ligands of LIT-PCBA.</p> <p>1_property_filtered.zip :</p> <ul> <li>TD set: the ligand file name, 2D T-sne vectors, Smiles, molecular weight (MW), Wildman-Crippen partition coefficient (log P), number of rotatable bonds (RB), number of hydrogen-bond acceptors (HBA), number of hydrogen-bond donors (HBD), number of halogens (HAL), topology similarities of decoys to the seed active ligands, active label (active or inactive) and training set label (whether belongs to training set or test set) <strong>OF active ligands and their topologically dissimilar decoys</strong></li> <li>CD set: the decoy conformations with low docking scores generated by docking active ligands into protein pockets using Glide, Schrödinger.</li> </ul> <p> </p>
Mean, standard deviation, and percentiles of the schema and domain scores of the German version of the Young Schema Questionnaire - Short Form 3 (YSQ-S3) in a German opportunity sample (n=1,150)
<p>Mean, standard deviation, and percentiles of the schema and domain scores of the German version of the Young Schema Questionnaire - Short Form 3 (YSQ-S3) in a German opportunity sample (n=1,150). Details are reported in: Kriston L, Schäfer J, Jacob GA, Härter M, Hölzel LP. Reliability and validity of the German version of the Young Schema Questionnaire - Short Form 3 (YSQ-S3). <em>Eur J Psychol Assess</em> 2013; 29: 205-212.</p> <p>IMPORTANT: This is an opportunity sample that is not representative of any well-defined population. Accordingly, the values should not be used as reference or norm values for the German version of the YSQ-S3.</p>
Mittelwert, Standardabweichung und Perzentile der Schema- und Domänen-Scores der deutschen Version des Young Schema Questionnaire - Short Form 3 (YSQ-S3) in einer deutschen Gelegenheitsstichprobe (n=1150)
<p>Mittelwert, Standardabweichung und Perzentile der Schema- und Domänen-Scores der deutschen Version des Young Schema Questionnaire - Short Form 3 (YSQ-S3) in einer deutschen Gelegenheitsstichprobe (n=1150). Details sind beschrieben in: Kriston L, Schäfer J, Jacob GA, Härter M, Hölzel LP. Reliability and validity of the German version of the Young Schema Questionnaire - Short Form 3 (YSQ-S3). <em>Eur J Psychol Assess</em> 2013; 29: 205-212.</p> <p>WICHTIG: Es handelt sich um eine Gelegenheitsstichprobe, die für keine gut definierbare Population repräsentativ ist. Dementsprechend sollten die Werte nicht als Referenz- oder Normwerte für die deutsche Version des YSQ-S3 verwendet werden.</p>
Reference videos for: Multi-track bottom-up synthesis from non-flattened AZee scores
<p>The upload contains videos referenced in the paper/poster "Multi-track bottom-up synthesis from non-flattened AZee scores"</p>
TCR-MHC Germline Interaction Scores Generated Using AIMS
<p>These data were generated using the AIMS interaction scoring function as outlined in the manuscript "A Systematic Characterization of Germline-Encoded Contacts Identifies the Source of Bias in TCR-MHC Interactions". They accompany the AIMS version 0.7 software available on GitHub: https://github.com/ctboughter/AIMS . These files are meant to be loaded into the mhc_germline_analysis.ipynb file, but are too large to be included on the GitHub page itself.</p>
pl-wnifc/humdrum-polish-scores: Digital scores from the "Polish Music Heritage in Open Access" project
<p>Digital scores from the Polish Music Heritage in Open Access project</p> <p>This repository contains transcriptions from the Heritage of Polish Music in Open Access project at the <a href="https://nifc.pl/en">Chopin Institute</a> in the Humdrum digital score format. Files are organized by the RISM siglum ID of the source archive.</p> <p>Frontend</p> <p>The main user interface for these digital scores is at <a href="https://polishscores.org/">https://polishscores.org/</a>. All transcriptions include scans of the physical scores.</p> <p>Staff:</p> <ul> <li><strong>Lead</strong>: Marcin Konik</li> <li><strong>Technical Lead</strong>: Craig Stuart Sapp</li> <li><strong>Project Manager</strong>: Jacek Iwaszko</li> <li><strong>Metadata Manager: </strong>Marcelina Chojecka</li> <li><strong>Assistant Project Manager</strong>: Emilia Ziętek</li> </ul> <p> </p> <p>Data created within EU funded project "Polish Music Heritage in Open Access"</p> <p>Co-financed by the Ministry of Culture and National Heritage</p> <p> </p> <table align="left"> <caption>Archives represented in repository</caption> <thead> <tr> <th scope="col">Siglum</th> <th scope="col">Library</th> <th scope="col">Scores</th> </tr> </thead> <tbody> <tr> <td><a href="https://github.com/pl-wnifc/humdrum-polish-scores/tree/main/pl-cz/kern">pl-cz</a></td> <td><a href="https://jasnagora.pl/en/o-sanktuarium/biblioteki/biblioteka-jasnogorska">Jasna Góra Monastery</a></td> <td>1294</td> </tr> <tr> <td><a href="https://github.com/pl-wnifc/humdrum-polish-scores/tree/main/pl-gd/kern">pl-gd</a></td> <td><a href="https://bgpan.gda.pl/?lang=en">Gdańsk Library PAoS</a></td> <td>244</td> </tr> <tr> <td><a href="https://github.com/pl-wnifc/humdrum-polish-scores/tree/main/pl-kc/kern">pl-kc</a></td> <td><a href="https://mnk.pl/branch/the-princes-czartoyski-library">Czartoryski Library, Cracow</a></td> <td>108</td> </tr> <tr> <td><a href="https://github.com/pl-wnifc/humdrum-polish-scores/tree/main/pl-kj/kern">pl-kj</a></td> <td><a href="https://bj.uj.edu.pl/en_GB/start-en">Jagiellonian Library, Cracow</a></td> <td>29</td> </tr> <tr> <td><a href="https://github.com/pl-wnifc/humdrum-polish-scores/tree/main/pl-kk/kern">pl-kk</a></td> <td><a href="http://akkk.com.pl/">Wawel Cathedral, Cracow</a></td> <td>1720</td> </tr> <tr> <td><a href="https://github.com/pl-wnifc/humdrum-polish-scores/tree/main/pl-kozmzk/kern">pl-kozmzk</a></td> <td><a href="https://www-muzeumzamoyskich-pl.translate.goog/?_x_tr_sch=http&_x_tr_sl=auto&_x_tr_tl=en&_x_tr_hl=en">Zamoyski Museum, Kozłówka</a></td> <td>168</td> </tr> <tr> <td><a href="https://github.com/pl-wnifc/humdrum-polish-scores/tree/main/pl-sa/kern">pl-sa</a></td> <td><a href="http://bc.bdsandomierz.pl/dlibra?language=en">Diocesan Library, Sandomierz</a></td> <td>1244</td> </tr> <tr> <td><a href="https://github.com/pl-wnifc/humdrum-polish-scores/tree/main/pl-stab/kern">pl-stab</a></td> <td><a href="https://rism.info/library_collections/2017/09/28/music-in-the-convent-of-st-adalberts-abbey-in.html">St. Adalbert Abbey, Staniątki</a></td> <td>154</td> </tr> <tr> <td><a href="https://github.com/pl-wnifc/humdrum-polish-scores/tree/main/pl-wn/kern">pl-wn</a></td> <td><a href="https://www.bn.org.pl/en">Polish National Library</a></td> <td>496</td> </tr> <tr> <td><a href="https://github.com/pl-wnifc/humdrum-polish-scores/tree/main/pl-wnifc/kern">pl-wnifc</a></td> <td><a href="https://nifc.pl/en">Chopin Institute, Warsaw</a></td> <td>390</td> </tr> <tr> <td><a href="https://github.com/pl-wnifc/humdrum-polish-scores/tree/main/pl-wtm/kern">pl-wtm</a></td> <td><a href="http://warszawskietowarzystwomuzyczne.pl/biblioteka/">Warsaw Music Society</a></td> <td>1163</td> </tr> <tr> <td><a href="https://github.com/pl-wnifc/humdrum-polish-scores/tree/main/pl-wumfc/kern">pl-wumfc</a></td> <td><a href="http://www.biblioteka.chopin.edu.pl/pl">Chopin University of Music</a></td> <td>151</td> </tr> </tbody> </table> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
Food Allergy Severity Score Diagram (Spanish)
<p>The Food Allergy Severity Score (FASS) is an instrument developed and validated to score the severity of allergic reactions elicited by foods. FASS has three formats that can be mapped to each other consistently: two ordinal scores with 3 (oFASS-3) or 5 grades (oFASS-5), and a numerical score (nFASS). The development and validation are reported in the manuscript of Fernández-Rivas et al. published in Allergy (https://doi.org/10.1111/all.15165). </p>
Food Allergy Severity Score Diagram (English)
<p>The Food Allergy Severity Score (FASS) is an instrument developed and validated to score the severity of allergic reactions elicited by foods. FASS has three formats that can be mapped to each other consistently: two ordinal scores with 3 (oFASS-3) or 5 grades (oFASS-5), and a numerical score (nFASS). The development and validation are reported in the manuscript of Fernández-Rivas et al. published in Allergy (https://doi.org/10.1111/all.15165). </p>
Figure 1: Frequency distribution & bar diagram of the combined score of ventilated patients
<p>In order to obtain the range of scores that represent 95% of the observations that ended in<br> respiratory failure, we used the frequency distribution curve (fig 9) with (2SD) above and below the<br> calculated mean. This gives a value of (16-24) as the limits of interval including the score of<br> patients at risk of developing respiratory failure</p>
Goat-CNN: A Lightweight Convolutional Neural Network for Pose-Independent Body Condition Score Estimation in Goats
<p>Here we introduce the dataset utilized in our published paper entitled "<a href="https://www.sciencedirect.com/science/article/pii/S2666154324002114">Goat-CNN: A Lightweight Convolutional Neural Network for Pose-Independent Body Condition Score Estimation in Goats</a>".</p> <p>Contained within the "bcs" folder are all the videos collected for this study. Each video file is named with a format denoting its respective details. The first number signifies the sequence of collection, the second denotes the ear tag, and the final figure represents the body condition score (BCS) value.</p> <p>For example: "1_158734_2.50" indicates the first sampling of an animal with the ear tag "158734" and a BCS value of "2.50".</p> <p>Additionally, we provide two Python scripts in this repository. The first script, "Video2Frame.py", facilitates the splitting of videos into individual frames. The second script, "Frames2npy.py", converts these frames into two numpy-friendly files with the extension ".npy". These files contain both the images ("X_train_bcs300.npy") and their corresponding labels ("Y_train_bcs300.npy").</p> <p>Furthermore, for the convenience of swift experimentation, we have included the desired .npy files within the repository.</p> <p>To load these files into your Python environment, you can use the following code snippet:</p> <div> <div>th4figs = '/content/drive/MyDrive/compag_2023/'</div> <br> <div>path4images = "/content/drive/MyDrive/CodeRefarm/datasets/BCS/X_train_bcs300.npy"</div> <div>Xtrain = np.load(path4images)</div> <br> <div>path4labels = "/content/drive/MyDrive/CodeRefarm/datasets/BCS/Y_train_bcs300.npy"</div> <div>Ytrain = np.load(path4labels).astype(float)</div> <br> <div>print("X train : ", Xtrain.shape)</div> <div>print("Y train : ", Ytrain.shape)</div> <div> <div> <div> <div> <div> <div> <div> </div> </div> <div> </div> </div> </div> </div> </div> <div> <div> <div> <div> <div> <div> <div> <div> <pre>X train : (5332, 300, 300, 3) Y train : (5332,)<br> </pre> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div>
iDRAMA-Scored-2024: A Dataset of the Scored Social Media Platform from 2020 to 2023
<p>ABSTRACT<br>---------------<br>Online web communities often face bans for violating platform policies, encouraging their migration to alternative platforms. This migration, however, can result in increased toxicity and unforeseen consequences on the new platform. In recent years, researchers have collected data from many alternative platforms, indicating coordinated efforts leading to offline events, conspiracy movements, hate speech propagation, and harassment. Thus, it becomes crucial to characterize and understand these alternative platforms. To advance research in this direction, we collect and release a large-scale dataset from Scored -- an alternative Reddit platform that sheltered banned fringe communities, for example, c/TheDonald (a prominent right-wing community) and c/GreatAwakening (a conspiratorial community). Over four years, we collected approximately 57M posts from Scored, with at least 58 communities identified as migrating from Reddit and over 950 communities created since the platform's inception. Furthermore, we provide sentence embeddings of all posts in our dataset, generated through a state-of-the-art model, to further advance the field in characterizing the discussions within these communities. We aim to provide these resources to facilitate their investigations without the need for extensive data collection and processing efforts.</p> <ul> <li>Scored platform: <a href="https://scored.co">https://scored.co</a></li> <li>Link to paper: <a href="https://arxiv.org/abs/2405.10233">https://arxiv.org/abs/2405.10233</a></li> <li>License: <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en">CC BY-NC-SA 4.0</a></li> </ul> <h1>Repository links</h1> <ul> <li><strong>Zenodo:</strong> From Zenodo, researchers can download `lite` version of this dataset, which includes only 57M posts from Scored (not the sentence embeddings).</li> <li><strong>Github:</strong> The main repository of this dataset, where we provide code-snippets to get started with this dataset. <ul> <li>Link here: <a href="https://github.com/idramalab/iDRAMA-scored-2024">https://github.com/idramalab/iDRAMA-scored-2024</a></li> </ul> </li> <li><strong>Huggingface:</strong> On Huggingface, we provide complete dataset with senetence embeddings.<br> <ul> <li>Link here: <a href="https://hf.co/datasets/iDRAMALab/iDRAMA-scored-2024">https://hf.co/datasets/iDRAMALab/iDRAMA-scored-2024</a></li> </ul> </li> </ul> <h1>Dataset Info</h1> <table> <tbody> <tr> <td><strong>File-name</strong></td> <td><strong>Data-points</strong></td> </tr> <tr> <td>comments-2020</td> <td>12,774,203</td> </tr> <tr> <td>comments-2021</td> <td>16,097,941</td> </tr> <tr> <td>comments-2022</td> <td>12,730,301</td> </tr> <tr> <td>comments-2023</td> <td>8,919,159</td> </tr> <tr> <td>submissions-2020-to-2023</td> <td>6,293,980</td> </tr> </tbody> </table> <h1>Authorship</h1> <p>This dataset is published at "AAAI ICWSM 2024 (INTERNATIONAL AAAI CONFERENCE ON WEB AND SOCIAL MEDIA)" hosted at Buffalo, NY, USA.</p> <ul> <li><strong>Academic Organization: </strong><a href="https://idrama.science/people/">iDRAMA Lab</a></li> <li><strong>Affiliation:</strong> Binghamton University, Boston University, University of California Riverside</li> </ul> <h1>Licensing</h1> <p>This dataset is available for free to use under terms of the non-commercial license <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en">CC BY-NC-SA 4.0</a>.</p> <h1>Citation</h1> <blockquote> <p>@inproceedings{patel2024idrama,<br> title={iDRAMA-Scored-2024: A Dataset of the Scored Social Media Platform from 2020 to 2023},<br> author={Patel, Jay and Paudel, Pujan and De Cristofaro, Emiliano and Stringhini, Gianluca and Blackburn, Jeremy},<br> booktitle={Proceedings of the International AAAI Conference on Web and Social Media},<br> volume={18},<br> pages={2014--2024},<br> year={2024},<br> issn = {2334-0770},<br> doi = {10.1609/icwsm.v18i1.31444},<br>}</p> </blockquote>
Figure 8. F1 scores for YOLOv5 in Use of open-source object detection algorithms to detect Palmer amaranth (Amoronthus polmeri) in soybean
Figure 8. F1 scores for YOLOv5 indicating the harmonic mean between precision and recall scores. Data indicated that detection results for both species would be best at a confidence threshold of 0.298.
Medical interview score data from PostCC-OSCE and programs for an extended many-facet IRT model
<p>Objective structured clinical examinations (OSCEs) are widely used performance assessments for medical and dental students. A common limitation of OSCEs is that the evaluation results depend on the characteristics of raters and the scoring rubric. To overcome this limitation, item response theory (IRT) models such as the many-facet models have been proposed to estimate examinee abilities while accounting for the characteristics of raters and evaluation items in a rubric. However, conventional IRT models have two impractical assumptions: constant rater severity across all evaluation items in a rubric and an equal interval rating scale among evaluation items, which can decrease model fitting and ability measurement accuracy.</p> <p>To resolve this problem, we propose a new IRT model that relaxes these assumptions. We demonstrate the effectiveness of the proposed model by applying it to actual data collected from a medical interview test conducted at Tokyo Medical and Dental University as part of a post-clinical clerkship (PostCC) OSCE. The experimental results showed that the proposed model fit our OSCE data well and measured ability accurately. Furthermore, it provided abundant information on rater and item characteristics that conventional models cannot, helping us to better understand rater and item properties.</p> <p>This dataset includes the actual score data collected from the above-mentioned medical interview test in a PostCC OSCE, as well as the program for estimating the parameters of the proposed IRT model.</p>
Dataset: FlexShares Credit-Scored US Corporate Bond Index Fund (SKOR) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Fig. 14. Multiple linear discriciminant score D in Ponera Testacea Emery, 1895 Stat. N. - A Sister Species Of P. Coarctata (Latreille, 1802) (Hymenoptera, Formicidae)
Fig. 14. Multiple linear discriciminant score D(7) = 0.068 FoDG +0.002 CS –0.43 PEL/NOH +0.02 PiMe –0.13 CL/CW –0.2 FR/CS –0.10 PEW/CS. for 126 individual workers of Ponera coarctata and testacea (based onSEIFERT's dataset)
Lost in Translation? Not for Large Language Models: Automated Divergent Thinking Scoring Performance Translates to Non-English Contexts (Datasets)
<p>Datasets for: Zielińska, A., Organisciak, P., Dumas, D., & Karwowski, M. (2023). Lost in translation? Not for large language models: Automated divergent thinking scoring performance translates to non-English contexts. <em>Thinking Skills and Creativity, 50</em>, 101414. https://doi.org/10.1016/j.tsc.2023.101414</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.