Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,250
datasets available to search
ShareScore release 0.7.1
Dataset results
6,250 results for “Classification”
Data for "On the Practice of Semantic Versioning for Ansible Galaxy Roles: An Empirical Study and a Change Classification Model"
<p>This dataset accompanies a replication package provided for a study on Semantic Versioning for Ansible Galaxy roles.</p> <p>The replication package is available at https://github.com/ROpdebee/ansible_semver_ext_replication</p>
Combined continuous nanoparticle synthesis with chromatographic size classification
<p>In this paper, we report a combination of the continuous flow synthesis of gold nanoparticles (AuNPs) with subsequent purification and narrowing of the particle size distribution (PSD) by size-exclusion chromatography (SEC) by adapting the flow rates of synthesis and classification. First, we show scalability of chromatographic classification with respect to column dimension and the absence of irreversible nanoparticle adhesion on the column material. Two different syntheses lead to a large and widely distributed and a small and narrowly distributed AuNP dispersion, which are classified by a semipreparative column. The PSDs of individual fractions are characterized by analytical SEC. The broadly distributed AuNP dispersion was classified into three fractions with distinct PSDs. For the narrowly distributed AuNPs, the separation is almost independent of the mobile phase flow rate: coarse and fine fractions with almost identical PSDs and separation efficiency curves are observed irrespective of the flow rate. Even NP samples with narrow PSDs can be classified into multiple fractions with tailored PSDs while simultaneously removing dissolved impurities from the dispersion. With our study, we demonstrate the potential of a direct combination of continuous NP synthesis with chromatographic classification for the optimization of final PSDs and the simultaneous purification of nanoparticulate dispersions.</p><p>Funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation)– Project-ID 416229255 – SFB 1411</p>
System transferability of a Raman-based oesophageal tissue classification
<p>This dataset contains 560819 Raman spectra (Uncorrected_MedianFiltered_SizeMatched_SMART_x) taken from 61 oesophageal samples (sampleID) representing 51 patients (patientID). The data was acquired across three independent sites (centre) using the same make of spectrometer - Renishaw RA816 Biological Analyser (Renishaw plc, Wotton-under-edge, UK). Ostensibly the same sample was collected by all three sites (three adjacent FFPE tissue slices were obtained and regions of interest were identified by a histopathologist on one of the centres). The samples belong to one of 5 pathology classes: NSQ (0), IM(1), LGD(2), HGD(3) and AC(4). </p> <p>Included in this version is the protocol used to collect this data, particularly focused on the aquisition of Raman spectra.</p>
Curlie Enhanced with LLM Annotations: Two Datasets for Advancing Homepage2Vec's Multilingual Website Classification
<h3>Advancing Homepage2Vec with LLM-Generated Datasets for Multilingual Website Classification</h3> <p>This dataset contains two subsets of labeled website data, specifically created to enhance the performance of Homepage2Vec, a multi-label model for website classification. The datasets were generated using Large Language Models (LLMs) to provide more accurate and diverse topic annotations for websites, addressing a limitation of existing Homepage2Vec training data.</p> <p><strong>Key Features:</strong></p> <ul> <li><strong>LLM-generated annotations:</strong> Both datasets feature website topic labels generated using LLMs, a novel approach to creating high-quality training data for website classification models.</li> <li><strong>Improved multi-label classification:</strong> Fine-tuning Homepage2Vec with these datasets has been shown to improve its macro F1 score from 38% to 43% evaluated on a human-labeled dataset, demonstrating their effectiveness in capturing a broader range of website topics.</li> <li><strong>Multilingual applicability:</strong> The datasets facilitate classification of websites in multiple languages, reflecting the inherent multilingual nature of Homepage2Vec.</li> </ul> <p><strong>Dataset Composition:</strong></p> <ul> <li><strong>curlie-gpt3.5-10k:</strong> 10,000 websites labeled using GPT-3.5, context 2 and 1-shot</li> <li><strong>curlie-gpt4-10k:</strong> 10,000 websites labeled using GPT-4, context 2 and zero-shot</li> </ul> <p><strong>Intended Use:</strong></p> <ul> <li>Fine-tuning and advancing Homepage2Vec or similar website classification models</li> <li>Research on LLM-generated datasets for text classification tasks</li> <li>Exploration of multilingual website classification</li> </ul> <p><strong>Additional Information:</strong></p> <ul> <li><strong>Project and report repository:</strong> https://github.com/CS-433/ml-project-2-mlp</li> </ul> <p><strong>Acknowledgments:</strong></p> <p>This dataset was created as part of a project at EPFL's Data Science Lab (DLab) in collaboration with <a href="https://people.epfl.ch/robert.west">Prof. Robert West</a> and <a href="https://tizianopiccardi.github.io/" rel="nofollow">Tiziano Piccardi.</a></p>
The effect of dynamical states on galaxy clusters populations. I. Classification of dynamical states
<p>This repository contains three figures mentioned in "The effect of dynamical states on galaxy clusters populations. I. Classification of dynamical states" <em>(DOI to follow on publication)</em>.</p> <p>We show the contours of the X-ray surface brightness distribution (solid green lines) and the distribution of galaxies belonging to the red sequence (solid gray lines). Black crosses symbolize the positions of the X-ray peaks, black "X" marks represent the positions of the X-ray centroids, and open red circles denote the positions of the BCGs. The blue circle corresponds to the R200 of each cluster.</p>
The classification of toric canonical Fano 3-folds
<p><strong>The classification of toric canonical Fano 3-folds</strong></p> <p>This dataset describes the classification of all toric canonical Fano 3-folds [1]. Equivalently, it describes the classification of all 3-dimensional convex lattice polytopes with exactly one interior lattice point.</p> <p>A toric Fano 3-fold <span class="math-tex">\(X\)</span> (that is, a 3-dimensional toric variety with ample anticanonical divisor <span class="math-tex">\(-K\)</span>) with at worst canonical singularities corresponds to a 3-dimensional convex polytope <span class="math-tex">\(P\)</span> with vertices that are primitive integer vectors, and such that <span class="math-tex">\(P\)</span> contains exactly one lattice point, the origin, in its strict interior. The fan of <span class="math-tex">\(X\)</span> is given by the spanning fan of <span class="math-tex">\(P\)</span>: that is, the fan whose cones are spanned by the faces of <span class="math-tex">\(P\)</span>. In the language of toric geometry, <span class="math-tex">\(P\)</span> is in the lattice <span class="math-tex">\(N \cong \mathbb{Z}^3\)</span>. Since two polytopes which are equal after a lattice change of basis give rise to isomorphic toric varieties, a polytope is regarded as being defined only up to the action of <span class="math-tex">\(\mathrm{GL}(\mathbb{Z}^3)\)</span>. There are 674688 isomorphism classes.</p> <p>For details, see the paper [1]. If you make use of this data, please consider citing [1] and the DOI for this data:</p> <p>doi:10.5281/zenodo.5866330</p> <p><strong>toricf3c.txt</strong></p> <p>The file "toricf3c.txt" contains key:value records with keys and values as described below, where each record is separated by a blank line. Each key:value record determines a toric Fano 3-fold <span class="math-tex">\(X\)</span>, or equivalently a lattice polytope <span class="math-tex">\(P\)</span>, in the classification. There are 674688 records in the file.</p> <p><strong>Example record</strong></p> <p>id: 1<br> num_vertices: 16<br> num_faces: 10<br> num_points: 22<br> is_terminal: false<br> is_simplicial: false<br> is_regular: false<br> is_reflexive: false<br> vertex_list: [[2,1,1],[-1,0,-1],[-2,-1,-1],[0,1,1],[0,-1,-1],[0,-1,-2],[-2,-1,0],[0,1,0],[2,1,2],[-1,0,1],[1,1,2],[1,0,1],[1,1,0],[1,0,-1],[-1,-1,0],[-1,-1,-2]]<br> point_list: [[2,1,1],[-1,0,-1],[-2,-1,-1],[0,1,1],[0,-1,-1],[0,0,-1],[0,-1,-2],[-1,0,0],[-2,-1,0],[1,0,0],[0,1,0],[0,0,1],[-1,-1,-1],[2,1,2],[1,1,1],[-1,0,1],[1,1,2],[1,0,1],[1,1,0],[1,0,-1],[-1,-1,0],[-1,-1,-2]]<br> dual_list: [[0,-1,1],[0,-1,0],[1,-1,0],[-1,1,0],[0,1,0],[0,1,-1],[1/2,-1/2,-1/2],[-1/2,-1/2,1/2],[-1/2,3/2,-1/2],[1/2,-1/2,1/2]]<br> ehrhart_delta: [1,19,23,1]<br> hilbert_delta: [1,7,25,47,47,25,7,1]<br> normal_form: [[1,0,0],[0,1,0],[1,1,2],[0,-1,-2],[-1,0,-2],[1,-1,0],[-1,1,0],[1,0,2],[0,1,2],[-1,-1,-2],[0,-1,0],[-1,0,0],[-1,1,1],[1,1,1],[-1,-1,-3],[1,-1,1]]<br> volume: 44<br> degree: 10<br> gorenstein_index: 2<br> h1: 7<br> h2: 29<br> h3: 75<br> h4: 157<br> h5: 283<br> h6: 465<br> h7: 711<br> h8: 1033<br> h9: 1439<br> h10: 1941<br> e1: 23<br> e2: 109<br> e3: 303<br> e4: 649<br> e5: 1191<br> e6: 1973<br> e7: 3039<br> e8: 4433<br> e9: 6199<br> e10: 8381<br> picard_rank: 1<br> automorphism_order: 24<br> is_barycentre_zero: true<br> is_dual_barycentre_zero: true</p> <p>We fix some notation. Let:</p> <ul> <li><span class="math-tex">\(N \cong \mathbb{Z}^3\)</span>be a lattice of rank 3;</li> <li><span class="math-tex">\(P\)</span> denote the lattice polytope in <span class="math-tex">\(N_\mathbb{Q}=N\otimes_\mathbb{Z}\mathbb{Q}\)</span> defined by the key:value record;</li> <li><span class="math-tex">\(F\)</span> denote the spanning fan in N of P;</li> <li><span class="math-tex">\(X\)</span> denote the toric Fano 3-fold corresponding to <span class="math-tex">\(F\)</span>;</li> <li><span class="math-tex">\(M = \mathrm{Hom}(N,\mathbb{Z})\cong\mathbb{Z}^3\)</span> denote the lattice dual to <span class="math-tex">\(N\)</span>;</li> <li><span class="math-tex">\(P^*\)</span> denote the (rational) polytope in <span class="math-tex">\(M_\mathbb{Q}\)</span> dual to <span class="math-tex">\(P\)</span> (also called the polar polytope);</li> <li><span class="math-tex">\(-K\)</span> denote the anticanonical divisor of <span class="math-tex">\(X\)</span>.</li> </ul> <p>The keys and values are as follows.</p> <p>id: A unique integer ID for this record, in the range 1 to 674688.<br> num_vertices: A positive integer. The number of vertices <span class="math-tex">\(\#\mathrm{vert}(P)\)</span> of <span class="math-tex">\(P\)</span>. Equivalently, the number of rays of <span class="math-tex">\(F\)</span>. By duality this is also equal to the number of 2-dimensional faces of <span class="math-tex">\(P^*\)</span>.<br> num_faces: A positive integer. The number of 2-dimensional faces of <span class="math-tex">\(P\)</span>. Equivalently, the number of top-dimensional cones of <span class="math-tex">\(F\)</span>. By duality this is also equal to the number of vertices of <span class="math-tex">\(P^*\)</span>.<br> num_points: A positive integer. The number of lattice points <span class="math-tex">\(\#(P \cap N)\)</span>.<br> is_terminal: A boolean. True if and only if <span class="math-tex">\(X\)</span> has at worst terminal singularities. Equivalently, true if and only if the only lattice points on the boundary of <span class="math-tex">\(P\)</span> are the vertices of <span class="math-tex">\(P\)</span>; that is, <span class="math-tex">\(P \cap N = \mathrm{vert}(P) \cup \{0\}\)</span>.<br> is_simplicial: A boolean. True if and only if <span class="math-tex">\(P\)</span> is simplicial. Equivalently, true if and only if <span class="math-tex">\(X\)</span> is <span class="math-tex">\(\mathbb{Q}\)</span>-factorial. By duality, this is true if and only if <span class="math-tex">\(P^*\)</span> is simple.<br> is_regular: A boolean. True if and only if <span class="math-tex">\(X\)</span> is smooth. Equivalently, true if and only if every 2-dimensional face of <span class="math-tex">\(P\)</span> is a triangle whose vertices <span class="math-tex">\(\mathbb{Z}\)</span>-generate the lattice <span class="math-tex">\(N\)</span>. If is_regular is true then both is_simplicial and is_terminal must be true.<br> is_reflexive: A boolean. True if and only if <span class="math-tex">\(X\)</span> is Gorenstein; that is, <span class="math-tex">\(-K\)</span> is Cartier. Equivalently, true if and only if <span class="math-tex">\(P^*\)</span> is a lattice polytope.<br> vertex_list: A sequence of lattice points in <span class="math-tex">\(N\)</span>. The vertices <span class="math-tex">\(\mathrm{vert}(P)\)</span> of <span class="math-tex">\(P\)</span>. Equivalently, the primitive lattice generators of the rays of <span class="math-tex">\(F\)</span>. The number of points is given by num_vertices.<br> point_list: A sequence of lattice points in <span class="math-tex">\(N\)</span>. The lattice points <span class="math-tex">\(P \cap N\)</span>. The number of points is given by num_points.<br> dual_list: A sequence of rational points in <span class="math-tex">\(M\)</span>. The vertices <span class="math-tex">\(\mathrm{vert}(P^*)\)</span> of <span class="math-tex">\(P^*\)</span>. These will be lattice points if and only if is_reflexive is true. The number of points is given (via duality) by num_faces.<br> ehrhart_delta: A sequence <span class="math-tex">\((1,a_1,a_2,1)\)</span> of four integers, the first and last of which are always 1. This is the Ehrhart <span class="math-tex">\(\delta\)</span>-vector (or <span class="math-tex">\(h^*\)</span>-vector) of <span class="math-tex">\(P\)</span>. The Ehrhart series <span class="math-tex">\(\mathrm{Ehr}(P)\)</span> of <span class="math-tex">\(P\)</span> is given by <span class="math-tex">\(\mathrm{Ehr}(P) = (1 + a_1t + a_2t^2 + t^3) / (1 - t)^4\)</span>.<br> hilbert_delta: A sequence <span class="math-tex">\((b_0,b_1,\ldots,b_N)\)</span> of integers such that <span class="math-tex">\(b_i = b_{N - i}\)</span>, and <span class="math-tex">\(b_0 = b_N = 1\)</span>. This is called the Ehrhart <span class="math-tex">\(\delta\)</span>-vector (or <span class="math-tex">\(h^*\)</span>-vector) of <span class="math-tex">\(P^*\)</span>. Write <span class="math-tex">\(N = 4r - 1\)</span>. Then <span class="math-tex">\(r\)</span> is the quasiperiod of <span class="math-tex">\(P^*\)</span>, and <span class="math-tex">\(r\)</span> divides gorenstein_index. The Ehrhart series of <span class="math-tex">\(P^*\)</span> is given by <span class="math-tex">\(\mathrm{Ehr}(P^*) = (b_0 + b_1t + \ldots + b_Nt^N) / (1 - t^r)^4\)</span>. The Ehrhart series <span class="math-tex">\(\mathrm{Ehr}(P^*)\)</span> of <span class="math-tex">\(P^*\)</span> is equal to the Hilbert series <span class="math-tex">\(\mathrm{Hilb}(X,-K)\)</span>.<br> normal_form: A sequence of lattice points in <span class="math-tex">\(N\)</span>. The PALP normal form of the vertices of <span class="math-tex">\(P\)</span>; see [2,3].<br> volume: A positive integer. The lattice-normalised volume <span class="math-tex">\(\mathrm{Vol}(P)\)</span> of <span class="math-tex">\(P\)</span>. This is equal to the sum of the ehrhart_delta: <span class="math-tex">\(\mathrm{Vol}(P) = 1 + a_1 + a_2 + 1\)</span>.<br> degree: A positive integer. The anticanonical degree <span class="math-tex">\((-K)^3\)</span> of <span class="math-tex">\(X\)</span>. Equivalently, the lattice-normalised volume <span class="math-tex">\(\mathrm{Vol}(P^*)\)</span> of <span class="math-tex">\(P^*\)</span>.<br> gorenstein_index: A positive integer. The Gorenstein index of <span class="math-tex">\(X\)</span>; that is, the smallest multiple <span class="math-tex">\(m>0\)</span> such that <span class="math-tex">\(-mK\)</span> is Cartier. Equivalently, the smallest multiple <span class="math-tex">\(m>0\)</span> such that <span class="math-tex">\(mP^*\)</span> is a lattice polytope. This is 1 if and only if is_reflexive is true.<br> h1,...,h10: Positive integers. The value hi <span class="math-tex">\(=h_i\)</span> is equal to the number of lattice points in the <span class="math-tex">\(i\)</span>-th dilation of <span class="math-tex">\(P^*\)</span>, that is, <span class="math-tex">\(h_i = \#(iP^* \cap M)\)</span>. Equivalently, <span class="math-tex">\(h_i = h^0(X,-iK)\)</span>. The values <span class="math-tex">\(h_i\)</span> can also be obtained from <span class="math-tex">\(\mathrm{Ehr}(P^*)\)</span>, or equivalently from <span class="math-tex">\(\mathrm{Hilb}(X,-K)\)</span>, via the power-series expansion <span class="math-tex">\((b_0 + b_1t + \ldots + b_Nt^N) / (1 - t^r)^4 = 1 + h_1t + h_2t^2 + h_3t^3 + \ldots\)</span>, where <span class="math-tex">\((b_0,b_1,\ldots,b_N)\)</span> is given by hilbert_delta.<br> e1,...,e10: Positive integers. The value ei <span class="math-tex">\(=e_i\)</span> is equal to the number of lattice points in the <span class="math-tex">\(i\)</span>-th dilation of <span class="math-tex">\(P\)</span>, that is, <span class="math-tex">\(e_i = \#(iP \cap N)\)</span>. In particular, <span class="math-tex">\(e_1\)</span> is equal to num_points. The values <span class="math-tex">\(e_i\)</span> can also be obtained from <span class="math-tex">\(\mathrm{Ehr}(P)\)</span> via the power-series expansion <span class="math-tex">\((1 + a_1t + a_2t^2 + t^3) / (1 - t)^4 = 1 + e_1t + e_2t^2 + e_3t^3 + \ldots\)</span>, where <span class="math-tex">\((1,a_1,a_2,1)\)</span> is given by ehrhart_delta.<br> picard_rank: Positive integer. The rank of the Picard group of <span class="math-tex">\(X\)</span>. When is_simplicial is true, this is equal to <span class="math-tex">\(\#\mathrm{vert}(P) - 3\)</span>.<br> automorphism_order: Positive integer. The order of the automorphism group <span class="math-tex">\(\mathrm{Aut}(P) \leq \mathrm{GL}(\mathbb{Z}^3)\)</span> of <span class="math-tex">\(P\)</span>.<br> is_barycentre_zero: A boolean. True if and only if the barycentre of <span class="math-tex">\(P\)</span> is equal to the origin in <span class="math-tex">\(N\)</span>.<br> is_dual_barycentre_zero: A boolean. True if and only if the barycentre of <span class="math-tex">\(P^*\)</span> is equal to the origin in <span class="math-tex">\(M\)</span>.</p> <p><strong>toricf3c.sql</strong></p> <p>The file "toricf3c.sql" contain an sqlite-formatted version of the data described above, and can be imported into an sqlite database via, for example:</p> <pre><code class="language-bash">$ cat toricf3c.sql | sqlite3 toricf3c.db</code></pre> <p>This can then be easily queried. For example:</p> <pre><code class="language-bash">$ sqlite3 toricf3c.db > SELECT COUNT(*) FROM toricf3c; 674688 > SELECT vertex_list FROM toricf3c WHERE degree = 72; [[-1,-4,-6],[1,0,0],[0,1,0],[0,0,1]] [[-1,-1,-3],[1,0,0],[0,1,0],[0,0,1]]</code></pre> <p> </p> <p><strong>References</strong></p> <p>[1] Alexander M. Kasprzyk. <em>Canonical toric Fano threefolds</em>. Canadian Journal of Mathematics, 62(6):1293–1309, 2010.<br> [2] Maximilian Kreuzer, Harald Skarke. <em>PALP, a package for analyzing lattice polytopes with applications to toric geometry</em>. Computer Phys. Comm., 157:87-106, 2004.<br> [3] Roland Grinis, Alexander M. Kasprzyk. <em>Normal forms of convex lattice polytopes</em>. arXiv:1301.6641 [math.CO], 2013</p>
8 years of dayside Magnetospheric Multiscale (MMS) unsupervised clustering plasma regions classifications
<p>These files contain the 1-minute resolution dataset (“labeled_sunside_data.csv”) and 15 minute or longer region list (“<region_name>_region_list.csv”) for Toy-Edens et al.'s Classifying 8 years of MMS Dayside Plasma Regions via Unsupervised Machine Learning. The 1-minute resolution file contains the rolled up 1-minute epoch, probe name (mms1, mms2, mms3, mms4), features that go into clustering and post-cleansing methods, spacecraft positions (in GSE, GSM, and magnetic latitude/local time), raw and cleansed clustering labels, and transition name. The 15+ minute region lists contain the name of the plasma region type, the probe name (mms1, mms2, mms3, mms4), and the start and stop epoch of >= 15 minute epoch where the probe is solidly within that region. NOTE: for the 15+ minute region lists we are only looking for changes in plasma regions, this means that missing data may artificially inflate the duration of the epoch, we suggest looking at the full 1-minute resolution dataset to confirm the region timing.</p> <p>We ask that if you use any parts of the dataset that you cite Toy-Edens et al.'s Classifying 8 years of MMS Dayside Plasma Regions via Unsupervised Machine Learning (DOI:10.1029/2024JA032431).</p> <p>This work was funded by grant 2225463 from the NSF GEM program.</p> <p> </p> <p>The following tables detail the contents of the described files:</p> <p><strong>labeled_sunside_data.csv description</strong></p> <table> <tbody> <tr> <td> <p><strong>Column Name</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p>Epoch</p> </td> <td> <p>Epoch in datetime</p> </td> </tr> <tr> <td> <p> probe</p> </td> <td> <p>MMS probe name</p> </td> </tr> <tr> <td> <p> ratio_max_width</p> </td> <td> <p>Ratio of the width of the most prominent ion spectra peak (in number of energy channels) to max number of energy channels. See paper for more information</p> </td> </tr> <tr> <td> <p> ratio_high_low</p> </td> <td> <p>Ratio of the mean of the log intensity of high energies in the ion spectra to the mean of the log intensity of low energies in the ion spectra. See paper for more information</p> </td> </tr> <tr> <td> <p> norm_Btot</p> </td> <td> <p>Magnitude of the total magnetic field normalized to 50nT. See paper for more information</p> </td> </tr> <tr> <td> <p> small_energy_mean</p> </td> <td> <p>The denominator in ratio_high_low</p> </td> </tr> <tr> <td> <p> large_energy_mean</p> </td> <td> <p>The numerator in ratio_high_low</p> </td> </tr> <tr> <td> <p> temp_total</p> </td> <td> <p>Total temperature from the DIS moments. See paper for more information</p> </td> </tr> <tr> <td> <p> r_gse_x</p> </td> <td> <p>x position of the spacecraft in GSE</p> </td> </tr> <tr> <td> <p> r_gse_y</p> </td> <td> <p>y position of the spacecraft in GSE</p> </td> </tr> <tr> <td> <p> r_gse_z</p> </td> <td> <p>z position of the spacecraft in GSE</p> </td> </tr> <tr> <td> <p> r_gsm_x</p> </td> <td> <p>x position of the spacecraft in GSM</p> </td> </tr> <tr> <td> <p> r_gsm_y</p> </td> <td> <p>y position of the spacecraft in GSM</p> </td> </tr> <tr> <td> <p> r_gsm_z</p> </td> <td> <p>z position of the spacecraft in GSM</p> </td> </tr> <tr> <td> <p> mlat</p> </td> <td> <p>magnetic latitude of spacecraft</p> </td> </tr> <tr> <td> <p> mlt</p> </td> <td> <p>magnetic local time of spacecraft</p> </td> </tr> <tr> <td> <p> raw_named_label</p> </td> <td> <p>Raw cluster assigned plasma region label (allowed values: magnetosheath, magnetosphere, solar wind, ion foreshock)</p> </td> </tr> <tr> <td> <p> modified_named_label</p> </td> <td> <p>Cleansed cluster assigned plasma region label (use these unless have a specific reason to use raw labels). See paper for more information</p> </td> </tr> <tr> <td> <p> transition_name</p> </td> <td> <p>Transition names (e.g. quasi-perpendicular bow shock, magnetopause). See paper for more information</p> </td> </tr> </tbody> </table> <p> </p> <p><strong><region_name>_region_list.csv description</strong></p> <table> <tbody> <tr> <td> <p><strong>Column Name</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p>start</p> </td> <td> <p>Starting Epoch in datetime</p> </td> </tr> <tr> <td> <p>stop</p> </td> <td> <p>Stopping Epoch in datetime</p> </td> </tr> <tr> <td> <p>probe</p> </td> <td> <p>MMS probe name</p> </td> </tr> <tr> <td> <p>region</p> </td> <td> <p>Cleansed cluster name associated with 1-minute resolution “modified_named_label”</p> </td> </tr> </tbody> </table> <p> </p> <p> </p> <p> </p>
Infection Inspection: Classifications and images of ciprofloxacin-treated Escherichia coli clinical isolates
<p>This dataset includes a .csv file with the image metadata and a folder of RGB images of <i>E. coli</i> grown from clinical isolates with varying concentrations of the antibiotic ciprofloxacin and varying minimum inhibitory concentrations. The <i>E. coli</i> cell membranes are stained with Nile Red and the DNA is stained with DAPI. The details of the image data collection are included in: https://doi.org/10.1038/s42003-023-05524-4. The classification data come from a Zooniverse citizen science project, Infection Inspection. (https://www.zooniverse.org/projects/conor-feehily/infection-inspection) Volunteers learned how to interpret ciprofloxacin response phenotypes as antibiotic-sensitive or antibiotic-resistant, and their classifications are included in the Metadata.csv file.</p><p>This dataset could be used for further analysis into the volunteer classifications, or the image data could be used for further image feature analysis of the ciprofloxacin response phenotypes.</p>
Classification of New Caledonian Forests According to Edge and Elevation Effects
<h1>Description</h1> <p>This map represents a classification of forest types based on the influence of the edge effect (distance to the forest edge) and elevation effect (temperature and area) on tree community richness.</p> <ul> <li>The edge effect influences tree diversity through an environmental aridity filter. In New Caledonia, the maximum temperature recorded at the forest edge is 41°C in February, while it never exceeds 24°C beyond 100 meters from the edge. This temperature difference induces a selection for species that tolerate the most arid conditions, leading to a reduction in the biological richness of tree communities (<a href="https://doi.org/10.1007/s10980-017-0534-7" target="_blank" rel="noopener">Ibanez et al., 2017</a>; <a href="https://cnrt.nc/wp-content/uploads/2022/12/CNRT-rappsc-RELIQUES_Tome-ENV-Edition-2022-cp.pdf" target="_blank" rel="noopener">Birnbaum et al., 2022</a>; <a href="https://doi.org/10.1111/1365-2745.14105" target="_blank" rel="noopener">Blanchard et al., 2023</a>).</li> <li>Altitude also affects tree diversity due to temperature variation and available area (<a href="https://doi.org/10.1111/avsc.12070" target="_blank" rel="noopener">Ibanez et al., 2014</a>; <a href="https://doi.org/10.1093/aobpla/plv075" target="_blank" rel="noopener">Birnbaum et al., 2015</a>; <a href="https://doi.org/10.1111/ddi.12374" target="_blank" rel="noopener">Pouteau et al., 2015</a>; <a href="https://doi.org/10.1111/jvs.12396" target="_blank" rel="noopener">Ibanez et al., 2016</a>; <a href="https://doi.org/10.1093/aob/mcx107" target="_blank" rel="noopener">Ibanez et al., 2018</a>). In New Caledonia, observed tree community richness ranges from 35 to 121 species per hectare within the NC-PIPPN network, peaking at mid-altitude ranges (refer to figure '<a title="1ha Plot Tree Richness Distribution Along Elevation" href="../records/12739730/files/amap_elevation_richness.png?download=1&preview=1" target="_blank" rel="noopener">amap_elevation_richness.png</a>'). Potential richness was assessed using the S-SDM model, with the 80th percentile used as a threshold to distinguish low and high potential richness across three elevation classes: [0 - 400m[, [400 - 900m[, and [900 - 1628m[.</li> </ul> <p>The classification of forest types combines distance from the forest edge and potential richness by elevation into three major categories, as illustrated in the figure '<a title="Illustration of the three forest types" href="../records/12739730/files/amap_forest_types_nc.png?download=1&preview=1" target="_blank" rel="noopener">amap_forest_types_nc.png</a>':</p> <ol> <li><strong>Edge Forest:</strong> Parts of the forest located less than 100 meters from the forest edge.</li> <li><strong>Mature Forest:</strong> Parts of the forest located beyond 100 meters from the edge with a lower potential richness of tree communities.</li> <li><strong>Core Forest:</strong> Parts of the forest located more than 300 meters from the edge with a higher potential richness of tree communities.</li> </ol> <h1>Content</h1> <p>The map is computed from the Forest Map of New Caledonia (v2024) and the Potential Tree Species Richness in the Forests of New Caledonia (v2024). This dataset was produced, analyzed, and verified using a combination of open-source software, including QGIS, PostgreSQL, PostGIS, Python, R, and the GDAL library, all running on Linux. </p> <ul> <li>amap_forest_types_nc.png is a picture illustrating the forest type classification </li> <li>amap_forest_types_nc.zip is a compressed file contains the six essential files for an ESRI-format GIS system, using the WGS84 international coordinate system, and can be uploaded to a spatial database such as PostgreSQL/PostGIS. Each row of the attribute table represents a forest type (a multi-polygon) with associated fields :</li> </ul> <table> <tbody> <tr> <td><strong>Field</strong></td> <td><strong>Type</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td><strong>type</strong></td> <td>TEXT</td> <td>One of the three forest types ("Edge Forest", "Mature Forest", "Core forest")</td> </tr> <tr> <td><strong>area_ha</strong></td> <td>NUMERIC (2 DECIMALS)</td> <td>Area of the multi-polygon in hectares</td> </tr> <tr> <td><strong>description<br></strong></td> <td>TEXT</td> <td>Description of the three forest types</td> </tr> <tr> <td><strong>geom</strong></td> <td>GEOMETRY (MULTIPOLYGON, 4326))</td> <td>Geometry with datum EPSG: 4326 (WGS 84 – World Geodetic System 1984)</td> </tr> </tbody> </table> <h1>Limitations</h1> <p>We caution users that the distinction between the three classes is based on an ecological interpretation and does not reflect directly perceptible breaks in the forest. The ecological transition from the edge to the core of the forest follows multiple gradient modulated by environmental conditions.</p> <p>Moreover, this classification is based on local observations and measurements, which are complex to generalize and extrapolate across a territory as environmentally diverse as New Caledonia. Nevertheless, it allows us to address the impact of fragmentation at the scale of New Caledonia.</p>
Viral Pneumonia Classification Using Machine and Transfer Learning Techniques
<p>Pneumonia is considered a deadly and harmful disease throughout the world. Pneumonia can be lethal if not treated promptly with antibiotics. As a result, early detection of pneumonia increases the likelihood of recovery and lowers mortality. X-rays are one of the most important diagnostic tools for pneumonia. Because of its lower diagnostic costs, the chest X-ray is routinely used to diagnose various lung illnesses. Indeed, diagnosis can be subjective for various reasons, including disease presentation, which might be confusing in chest X-ray images or misdiagnosed as another condition. As a result, the employment of chest X-rays for the diagnoses of pneumonia disease is considered a way forward to fight the challenges being faced with during the examination process and expert readings of results. The dataset comprises 1,067 Pneumonia Chest X-ray images that were curated from the Hopskin Diagnostic Center Nigeria for Research Purposes. This was used to classify Pneumonia disease for pneumonia class encoding. The result yield Pneumonia Disease with High Accuracy, precision and Recall. </p>
Data of the article Analysis of the self-archiving policies of journals in the highest rank category of the Finnish journal classification system within computer science, physics and electronic engineering
<p>The publication forum level three journals representing the three fields of science of computer science, computer science and electrical engineering were identified by utilizing the MinEdu field search filter while searching for the top-ranked journals from the publication channel search (https://www.tsv.fi/julkaisufoorumi/haku.php?lang=en), which is based on Field of Science, Statistics Finland classification (https://www.stat.fi/meta/luokitukset/tieteenala/001-2010/index_en.html). The data were extracted during august 2017 consists of total of 127 individual journals. It is worth noting that circa 30 journals were classified into more than one fields of sciences under scrutiny. First, the journals were divided into representing gold and hybrid model journals. Second, green open access policies of the identified hybrid journals were analyzed using Laakso’s (2014) publisher policy coding framework. Also publishers of the individual journals were identified and subsequently added to the data.</p> <p>NOTE! The data includes the shortest embargo to either institutional or subject repositories. For example, Elsevier had no embargo to opening accepted manuscripts from arXiv subject repository and thus no embargoes to Elsevier's journals are included within this datasheet.</p> <p>Data is in CSV. format</p> <p> </p> <p> </p>
Machine learning for Gravity Spy: Glitch classification and dataset
<p>We present the first version of the training set used in the Gravity Spy citizen science project. This training set, discussed in detail <a href="https://www.sciencedirect.com/science/article/pii/S0020025518301634">here</a>, was utilized to train the convolutional neural network employed in the Gravity Spy project. We anticipate moving forward to release more labelled Gravity Spy data sets, including a refined version of this training set which can be found here <a href="https://doi.org/10.5281/zenodo.1476551">10.5281/zenodo.1476551</a>, and data sets containing the annotations provided by our citizen science volunteers.</p> <p><strong>Data Set Information</strong></p> <p>There are three files provided in this data set</p> <ul> <li><strong>trainingset_v1d0_metadata.csv</strong> <ul> <li>This file has three columns, <em>gravityspy_id, label, </em>and <em>sample_type.</em><em> gravityspy_id </em>is the unique 10 character hash given to every Gravity Spy sample. <em>label</em> is the string label of the sample. <em>sample_type </em>indicates whether this sample was used in the paper for testing training or validating the models. This is provided for those who would like to do direct comparisons to the network described in the paper.</li> </ul> </li> <li><strong>trainingsetv1d0.h5</strong> <ul> <li>This file contains the exact arrays used in the paper for every Gravity Spy sample. Each Gravity Spy sample is defined by four different images with varying temporal duration, <em>0.5, 1.0, 2.0, and 4.0</em> second, respectively. This also determines the naming conventions of the PNGs: <em>interferometer_gravityspyid_spectrogram_duration.png (e.g. H1_Fv3p6eROvA_spectrogram_0.5.png, H1_Fv3p6eROvA_spectrogram_1.0.png, H1_Fv3p6eROvA_spectrogram_2.0.png, H1_Fv3p6eROvA_spectrogram_4.0.png</em>).</li> <li>This file contains all the information needed for each sample in the Gravity Spy dataset (i.e. the label, the sample type of the sample, the unique id of the sample, and the image data for that sample. <ul> <li>/1080Lines/validation/xUEyaWr34c Group<br> /1080Lines/validation/xUEyaWr34c/0.5.png Dataset {1, 140, 170}<br> /1080Lines/validation/xUEyaWr34c/1.0.png Dataset {1, 140, 170}<br> /1080Lines/validation/xUEyaWr34c/2.0.png Dataset {1, 140, 170}<br> /1080Lines/validation/xUEyaWr34c/4.0.png Dataset {1, 140, 170}</li> </ul> </li> </ul> </li> <li><strong>trainingsetv1d0.tar.gz</strong> <ul> <li>Contains the raw PNGs of the Gravity Spy training set.</li> <li>The structure of the folder is <em>/"label"/"sample_type"/"pngs"</em></li> </ul> </li> </ul> <p><strong>Data Set Parsing Information</strong></p> <p>To read and crop out the plot axis and labels of the provided PNGs, the following small python code using scikit-image should work.</p> <p>from skimage import io</p> <p>image_data = io.imread("filename_of_image")</p> <p>x=[66, 532]; y=[105, 671]</p> <p>image_data = image_data[x[0]:x[1], y[0]:y[1], :3]</p>
Unblinded Data for PLAsTiCC Classification Challenge
<p>For classification challenge (https://plasticc.org), the unblinded data files are included here. See PDF note above for more information. The original challenge (Sep 28, 2018 - Dec 17, 2018) was hosted at https://www.kaggle.com/c/PLAsTiCC-2018 .</p>
Meteorological drought lacunarity around the world and its classification
<p>Drought duration strongly depends on the definition thereof. In meteorology, dryness is habitually measured by means of fixed thresholds (e.g. 0.1 or 1 mm usually define dry spells) or climatic mean values (as is the case of the Standardised Precipitation Index), but this also depends on the aggregation time interval considered. However, robust measurements of drought duration are required for analysing the statistical significance of possible changes. Herein we have climatically classified the drought duration around the world according to their similarity to the voids of the Cantor set. Dryness time structure can be concisely measured by the n-index (from the regular/irregular alternation of dry/wet spells), which is closely related to the Gini index and to a Cantor-based exponent. This enables the world’s climates to be classified into six large types based upon a new measure of drought duration. We performed the dry-spell analysis using the full global gridded daily Multi-Source Weighted-Ensemble Precipitation (MSWEP) dataset. The MSWEP combines gauge-, satellite-, and reanalysis-based data to provide reliable precipitation estimates. The study period comprises the years 1979-2016 (total of 45165 days), and a spatial resolution of 0.5º, with a total of 259,197 grid points.</p> <p>FILES </p> <p>1. "drought_class" (geotiff)</p> <p>2. "legend_drought_class" (csv): legend values for drought classification. </p> <p>3. "rasterbrick_index_HurstCantorGini" (geotiff): raster with three layers (Hurst, Cantor and Gini Index applied to dry spells). </p> <p>4. "rasterbrick_nindex_spells" (geotiff): raster with four layers (Dry Spell Spells n-index, maximum expected dry spell<em>Y</em><sub>1 </sub>, mean dry spell and mean wet spell).</p> <p> </p> <p>Projection: "+proj=longlat +ellps=WGS84 +datum=WGS84 +no_defs" (EPSG.4326)</p>
Training data for: CoastSat image classification
<p><strong>CoastSat image classification training data </strong></p> <p>CoastSat is an open-source global shoreline mapping toolbox, available at https://github.com/kvos/CoastSat, which enables users to extract time-series of shoreline change from 30+ years of publicly available satellite imagery (Landsat 5, 7, 8 and Sentinel-2).</p> <p>The automated shoreline extraction relies on a classifier (Multilayer Perceptron from scikit-learn) which labels each pixels on the images with one of four classes: sand, water, white-water and other land features.</p> <p>The data used to train the classifier is stored here, the README.md file provides information on the data organisation and content of each file.</p>
Classification of majority opinions and headnotes written by U.S. Supreme Court Justice Antonin Scalia
<p>This is data to accompany an article by Linda L. Berger and Eric C. Nystrom, "A Rhetorical-Computational Analysis of Justice Antonin Scalia's 'Remarkable Influence': The Unexpected Importance of Deceptively Unanimous and Contested Majority Opinions," <em>Journal of Appellate Practice and Process</em> 20, no. 2 (2020).</p> <p>In "scalia-HN-with-ruletype.tsv," Berger classified each headnote from a Scalia-authored majority opinion as one of the following rhetorical types: argument, scalia rule, or preexisting rule. (See article for further explanation of these categories.) Organized by SCDB ID and headnote number.</p> <p>In "unanimity.tsv," Berger addressed each case with a Scalia-authored majority opinion, assessing the degree of unanimity, which may or may not be the same as that implied by the for/against vote in the case. Fields include case SCDB ID, majority-minority vote, and degree of unanimity.</p> <p>Both data files are in Tab-separated format. For further information, please contact the authors.</p>
IDLAB-UA Dataset for Traffic Classification using Spectrum Data
<p>This dataset contains IQ values of physical layer (L1) packets associated with WLAN transmission and the set of labels that associated each of the packets to properties/features at different radio stack layer (from L1 to L7). </p>
Swahili : News Classification Dataset
<p>Swahili is spoken by 100-150 million people across East Africa. In Tanzania, it is one of two national languages (the other is English) and it is the official language of instruction in all schools. News in Swahili is an important part of the media sphere in Tanzania.</p> <p>News contributes to education, technology, and the economic growth of a country, and news in local languages plays an important cultural role in many Africa countries. In the modern age, African languages in news and other spheres are at risk of being lost as English becomes the dominant language in online spaces.<br> <br> The Swahili news dataset was created to reduce the gap of using the Swahili language to create NLP technologies and help AI practitioners in Tanzania and across Africa continent to practice their NLP skills to solve different problems in organizations or societies related to Swahili language. Swahili News were collected from different websites that provide news in the Swahili language. I was able to find some websites that provide news in Swahili only and others in different languages including Swahili.<br> <br> The dataset was created for a specific task of text classification, this means each news content can be categorized into six different topics (Local news, International news , Finance news, Health news, Sports news, and Entertainment news). The dataset comes with a specified train/test split. The train set contains 75% of the dataset and test set contains 25% of the dataset.</p> <p><strong>Acknowledgment</strong>: This project was supported by the <a href="https://www.k4all.org/project/language-dataset-fellowship/">AI4D language dataset fellowship</a> through K4All and <a href="https://zindi.africa/">Zindi Africa</a>.</p>
Color Classification of Extrasolar Giant Planets: Prospects and Cautions
<p>This dataset contains the reflected light models used in the analysis of<a href="http://adsabs.harvard.edu/abs/2018AJ....156..158B"> Batalha et al. 2018 (Color Classification of Extrasolar Giant Planets: Prospects and Cautions) </a>. The paper explains the full calculation of the models and the parameter space. </p> <p>The public repository <a href="https://github.com/natashabatalha/colorcolor">colorcolor</a> contains <a href="https://github.com/natashabatalha/colorcolor/tree/master/notebooks">notebooks</a> that explain how extract spectra from the database. To show it's simplicity, we post a small code snippet below. </p> <pre><code class="language-python">import colorcolor as c planet_dict = {'cloud': 0.03, 'distance': 0.85, 'gravity': 25, 'metallicity': 0.0, 'phase': 100.0, 'temp': 150} planet = c.select_model(planet_dict) wave, albedo = planet['WAVELN'], planet['GEOMALB']</code></pre> <p> </p> <p>Further explanation can be found at the GitHub repository and within the paper. </p> <p> </p> <p> </p>
Classification of GTP-dependent K-Ras4B active and inactive conformational states
<p>Dataset for the paper: Classification of GTP-dependent K-Ras4B active and inactive conformational states.</p> <p>Cite as: J. Chem. Phys. 158, 000000 (2023); DOI: 10.1063/5.0139181<br> Submitted: 18 December 2022; Accepted: 13 February 2023; Published Online: 13 February 2023.</p> <p>All molecular dynamics and molecular docking data presented, analyzed, and discussed in this paper are available at reasonable<br> requests submitted to the corresponding author. The following data are also available online: (i) file KRas4B_pdbs.zip, a compressed<br> archive including coordinate (.pdb) and structure (.psf) files for KRas-4B WT and D33E proteins, (ii) KRas4B_WT_traj.zip,<br> a compressed archive including trajectory files (.trr) for the WT KRas-4B runs, including 120 "*.trr"-formatted trajectories each corresponding to 40 ns of MD simulation time, and (iii) a sample Python script to generate a free energy plot as shown in Fig. 2.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.