Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
64
datasets available to search
ShareScore release 0.7.1
Dataset results
64 results for “1k”
Image Data Part 3 of AbdomenCT-1K: Is Abdominal Organ Segmentation A Solved Problem
<p>Image Data Part 3 of AbdomenCT-1K: Is Abdominal Organ Segmentation A Solved Problem</p> <p><a href="https://ieeexplore.ieee.org/document/9497733/">Paper</a>: https://ieeexplore.ieee.org/document/9497733/</p> <p> </p> <p>Other two parts: <a href="https://zenodo.org/record/5903099">Part 1</a>, <a href="https://zenodo.org/record/5903846">Part 2</a></p> <p> </p> <p> </p>
AbdomenCT-1K: Continual Learning Benchmark
<p>This is the dataset of AbdomenCT-1K: Continual Learning Benchmark.</p> <p>Related paper: <a href="https://ieeexplore.ieee.org/document/9497733/">https://ieeexplore.ieee.org/document/9497733/</a></p> <p>Benchmark homepage: https://abdomenct-1k-continual-learning.grand-challenge.org/</p>
LD Estimated from 1k Genomes CEU Population
<p>These data contain estimated pairwise r^2 for variants with allele frequency greater than 0.05 in the 1000 Genomes CEU population. They were estimated using LDshrink (https://github.com/stephenslab/LDshrink). R^2 is only reported when the estimate is greater than 0.1.</p> <p> </p> <p>For each chromosome there are two files:</p> <p>chr<chr>_AF0.5_0.1.RDS is an R object containing a data frame with three columns: rowsnp, colsnp, and r2</p> <p>chr<chr>_AF0.5_snpdata.RDS is an R object containing a data frame with information for every SNP meeting the allele frequency cutoff.</p>
Myanmar Agriculture 1K
<p>The "Myanmar Agriculture 1K" Dataset is curated to build a knowledge bank for further studies in Natural Language Processing in the Burmese Language and to train instruction fine-tuned language model for the Burmese Language.</p> <p>Moreover, this dataset is a motivation to move the Burmese language from a low-resource language to a high-resource language.</p>
Blog-1K
<p>The Blog-1K corpus is a redistributable authorship identification testbed for contemporary English prose. It has 1,000 candidate authors, 16K+ posts, and a pre-defined data split (train/dev/test proportional to ca. 8:1:1). It is a subset of the <a href="http://www.kaggle.com/datasets/rtatman/blog-authorship-corpus">Blog Authorship Corpus</a> from Kaggle. The MD5 for Blog-1K is '0a9e38740af9f921b6316b7f400acf06'.</p> <p>1. Preprocessing</p> <p>We first filter out texts shorter than 1,000 characters. Then we select one thousand authors whose writings meet the following criteria:<br> - accumulatively at least 10,000 characters,<br> - accumulatively at most 49,410 characters,<br> - accumulatively at least 16 posts,<br> - accumulatively at most 40 posts, and <br> - each text has at least 50 function words found in the Koppel512 list (to filter out non-English prose).</p> <p>Blog-1K has three columns: 'id', 'text', and 'split', where 'id' corresponds to its parent corpus.</p> <p>2. Statistics</p> <p>Its creation and statistics can be found in <a href="https://codeberg.org/haining/Blog-1K/src/branch/main/blog-1k_generation_and_stats.ipynb">the Jupyter Notebook</a>.</p> <table> <tbody> <tr> <td>Split</td> <td># Authors</td> <td># Posts</td> <td># Characters</td> <td>Avg. Characters Per Author (Std.)</td> <td>Avg. Characters Per Post (Std.)</td> </tr> <tr> <td>Train</td> <td>1,000</td> <td>16,132</td> <td>30,092,057</td> <td>30,092 (5,884)</td> <td>1,865 (1,007)</td> </tr> <tr> <td>Validation</td> <td>935</td> <td>2,017</td> <td>3,755,362</td> <td>4,016 (2,269)</td> <td>1,862 (999)</td> </tr> <tr> <td>Test</td> <td>924</td> <td>2,017</td> <td>3,732,448</td> <td>4,039 (2,188)</td> <td>1,850 (936)</td> </tr> </tbody> </table> <p><br> 3. Usage</p> <pre><code class="language-python">import pandas as pd df = pd.read_csv('blog1000.csv.gz', compression='infer') # read in training data train_text, train_label = zip(*df.loc[df.split=='train'][['text', 'id']].itertuples(index=False))</code></pre> <p> </p> <p>4. License<br> All the materials is licensed under the ISC License.</p> <p><br> 5. Contact<br> Please contact <a href="mailto:hw56@indiana.edu">its maintainer</a> for questions.</p>
Training datasets of size 1K, 3K, and 9K for four utility function variants
<p>The datasets include 1K, 3K, and 9K training data points that are used for training utility-change prediction models in systems with ground truth from four different mathematical complexities:</p> <ul> <li>Linear utility function</li> <li>Saturating utility function</li> <li>Discontinuous utility function</li> <li>Combined utility function ( combining the above three variants)</li> </ul>
VisualAtom-1k
<p>VisualAtom is a cutting-edge artificial image dataset, specifically designed for pre-training deep learning models for image recognition tasks, such as Vision Transformers. Generated through the innovative synthesis of geometric contours, VisualAtom offers a rich and diverse synthetic images, achieved by assigning various stationary waveforms to the contour lines.<br> The primary goal of VisualAtom is to provide pre-training effect that rivals large real image datasets, such as ImageNet and JFT. By offering a wide variety of synthesized geometric contours, VisualAtom allows deep learning models to develop a robust understanding of diverse visual structures, thus enabling them to perform at comparable levels to models pre-trained on real images. Furthermore, the datasets and models are licensed for commercial use and are not restricted to educational or academic use only.<br> To facilitate easy access and customization, the generation scripts and usage instructions for VisualAtom are available on our GitHub page at <a href="https://github.com/masora1030/CVPR2023-FDSL-on-VisualAtom">https://github.com/masora1030/CVPR2023-FDSL-on-VisualAtom</a>. Users are encouraged to explore the repository and generate and pre-train on VisualAtom to their specific needs, further expanding the possibilities of VisualAtom.</p>
Proarna montevidensis Berg, 1882(Figura 1k) in Cicádidos (Insecta: Hemiptera: Cicadidae) Del Museo Provincial De Ciencias Naturales "Florentino Ameghino" Santa Fe, Argentina
Proarna montevidensis Berg, 1882(Figura 1k)
2010_2024_ERA5_Precipitation_Rainfall_FourierProcessed_1k_WG
<p>This is a set of images produced by Temporal Fourier Analysis (TFA) of ERA5 data:</p> <p>ERA5: Total Precipitation </p> <p>The imagery summarises some key environmental indicators, incorporating seasonal dynamics, for the whole world<br>This series of ERA5 data, processed according to Scharlemann et al (2008), has been updated to include imagery from 2010 to 2024. </p> <p> </p> <p>Precipitation from the ERA5 reanalysis archive supplied by the European Centre for Medium Range Weather Forecasting for 2010 - 2022. Abstract: Precipitation from the ERA5 reanalysis archive supplied by the European Centre for Medium Range Weather Forecasting. The original data is at 0.25 degree resolution and was downscaled by ERA extraction algorithms to 1km resolution, then downloaded. The daily data have been aggregated to dekadal, monthly, and annual datasets to match the outputs produced by NASA from the MODIS imagery temperature and vegetation Index datasets. The resolution was also chosen to match these MODIS datasets.</p> <h4>Process:</h4> <p>Image values were extracted from ERA5 ( Total precipitation) at 1 km resolution imagery from 2010 to 2024. Each parameter extract dataset was then processed by a temporal Fourier processing algorithm. A stepwise system of thresholds and interpolations screened erroneous values and bridged gaps in the time series. The smoothed series was sampled at 5-day intervals and transformed into a set of sine curves describing annual, bi-annual, and tri-annual fluctuations. For each of these curves, the Fourier algorithm generated images expressing the amplitude, phase, and variance. Other output recorded the mean, minimum, and maximum of the time series, and errors measured during the Fourier transform. For a detailed description of the Fourier algorithm and its output, please see the article by Scharlemann et al., 2008 (<a href="https://doi.org/10.1371/journal.pone.0001408">https://doi.org/10.1371/journal.pone.0001408</a>) <br>Idrisi rasters were converted to GeoTIFF format in order to give data users more flexibility. Sea pixels were masked with a VIIRS land/sea layer. </p> <p> </p> <p>This new ERA5 Dataset is used as an update and continuation of our MODIS TFA product and can be utilised in the same way. </p> <p>Projection + EPSG code:</p> <p>Latitude-Longitude/WGS84 (EPSG: 4326)</p> <p>Extent -180.0000000000000000,-90.0000000000000000 : 179.9999999999998295,89.9999999999999147</p> <h4>File names:</h4> <p><br>The wg at the start of each file name indicates that the image covers the whole world in the E4warning and is in geographic projection. 04 refers to the year timeline of 2010-2024.<br><br>The next two characters identify the channel:<br>20 - Monthly Total Precipitation<br><br>The last two characters of each file name denote the output from Fourier processing:<br>a0 - mean<br>mn - minimum<br>mx - maximum<br>a1 - amplitude of annual cycle<br>a2 - amplitude of bi-annual cycle<br>a3 - amplitude of tri-annual cycle<br>p1 - phase of annual cycle<br>p2 - phase of bi-annual cycle<br>p3 - phase of tri-annual cycle<br>d1 - variance in annual cycle<br>d2 - variance in bi-annual cycle<br>d3 - variance in tri-annual cycle<br>da - combined variance in annual, bi-annual, and tri-annual cycles<br>vr - variance in raw data<br><br>Parameter Fourier Variable Image values are<br>ERA5 A0, A1, A2, A3, Min, Max, Vr Reflectance values monthly total precipitation in mm<br>ALL D1,D2,D3,Da Percentages<br>ALL E1,E2,E3 Percentages<br>ALL P1,P2.P3 Months*100. (Jan=100)</p>
[1k] Head of the Buddha
Heavily optimised version of this scan https://skfb.ly/FtIw from Rijksmuseum, Amsterdam for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab
[1k] Bust of Virginia Woolf
Heavily optimised version of this scan https://skfb.ly/GZTw from Tavistock Square, London for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab
[1k] Marble torso from a statue of Dionysos
Heavily optimised version of this scan https://skfb.ly/KwpI from the Fitzwilliam Musuem, Cambridge, UK for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab
[1k] Head of Amenhotep III
Heavily optimised version of this scan https://skfb.ly/67vFW from the Cleveland Museum of Art for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab
Processed KuaiRand-1K dataset for the paper: Large-Scale Multi-Domain Recommendation: an Automatic Domain Feature Extraction and Personalized Integration Framework
<p>The original public dataset is published in https://zenodo.org/records/10439422, we processed the KuaiRand-1K dataset for the paper: Large-Scale Multi-Domain Recommendation: an Automatic Domain Feature Extraction and Personalized Integration Framework.</p>
[1k] Head of an Ox on a Tree Trunk
Heavily optimised version of this scan https://skfb.ly/68X9V from the Victoria and Albert Museum, London for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab
[1k] Ninfa (Mulva)
Heavily optimised version of this scan https://skfb.ly/ENKR for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab
[1k] Fossilized Ammonite
Heavily optimised version of this scan https://skfb.ly/6qrWs for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab
1K_genomes_reference_panel
<p>The 1K Genomes Reference Panel for Europeans was downloaded from the GitHub repository of LOGODetect (https://github.com/ghm17/LOGODetect).</p>
Variant metadata for 1K genomes reference panel
<p>1K genomes reference panel - variant metadata</p>
[1k] Basalt Column Base Depicting a Sphinx
Heavily optimised version of this scan https://skfb.ly/6tQEB from the Pergamon Museum, Berlin for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.