Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

64

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

64 results for “1k”

Learn how ShareScore rates datasets ↗
zenodo36/100

Image Data Part 3 of AbdomenCT-1K: Is Abdominal Organ Segmentation A Solved Problem

<p>Image Data Part 3 of AbdomenCT-1K: Is Abdominal Organ Segmentation A Solved Problem</p> <p><a href="https://ieeexplore.ieee.org/document/9497733/">Paper</a>: https://ieeexplore.ieee.org/document/9497733/</p> <p>&nbsp;</p> <p>Other two parts: <a href="https://zenodo.org/record/5903099">Part 1</a>,&nbsp;<a href="https://zenodo.org/record/5903846">Part 2</a></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

AbdomenCT-1K: Continual Learning Benchmark

<p>This is the dataset of AbdomenCT-1K: Continual Learning Benchmark.</p> <p>Related paper: <a href="https://ieeexplore.ieee.org/document/9497733/">https://ieeexplore.ieee.org/document/9497733/</a></p> <p>Benchmark homepage: https://abdomenct-1k-continual-learning.grand-challenge.org/</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

LD Estimated from 1k Genomes CEU Population

<p>These data contain estimated pairwise r^2 for variants with allele frequency greater than 0.05 in the 1000 Genomes CEU population. They were estimated using LDshrink (https://github.com/stephenslab/LDshrink). R^2 is only reported when the estimate is greater than 0.1.</p> <p>&nbsp;</p> <p>For each chromosome there are two files:</p> <p>chr&lt;chr&gt;_AF0.5_0.1.RDS is an R object containing a data frame with three columns: rowsnp, colsnp, and r2</p> <p>chr&lt;chr&gt;_AF0.5_snpdata.RDS is an R object containing a data frame with information for every SNP meeting the allele frequency cutoff.</p>

opencc-by-4.0Oct 2018View details →
zenodo36/100

Myanmar Agriculture 1K

<p>The "Myanmar Agriculture 1K" Dataset is curated to build a knowledge bank for further studies in Natural Language Processing in the Burmese Language and to train instruction fine-tuned language model for the Burmese Language.</p> <p>Moreover, this dataset is a motivation to move the Burmese language from a low-resource language to a high-resource language.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Blog-1K

<p>The Blog-1K corpus is a redistributable authorship identification testbed for contemporary English prose.&nbsp;It has 1,000 candidate authors, 16K+ posts, and a pre-defined data split (train/dev/test proportional to ca. 8:1:1).&nbsp;It is a subset of the <a href="http://www.kaggle.com/datasets/rtatman/blog-authorship-corpus">Blog Authorship Corpus</a>&nbsp;from Kaggle. The MD5 for Blog-1K is &#39;0a9e38740af9f921b6316b7f400acf06&#39;.</p> <p>1. Preprocessing</p> <p>We first filter out texts shorter than 1,000 characters. Then we select one thousand&nbsp;authors whose writings meet the following criteria:<br> - accumulatively at least 10,000 characters,<br> - accumulatively at most 49,410 characters,<br> - accumulatively at least 16 posts,<br> - accumulatively at most 40 posts, and&nbsp;<br> - each text has at least 50 function words found in the Koppel512 list (to filter out non-English prose).</p> <p>Blog-1K has three columns: &#39;id&#39;, &#39;text&#39;, and &#39;split&#39;, where &#39;id&#39;&nbsp;corresponds to its parent corpus.</p> <p>2. Statistics</p> <p>Its creation and statistics can be found in <a href="https://codeberg.org/haining/Blog-1K/src/branch/main/blog-1k_generation_and_stats.ipynb">the Jupyter Notebook</a>.</p> <table> <tbody> <tr> <td>Split</td> <td># Authors</td> <td># Posts</td> <td># Characters</td> <td>Avg. Characters Per Author (Std.)</td> <td>Avg. Characters Per Post (Std.)</td> </tr> <tr> <td>Train</td> <td>1,000</td> <td>16,132</td> <td>30,092,057</td> <td>30,092 (5,884)</td> <td>1,865 (1,007)</td> </tr> <tr> <td>Validation</td> <td>935</td> <td>2,017</td> <td>3,755,362</td> <td>4,016 (2,269)</td> <td>1,862 (999)</td> </tr> <tr> <td>Test</td> <td>924</td> <td>2,017</td> <td>3,732,448</td> <td>4,039 (2,188)</td> <td>1,850 (936)</td> </tr> </tbody> </table> <p><br> 3. Usage</p> <pre><code class="language-python">import pandas as pd df = pd.read_csv('blog1000.csv.gz', compression='infer') # read in training data train_text, train_label = zip(*df.loc[df.split=='train'][['text', 'id']].itertuples(index=False))</code></pre> <p>&nbsp;</p> <p>4. License<br> All the materials is licensed under the ISC License.</p> <p><br> 5. Contact<br> Please contact <a href="mailto:hw56@indiana.edu">its maintainer</a>&nbsp;for questions.</p>

openisc-licenseDec 2022View details →
zenodo36/100

Training datasets of size 1K, 3K, and 9K for four utility function variants

<p>The datasets include 1K, 3K, and 9K training data points that are used for training utility-change prediction models in systems with ground truth from four different mathematical complexities:</p> <ul> <li>Linear utility function</li> <li>Saturating utility function</li> <li>Discontinuous utility function</li> <li>Combined utility function ( combining the above three variants)</li> </ul>

opencc-by-4.0Feb 2023View details →
zenodo36/100

VisualAtom-1k

<p>VisualAtom is a cutting-edge artificial image dataset, specifically designed for pre-training deep learning models for image recognition tasks, such as Vision Transformers. Generated through the innovative synthesis of geometric contours, VisualAtom offers a rich and diverse synthetic images, achieved by assigning various stationary waveforms to the contour lines.<br> The primary goal of VisualAtom is to provide pre-training effect that rivals large real image datasets, such as ImageNet and JFT. By offering a wide variety of synthesized geometric contours, VisualAtom allows deep learning models to develop a robust understanding of diverse visual structures, thus enabling them to perform at comparable levels to models pre-trained on real images. Furthermore, the datasets and models are licensed for commercial use and are not restricted to educational or academic use only.<br> To facilitate easy access and customization, the generation scripts and usage instructions for VisualAtom are available on our GitHub page at <a href="https://github.com/masora1030/CVPR2023-FDSL-on-VisualAtom">https://github.com/masora1030/CVPR2023-FDSL-on-VisualAtom</a>. Users are encouraged to explore the repository and generate and pre-train on VisualAtom to their specific needs, further expanding the possibilities of VisualAtom.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Proarna montevidensis Berg, 1882(Figura 1k) in Cicádidos (Insecta: Hemiptera: Cicadidae) Del Museo Provincial De Ciencias Naturales "Florentino Ameghino" Santa Fe, Argentina

Proarna montevidensis Berg, 1882(Figura 1k)

opencc-by-4.0Jul 2007View details →
zenodo36/100

2010_2024_ERA5_Precipitation_Rainfall_FourierProcessed_1k_WG

<p>This is a set of images produced by Temporal Fourier Analysis (TFA) of ERA5 data:</p> <p>ERA5: Total Precipitation&nbsp;</p> <p>The imagery summarises some key environmental indicators, incorporating seasonal dynamics, for the whole world<br>This series of ERA5 data, processed according to Scharlemann et al (2008), has been updated to include imagery from 2010 to 2024.&nbsp;</p> <p>&nbsp;</p> <p>Precipitation from the ERA5 reanalysis archive supplied by the European Centre for Medium Range Weather Forecasting for 2010 - 2022. Abstract: Precipitation from the ERA5 reanalysis archive supplied by the European Centre for Medium Range Weather Forecasting. The original data is at 0.25 degree resolution and was downscaled by ERA extraction algorithms to 1km resolution, then downloaded. The daily data have been aggregated to dekadal, monthly, and annual datasets to match the outputs produced by NASA from the MODIS imagery temperature and vegetation Index datasets. The resolution was also chosen to match these MODIS datasets.</p> <h4>Process:</h4> <p>Image values were extracted from ERA5 ( Total precipitation) at 1 km resolution imagery from 2010 to 2024.&nbsp; Each parameter extract dataset was then processed by a temporal Fourier processing algorithm. A stepwise system of thresholds and interpolations screened erroneous values and bridged gaps in the time series. The smoothed series was sampled at 5-day intervals and transformed into a set of sine curves describing annual, bi-annual, and tri-annual fluctuations. For each of these curves, the Fourier algorithm generated images expressing the amplitude, phase, and variance. Other output recorded the mean, minimum, and maximum of the time series, and errors measured during the Fourier transform. For a detailed description of the Fourier algorithm and its output, please see the article by Scharlemann et al., 2008 (<a href="https://doi.org/10.1371/journal.pone.0001408">https://doi.org/10.1371/journal.pone.0001408</a>)&nbsp;&nbsp;<br>Idrisi rasters were converted to GeoTIFF format in order to give data users more flexibility. Sea pixels were masked with a VIIRS land/sea layer.&nbsp;</p> <p>&nbsp;</p> <p>This new ERA5 Dataset is used as an update and continuation of our MODIS TFA product and can be utilised in the same way.&nbsp;</p> <p>Projection + EPSG code:</p> <p>Latitude-Longitude/WGS84 (EPSG: 4326)</p> <p>Extent &nbsp; &nbsp;-180.0000000000000000,-90.0000000000000000 : 179.9999999999998295,89.9999999999999147</p> <h4>File names:</h4> <p><br>The wg at the start of each file name indicates that the image covers the whole world in the E4warning&nbsp; and is in geographic projection. 04 refers to the year timeline of 2010-2024.<br><br>The next two characters identify the channel:<br>20 - Monthly Total Precipitation<br><br>The last two characters of each file name denote the output from Fourier processing:<br>a0 - mean<br>mn - minimum<br>mx - maximum<br>a1 - amplitude of annual cycle<br>a2 - amplitude of bi-annual cycle<br>a3 - amplitude of tri-annual cycle<br>p1 - phase of annual cycle<br>p2 - phase of bi-annual cycle<br>p3 - phase of tri-annual cycle<br>d1 - variance in annual cycle<br>d2 - variance in bi-annual cycle<br>d3 - variance in tri-annual cycle<br>da - combined variance in annual, bi-annual, and tri-annual cycles<br>vr - variance in raw data<br><br>Parameter Fourier Variable Image values are<br>ERA5&nbsp; A0, A1, A2, A3, Min, Max, Vr Reflectance values&nbsp; monthly total precipitation in mm<br>ALL D1,D2,D3,Da Percentages<br>ALL E1,E2,E3 Percentages<br>ALL P1,P2.P3 Months*100. (Jan=100)</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

[1k] Head of the Buddha

Heavily optimised version of this scan https://skfb.ly/FtIw from Rijksmuseum, Amsterdam for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab

opencc-by-nc-sa-2.0Nov 2018View details →
zenodo32/100

[1k] Bust of Virginia Woolf

Heavily optimised version of this scan https://skfb.ly/GZTw from Tavistock Square, London for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab

opencc-by-nc-sa-2.0Nov 2018View details →
zenodo32/100

[1k] Marble torso from a statue of Dionysos

Heavily optimised version of this scan https://skfb.ly/KwpI from the Fitzwilliam Musuem, Cambridge, UK for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab

opencc-by-nc-sa-2.0Nov 2018View details →
zenodo32/100

[1k] Head of Amenhotep III

Heavily optimised version of this scan https://skfb.ly/67vFW from the Cleveland Museum of Art for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab

opencc-by-nc-sa-2.0Nov 2018View details →
zenodo32/100

Processed KuaiRand-1K dataset for the paper: Large-Scale Multi-Domain Recommendation: an Automatic Domain Feature Extraction and Personalized Integration Framework

<p>The original public dataset is published in https://zenodo.org/records/10439422, we processed the KuaiRand-1K dataset for the paper: Large-Scale Multi-Domain Recommendation: an Automatic Domain Feature Extraction and Personalized &nbsp;Integration Framework.</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

[1k] Head of an Ox on a Tree Trunk

Heavily optimised version of this scan https://skfb.ly/68X9V from the Victoria and Albert Museum, London for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab

opencc-by-nc-sa-2.0Nov 2018View details →
zenodo32/100

[1k] Ninfa (Mulva)

Heavily optimised version of this scan https://skfb.ly/ENKR for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab

opencc-by-nc-sa-2.0Nov 2018View details →
zenodo32/100

[1k] Fossilized Ammonite

Heavily optimised version of this scan https://skfb.ly/6qrWs for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab

opencc-by-nc-sa-2.0Nov 2018View details →
zenodo32/100

1K_genomes_reference_panel

<p>The 1K Genomes Reference Panel for Europeans was downloaded from the GitHub repository of LOGODetect (https://github.com/ghm17/LOGODetect).</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Variant metadata for 1K genomes reference panel

<p>1K genomes reference panel - variant metadata</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

[1k] Basalt Column Base Depicting a Sphinx

Heavily optimised version of this scan https://skfb.ly/6tQEB from the Pergamon Museum, Berlin for the Mozilla Hubs Challenge: https://blog.sketchfab.com/vr-design-challenge-mozilla-hubs-clubhouse 1000 faces, 1024x1024 colour and normal maps. Source: Objaverse 1.0 / Sketchfab

opencc-by-nc-sa-2.0Nov 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record