Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

11 results for “bangla”

Learn how ShareScore rates datasets ↗
zenodo44/100

Bangla Information Retrieval Test Collection | Revisiting Anwesha

<p>There are several IR test collections available in English (e.g. http://ir.dcs.gla.ac.uk/resources/test_collections/). Unfortunately, no Gold standard dataset existed for Bangla IR until recently (https://zenodo.org/record/6583149). Our work expands the existing Gold standard dataset by creating 100 query document relevance pairs across a new test collection of 1000 documents. The corpus contains news articles&nbsp;from&nbsp;Ebela, Zee News and Anandabazar&nbsp;Patrika, Vikaspedia and various Bangla travel blogs.&nbsp;The definition of&nbsp;the complexity level of a query is described below:</p> <p>Complexity Level 1:&nbsp;The query contains exact words, phrases or sentence from the document.</p> <p>Complexity Level 2:&nbsp;The query is not present as it is in the document. There is a slight deviation.</p> <p>Complexity Level 3:&nbsp;The query is a generalised phrase capturing the overall story or the document&rsquo;s theme.</p> <p>Complexity Level 4: It is a general query not related to any specific document.</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

BanglaWriting Words Dataset: A Collection of Isolated Word Images from the BanglaWriting multi-purpose Bangla offline-handwriting dataset (WoBW)

<p>The WoBW (Words from BanglaWriting) dataset is a curated collection of isolated word images, adapted from the original BanglaWriting dataset (url:&nbsp;https://data.mendeley.com/datasets/r43wkvdk4w/1).</p> <p>WoBW focuses on individual words extracted from handwritten Bangla text samples in the BanglaWriting corpus, making it a valuable resource for research in word-level Bangla handwriting recognition and related natural language processing tasks.</p> <p>Mridha, Dr. M. F.; Quwsar Ohi, Abu; Ali, M. Ameer; Emon, Mazedul Islam; Kabir, Md Mohsin (2020), &ldquo;BanglaWriting: A multi-purpose offline Bangla handwriting dataset&rdquo;, Mendeley Data, V1, doi: 10.17632/r43wkvdk4w.1</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Bangla Information Retrieval Test Collection

<p>There are several IR test collections available in English (e.g. http://ir.dcs.gla.ac.uk/resources/test_collections/). Unfortunately, there is no Gold standard dataset available to test the effectiveness of Bangla IR. So, we have created a document collection containing 182 short stories, novels, and essays written by Rabindranath Tagore11 and 1000 newspaper articles published in 2013 crawled from the Bangla newspaper Prothom Alo12. The collection contains 100 newspaper articles each from one of the ten categories: বাংলাদেশ/ Bānlādēśa(EN: `Bangladesh&#39;), খেলা/ khēlā(EN: `sports&#39;), বিজ্ঞান ও প্রযুক্তি/ bij&ntilde;āna ō prayukti(EN: `technology&#39;), বিনোদন/ binōdana(EN: `entertainment&#39;), আন্তর্জাতিক/ āntarjātika(EN: `international&#39;), অর্থনীতি/ arthanīti(EN: `economy&#39;), জীবনযাপন/ jībanayāpana(EN: `life-style&#39;), মতামত/ &nbsp;matāmata(EN: `opinion&#39;), শিক্ষা/ śikṣā(EN: `education&#39;) and আমরা/ āmarā(EN:`we-are&#39;). There are 94 queries in the dataset, 26 queries belonging to complexity levels 1 and 2, 19 queries&nbsp;in complexity level 3 and 23 queries in complexity level 4.&nbsp;The definition of&nbsp;the complexity level of a query is described below:</p> <p>Complexity Level 1:&nbsp;The query contains exact words, phrases or sentence from the document.</p> <p>Complexity Level 2:&nbsp;The query is not present as it is in the document. There is a slight deviation.</p> <p>Complexity Level 3:&nbsp;The query is a generalised phrase capturing the overall story or the document&rsquo;s theme.</p> <p>Complexity Level 4: It is a general query not related to any specific document.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Bangla Natural Language Image to Text (BNLIT)

<p>We represented a new Bangla dataset with a Hybrid Recurrent Neural Network model which generated Bangla natural language description of images. This dataset achieved by a large number of images with classification and containing natural language process of images. We conducted experiments on our self-made Bangla Natural Language Image to Text (BNLIT) dataset. Our dataset contained 8,743 images. We made this dataset using Bangladesh perspective images. We used one annotation for each image. In our repository, we added two types of pre-processed data which is 224 &times; 224 and 500 &times; 375 respectively alongside annotations of full dataset. We also added CNN features file of whole dataset in our repository which is features.pkl.</p>

opencc-by-4.0Dec 2018View details →
zenodo36/100

Bangla Text Normalization Benchmark Dataset

<p>This is a small benchmark dataset for Bangla Text Normalization.&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

SignBD-Word: Video-Based Bangla Word-Level Sign Language Dataset

<p>Bangla sign language (BdSL) is a complete and independent natural sign language with its own linguistic characteristics. While there exists video datasets for well-known sign languages, there is currently no available dataset for word-level BdSL. In this study, we present a video-based word-level dataset for Bangla sign language, called SignBD-Word, consisting of 6000 sign videos representing 200 unique words. The dataset includes full and upper-body views of the signers, along with 2D body pose information. This dataset can also be used as a benchmark for testing sign video classification algorithms.<br><br>Official Train Test Spllit (for both RGB and bodypose) can be found from the following link:&nbsp;<br>https://sites.google.com/view/signbd-word/dataset<br><br>This dataset is part of the following paper:<br>A. Sams, A. H. Akash and S. M. M. Rahman, "SignBD-Word: Video-Based Bangla Word-Level Sign Language and Pose Translation," 2023 14th International Conference on Computing Communication and Networking Technologies (ICCCNT), Delhi, India, 2023, pp. 1-7, doi: 10.1109/ICCCNT56998.2023.10306914.<br><br>Download the corresponding paper from this link:<br>https://asnsams.github.io/Publications.html</p>

opencc-by-sa-4.0Jun 2022View details →
dryad36/100

BdSL47: A complete dataset of sign alphabet and digits of Bangla Sign Language (BdSL) using depth information via MediaPipe

<p><strong>BdSL47</strong> is the first open-access complete dataset in Bangla Sign Language that contains hand signs from both 10 sign digits (from sign ০ to sign ৯) and 37 sign alphabet (from sign অ to sign ँ).</p> <p>Dataset summary :</p> <ul> <li>100 RGB images per sign (total 47 signs) from each of 10 users</li> <li>Total input images : 100×47×10 = 47000</li> <li>Input images are processed via MediaPipe, which provided <ul> <li>an output image with hand key-points being detected</li> <li>3D coordinate values of 21 predefined key-points</li> <li>Total 63 coordinate values for each sample</li> </ul> </li> <li>The values are stored in csv files</li> <li>1 CSV file contains values from 100 samples of 1 sign from 1 user</li> <li>Total CSV files : 47×10 = 470</li> </ul> <p>The dataset has been made public for further research purposes. It is also available upon request <a href="https://drive.google.com/drive/u/8/folders/1wmJUlgWUrWNnOvzuL8Ci82Hm3zUx4wS-" rel="noopener">here</a>.</p>

opencc-zeroDec 2021View details →
zenodo36/100

Bangla License Plate Dataset 2.5k

<p>This comprehensive dataset of 2519 Bangladeshi vehicle images with clearly legible Bangla license plates. The dataset contains preprocessed license plate images for detection and recognition systems.</p> <p>Upon downloading, you will get four directories:</p> <p>1. training: Contains 2211 high-resolution Bangla license plate images of variable sizes cropped from pictures with license plates. All the files are in jpg format.</p> <p>2. training_data: Contains 2211&nbsp;Bangla license plate images of ‪256 x 192‬ size, resized from images from the training directory. All the files are in png format.</p> <p>3. testing: Contains 200 high-resolution Bangla license plate images of variable sizes cropped from pictures with license plates. All the files are in jpg format.</p> <p>4. training_data: Contains 200 Bangla license plate images of ‪256 x 192‬ size, resized from images from the testing directory. All the files are in png format.&nbsp;</p> <p>5. new data: Contains extra&nbsp;308 images of variable size,&nbsp;</p>

opencc-by-4.0Sep 2022View details →
dryad36/100

BdSL47: A complete dataset of sign alphabet and digits of Bangla Sign Language (BdSL) using depth information via MediaPipe

Open the record for dataset details and reuse information.

publicSep 2022View details →
zenodo28/100

A Computer Graphics Approach to Creating New Method for Generating 3D Gesture Animations in Bangla Sign Language via HamNoSys to SiGML Conversion

<p>To prepare the system, we employed 94 classes of data. In this dataset, there are 13 Bangla numerical data classes, 36 Bangla alphabet data classes, and 41 Bangla word data classes. Every class of data contains different data types like Hand Shape, Hand Orientation, Hand Movement, Notations, etc. These data were created in SiGML tags. Every class of data is unique and different from others. The system was prepared using BdSL. And BdSL is an uncommon and unique sign language, among others. That's why every class of data is unique and created by us. We search HamNoSys notations for Bangla alphabets, words, and numbers in English HamNoSys datasets (almost 6,000 data). However, we find only 20% of the data, which is quite similar to BdSL. We create 80% HamNoSys notation for BdSL and we modify the matching 20% notations. Then, we converted them into SiGML and made the data classes. This was a big challenge in our research.&nbsp;&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo28/100

Bangla-REX: A Distinct Dataset for Bangla Relation Extraction

<p>The dataset is grounded in theoretical and methodological frameworks that emphasize the importance of structured knowledge bases and annotated corpora for effective relation extraction. To generate this dataset, we compiled a comprehensive Bangla Knowledge Base (KB) consisting of 63,256 entries, which serves as a foundation for automating the labeling process with relation tags. The corpus itself is extensive, comprising 90,441 text entries that have been meticulously processed to include Named Entity Recognition (NER) and Part-of-Speech (POS) tagging, ensuring that it is ready for immediate use in relation extraction tasks.<br>Additionally, we developed mnemonics for 440 distinct locations in Bangla, specifically tailored to enhance performance in location-based relation extraction. These mnemonics are particularly beneficial in the context of distant supervision-based relation extraction, where they help in establishing clear associations between locations and their corresponding entities or contexts.</p>

opencc-by-4.0Jul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record