TBX 11

TBX 11

Datasets

TBX 11

File

TBX 11

Use Case

TBX 11

Description

Explore the TBX 11 dataset, featuring 11,200 X-ray images with bounding box annotations for tuberculosis detection. Categorized into Healthy, Sick but Non-TB, Active TB, Latent TB, and Uncertain TB, this dataset supports training and evaluation of TB detection algorithms. Includes detailed training, validation, and testing subsets, along with JSON and XML annotations. Ideal for advancing computer-aided TB diagnosis.

TBX 11

Description:

The TBX 11 dataset comprises 11,200 X-ray images annotated with bounding boxes highlighting tuberculosis (TB) areas. Each image has a resolution of 512×512 pixels and falls into one of five categories: Healthy, Sick but Non-TB, Active TB, Latent TB, and Uncertain TB. The dataset is divided into training, validation, and testing subsets with 6,600, 1,800, and 2,800 images, respectively. We provide the image lists for these sets in ‘TBX11K_train.txt’, ‘TBX11K_val.txt’, and ‘TBX11K_trainval.txt’.

Download Dataset

Due to the high cost of obtaining TB data, we have combined data from four smaller TB datasets: DA, DB, Montgomery, and Shenzhen. These datasets contain 156, 150, 138, and 662 X-ray images, respectively. We use parts of each dataset for training and validation, as indicated in the ‘imgs/extra/’ folder. The remaining images from these datasets, along with our 2,800 testing X-rays, form the new testing set listed in ‘all_test.txt’. We also merge our data with the training and validation X-rays from these four datasets to create updated training, validation, and combined training+validation sets, available as ‘all_train.txt’, ‘all_val.txt’, and ‘all_trainval.txt’.

Please note that the ground truth for the testing set will not be released, as it is used for an online competition in computer-aided tuberculosis diagnosis. For model development, we recommend using the ‘all_train’ set for training and ‘all_val’ set for validation. When submitting results, train your model on the ‘all_trainval’ set and test on the ‘all_test’ set.

Dataset Composition and Structure

The dataset includes a total of 1,000 CAPTCHA images, each carefully generated with the following characteristics:

  • Image Format: PNG
  • Resolution: 180 × 50 pixels
  • Character Length: 5 characters per image
  • Character Set: Uppercase letters (A–Z) and digits (0–9)

Key Features of the Dataset

  • 1,000 labeled CAPTCHA images
  • Fixed 5-character structure for consistency
  • Built-in labeling via filenames
  • Suitable for OCR and deep learning tasks
  • Lightweight and easy to use for experimentation

Applications and Use Cases

The 5-Character CAPTCHA Dataset can be used in several important applications. For example:

  • CAPTCHA Recognition Systems: Train models to decode CAPTCHA images
  • Security Testing: Evaluate the strength of CAPTCHA mechanisms
  • OCR Development: Improve recognition of distorted text
  • Web Automation: Enable automated workflows involving CAPTCHA solving

Conclusion


The 5-Character CAPTCHA Dataset is a valuable resource for developing and testing CAPTCHA recognition systems. Overall, it provides a structured and practical dataset for OCR and security-related applications. More importantly, it helps improve the performance of machine learning models in handling complex visual text challenges.

This dataset is sourced from Kaggle.

Contact Us

Technology

Quality Data Creation

Technology

Guaranteed TAT

Technology

ISO 9001:2015, ISO/IEC 27001:2013 Certified

Technology

HIPAA Compliance

Technology

GDPR Compliance

Technology

Compliance and Security

Let's Discuss your Data collection Requirement With Us

To get a detailed estimation of requirements please reach us.

Scroll to Top

Please provide your details to download the Dataset.