TBX 11
Home » Dataset Download » TBX 11
TBX 11
Datasets
TBX 11
File
TBX 11
Use Case
TBX 11
Description
Explore the TBX 11 dataset, featuring 11,200 X-ray images with bounding box annotations for tuberculosis detection. Categorized into Healthy, Sick but Non-TB, Active TB, Latent TB, and Uncertain TB, this dataset supports training and evaluation of TB detection algorithms. Includes detailed training, validation, and testing subsets, along with JSON and XML annotations. Ideal for advancing computer-aided TB diagnosis.
Description:
The TBX 11 dataset comprises 11,200 X-ray images annotated with bounding boxes highlighting tuberculosis (TB) areas. Each image has a resolution of 512×512 pixels and falls into one of five categories: Healthy, Sick but Non-TB, Active TB, Latent TB, and Uncertain TB. The dataset is divided into training, validation, and testing subsets with 6,600, 1,800, and 2,800 images, respectively. We provide the image lists for these sets in ‘TBX11K_train.txt’, ‘TBX11K_val.txt’, and ‘TBX11K_trainval.txt’.
Download Dataset
Due to the high cost of obtaining TB data, we have combined data from four smaller TB datasets: DA, DB, Montgomery, and Shenzhen. These datasets contain 156, 150, 138, and 662 X-ray images, respectively. We use parts of each dataset for training and validation, as indicated in the ‘imgs/extra/’ folder. The remaining images from these datasets, along with our 2,800 testing X-rays, form the new testing set listed in ‘all_test.txt’. We also merge our data with the training and validation X-rays from these four datasets to create updated training, validation, and combined training+validation sets, available as ‘all_train.txt’, ‘all_val.txt’, and ‘all_trainval.txt’.
Please note that the ground truth for the testing set will not be released, as it is used for an online competition in computer-aided tuberculosis diagnosis. For model development, we recommend using the ‘all_train’ set for training and ‘all_val’ set for validation. When submitting results, train your model on the ‘all_trainval’ set and test on the ‘all_test’ set.
Dataset Composition and Structure
The dataset includes a total of 1,000 CAPTCHA images, each carefully generated with the following characteristics:
- Image Format: PNG
- Resolution: 180 × 50 pixels
- Character Length: 5 characters per image
- Character Set: Uppercase letters (A–Z) and digits (0–9)
Key Features of the Dataset
- 1,000 labeled CAPTCHA images
- Fixed 5-character structure for consistency
- Built-in labeling via filenames
- Suitable for OCR and deep learning tasks
- Lightweight and easy to use for experimentation
Applications and Use Cases
The 5-Character CAPTCHA Dataset can be used in several important applications. For example:
- CAPTCHA Recognition Systems: Train models to decode CAPTCHA images
- Security Testing: Evaluate the strength of CAPTCHA mechanisms
- OCR Development: Improve recognition of distorted text
- Web Automation: Enable automated workflows involving CAPTCHA solving
Conclusion
The 5-Character CAPTCHA Dataset is a valuable resource for developing and testing CAPTCHA recognition systems. Overall, it provides a structured and practical dataset for OCR and security-related applications. More importantly, it helps improve the performance of machine learning models in handling complex visual text challenges.
This dataset is sourced from Kaggle.
Contact Us

Quality Data Creation

Guaranteed TAT

ISO 9001:2015, ISO/IEC 27001:2013 Certified

HIPAA Compliance

GDPR Compliance

Compliance and Security
Let's Discuss your Data collection Requirement With Us
To get a detailed estimation of requirements please reach us.
