TikHarm Dataset

TikHarm Dataset

Datasets

TikHarm Dataset

File

TikHarm Dataset

Use Case

TikHarm Dataset

Description

Explore the TikHarm dataset, designed to train models for classifying harmful content on TikTok. This curated collection focuses on videos accessible to children, categorized into harmful content, adult content, safe content, and suicide-related material

TikHarm Dataset

Description:

The TikHarm dataset is a curated collection of TikTok videos aimed at training models to identify harmful content, specifically focusing on material accessible to children. The dataset is formatted similarly to UCF101 and includes videos categorized into four distinct types: harmful content, adult content, safe content, and suicide-related content.

Download Dataset

Videos were gathered from TikTok with a focus on content that children might encounter. Each video was manually labeled to reflect its category, ensuring that harmful behaviors, inappropriate material, and safe content are clearly identified. This enables the development of models that can accurately classify and filter potentially harmful material for child-friendly environments.

Dataset Structure and Statistics

The dataset is divided into three main subsets:

  • Training Set: Used for model training and learning patterns
  • Validation (Dev) Set: Used for tuning and evaluation during training
  • Test Set: Used for final performance evaluation

Moreover, each subset contains videos of varying durations, ranging from a few seconds to several minutes. Consequently, this ensures diversity and realism in the dataset.

Key Features of the TikHarm Dataset

  • Manually labeled video dataset for high accuracy
  • Four distinct content categories for classification
  • Real-world social media data from TikTok
  • Suitable for video classification and deep learning models
  • Designed with child safety and content moderation in mind

Applications and Use Cases

The TikHarm Dataset can be used in a wide range of AI and safety-focused applications. For example:

  • Content Moderation Systems: Automatically filter harmful or inappropriate videos
  • Child Safety Platforms: Identify unsafe content for younger audiences
  • Social Media Monitoring: Improve detection of sensitive or dangerous content
  • AI Research: Develop advanced video classification and behavior detection models

Conclusion

 

The TikHarm Dataset is a valuable resource for building AI systems that detect and classify harmful content in videos. With its structured labeling and real-world data, it supports the development of safer digital platforms, particularly for protecting children from inappropriate or dangerous content.

This dataset is sourced from Kaggle.

Contact Us

Technology

Quality Data Creation

Technology

Guaranteed TAT

Technology

ISO 9001:2015, ISO/IEC 27001:2013 Certified

Technology

HIPAA Compliance

Technology

GDPR Compliance

Technology

Compliance and Security

Let's Discuss your Data collection Requirement With Us

To get a detailed estimation of requirements please reach us.

Scroll to Top

Please provide your details to download the Dataset.