CCTV Shoplifting Detection Dataset (YOLO + VLM)
Sample video clip from the dataset — synthetic retail CCTV with person + pose detection.

Retail shrink is a $100B+ problem and most off-the-shelf vision models can't tell a shopper from a shoplifter — because both look identical until the moment of concealment. This dataset is purpose-built for that gap. Scenes are rendered from realistic high-angle in-store CCTV perspectives across electronics stores, grocery aisles, apparel sections, and convenience layouts, with person bounding boxes plus 17-keypoint pose annotations capturing the subtle body language of pocketing, bag-stashing, and shelf-sweep behaviors. The package ships in two formats: YOLO labels for spatial detection and VLM-style natural language captions for fine-tuning vision-language models on behavioral context. Video sequences are included so temporal models (3D CNNs, video transformers) can learn the time signature of a concealment event — the pause, the glance, the hand returning empty. 100% synthetic — no real customers, no biometrics, no GDPR exposure. The open-source sample on Kaggle includes 400 frames and 8 video clips; the full package contains 5,000+ frames and 100+ video sequences covering diverse store types, lighting conditions, and shoplifting techniques.
5,000+ frames + 100+ videos
Full Package
400
Open Source Samples
YOLO + VLM
Annotation Format
100%
Privacy Compliant
Dataset Features
Intended Use Cases
Free Sample vs. Commercial Package
Free Open-Source Sample
- 400 annotated images
- Format: YOLO + VLM
- Hosted on Kaggle
- Licence: See the hosting platform's terms of use
Commercial Package
- 5,000+ frames + 100+ videos
- Format: YOLO + VLM
- 8 evaluation videos included
- Licence: Student & Research or Business licence
The Creative Commons licence above applies only to the free sample, not to the full commercial package.
Limitations & Recommended Validation
This dataset is 100% synthetic. While it is designed to closely match real-world sensor and camera conditions, synthetic imagery can still differ from live footage in ways that affect model accuracy (a "domain gap"). Validate a trained model against real-world footage from your specific deployment environment before production use.
Not intended as a sole basis for biometric identification, legal evidence, or safety-critical decisions without independent human review and real-world testing.