Pedestrian Attribute Recognition Dataset (Pedestrian-1K)

Pedestrian-1K is a specialized synthetic dataset for Pedestrian Attribute Recognition (PAR) and Person Re-Identification (Re-ID) from overhead CCTV perspectives. While existing datasets like PA-100K rely on aging, low-resolution real-world footage with inconsistent labels, Pedestrian-1K provides high-fidelity synthetic subjects from a consistent high-angle security camera geometry. Each image is uniquely annotated with both a clean Natural Language Description (e.g., 'A young adult man with black hair, seen from the side, wearing a navy blue hoodie and black jeans') and a Structured Attribute Map covering gender, age, hair color, clothing type/color, camera angle, and accessories. The dataset features realistic digital noise, scanlines, and compression artifacts to ensure models generalize to real CCTV hardware. Attribute distributions use weighted demographic logic (e.g., elderly subjects lean 80% towards grey hair) for realistic sampling. The annotation format is JSONL with both descriptive strings and categorized attribute labels per image. The open-source sample includes 1,000 images; full packages are available with 10,000 and 50,000 images featuring descriptive annotations and Vision Transformer annotations. Licensed under CC BY 4.0.
50,000 images
Full Package
1000
Open Source Samples
JSONL (Descriptive + Structured)
Annotation Format
100%
Privacy Compliant
Dataset Features
Intended Use Cases
Free Sample vs. Commercial Package
Free Open-Source Sample
- 1000 annotated images
- Format: JSONL (Descriptive + Structured)
- Hosted on Kaggle
- Licence: CC BY 4.0
Commercial Package
- 50,000 images
- Format: JSONL (Descriptive + Structured)
- Licence: Student & Research or Business licence
The Creative Commons licence above applies only to the free sample, not to the full commercial package.
Limitations & Recommended Validation
This dataset is 100% synthetic. While it is designed to closely match real-world sensor and camera conditions, synthetic imagery can still differ from live footage in ways that affect model accuracy (a "domain gap"). Validate a trained model against real-world footage from your specific deployment environment before production use.
Not intended as a sole basis for biometric identification, legal evidence, or safety-critical decisions without independent human review and real-world testing.