AWS described a synthetic data augmentation pipeline for detecting people near heavy machinery
AWS described a pipeline on Amazon SageMaker AI that uses the diffusion model Qwen-Image-Edit-2509 to insert synthetic figures into real images of machinery and automatically labels them using Amazon Rekognition. According to AWS, this improved the detection of people near machinery (mAP50) by up to 160 % without manual annotation.
AWS described a synthetic data augmentation pipeline on its technical blog for training models to detect people dangerously close to heavy machinery. This addresses the problem that real images of the most dangerous situations (a person in a machine's blind spot, a child near moving equipment) are rare in datasets, and deliberately photographing them would be risky.
Instead of generating entirely synthetic scenes, the pipeline edits real images containing machinery but no people, using the diffusion model Qwen-Image-Edit-2509 to insert synthetic figures. According to AWS, this approach preserves the fidelity of the background, lighting and scale, avoiding the so-called domain gap, meaning a decline in model performance when switching to entirely synthetic data. The model runs on an Amazon SageMaker AI instance of type ml.g5.12xlarge with four NVIDIA A10G GPUs and a total of 96 GB VRAM. The generated images are then automatically labeled using Amazon Rekognition DetectLabels API (minimum confidence 80 %), with duplicate bounding boxes removed using NMS, and converted to YOLO format, so the entire process involves no manual annotation.
According to experiments by AWS on models in the YOLO11 family, augmentation delivered up to a 160% improvement in the mAP50 metric (detection accuracy at an overlap threshold of 0.5) compared with training only on real data. The key finding was that the wording of the prompt specifying the placement of the inserted person has the greatest influence on the result: placement in relevant high-risk positions doubled mAP50, while placement in the background actually worsened performance. For the tests, AWS used a publicly available subset of the OpenImages dataset as a substitute for customer data from industrial environments.
AWS also estimates that inference costs per image could decrease by approximately a factor of ten when deployed on a single NVIDIA H100 GPU (instance ml.p5.4xlarge) instead of the current configuration with four A10G GPUs, but according to AWS, this estimate has not yet been validated. The rest of the article, containing further details of the experiments, was not available.
Why it matters
The approach shows a way to train safety detection systems for rare and dangerous scenarios without risky collection of real data and without costly manual annotation — particularly relevant to companies deploying autonomous machinery in agriculture, construction, mining or manufacturing. For ML practitioners, the key insight is that the result depends primarily on correctly wording the prompt that specifies the person's placement, rather than just the amount of generated data.
Two audiences, two different impacts
What this means
For individuals
For data scientists and ML engineers dealing with insufficient training data for rare, high-risk scenarios, the approach from AWS shows that the key success factor is the placement of the inserted person as specified by the prompt, rather than just the volume of generated data.
For a business
According to AWS, the pipeline enables industrial companies using autonomous machinery (agriculture, construction, mining, manufacturing) to fill gaps in training data for detecting dangerous situations without risky photography in the field and without manual annotation, reducing both costs and safety risks during data collection.
Risks and complianceCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.