Skip to content
worth noting Security

AWS released an open-source personal data detector that works with any LLM in Amazon Bedrock

only one source so far

AWS released an open-source PII detector (the pii-detector package) that uses prompt-based instructions to work with any LLM in Amazon Bedrock and was tested against nine other detectors, including OpenAI PrivacyFilter.

Amazon Web Services (AWS) published an open-source detector for personally identifiable information (PII) that, according to the company, works with any large language model (LLM) available through Amazon Bedrock. The tool is released as the pii-detector package in the sample-llm-pii-detection repository. Unlike traditional approaches based on token-classification models, where the set of recognized data types is fixed during training, this detector defines entities and the output format using instructions (a prompt) supplied to the model at runtime. According to the company, adding a new data type, such as an internal employee ID or a crypto wallet address, therefore requires only a prompt adjustment, rather than retraining.

According to AWS, the detector was evaluated on five public datasets containing PII and compared with nine other LLM-based detectors, including the OpenAI PrivacyFilter tool. The detector uses a schema of fifteen entity categories defined in the system prompt, the model returns a structured list of detected data in JSON format, and subsequent post-processing calculates the exact text positions (character offsets), because, according to the article, the LLM cannot reliably return them directly. According to the source, the tool supports detection in eight languages and allows custom entity types to be defined.

Architecturally, the detector consists of four parts: a prompt template defining the entity schema, an interchangeable backend for running the model (with a supplied implementation for Amazon Bedrock through the Converse API), a layer for processing the response and calculating the positions of detected data, and logic connecting these components. According to AWS, separating model selection from entity definitions makes it possible to choose either a model on Amazon Bedrock or a smaller open model running locally on a single GPU.

The source article describing how to install and run the tool on your own data was only partially available. Details can be found in the source article.

What changed

Why it matters

According to the company, teams fine-tuning models on corporate text data (support, HR, chat logs) have access to a tool that can detect personal data without having to train their own classifier and without being locked into a single model or deployment — new types of sensitive data can be added by adjusting the prompt instead of retraining. This reduces the risk of real names, addresses or identity document numbers entering the training data and later being unintentionally reproduced by the model.

Relevant practical impact

What this means

01

For a business

Companies processing customer data (support, HR records, chat logs) for fine-tuning or analysis gain an open-source tool for PII detection that can be deployed in their own VPC without having to train a separate model, which may reduce the risk of personal data leaking through models.

Risks and compliance
What to decide Evaluate the use of the pii-detector package as an additional check when preparing training or corporate datasets containing personal data.
More business impacts →
AWS Bedrock data security LLM PII detection privacy

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected Model-agnostic PII detection with LLMs