Skip to content
context New models

AWS introduced an architecture for automated metadata harmonization using Amazon Bedrock

only one source so far

AWS described a system for automated metadata harmonization built on Amazon Bedrock: it combines LLMs, embeddings and traditional NLP to correct and align schemas, but final approval of changes remains with the user.

AWS published a case study on its Machine Learning Blog describing a system for automated metadata correction and harmonization, built on Amazon Bedrock. The system uses large language models available through Amazon Bedrock for semantic schema alignment and generating recommended corrections, as well as Amazon S3 for storing schemas and results, Amazon DynamoDB for task tracking, Amazon Cognito for authentication and Amazon ECS as the compute layer.

The workflow operates as a cyclical process with a human as the final approver (human-in-the-loop): the user uploads a metadata file, and the system performs two parallel validation steps – schema alignment (checking whether columns and their structure match) and validation of individual fields (checking values against rules). Schema alignment addresses issues such as inconsistent naming, missing or extra columns, or the need to split or merge columns. In addition to fuzzy matching, it uses the semantic understanding of an LLM, which can recognize domain-specific synonyms and infer meaning from the context of surrounding columns.

Field validation divides errors into three categories: missing required fields, values outside a defined controlled vocabulary (e.g. a disallowed instrument type) and failure to match a required format checked using regular expressions (e.g. a date format). According to AWS, when generating corrections, the system first prioritizes less expensive traditional NLP and embedding-based similarity matching, turning to an LLM only for ambiguous or new cases – the aim is to keep inference costs predictable while maintaining accuracy. Vector embeddings (available through Amazon Bedrock) enable semantic mapping of values, for example mapping “Human” to “Homo sapiens” or “NYC” to “New York City”, and their outputs can be cached to reduce costs. The company states that it tested multiple embedding models, including models specialized in biomedicine.

The source does not describe the other models tested or the rest of the architecture. Details can be found in the source article.

What changed

Why it matters

For teams managing large or multi-source datasets (typically in research or science), this provides a concrete reference approach for speeding up metadata standardization without fully losing human control – layering less expensive methods (NLP, embeddings) before using LLMs also helps keep inference costs under control.

Two audiences, two different impacts

What this means

01

For individuals

Data engineers and researchers working with heterogeneous metadata gain a concrete architectural pattern – an LLM accessed through Amazon Bedrock combined with embeddings and a human approval step – that can be applied to their own pipelines without requiring fully autonomous automation.

What to do Study the described approach (layering traditional NLP, embeddings and LLMs according to confidence level) as inspiration for your own data cleaning and harmonization projects.
More practical updates →
02

For a business

According to AWS, organizations managing large volumes of data from different sources (e.g. research institutions) can make metadata standardization faster and less expensive by combining traditional NLP, embeddings and LLMs, without losing control over the final approval of changes.

Productivity
What to decide Consider piloting a similar architecture (LLM on Amazon Bedrock + embeddings + workflow with human approval) to harmonize metadata in your own large or multi-source datasets.
More business impacts →
Amazon Bedrock AWS data standardization LLM metadata harmonization schema alignment

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected AI-powered metadata correction and harmonization