Skip to content
worth noting Coding

AWS described the KnowledgeForge architecture for automated knowledge base generation and curation from ITSM tickets

only one source so far

AWS described KnowledgeForge, a system that generates new knowledge base articles from resolved ITSM tickets using Claude Sonnet 4.5 in Amazon Bedrock and uses Amazon S3 Vectors to remove duplicates and assess content quality. A human approves the outputs.

AWS described KnowledgeForge on its technical blog, a system for automatically expanding and maintaining a knowledge base from IT service management (ITSM) tickets. The goal is to use knowledge that remains locked in the history of resolved tickets - the symptoms of a problem, its cause and the solution used - and that otherwise never makes it into the knowledge base. At the same time, the system is intended to address the opposite problem in the existing knowledge base: the accumulation of duplicate articles, outdated content and inconsistent quality depending on who wrote an article and when.

The system consists of two interconnected parts. Generation clusters related resolved tickets by topic and uses Claude Sonnet 4.5 in Amazon Bedrock to create a draft of a new knowledge base article and a document describing the root cause (root cause analysis). Before writing, the system uses RAG (retrieval augmented generation) to retrieve the five most similar existing articles for the given customer from the Amazon S3 Vectors index and uses them as reference context, which, according to AWS, is intended to keep terminology consistent and limit the invention of nonexistent procedures. Generation runs on Amazon ECS with AWS Fargate because processing a single topic can take several minutes and the workload arrives in bursts.

Curation processes every article - both new and existing - in four steps: type classification, duplicate detection, quality assessment and rewriting weak content. Duplicate detection uses 1024-dimensional embeddings from Amazon Titan Text Embeddings V2 stored in a separate Amazon S3 Vectors index for each customer; articles with a cosine distance below the threshold of 0.05 (meaning a similarity of 0.95 or higher) among the five nearest neighbors are flagged as duplicates. According to AWS, AWS Step Functions orchestrates the entire process. Completed articles go to ServiceNow, where a knowledge manager must approve them before publication, so the output always undergoes human review.

You can find details in the source article.

What changed

Why it matters

For developers and architects, this is a ready-made, documented pattern combining Amazon Bedrock, Amazon S3 Vectors and AWS Step Functions for RAG and duplicate detection in a large-scale document processing pipeline - it can also be applied outside ITSM. For companies with customer support or internal IT operations, it offers a way to automatically turn resolved tickets into useful documentation while removing duplicates and outdated content from the existing knowledge base, without dropping human review before publication.

Two audiences, two different impacts

What this means

01

For individuals

Developers and architects building document processing pipelines with generative AI get a concrete, reusable pattern from AWS: duplicate detection using embeddings in Amazon S3 Vectors, content generation with RAG through Amazon Bedrock, and orchestration through AWS Step Functions.

What to do Study the described pattern (embeddings for duplicate detection in Amazon S3 Vectors, generation and RAG over existing content through Amazon Bedrock, orchestration through AWS Step Functions) as a starting point for your own pipeline…
More practical updates →
02

For a business

Companies with a large volume of ITSM tickets can use the described architecture to automate the expansion and cleanup of their support knowledge base (duplicates, outdated content, inconsistent quality), with a knowledge manager always approving the output before publication in ServiceNow.

Development
What to decide Evaluate whether a similar system for automated knowledge base curation (duplicates, quality, generation from tickets) would reduce the burden on IT support and documentation obsolescence.
More business impacts →
Amazon Bedrock AWS AWS Step Functions ITSM KnowledgeForge S3 Vectors

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected KnowledgeForge: mining gold from the ITSM ticket graveyard