AWS described the KnowledgeForge architecture for automated knowledge base generation and curation from ITSM tickets
AWS described KnowledgeForge, a system that generates new knowledge base articles from resolved ITSM tickets using Claude Sonnet 4.5 in Amazon Bedrock and uses Amazon S3 Vectors to remove duplicates and assess content quality. A human approves the outputs.
AWS described KnowledgeForge on its technical blog, a system for automatically expanding and maintaining a knowledge base from IT service management (ITSM) tickets. The goal is to use knowledge that remains locked in the history of resolved tickets - the symptoms of a problem, its cause and the solution used - and that otherwise never makes it into the knowledge base. At the same time, the system is intended to address the opposite problem in the existing knowledge base: the accumulation of duplicate articles, outdated content and inconsistent quality depending on who wrote an article and when.
The system consists of two interconnected parts. Generation clusters related resolved tickets by topic and uses Claude Sonnet 4.5 in Amazon Bedrock to create a draft of a new knowledge base article and a document describing the root cause (root cause analysis). Before writing, the system uses RAG (retrieval augmented generation) to retrieve the five most similar existing articles for the given customer from the Amazon S3 Vectors index and uses them as reference context, which, according to AWS, is intended to keep terminology consistent and limit the invention of nonexistent procedures. Generation runs on Amazon ECS with AWS Fargate because processing a single topic can take several minutes and the workload arrives in bursts.
Curation processes every article - both new and existing - in four steps: type classification, duplicate detection, quality assessment and rewriting weak content. Duplicate detection uses 1024-dimensional embeddings from Amazon Titan Text Embeddings V2 stored in a separate Amazon S3 Vectors index for each customer; articles with a cosine distance below the threshold of 0.05 (meaning a similarity of 0.95 or higher) among the five nearest neighbors are flagged as duplicates. According to AWS, AWS Step Functions orchestrates the entire process. Completed articles go to ServiceNow, where a knowledge manager must approve them before publication, so the output always undergoes human review.
You can find details in the source article.
Why it matters
For developers and architects, this is a ready-made, documented pattern combining Amazon Bedrock, Amazon S3 Vectors and AWS Step Functions for RAG and duplicate detection in a large-scale document processing pipeline - it can also be applied outside ITSM. For companies with customer support or internal IT operations, it offers a way to automatically turn resolved tickets into useful documentation while removing duplicates and outdated content from the existing knowledge base, without dropping human review before publication.
Two audiences, two different impacts
What this means
For individuals
Developers and architects building document processing pipelines with generative AI get a concrete, reusable pattern from AWS: duplicate detection using embeddings in Amazon S3 Vectors, content generation with RAG through Amazon Bedrock, and orchestration through AWS Step Functions.
For a business
Companies with a large volume of ITSM tickets can use the described architecture to automate the expansion and cleanup of their support knowledge base (duplicates, outdated content, inconsistent quality), with a knowledge manager always approving the output before publication in ServiceNow.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.