Skip to content
context Tools and apps

Amazon described a combination of Amazon Textract and Amazon Bedrock for RAG over complex documents

only one source so far

Amazon described a procedure combining Amazon Textract and Amazon Bedrock for RAG processing of complex multi-page documents (PDF, DOCX, TXT, HTML, PNG, XLSX), intended to reduce errors and hallucinations compared with uploading raw files. The code is on GitHub.

Amazon published a guide on its AWS Machine Learning blog explaining how to combine Amazon Textract with Amazon Bedrock for the automated processing and analysis of complex, multi-page documents. The example centers on customer support processing thousands of utility bills (energy invoices) each month in varied formats and layouts, where manual information searches led to delays, errors, and customer dissatisfaction.

According to Amazon, an initial attempt to deploy a RAG (Retrieval Augmented Generation) solution directly over raw documents proved unreliable — according to the article, the language model omitted key data and, in some cases, invented information (hallucinated). The solution, as described, is to preprocess documents using Amazon Textract, which extracts structured and unstructured content from them before they are passed to an Amazon Bedrock knowledge base. Supported file formats are PDF, DOCX, TXT, HTML, XLSX, and PNG.

The deployment described in the guide uses a shell script that creates a CloudFormation stack in AWS with Lambda functions, an S3 bucket, an OpenSearch Serverless cluster, and an Amazon Bedrock knowledge base. Uploaded documents are automatically processed and converted into a format suitable for the knowledge base. The solution code is available on GitHub. For production deployment, Amazon recommends enabling Amazon Bedrock Guardrails, which filter harmful content, block prohibited topics, redact sensitive information, and reduce hallucinations through source verification (grounding validation).

The rest of the article, including the detailed setup procedure and specific configuration steps, is not available — you can find the details in the source article.

What changed

Why it matters

For teams building RAG applications over unstructured documents (invoices, contracts, statements), it offers a documented example procedure for using OCR/extraction preprocessing to reduce a typical problem with direct RAG over raw files — missing data and model hallucinations. The described CloudFormation deployment and the code available on GitHub allow the procedure to be replicated without having to build an extraction pipeline from scratch.

Two audiences, two different impacts

What this means

01

For individuals

Developers and teams working with RAG solutions over documents gain a documented procedure for preprocessing complex PDF and image files using Amazon Textract to reduce the hallucinations and missing data that occurred when raw documents were uploaded directly to a language model.

What to do Review the published code and procedure on GitHub if the document processing task is relevant.
More practical updates →
02

For a business

Companies processing large volumes of unstructured documents (invoices, statements) can, according to the guide from Amazon, replace manual information searches with automated querying, which may shorten customer support response times and reduce errors caused by manual processing.

Processes
What to decide Evaluate whether combining Amazon Textract and Amazon Bedrock (using the described procedure with CloudFormation) could simplify the processing of similar multi-page documents within the company.
More business impacts →
Amazon Bedrock Amazon Textract Automation Document Processing LLM RAG

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected Customizing your knowledge base on Amazon Bedrock for large and complex documents using Amazon Textract