TReNDS Center automated error cause analysis using Amazon Bedrock and AWS Lambda
The research center TReNDS (Georgia State University, Georgia Institute of Technology, Emory University) deployed a system in production that uses CloudWatch, Lambda, Strands Agents SDK and Amazon Bedrock to automatically analyze the causes of errors in applications and send the results to the team through SNS.
The research center TReNDS (a joint center of Georgia State University, Georgia Institute of Technology and Emory University focused on neuroimaging and data science) described in a post on AWS Machine Learning Blog an architecture it deployed in production to automatically analyze the causes of errors in its applications. The center has been running its infrastructure on AWS since 2019, its applications run on Amazon EKS, and logs are sent through FluentBit to Amazon CloudWatch.
The system works by using a subscription filter in CloudWatch to detect error patterns (ERROR, Exception, FATAL, CRITICAL) in real time and trigger an AWS Lambda function when a match occurs. The function uses Strands Agents SDK to run an agent built on an Amazon Bedrock foundation model that investigates the error — it retrieves the surrounding log context and the relevant source code from a GitHub repository through a custom tool defined in Python, then traces the execution path through the code that led to the error. According to the company TReNDS, the model not only summarizes the error message but actively investigates the cause and decides on its own when to use a tool and which tool to use. The resulting structured analysis is sent to the team through Amazon SNS.
According to TReNDS, manually analyzing a simple error took 15 to 30 minutes, and significantly longer for complex problems spanning multiple services. The center states that, given its work with medical research data (potentially subject to HIPAA), it was important that Amazon Bedrock processes requests within its own AWS account and that neither log data nor source code leaves the managed environment. According to the company, the same pattern also works for other log sources sent to CloudWatch, such as ECS, Lambda, EC2 or on-premises workloads with CloudWatch Agent.
The source article is a technical guide with architecture and code examples; the complete text, including details of the four-stage pipeline and the full implementation, is not available in the accessible portion of the source. Details can be found in the source article.
Why it matters
It shows a concrete, documented implementation of incident response automation with a foundation model connected to production logs and source code, including how it addresses where data remains (within its own AWS account) when handling sensitive data subject to regulations such as HIPAA. For teams running applications on AWS, it provides a ready-made pattern that can be adopted without having to send logs or code outside their own infrastructure.
Two audiences, two different impacts
What this means
For individuals
A developer or DevOps engineer working with AWS gets a concrete, replicable procedure for using Bedrock and Strands Agents SDK to connect logs from CloudWatch with source code and obtain an automatically proposed cause of an error instead of performing a manual analysis lasting 15-30 minutes.
For a business
Companies running infrastructure on AWS can use this pattern to reduce the time spent manually diagnosing errors, lowering operational support costs (SRE/DevOps) and speeding up incident response.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.