Skip to content
worth noting Coding

AWS described automated agent evaluation on AWS Bedrock AgentCore in CI/CD using GitHub Actions

only one source so far

AWS described a GitHub Actions pipeline that deploys an agent to AWS Bedrock AgentCore, evaluates it using AgentCore Evaluate API, and blocks a pull request when quality deteriorates; it also addresses OAuth authentication for CI against an MCP server with roles.

Amazon Web Services published a tutorial on AWS Machine Learning Blog describing a pipeline in GitHub Actions that automatically deploys an agent to the AWS Bedrock AgentCore runtime, evaluates its performance using AgentCore Evaluate API, and blocks a pull request if agent quality drops below a set threshold. Evaluation uses LLM-as-a-judge or, alternatively, evaluation through AWS Lambda, and operates on OpenTelemetry trace data that the agent routinely sends through AgentCore Observability.

The described architecture includes an agent built on the Strands framework that communicates with an MCP server with role-based access to tools (role-based access control). Authentication is handled by a shared Cognito user pool with two types of flows – machine access using client_credentials for the CI pipeline and user access using authorization_code with roles in the token for interactive use. The infrastructure is defined using CDK, and according to the company, the complete reference implementation is available as open source in the accompanying repository.

The main problem addressed by the tutorial is that a headless CI pipeline lacks the user context needed for interactive OAuth consent, while the MCP server expects a token with role claims. The article describes three possible approaches: evaluating stored trace data from fixtures instead of making live calls to the agent (using the included evaluate_stored_traces.py script), using a dedicated test user with a refresh token stored in AWS Secrets Manager (with the need to handle its renewal), and configuring the MCP server to support both machine and user login types simultaneously.

The description of the third approach and further technical details of the evaluation categories are incomplete in the available text. You can find details in the source article.

What changed

Why it matters

Without automated evaluation, agent quality is assessed subjectively, and regressions following a prompt or model change are often discovered only through user complaints. The described process makes it possible to catch such deterioration during code review (pull request) and also addresses a specific technical obstacle – how to authenticate a headless CI pipeline against an OAuth-protected MCP server with role-restricted access to tools, a problem relevant to anyone running agents with MCP tools in a production environment.

Two audiences, two different impacts

What this means

01

For individuals

Developers and DevOps engineers working with AWS Bedrock AgentCore and MCP servers gain concrete guidance on how to handle authentication in GitHub Actions for a headless CI pipeline against an OAuth-protected server with roles, and how to connect evaluation to an existing OpenTelemetry trace.

What to do Study the reference implementation in the accompanying repository and assess which of the three described approaches to authenticating CI against an OAuth-protected MCP server fits your pipeline.
More practical updates →
02

For a business

Companies running AI agents in production can add an automated quality check to their development process that blocks merging a change (prompt, model, tool configuration) if it worsens agent behavior – reducing the risk of regressions reaching users.

Development
What to decide Consider introducing automated agent evaluation (LLM-as-a-judge or Lambda) as a mandatory quality check in CI/CD before deployment to production.
More business impacts →
AgentCore AWS Bedrock CI/CD evaluation GitHub Actions MCP

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected Automated agent evaluation with Amazon Bedrock AgentCore and GitHub Actions