Skip to content
worth noting AI agents

Amazon Bedrock AgentCore adds evaluators for agents with skills

clearly official source

Amazon Web Services added three evaluators (Skill Selection Accuracy, Skill Instruction Following, SkillInvoked) to Amazon Bedrock AgentCore Evaluations and the Strands Evals SDK that measure whether an agent selected the right skill and followed its instructions.

Amazon Web Services introduced three new evaluators in Amazon Bedrock AgentCore Evaluations and the Strands Evals SDK, focused on agents equipped with so-called skills – modular sets of instructions (typically in a SKILL.md file) that teach an agent a specific task, such as processing an invoice or following company rules for pull requests. According to AWS, skills are part of the open Agent Skills standard and are therefore portable between compatible harnesses; at runtime, the agent loads only the skill it currently needs.

The new evaluators address two types of errors that, according to the article, standard output quality metrics overlook: the agent selects an unsuitable skill for the task, or selects the right skill but follows its instructions only partially. Skill Selection Accuracy is a binary evaluator that assesses whether each invoked skill was suitable for the task. Skill Instruction Following rates, on five levels, how consistently the agent followed the prescribed steps of the skill. The third, deterministic evaluator, SkillInvoked (available only in Strands Evals), checks whether a named skill was actually loaded.

As an example, the article describes an HR assistant with skills for vacation planning and employee benefits: if the agent uses the benefits skill to answer a vacation question, Skill Selection Accuracy detects this incorrect choice; if it correctly invokes the vacation skill but skips the step of checking the rules for carrying over unused vacation, Skill Instruction Following catches the error. According to AWS, the two types of errors require different fixes – an unsuitable skill choice often points to overlapping or unclear skill descriptions, while incomplete adherence to instructions is more likely to indicate a need for clearer steps, a different skill structure, or a more capable agent model.

The evaluators work with a record of the agent run – either a session from Strands Evals or an OpenTelemetry trace from the user's observability layer. Skill recognition during extraction works with signals from, among others, the Strands AgentSkills plugin, Claude Code, Claude Agent SDK, OpenAI Agents SDK, Codex, Gemini CLI, OpenHands, Google ADK, and generic reading of SKILL.md files. To run the evaluations, AWS uses AgentCore CLI in the article and requires, among other things, InvokeModel permission for the evaluation model. Details can be found in the source article.

What changed

Why it matters

The evaluators make it possible to distinguish between two different causes of agent failure – choosing the wrong skill versus incompletely following its steps – which, according to AWS, require different fixes (clearer skill descriptions versus clearer instructions or a more capable model). This gives teams building agents with procedural skills a tool for targeted diagnosis instead of assessing only the final answer, which can sound credible even when the prescribed procedure has been violated.

Two audiences, two different impacts

What this means

01

For individuals

Developers building agents with modular skills can now automatically test whether an agent selects the right skill and fully follows its steps, including in CI, instead of manually assessing the final answers.

What to do Try the Skill Selection Accuracy and Skill Instruction Following evaluators on your own skill-equipped agent through Strands Evals SDK or AgentCore CLI and add them to your test suite.
More practical updates →
02

For a business

Companies deploying agents with skills for processes such as compliance checks or HR workflows gain a tool that verifies that the agent selected the right procedure and followed its steps, instead of relying solely on the final answer sounding credible.

Development
What to decide Consider introducing skill evaluation (Skill Selection Accuracy, Skill Instruction Following) into the testing process for agents built on AWS Bedrock AgentCore, especially where the agent must follow company procedures such as…
More business impacts →
Amazon Bedrock Bedrock AgentCore evaluace agentů OpenTelemetry skills Strands Evals

Check the original

Event sources

clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore