Amazon Bedrock AgentCore adds evaluators for agents with skills
Amazon Web Services added three evaluators (Skill Selection Accuracy, Skill Instruction Following, SkillInvoked) to Amazon Bedrock AgentCore Evaluations and the Strands Evals SDK that measure whether an agent selected the right skill and followed its instructions.
Amazon Web Services introduced three new evaluators in Amazon Bedrock AgentCore Evaluations and the Strands Evals SDK, focused on agents equipped with so-called skills – modular sets of instructions (typically in a SKILL.md file) that teach an agent a specific task, such as processing an invoice or following company rules for pull requests. According to AWS, skills are part of the open Agent Skills standard and are therefore portable between compatible harnesses; at runtime, the agent loads only the skill it currently needs.
The new evaluators address two types of errors that, according to the article, standard output quality metrics overlook: the agent selects an unsuitable skill for the task, or selects the right skill but follows its instructions only partially. Skill Selection Accuracy is a binary evaluator that assesses whether each invoked skill was suitable for the task. Skill Instruction Following rates, on five levels, how consistently the agent followed the prescribed steps of the skill. The third, deterministic evaluator, SkillInvoked (available only in Strands Evals), checks whether a named skill was actually loaded.
As an example, the article describes an HR assistant with skills for vacation planning and employee benefits: if the agent uses the benefits skill to answer a vacation question, Skill Selection Accuracy detects this incorrect choice; if it correctly invokes the vacation skill but skips the step of checking the rules for carrying over unused vacation, Skill Instruction Following catches the error. According to AWS, the two types of errors require different fixes – an unsuitable skill choice often points to overlapping or unclear skill descriptions, while incomplete adherence to instructions is more likely to indicate a need for clearer steps, a different skill structure, or a more capable agent model.
The evaluators work with a record of the agent run – either a session from Strands Evals or an OpenTelemetry trace from the user's observability layer. Skill recognition during extraction works with signals from, among others, the Strands AgentSkills plugin, Claude Code, Claude Agent SDK, OpenAI Agents SDK, Codex, Gemini CLI, OpenHands, Google ADK, and generic reading of SKILL.md files. To run the evaluations, AWS uses AgentCore CLI in the article and requires, among other things, InvokeModel permission for the evaluation model. Details can be found in the source article.
Why it matters
The evaluators make it possible to distinguish between two different causes of agent failure – choosing the wrong skill versus incompletely following its steps – which, according to AWS, require different fixes (clearer skill descriptions versus clearer instructions or a more capable model). This gives teams building agents with procedural skills a tool for targeted diagnosis instead of assessing only the final answer, which can sound credible even when the prescribed procedure has been violated.
Two audiences, two different impacts
What this means
For individuals
Developers building agents with modular skills can now automatically test whether an agent selects the right skill and fully follows its steps, including in CI, instead of manually assessing the final answers.
For a business
Companies deploying agents with skills for processes such as compliance checks or HR workflows gain a tool that verifies that the agent selected the right procedure and followed its steps, instead of relying solely on the final answer sounding credible.
DevelopmentCheck the original
Event sources
clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.