AWS described training a search agent with the Qwen3.6-27B model using MTRL
AWS described a procedure for training a search agent with the Qwen3.6-27B model in the Amazon SageMaker AI MTRL service. The reward evaluates the final ranking of the documents found using nDCG@10; reaching interaction or generation limits results in a penalty.
AWS described a procedure for training a search agent based on the Qwen3.6-27B model using the Amazon SageMaker AI MTRL service. Over several rounds, the agent uses search tools and adjusts further searching based on the results obtained. Training data is generated using multi-turn rollouts, and the model is optimized using policy gradient algorithms. Support for the stated model is documented for the US West (Oregon) region, designated us-west-2.
Training uses the FRAMES, BRIGHT, Enterprise RAG, ESCI, Musique and MLQA datasets. 5% of the examples from each training set is set aside for validation. For testing, the FreshStack, WixQA, BrowseComp-Plus and Wands datasets are listed, covering technical queries, customer support, web search and product search.
The reward is directly the nDCG@10 metric, which evaluates the relevance and ranking of the first ten documents found. It is evaluated only after the entire search run has been completed. If the maximum number of interactions or the generation length limit in a single round is reached, the agent receives a reward of −1. See the source article for details.
Why it matters
The procedure gives developers a concrete way to train a search agent's decision-making based on the outcome of an entire sequence of searches. For enterprise search, it also offers a way to separate training, validation and testing, which can be used to verify the quality of retrieved documents before deployment.
Two audiences, two different impacts
What this means
For individuals
A developer of a search agent can use nDCG@10 as the reward for the entire completed run and penalize reaching interaction or generation limits.
For a business
Teams developing enterprise search gain a concrete procedure for quality control: separate test sets and 5% of training examples set aside for validation.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.