T-Search released with open weights for multi-step retrieval of evidence passages
The authors released T-Search, a testing environment and a live demo. The model searches for evidence passages in a given corpus over several rounds. According to the authors, it achieves an average Recall@10 of 56.0 in a single run and 61.3 when combining three runs.
The authors introduced T-Search, a model with open weights that performs a limited number of search rounds within a fixed corpus for a given question. It returns ranked evidence passages with brief explanations and leaves answer generation to a downstream model. According to the authors, this separation allows the search backend and the answer-generation model to be changed without retraining. The release includes a testing environment, a live demo and three sets of test tasks.
T-Search is based on Qwen3.6-35B-A3B and was trained on synthetic search tasks using supervised fine-tuning followed by GSPO, with a reward for retrieving evidence. The authors report an average Recall@10 of 56.0 across seven English and Russian datasets with reference evidence passages, which is 14.4 points higher than the base model. When combining the results of three runs, they report a value of 61.3. The released datasets include TRuST, which the authors describe as the first dataset of challenging search tasks created directly in Russian.
Why it matters
For questions that require several search steps, it is possible to assess separately whether the system retrieved the supporting material needed for an answer. The published testing environment and task datasets allow comparisons of this capability; the reported results are measurements by the authors.
Relevant practical impact
What this means
For individuals
A developer or researcher can try multi-step search in the live demo and inspect the retrieved passages and the rationale for their selection.
Check the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.