AIToday
Large Language ModelsAI Coding AssistantsAmazon AI BlogPublished: Oct 3, 2026, 01:00 JST

AWS fine-tunes Qwen3.6-27B search agent on Amazon SageMaker AI

AWS fine-tunes Qwen3.6-27B search agent on Amazon SageMaker AI

3 Key Points

  1. What happened

    AWS fine-tuned a Qwen3.6-27B search agent using multi-turn reinforcement learning on Amazon SageMaker AI, with nDCG@10 gains of +23.7 percent on BrowseComp-Plus and +18.4 percent on WixQA.

  2. Why it matters

    The results suggest a smaller, specialized search agent can approach frontier-model reliability at lower latency and cost, though the body notes a slight regression on the FreshStack benchmark.

  3. What to watch

    With only three hyperparameters changed and everything else on default, the test is whether this low-configuration recipe holds on a company's own tools and data.

WHO IT HITSEnterprise teams building internal search or RAG systems can train a smaller model on their own tools and environment, potentially reducing reliance on larger frontier models.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

AWS's post walks through a practical fine-tuning pipeline for a search agent, starting from the observation that base models do not arrive knowing a company's tools or environment. The write-up frames multi-turn reinforcement learning as a middle path between prompting small models, which rarely produces dependable multi-turn behavior, and prompting frontier models, which works more often but at higher latency and cost. It also positions MTRL against supervised fine-tuning, which depends on costly expert demonstrations, and single-turn RL, which scores one response at a time and misses the dependencies between turns.

The setup used a Qwen3.6-27B model and a training configuration that changed only three hyperparameters, leaving the algorithm, advantage estimator, and off-policy staleness bounds on defaults. Training data came from several public datasets, with five percent of training instances reserved for validation, and the reward was the retrieval metric nDCG@10. AWS says the fine-tuned model improved on WixQA, Wands, and BrowseComp-Plus, with a slight decrease on FreshStack, and that the number of turns per question fell on BrowseComp-Plus and FreshStack.

The stakes for readers evaluating internal search or RAG systems appear to hinge on whether this low-configuration recipe transfers to their own tools, datasets, and reward metrics, since the results come from AWS's selected benchmarks rather than a customer deployment. The body notes that MTRL training can span multiple days and that the default time limit is 24 hours, which may matter for teams planning longer runs.

FAQ
What model did AWS fine-tune?
AWS fine-tuned a Qwen3.6-27B model, supported in the US West (Oregon) Region (us-west-2), for a search agent.
How much configuration did the training need?
Only three hyperparameters were changed: max_epochs: 1, global_batch_size: 128, and rollout_max_concurrency: 32. Everything else ran on defaults.
What metric did AWS use for the reward?
The main metric was nDCG@10 (Normalized Discounted Cumulative Gain at rank 10), used directly as the reward function in MTRL.
Amazon AI BlogRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleLanguage discrimination lifts multilingual speech AI to 10.4% error