· via dev.to (home feed)
AWS publishes guidance for custom reward rules on Nova Forge multi-turn agents
AWS has published technical guidance for writing custom multi-turn reward functions in Amazon Nova Forge, letting developers encode their own definition of agent success into fine-tuning.

AWS has published new technical guidance on writing custom reward functions for multi-turn reinforcement learning in Amazon Nova Forge, the platform it offers for tailoring its Nova family of models. As dev.to reports, the post appeared on the AWS Machine Learning Blog and walks developers through defining reward signals that grade what an agent does across a whole conversation or task, not each individual reply.
Scoring whole conversations, not single replies
Most reinforcement learning setups have historically been tuned for one-shot quality — whether a single answer looks accurate, useful, or on brand. According to dev.to, AWS's guidance argues that this is the wrong yardstick for agents working through long, realistic interactions. The outcomes that matter in a multi-turn setting — whether the exchange reached a resolution, whether the agent stayed consistent with what it said earlier, whether it avoided redundant steps or needless hand-offs — only become visible when you look at the sequence as a whole. Reward functions built for multi-turn training score the agent on those session-level results instead.
Reward rules as code
The concrete change described in the coverage is that Nova Forge lets developers express these reward functions as custom code that runs inside the training loop, giving them direct influence over what the model is optimized to accomplish across a session. Per dev.to, AWS positions this as an alternative to relying exclusively on generic preference models or feedback labels collected from people, both of which are costly to gather at the scale reinforcement learning demands and often miss the success criteria specific to a given domain.
The practical example the coverage highlights is customer support: rather than rewarding a bot for answers that merely sound plausible, a team could define success as closing a ticket in fewer exchanges while still following its escalation policy. The same pattern extends to technical support and task automation, where "done correctly" is easier to codify than "sounds convincing."
Limits and open questions
dev.to flags several caveats. The capability is tied to Nova Forge on AWS infrastructure, so it is not a portable feature across other reinforcement learning tooling. The AWS post includes no independent benchmark comparisons against other reward-modeling approaches, so performance claims should be limited to what AWS itself describes. And pricing details for the reward-function workflow have not been spelled out beyond what the existing Nova Forge documentation already covers.
Why it matters
For teams fine-tuning Nova models, the guidance turns an abstract problem — how to train an agent for long tasks — into a concrete mechanism: write down what success looks like, in code, and let training optimize for it. That narrows the gap between how an agent is judged in evaluation and what it was actually optimized to do while learning.
The broader takeaway travels beyond AWS. Evaluating agents on full-task outcomes rather than individual responses is a principle any team can apply, whether its agent runs on Nova, a competing foundation model, or an in-house automation stack. As agents take on longer, multi-step work, single-turn quality scores look increasingly like the wrong unit of measurement, and this is a signal that a major cloud vendor is building tooling around that shift.
- #aws
- #amazon-nova
- #reinforcement-learning
- #fine-tuning
- #ai-agents