Choose the next test from the unresolved decision
After a comparison, identify what still prevents a decision. If the candidate has a known critical failure, the next work is to understand and repair that behaviour or reject the candidate. Running a much larger ordinary sample may improve the headline score without addressing the stopping condition. Keep the test programme focused on the evidence needed to choose keep, revise or release under the agreed rules.
If the candidate appears equivalent on a small sample, ask which uncertainty matters commercially. You may need more examples from an unrepresented document group, a longer task with state changes or better measurement of review effort. Do not simply repeat the same cases until the average looks stable. Record why each additional test was selected and how it could change the decision. This prevents evaluation effort from drifting into open-ended model comparison.
If the proposed benefit is lower cost, verify the relevant usage and the surrounding work before expanding testing. A small difference in model charges may be outweighed by conversion, tool calls or human review. Keep actual measured charges separate from assumptions about future volume. The decision may be to preserve the current model and improve an expensive retry or retrieval path, which is a different change that should receive its own evidence.
If a provider migration makes some change necessary, the comparison still needs clear acceptance boundaries. The available choice may be a narrower replacement, additional review or a staged transition rather than a direct swap. Document that constraint without claiming the candidate is superior. A release decision can be justified by continuity requirements while still stating the extra work, limitations and monitoring needed to keep the business process controlled.