A model must translate a short story problem into arithmetic reasoning and produce the expected answer. Evaluations usually parse the final numerical result while prompts may include or omit worked examples. Exact setup matters because intermediate reasoning prompts can change performance.