After shortening a prompt, check the savings using the usage for each response and the cost of the entire task. A short answer may come with substantial reasoning usage, and continuing a conversation may count its history again. To monitor your budget, separate token volume, pricing categories, and the number of calls needed to get the result.

Four counters and how they nest
In its usage object example, OpenAI shows input, output, and the details for both counters. Read them as follows:
input_tokens— the total input tokens for this call, including context used to continue a conversation.input_tokens_details.cached_tokens— the portion of input served from the cache. It is already included ininput_tokens.output_tokens— the total number of generated tokens, including reasoning.output_tokens_details.reasoning_tokens— the portion of output used for reasoning.
Reasoning tokens are already included in output_tokens; do not add them to the output a second time. OpenAI bills them at the output token rate. Similarly, adding cached_tokens to input counts the same portion of the request twice. For the total, use total_tokens or the sum of input_tokens + output_tokens.
Cost: break input down by category
The current pricing page lists standard input, cached input, cache writes, and output separately. Choose rates for the actual model, processing mode, and applicable context length. Store the exact values in a version-controlled calculation configuration.
An important detail in the current caching documentation: for GPT-5.6 and newer models, input_tokens_details.cache_write_tokens is counted. If your pricing scheme charges cache writes separately, subtract this category from standard input and apply its rate.
I = input_tokens
C = input_tokens_details.cached_tokens
W = input_tokens_details.cache_write_tokens
O = output_tokens
ordinary_input = I - C - W
cost = ((I - C - W) * P_input
+ C * P_cached
+ W * P_write
+ O * P_output) / 1_000_000
Prices here are per million tokens; for a scheme without a separate cache-write category, use W = 0. The formula covers the token categories listed above. Add paid tools as separate line items according to their pricing terms.
Hypothetical example: input is 4,000, cached input is 3,000, cache writes are 500, and output is 1,000, of which 600 are reasoning tokens. This gives 500 standard input tokens and 5,000 tokens in total. The billed output is 1,000 tokens; retain the value of 600 for diagnostics.
Continuing a conversation uses input again
When managing history manually, the application sends previous messages along with the new input. Messages included in the next request become input context again. Therefore, measuring only the latest user message misses the cost of conversation history.
With previous_response_id, the application passes a reference to the previous response, and the API links the context. According to the rules for accounting for continuations, previous input tokens in the chain are billed again as input. Caching may change the pricing category for some of this input; check cached_tokens in each response to see what was actually cached.
How to verify optimization results
- Save a baseline sample. Compare the same tasks, model, success criterion, and comparable conversation lengths.
- Log every call. Record response and task IDs, model, reasoning settings, status, latency, and the complete
usage. Save the prompt version and pricing configuration. - Sum task costs. Include all calls and responses from retries. The primary metric is the cost per successfully completed task.
- Separate the reasons for changes. Track standard input, cache reads and writes, output, and the reasoning share. For the aggregate cache-hit rate, divide the sum of
Cby the sum ofI. - Check quality and the distribution tail. Along with average cost, compare p95, the number of attempts, and the task success rate.
The max_output_tokens parameter limits generation together with reasoning. If the limit is reached, the response may have an incomplete status, including cases with no visible text but with token usage. Include such responses when checking savings: optimization has met its goal when comparable tasks are completed successfully at a lower total cost.