Inputs the calculator needs
Use provider-reported uncached input tokens, output tokens and the number of similar sessions. If cached input has a different rate, calculate it separately or include only the uncached portion here. Rates should be entered per one million tokens.
How to improve the estimate
Average several normal sessions, including failed attempts and retries. Calculate cost per completed task. Update model rates whenever your provider changes pricing or when you switch models. Keep a low, expected and high scenario for planning.
Cost calculator
Enter your current rates and measured average. Nothing is uploaded.
Testing a lower-input scenario
After calculating the baseline, reduce only obvious repetition or irrelevant context and enter the measured provider-bound input. The difference is an estimate until the provider confirms actual usage and the task outcome remains acceptable.
Measurement checklist
- Choose a representative completed task, not an artificial one-line prompt.
- Record the selected model, provider input, cached input, output, retries and final result.
- Change one optimization mechanism at a time so the cause remains visible.
- Verify required identifiers, tool calls, code changes or business fields.
- Keep passthrough available when the reduced request does not pass.
How Varion fits
Varion Token Engine is a gateway and testing platform for reducing eligible input-token waste across supported AI traffic. It reports original and provider-bound input, keeps provider charges separate, and does not claim that every request can be reduced. New verified users receive 100,000 processed input tokens and 50 local test runs.
Frequently asked questions
Does this calculator know my Claude plan?
No. It runs locally in your browser and uses only the values you enter.
Are the results sent to Varion?
No. The calculator runs in the browser. Testing through Varion is a separate action you choose.