Integration model
OpenAI-compatible applications can use a Varion base URL and Varion API key. Upstream provider credentials are supplied according to the documented integration and are not stored in normal request logs. Native provider integrations may use separate gateway endpoints.
What gets measured
Varion records original input, provider-bound input, cache status, provider usage when available and verification fields. Prompt and generated-answer content are excluded from normal request logs. Customers can test before production.
Local token estimate
This is an approximate comparison, not provider billing data.
Safety behavior
Critical instructions should remain exact. Strategies are selected by workflow, and unverified requests can pass through unchanged. Failed provider calls do not consume Varion processing tokens under the current package model.
Measurement checklist
- Choose a representative completed task, not an artificial one-line prompt.
- Record the selected model, provider input, cached input, output, retries and final result.
- Change one optimization mechanism at a time so the cause remains visible.
- Verify required identifiers, tool calls, code changes or business fields.
- Keep passthrough available when the reduced request does not pass.
How Varion fits
Varion Token Engine is a gateway and testing platform for reducing eligible input-token waste across supported AI traffic. It reports original and provider-bound input, keeps provider charges separate, and does not claim that every request can be reduced. New verified users receive 100,000 processed input tokens and 50 local test runs.
Frequently asked questions
Does Varion replace my AI provider?
No. Customers keep their provider account and pay provider charges separately.
Is every request reduced?
No. Some requests are already compact or unsafe to change and should pass through.