OpenAI-compatible API gateway
Change the base URL, authenticate with a Varion key and pass your provider credential securely per request.
- Chat Completions and Responses API
- Streaming and tool calls
- Provider-reported usage measurement
Varion uses proprietary optimisation technology to reduce avoidable input-token usage across supported AI providers while preserving request intent and expected quality.
One verified example, not a guaranteed result. Up to 60% applies to eligible workloads.
Choose a supported integration today, with more AI providers planned for the future.
Change the base URL, authenticate with a Varion key and pass your provider credential securely per request.
Keep Claude Code installed as it is. Set three local environment variables and start Claude normally.
/v1/messages gatewayChoose a supported integration, test your own workload and connect production with minimal changes.
Receive 100,000 trial input tokens and create a separate Varion API key.
Compare original and optimised requests before routing production traffic.
Connect a supported workflow, then track original usage, provider-bound usage and measured savings.
Varion focuses on measurable customer outcomes while its optimisation methods remain proprietary.
Reduce avoidable input-token usage across supported AI providers.
Designed to maintain request intent, critical instructions and expected output quality.
Connect existing applications and AI workflows with minimal changes.
View original usage, provider-bound usage and verified token savings.
Use current integrations and add future providers through one scalable platform.
Requests remain protected whenever optimisation is not suitable.
Choose a non-expiring package that matches your usage.
Support real applications, development tools and high-volume AI workflows.
Varion reports original input, provider-bound input and measured token savings without exposing proprietary optimisation methods.
These are measured examples, not guaranteed averages. Savings vary by workload and usage.
Prepay Varion processing tokens. Charges from your selected AI provider remain separate.
10 million Varion input tokens
50 million Varion input tokens
200 million Varion input tokens
1 billion Varion input tokens
No. “Up to 60%” describes eligible workloads. Results depend on repetition, conversation length, tool schemas and safety classification.
Varion is designed to preserve critical instructions and coding intent. Uncertain or high-risk requests default to safe passthrough, but customers should test their own workloads.
Yes. OpenAI-compatible applications use the Varion /v1 base URL. Claude Code uses the native Anthropic Messages gateway at https://api.varion.tech.
Local testers keep provider keys on the customer device. In direct production gateway mode, the credential is forwarded transiently to the selected provider and is not stored in normal request logs.
One original customer input token deducts one Varion token after a successful provider request. Failed provider requests deduct zero.
Create an account and receive 100,000 trial input tokens after email verification.