Token Optimisation

Spend less on eligible AI input traffic.

Measure the result on your own workload before committing production traffic.

100,000 free processed input tokens + 50 local test runs after email verification · Results vary by workload
Measure first

Prove the value on your own traffic.

Use the free testing allowance before a production decision.

Lower eligible input spend

Focus on measurable input-token reduction rather than a universal savings promise.

Validate before rollout

Compare results on representative requests before production adoption.

Independent product access

Use Token Optimisation without purchasing Forge or Compute.

NVIDIA CUDA acceleration

Available for supported high-volume retrieval workloads, with automatic CPU fallback when a compatible NVIDIA CUDA runtime is not available.

Packages

Token Optimisation pricing

Prepaid processed-input-token packages.

Loading packages…