Varion Compute · Commercial Early Access

More AI throughput. Less energy per generated token.

Breakthrough compute-efficiency results validated on real LLM inference workloads using NVIDIA A40 hardware.

Current validation is on NVIDIA A40. Results vary by model, hardware and workload; broader cross-GPU validation is ongoing.
Validated results

Measured on real LLM workloads.

Selected completed validation runs from the current NVIDIA A40 evidence set.

36.4%lower energy per generated token · Qwen 2.5 1.5B
131.4%higher generation throughput · Qwen 2.5 1.5B
100%identical greedy-generation output in validated KV-cache tests
1.5Bparameter Qwen model validated
Evidence set

Results that survived validation.

No universal percentage is claimed.

DistilGPT-2

31.3% lower energy per generated token
54%+ higher generation throughput
100% identical validated greedy output

Qwen 2.5 0.5B

35.1% lower energy per generated token
70.3% higher generation throughput
100% identical validated greedy output

Qwen 2.5 1.5B

36.4% lower energy per generated token
131.4% higher generation throughput
100% identical validated greedy output

Test before buying

Run Varion Compute against your own workload.

Commercial early access includes a free self-test path and enterprise evaluation options.

Current validation boundary

NVIDIA A40 results are validated. Cross-GPU claims are not made until additional hardware results are completed.