Varion Compute · Commercial Release

More AI throughput. Less energy per generated token.

Breakthrough compute-efficiency results validated on real LLM inference workloads using NVIDIA A40 hardware.

Current validation is on NVIDIA A40. Results vary by model, hardware and workload; broader cross-GPU validation is ongoing.
Validated results

Measured on real LLM workloads.

Selected completed validation runs from the current NVIDIA A40 evidence set.

36.4%lower energy per generated token · Qwen 2.5 1.5B
131.4%higher generation throughput · Qwen 2.5 1.5B
100%identical greedy-generation output in validated KV-cache tests
1.5Bparameter Qwen model validated
Evidence set

Results that survived validation.

No universal percentage is claimed.

DistilGPT-2

31.3% lower energy per generated token
54%+ higher generation throughput
100% identical validated greedy output

Qwen 2.5 0.5B

35.1% lower energy per generated token
70.3% higher generation throughput
100% identical validated greedy output

Qwen 2.5 1.5B

36.4% lower energy per generated token
131.4% higher generation throughput
100% identical validated greedy output

10 free GPU-hours

Run Varion Compute against your own workload.

Commercial release includes 10 free GPU-hours for metered testing, plus the offline self-test path. No credit card is required for the free trial.

Current validation boundary

NVIDIA A40 results are validated. Cross-GPU claims are not made until additional hardware results are completed.

Open benchmark data

Inspect the published validation snapshot.

The machine-readable JSON records the current NVIDIA A40 validation boundary and headline results used on this page. Results remain workload-specific and are not a universal savings guarantee.