More AI throughput. Less energy per generated token.
Breakthrough compute-efficiency results validated on real LLM inference workloads using NVIDIA A40 hardware.
Measured on real LLM workloads.
Selected completed validation runs from the current NVIDIA A40 evidence set.
Results that survived validation.
No universal percentage is claimed.
DistilGPT-2
31.3% lower energy per generated token
54%+ higher generation throughput
100% identical validated greedy output
Qwen 2.5 0.5B
35.1% lower energy per generated token
70.3% higher generation throughput
100% identical validated greedy output
Qwen 2.5 1.5B
36.4% lower energy per generated token
131.4% higher generation throughput
100% identical validated greedy output
Run Varion Compute against your own workload.
Commercial release includes 10 free GPU-hours for metered testing, plus the offline self-test path. No credit card is required for the free trial.
Current validation boundary
NVIDIA A40 results are validated. Cross-GPU claims are not made until additional hardware results are completed.
Inspect the published validation snapshot.
The machine-readable JSON records the current NVIDIA A40 validation boundary and headline results used on this page. Results remain workload-specific and are not a universal savings guarantee.