More AI throughput. Less energy per generated token.
Breakthrough compute-efficiency results validated on real LLM inference workloads using NVIDIA A40 hardware.
Measured on real LLM workloads.
Selected completed validation runs from the current NVIDIA A40 evidence set.
Results that survived validation.
No universal percentage is claimed.
DistilGPT-2
31.3% lower energy per generated token
54%+ higher generation throughput
100% identical validated greedy output
Qwen 2.5 0.5B
35.1% lower energy per generated token
70.3% higher generation throughput
100% identical validated greedy output
Qwen 2.5 1.5B
36.4% lower energy per generated token
131.4% higher generation throughput
100% identical validated greedy output
Run Varion Compute against your own workload.
Commercial early access includes a free self-test path and enterprise evaluation options.
Current validation boundary
NVIDIA A40 results are validated. Cross-GPU claims are not made until additional hardware results are completed.