GPU compute optimization

More useful AI from every GPU.

Tensor Machines is building a physics-informed optimization layer that predicts what each GPU can sustain and guides how it should run. The objective: more useful output, lower cost per result, and longer productive hardware life.

GPU + SERVER SCHEMATIC CONCEPTUAL · NO MEASURED VALUES

[01] The problem

More watts don’t guarantee more work.

More power can raise energy use faster than it raises useful output. When memory movement or thermal limits constrain execution, increasing the power budget may deliver little additional work.

The gap between peak compute and useful output is expensive. Operators need to know what is limiting performance and which operating change can recover useful capacity.

ILLUSTRATION
MEMORY
DATA MOVEMENT
COMPUTE
MOST COMPUTE WAITING ON DATA

Memory bottlenecks

When data movement limits execution, more compute alone cannot keep requests moving.

ILLUSTRATION THERMAL LIMIT USEFUL OUTPUTTEMPERATURECLIFF

Performance cliffs

A workload can start fast, then lose throughput as power or thermal constraints take hold.

ILLUSTRATION
SAME MODEL NAME · SUSTAINED OUTPUT
GPU A
GPU B
CONCEPTUAL · NOT MEASURED VALUES

Changing hardware capabilities

Two GPUs with the same model name can sustain different results. Operating conditions and hardware history matter.

[02] Optimization

Run each GPU where it delivers the most.

Tensor’s optimization work connects a workload’s demands to the GPU’s physical response. Our physics-informed model can identify the operating configurations that sustain the most useful work within your service, power, thermal, and hardware limits.

ASSESSMENT DASHBOARD · DEVICE COMPARISON
WORKLOADInference APHASESustained loadDEVICES8 · same model
Performance alongside power, temperature and clocksNo values shown

The measurement foundation: workload performance alongside power, temperature, and clocks.


01/

Tune performance to the workload

Evaluate supported GPU core and memory frequencies against sustained throughput, latency, and energy use. Find where additional frequency helps and where another constraint takes over.


02/

Make cooling serve useful output

Connect thermal response to performance. Evaluate supported server cooling controls against the heat a workload produces and the output it can sustain.


03/

Put power where it pays back

Compare useful output across supported power settings. Identify where another watt adds capacity and where it mostly adds energy cost.


01Backed by
  • Omni
  • Avesta Fund
  • RV
  • Draper University
02Working with
  • Cato Digital
  • Texas A&M University
03Part of
  • NVIDIA Inception Program

[03] Open-Source Benchmark

Start with what your hardware actually delivers.

The Tensor Machines GPU benchmark applies controlled workloads and records how the system performs, draws power, heats up, and recovers. It connects workload results to the physical behavior of the hardware producing them.

Use the assessment to compare devices, investigate performance differences, and establish a repeatable baseline for operating changes. Published methodology makes the test conditions part of the result.

01 BASELINE02 LOAD03 RECOVERY

[04] Physics-Informed Models [Private Beta]

Predict the performance a workload can sustain.

Every execution choice places a different demand on memory, power, and cooling. The hardware’s response then shapes how much work gets done.

Our models study that relationship. We are developing them to estimate each GPU’s performance envelope: the throughput, latency, and energy outcomes attainable for a workload as conditions change.

That prediction can guide a more useful question: which supported configuration or placement is expected to deliver the best sustained result?

PERFORMANCE RESPONSE OPTIMIZATION ARCHITECTURE IN DEVELOPMENT
01 INPUTS
OBSERVEDWorkload contextModel · request length · concurrency · phase
ESTIMATEDPhysical stateThermal and electrical state, inferred from telemetry
CANDIDATESOperating choicesSupported frequency, power and cooling settings
02 RESPONSE MODELPhysics-informed response model PRIVATE BETA
03 PREDICTED OUTCOMES
Throughput
Latency
Energy
Thermal response
04 DECISION Best supported configuration or placement, within service, power and thermal limits
OBSERVEDESTIMATED / PREDICTED

Operator Use Case

Bring the decision you need to make.


A/

GPU clouds and inference providers

The performance you can sustain determines the capacity you can sell. Evaluate supported frequency, power, and cooling settings against the throughput, latency, and cost your service needs.


B/

Enterprise and research infrastructure teams

Establish a performance baseline, investigate variation across systems, and compare configurations against the requirements of your workloads.


C/

Hardware lifecycle teams

Build an operating record to inform continued use, maintenance, and redeployment as hardware ages.


How much useful AI can your infrastructure deliver?

Bring us the workload, service target, and constraint. Start with a repeatable assessment, then define an optimization evaluation around the result you need to improve.