Physics-informed GPU optimization

Predict what each GPU can sustain. Choose how it should run.

Execution changes a GPU’s physical state. That changing state affects execution. Tensor Machines models this relationship to connect the workload you want to run with the performance your hardware can deliver.

We are building a predictive optimization layer over existing GPU infrastructure, combining workload context, physical modeling, and measured feedback to guide operating choices.

GPU SCHEMATIC
01ComputeWhere work executes, at a frequency the system can sustain
02MemoryBandwidth and capacity that feed the compute
03PowerDelivery and limits shaping each operating point
04CoolingThe path heat takes out to the environment

[01] Hard inference problems

Understand the constraint. Recover the useful work.

What is limiting execution?

Relieve memory bottlenecks

Inference can spend valuable GPU time waiting for data. Longer contexts and greater concurrency also increase pressure on memory capacity. The right response depends on what is limiting execution: bandwidth, capacity, compute, communication, or the system’s operating state.

Our modeling direction combines workload and runtime measurements with physical behavior to distinguish these constraints and evaluate changes that can relieve them.

SERVICE TARGET NOW OBSERVED THROUGHPUT FORECAST ± UNCERTAINTY
Will this rate hold?

Anticipate performance cliffs

Every performance cliff has a cost. When throughput drops or latency spikes, the same fleet can complete less work within its service targets. We aim to forecast how power and temperature evolve during execution and when they are likely to reduce useful capacity.

Which system can sustain this workload?

Adapt to the hardware you actually have

Devices with the same specification can respond differently to a workload. We aim to model each system’s measured capability as operating conditions and hardware behavior change, so configuration and placement can reflect what it can sustain.

PERFORMANCE ENVELOPE THERMAL LIMIT OPERATING POINT (POWER / FREQUENCY) → USEFUL OUTPUT → MEMORY-BANDWIDTH CEILING COMPUTE-BOUND WORKLOAD MEMORY-BOUND WORKLOAD BEST SUSTAINED POINT
Compute-bound Memory-bound Memory-bandwidth ceiling Thermal limit

[02] Performance response models

Model the response before choosing the next move.

PERFORMANCE RESPONSE OPTIMIZATION ARCHITECTURE IN DEVELOPMENT
01 INPUTS
OBSERVEDWorkload contextModel · request length · concurrency · phase
ESTIMATEDPhysical stateThermal and electrical state, inferred from telemetry
CANDIDATESOperating choicesSupported frequency, power and cooling settings
02 RESPONSE MODELPhysics-informed response model PRIVATE BETA
03 PREDICTED OUTCOMES
Throughput
Latency
Energy
Thermal response
04 DECISION Best supported configuration or placement, within service, power and thermal limits
OBSERVEDESTIMATED / PREDICTED
OPTIMIZATION AND VERIFICATION LOOP
01 →Selected supported actionWithin validated limits
02 →Operator approvalInitial evaluation workflow
03 →RuntimeWorkload runs under the new setting
04 →Measured outputThroughput · latency · energy
05 →Model updatePrediction compared with the observed result
↺ MEASURED FEEDBACK UPDATES THE MODEL
Intended architecture, not shipped autonomous control. No improvement figures shown.

A useful performance model needs three things: the work being requested, the system’s current state, and the execution choice being considered. Tensor’s architecture brings them together to estimate how much useful output a GPU can sustain over a defined interval.


01/

Understand the workload

Characterize the model, request lengths, concurrency, precision, and execution phase. Prefill and decode can place different demands on compute and memory. Runtime measurements provide the context needed to interpret hardware behavior.


02/

Estimate the physical response

Combine telemetry and operating history with models of thermal and electrical behavior. Estimate how the system heats, cools, and responds to load, and how that response affects attainable performance.


03/

Compare supported operating choices

Predict throughput, latency, energy, and thermal response under candidate settings or placements. Evaluate them against the service target and validated operating limits, accounting for prediction uncertainty and the cost of changing plans.


04/

Measure what actually happened

Compare the observed result with the prediction. Use repeated, controlled measurements to refine the model and determine whether a change improved useful output.


The performance envelope describes what a workload can achieve. The operating limits define the conditions within which it must run. Tensor’s objective is to find the most productive combination.

TECHNICAL DETAIL · RESEARCH AND MODEL DEVELOPMENTFrom hardware condition to attainable performance

Build a model of the system behind the signals.

Our early work focused on hardware condition and remaining productive life. The same measurement foundation supports a broader objective: predicting slowdown risk, attainable throughput, and the response to a different execution plan.

Condition contributes to that prediction alongside workload demand, software behavior, and operating limits. Learning which intervention will help requires measurements of the system under different operating choices.

Observe the system over time

Combine available power, temperature, clock, utilization, voltage, fan, throttling, error, and interconnect measurements with workload context. Compare each device with its own baseline and relevant peers.

Combine dynamics with physical relationships

One view describes how signals evolve, including recurrence patterns and transitions between operating regimes. Methods under study include recurrence analysis, dynamic mode decomposition, and topological data analysis. A second view describes thermal and electrical response, including heating, recovery, and estimated stress over time. Our research combines these views to estimate the hardware state underlying the observations.

Constrain the predictions

Physical relationships and validated operating limits inform model training and candidate selection. The objective is a workload-specific estimate of performance, energy, and future system behavior, with uncertainty attached to the forecast.

Forecast usable capacity

Estimate the work a device or fleet can complete within a defined interval and service target. Account for queueing, shared bottlenecks, and changing device behavior when translating individual predictions into capacity an operator can plan around.

G9 · STATE ESTIMATION → PERFORMANCE PREDICTION
OBSERVEDINFERREDPREDICTEDRESEARCH · IN DEVELOPMENT

[03] Performance, thermal and power optimization

Turn physical understanding into operating decisions.

MODULE 01

Performance and frequency

Evaluate increases or reductions in supported core and memory clock settings against the workload’s limiting resource. Measure sustained throughput and latency alongside power and temperature. Select settings within validated manufacturer and operator limits, including when a lower setting produces the better cost-performance tradeoff.

MODULE 02

Thermal management

Connect cooling response to sustained performance. Where server controls are available, evaluate fan or cooling policy changes against thermal headroom, workload output, and measured energy use.

MODULE 03

Power management

Characterize the relationship between power settings and useful output. Compare configurations against the available power budget and service target, including whether a lower-power operating point offers better economics.

Optimization modules are qualified for the hardware and accepted scope of each engagement. Supported controls, operating limits, and measurement coverage determine what can be evaluated.

Our commercial optimization program focuses on three connected controls. Each evaluation begins with the workload, the objective, and the controls supported by the actual server.

[04] Workload-aware execution [Execution Roadmap]

Make execution respond to the physical system.


Route work to the capacity that can sustain it

Evaluate which GPU and configuration can meet a request’s model, context-length, and latency requirements. Compare predicted capacity with queueing, data locality, and reusable KV cache. Moving to another GPU should earn back the transfer and cache-rebuilding costs.


Recover time lost to stalls and recomputation

Use serving metrics to evaluate batching, concurrency, prefix reuse, and KV-cache pressure. Excess admission can trigger preemption and recomputation in serving engines. The objective is enough concurrent work to use the GPU effectively while meeting the service target.

EXECUTION TRADEOFF
COMPAREPLAN A · STAYPLAN B · MOVE PlacementCurrent GPU, reuse cacheAnother device or config Qualifying outputpredictedpredicted Latency vs targetpredictedpredicted Energypredictedpredicted Change overheadnonetransfer + cache rebuild

Keep the decision process economical: evaluate a bounded set of supported plans, minimize time in the token-generation path, and recalibrate when observed behavior departs from the model. Uncertainty, switching costs, and Tensor’s own resource use belong in every comparison.

Our longer-term direction extends the response model into placement and serving decisions. For each candidate plan, estimate the useful output it can sustain, the resources it will consume, and whether a change is worth its overhead.


Reduce the work and data movement per result

Evaluate supported plans that improve data reuse, reduce unnecessary memory transfers, or use validated lower precision, including KV-cache quantization. Preserve the declared quality requirement. Memory bandwidth and capacity remain distinct constraints in the model.


Choose execution paths for the actual device

Characterize supported kernel variants, including their throughput, power, and thermal response. A future dispatcher could compare tile sizes, fusion choices, or parallelism configurations against the device’s predicted sustained behavior. Optimized kernels remain the execution foundation; physical-response modeling adds context for choosing among them.


Preserve performance through sustained load

Evaluate changes to placement, workload intensity or timing, or supported operating settings before a predicted loss of capacity. Use measured feedback to determine whether the intervention helped.

[05] Performance passports [Product Direction]

Give every GPU a record of what it can deliver.


CONSTRAINT LOCALIZATION
EVIDENCE FOR THE NEXT TEST
OBSERVEDLink error counters rising on the GPU 5 link
OBSERVEDGPU 5 clocks below its peers under the same load
CONTEXTShares cooling zone B with GPUs 4–7
SUSPECTEDLink, device or cooling: cause not yet confirmed
NEXTTargeted link test before any reassignment

We are developing the performance-passport concept as a record of each GPU’s measured behavior, operating history, and workload-specific capabilities. The aim is to help operators assess workload fit as hardware remains in service.

A useful passport combines test conditions and observed results with clearly identified model estimates. It can track thermal response, stress history, performance variation, and the configurations evaluated for a particular workload.


Match capability to service requirements

Use demonstrated performance to qualify devices for latency-sensitive serving, long-context requests, batch inference, or embeddings. Devices that need further investigation can be flagged for review before reassignment, maintenance, or retirement.

Device and system diagnosis

Find where useful performance is being lost.

A slowdown may originate in a GPU, an interconnect, or a shared power or cooling constraint. Our diagnostic direction combines device comparisons, server context, and available topology and runtime measurements to narrow where the loss begins.

The objective is to distinguish a local device issue from a link, server, or rack-level constraint and guide the next targeted test or operating action. Affected devices may need reduced load, isolation, or maintenance review; the decision should follow the evidence and operating policy.

Device record

GPU-07 · CONCEPT
ASSESSMENT DATE
YYYY-MM-DD
HARDWARE
GPU model · board
SOFTWARE
Driver · runtime
MODEL VERSION
Estimator vX.Y
MEASURED
Sustained throughput · workload Ameasured
Thermal response · load / recoverymeasured
Power under tested capsmeasured
MODEL ESTIMATE
Performance envelope · workload Aestimate ± uncertainty
Estimated stress historyestimate
WORKLOAD QUALIFICATIONS
✓ Batch inference✓ Embeddings Latency-sensitive · not tested Long-context · review

[06] Useful output and economics [Framework in development]

Optimize for work that meets the requirement.

The target is more completed, useful work per GPU-hour at a lower cost per result. For inference, the model, output quality, request length, and service deadline define what counts as useful.

We are developing a measurement framework that connects the output meeting those requirements with the time, energy, and cost consumed to produce it.

accepted / eligible

Service attainment

What share of eligible requests meets the declared deadline and latency requirements? Publish the service target and workload conditions with the result.

J / accepted token

Energy per accepted output token

How much measured energy is consumed for each output token from requests that meet the acceptance criteria? Include energy spent on failed or late work within the stated measurement boundary.

$ / accepted token

Cost per accepted output token

What is the stated cost of the test interval divided by the accepted output tokens produced? Specify the included hardware, energy, and operating costs.

Device and system diagnosis

Maximum useful performance is an objective with constraints: service levels, quality, available power, thermal limits, and hardware operating policy. We aim to learn a cost-performance curve for each GPU and workload, so an operator can compare useful output with the total cost of sustaining it.

[07] Benchmark and validation

Make every optimization answer to a measured result.

The open-source benchmark recipe provides a repeatable measurement foundation: establish an idle baseline, apply controlled workloads, observe recovery, and compare like-for-like runs.

To evaluate an optimization, define the workload and service target, record the baseline, change a supported setting, and repeat under comparable conditions. Sustained output, energy, and latency belong in the result together.

LOAD AND RECOVERY PLOT
01 IDLE BASELINE02 CONTROLLED WORKLOAD03 RECOVERY
t₀TIME →

[08] Integration and deployment

Start with the system you operate.

Our assessment software collects available GPU telemetry through NVIDIA NVML and server telemetry through the BMC. These sources connect device performance with the power and thermal behavior of the surrounding system.

The open benchmark recipe makes the assessment method inspectable. Tensor’s proprietary models and optimization work build on that measurement foundation through the partner program.

COLLECTIONNVML · GPU telemetryBMC · server telemetry
CONTROLSupported settings only · qualified per server

Hardware qualification

Confirm telemetry access, the hardware and driver profile, and supported power, frequency, and cooling controls. Evaluate changes within agreed manufacturer and operator limits.


Runtime context

Workload-aware prediction needs serving metrics as well as hardware signals. Execution controls and automated feedback require integration with the relevant runtime or scheduler.


Operational fit

Begin with a defined evaluation scope and agreed test windows. Validate the result before expanding the workload, hardware coverage, or degree of automation.


[09] Getting started

Bring one workload and a result worth improving.

01

Define useful output

Choose the throughput, latency, energy, or cost objective and the service and quality requirements that must hold.

02

Establish the baseline

Confirm hardware support and run a controlled assessment. Preserve workload, software, power, and cooling conditions with the result.

03

Qualify the operating choices

Work with Tensor to select supported settings and the scope of an optimization evaluation.

04

Measure the difference

Repeat the workload, compare sustained results, and use the evidence to decide what to deploy or test next.

[10] FAQ

Questions operators ask