Physics-informed GPU optimization
Predict what each GPU can sustain. Choose how it should run.
Execution changes a GPU’s physical state. That changing state affects execution. Tensor Machines models this relationship to connect the workload you want to run with the performance your hardware can deliver.
We are building a predictive optimization layer over existing GPU infrastructure, combining workload context, physical modeling, and measured feedback to guide operating choices.
[01] Hard inference problems
Understand the constraint. Recover the useful work.
Relieve memory bottlenecks
Inference can spend valuable GPU time waiting for data. Longer contexts and greater concurrency also increase pressure on memory capacity. The right response depends on what is limiting execution: bandwidth, capacity, compute, communication, or the system’s operating state.
Our modeling direction combines workload and runtime measurements with physical behavior to distinguish these constraints and evaluate changes that can relieve them.
Anticipate performance cliffs
Every performance cliff has a cost. When throughput drops or latency spikes, the same fleet can complete less work within its service targets. We aim to forecast how power and temperature evolve during execution and when they are likely to reduce useful capacity.
Adapt to the hardware you actually have
Devices with the same specification can respond differently to a workload. We aim to model each system’s measured capability as operating conditions and hardware behavior change, so configuration and placement can reflect what it can sustain.
[02] Performance response models
Model the response before choosing the next move.
A useful performance model needs three things: the work being requested, the system’s current state, and the execution choice being considered. Tensor’s architecture brings them together to estimate how much useful output a GPU can sustain over a defined interval.
01/
Understand the workload
Characterize the model, request lengths, concurrency, precision, and execution phase. Prefill and decode can place different demands on compute and memory. Runtime measurements provide the context needed to interpret hardware behavior.
02/
Estimate the physical response
Combine telemetry and operating history with models of thermal and electrical behavior. Estimate how the system heats, cools, and responds to load, and how that response affects attainable performance.
03/
Compare supported operating choices
Predict throughput, latency, energy, and thermal response under candidate settings or placements. Evaluate them against the service target and validated operating limits, accounting for prediction uncertainty and the cost of changing plans.
04/
Measure what actually happened
Compare the observed result with the prediction. Use repeated, controlled measurements to refine the model and determine whether a change improved useful output.
The performance envelope describes what a workload can achieve. The operating limits define the conditions within which it must run. Tensor’s objective is to find the most productive combination.
TECHNICAL DETAIL · RESEARCH AND MODEL DEVELOPMENTFrom hardware condition to attainable performance
Build a model of the system behind the signals.
Our early work focused on hardware condition and remaining productive life. The same measurement foundation supports a broader objective: predicting slowdown risk, attainable throughput, and the response to a different execution plan.
Condition contributes to that prediction alongside workload demand, software behavior, and operating limits. Learning which intervention will help requires measurements of the system under different operating choices.
Observe the system over time
Combine available power, temperature, clock, utilization, voltage, fan, throttling, error, and interconnect measurements with workload context. Compare each device with its own baseline and relevant peers.
Combine dynamics with physical relationships
One view describes how signals evolve, including recurrence patterns and transitions between operating regimes. Methods under study include recurrence analysis, dynamic mode decomposition, and topological data analysis. A second view describes thermal and electrical response, including heating, recovery, and estimated stress over time. Our research combines these views to estimate the hardware state underlying the observations.
Constrain the predictions
Physical relationships and validated operating limits inform model training and candidate selection. The objective is a workload-specific estimate of performance, energy, and future system behavior, with uncertainty attached to the forecast.
Forecast usable capacity
Estimate the work a device or fleet can complete within a defined interval and service target. Account for queueing, shared bottlenecks, and changing device behavior when translating individual predictions into capacity an operator can plan around.
+ Candidate plans
[03] Performance, thermal and power optimization
Turn physical understanding into operating decisions.
Performance and frequency
Evaluate increases or reductions in supported core and memory clock settings against the workload’s limiting resource. Measure sustained throughput and latency alongside power and temperature. Select settings within validated manufacturer and operator limits, including when a lower setting produces the better cost-performance tradeoff.
Thermal management
Connect cooling response to sustained performance. Where server controls are available, evaluate fan or cooling policy changes against thermal headroom, workload output, and measured energy use.
Power management
Characterize the relationship between power settings and useful output. Compare configurations against the available power budget and service target, including whether a lower-power operating point offers better economics.
Optimization modules are qualified for the hardware and accepted scope of each engagement. Supported controls, operating limits, and measurement coverage determine what can be evaluated.
Our commercial optimization program focuses on three connected controls. Each evaluation begins with the workload, the objective, and the controls supported by the actual server.
[04] Workload-aware execution [Execution Roadmap]
Make execution respond to the physical system.
Route work to the capacity that can sustain it
Evaluate which GPU and configuration can meet a request’s model, context-length, and latency requirements. Compare predicted capacity with queueing, data locality, and reusable KV cache. Moving to another GPU should earn back the transfer and cache-rebuilding costs.
Recover time lost to stalls and recomputation
Use serving metrics to evaluate batching, concurrency, prefix reuse, and KV-cache pressure. Excess admission can trigger preemption and recomputation in serving engines. The objective is enough concurrent work to use the GPU effectively while meeting the service target.
Keep the decision process economical: evaluate a bounded set of supported plans, minimize time in the token-generation path, and recalibrate when observed behavior departs from the model. Uncertainty, switching costs, and Tensor’s own resource use belong in every comparison.
Our longer-term direction extends the response model into placement and serving decisions. For each candidate plan, estimate the useful output it can sustain, the resources it will consume, and whether a change is worth its overhead.
Reduce the work and data movement per result
Evaluate supported plans that improve data reuse, reduce unnecessary memory transfers, or use validated lower precision, including KV-cache quantization. Preserve the declared quality requirement. Memory bandwidth and capacity remain distinct constraints in the model.
Choose execution paths for the actual device
Characterize supported kernel variants, including their throughput, power, and thermal response. A future dispatcher could compare tile sizes, fusion choices, or parallelism configurations against the device’s predicted sustained behavior. Optimized kernels remain the execution foundation; physical-response modeling adds context for choosing among them.
Preserve performance through sustained load
Evaluate changes to placement, workload intensity or timing, or supported operating settings before a predicted loss of capacity. Use measured feedback to determine whether the intervention helped.
[05] Performance passports [Product Direction]
Give every GPU a record of what it can deliver.
We are developing the performance-passport concept as a record of each GPU’s measured behavior, operating history, and workload-specific capabilities. The aim is to help operators assess workload fit as hardware remains in service.
A useful passport combines test conditions and observed results with clearly identified model estimates. It can track thermal response, stress history, performance variation, and the configurations evaluated for a particular workload.
Match capability to service requirements
Use demonstrated performance to qualify devices for latency-sensitive serving, long-context requests, batch inference, or embeddings. Devices that need further investigation can be flagged for review before reassignment, maintenance, or retirement.
Device and system diagnosis
Find where useful performance is being lost.
A slowdown may originate in a GPU, an interconnect, or a shared power or cooling constraint. Our diagnostic direction combines device comparisons, server context, and available topology and runtime measurements to narrow where the loss begins.
The objective is to distinguish a local device issue from a link, server, or rack-level constraint and guide the next targeted test or operating action. Affected devices may need reduced load, isolation, or maintenance review; the decision should follow the evidence and operating policy.
Device record
GPU-07 · CONCEPT[06] Useful output and economics [Framework in development]
Optimize for work that meets the requirement.
The target is more completed, useful work per GPU-hour at a lower cost per result. For inference, the model, output quality, request length, and service deadline define what counts as useful.
We are developing a measurement framework that connects the output meeting those requirements with the time, energy, and cost consumed to produce it.
Service attainment
What share of eligible requests meets the declared deadline and latency requirements? Publish the service target and workload conditions with the result.
Energy per accepted output token
How much measured energy is consumed for each output token from requests that meet the acceptance criteria? Include energy spent on failed or late work within the stated measurement boundary.
Cost per accepted output token
What is the stated cost of the test interval divided by the accepted output tokens produced? Specify the included hardware, energy, and operating costs.
Device and system diagnosis
Maximum useful performance is an objective with constraints: service levels, quality, available power, thermal limits, and hardware operating policy. We aim to learn a cost-performance curve for each GPU and workload, so an operator can compare useful output with the total cost of sustaining it.
[07] Benchmark and validation
Make every optimization answer to a measured result.
The open-source benchmark recipe provides a repeatable measurement foundation: establish an idle baseline, apply controlled workloads, observe recovery, and compare like-for-like runs.
To evaluate an optimization, define the workload and service target, record the baseline, change a supported setting, and repeat under comparable conditions. Sustained output, energy, and latency belong in the result together.
[08] Integration and deployment
Start with the system you operate.
Our assessment software collects available GPU telemetry through NVIDIA NVML and server telemetry through the BMC. These sources connect device performance with the power and thermal behavior of the surrounding system.
The open benchmark recipe makes the assessment method inspectable. Tensor’s proprietary models and optimization work build on that measurement foundation through the partner program.
Hardware qualification
Confirm telemetry access, the hardware and driver profile, and supported power, frequency, and cooling controls. Evaluate changes within agreed manufacturer and operator limits.
Runtime context
Workload-aware prediction needs serving metrics as well as hardware signals. Execution controls and automated feedback require integration with the relevant runtime or scheduler.
Operational fit
Begin with a defined evaluation scope and agreed test windows. Validate the result before expanding the workload, hardware coverage, or degree of automation.
[09] Getting started
Bring one workload and a result worth improving.
Define useful output
Choose the throughput, latency, energy, or cost objective and the service and quality requirements that must hold.
Establish the baseline
Confirm hardware support and run a controlled assessment. Preserve workload, software, power, and cooling conditions with the result.
Qualify the operating choices
Work with Tensor to select supported settings and the scope of an optimization evaluation.
Measure the difference
Repeat the workload, compare sustained results, and use the evidence to decide what to deploy or test next.
[10] FAQ
Questions operators ask
-
A physics-informed optimization layer for GPU infrastructure. Our models study how workloads interact with each system’s physical state, with the objective of predicting attainable performance and guiding operating choices. Assessments provide the measurement foundation; models are in private beta and the optimization layer is in development.
-
Work completed within the requirements of the workload. For inference, throughput matters alongside latency and output quality. We aim to increase that qualifying output per GPU-hour and per unit of energy.
-
Memory bandwidth and capacity remain real limits. Physical modeling contributes an estimate of what the system can sustain under a given execution plan. Combined with workload and runtime measurements, that estimate can help evaluate changes to data movement, concurrency, placement, or supported operating settings.
-
Tensor’s proposed execution layer would work with existing serving engines and optimized kernels. Its contribution is a model of the physical system’s response, helping evaluate which execution and operating choices suit the actual hardware and workload.
-
Performance, thermal, and power modules are evaluated after hardware qualification. Supported controls and activation scope depend on the server, driver, permissions, and agreed operating limits. Contact us to discuss your environment.
-
GPU workloads are the broader scope. Training throughput and simulation completion time are potential applications of the same physical-response approach, with their own workload models, objectives, and validation.
-
The benchmark stress recipe and published methodology. The physics-informed models and surrounding platform tooling remain proprietary. The repository defines the exact release contents and license.
-
No. Workload behavior, software, power settings, and cooling can all affect a comparison. Diagnosing degradation requires additional evidence and repeated observations.
-
The findings section describes an eight-GPU A100 assessment, with workload and operating conditions attached. These are baseline comparisons. Optimization gains will be reported with the intervention, comparison method, and service requirements used to measure them.