Dynamic workloads
Workloads whose shapes or execution paths vary at runtime, such as sparse LiDAR perception with changing occupied-voxel counts.
Get more performance from your autonomy/perception stack than you thought possible -
with clarity on how much more your onboard compute can actually give.
Supported platforms today. More coming soon.
The steps below let RunLocal optimize your proprietary stack while minimizing IP sharing. For open-source models, just share the repository and target HW/SW versions.
Your coding agent uses the RunLocal Skill to locally prepare a Request Package from your internal model stack. You review it, then your agent uploads it to the RunLocal Platform.
No weights or training data are uploaded; full source code is not required. See the docs for details.
Our agentic ML compiler autonomously optimizes your model stack for your target compute platform (running on RunLocal’s infrastructure).
It experiments with inference optimizations, including custom kernels, quantization, operator fusion, and scheduling. It doesn’t perform distillation or broader model architecture changes (currently).
Optimized code, results, and insights are uploaded to the RunLocal Platform as a Deliverable Package. Your coding agent downloads it and integrates the changes locally. You then validate the results locally.
When standard compilers and libraries don’t deliver the performance you need
Workloads whose shapes or execution paths vary at runtime, such as sparse LiDAR perception with changing occupied-voxel counts.
Multiple models sharing hardware or feeding into one another, where coordinating execution and data movement can reduce overall latency.
Custom and emerging model operations that need more than standard library kernels.
Improve performance, reduce costs, and plan your next move with more confidence.
Run your stack faster on the same compute
Fit larger models on the same compute
Do more without hiring more rare ML compiler talent
Meet performance targets with cheaper compute
Better estimate whether models will fit your compute
Better estimate performance headroom before committing
Helping the agent choose better optimizations, understand the performance headroom remaining, and squeeze every drop out of your target compute.
For supported workloads, demand modeling and calibrated device models help estimate how fast your workload could run and where performance headroom remains. These predictions help the agent choose which implementation changes to try and understand their combined effect on latency.
Calculates the instructions, data movement and resources an implementation requires. Our MLIR-based compiler gives agents control over kernels, fusion, layouts and scheduling beyond the standard vendor path. Custom dialects connect those choices to their hardware demand through compiler stages.
Benchmarks across opcodes, memory access patterns, parallelism and operating conditions calibrate models of hardware speeds and limits.
Combines compiler demand with calibrated device models to predict latency, identify modeled bottlenecks and explain the combined effects of implementation changes.
Instructions, data movement and resource demand for a proposed implementation
Hardware speeds and limits, calibrated against real benchmarks
Combines workload demand with the device’s calibrated behavior
Estimated time to run the workload
How proposed changes affect performance
Where further optimization could help
Uses predictions to choose what to try next
Predictions guide which implementations to test on the device.
Different models share operations such as matrix multiplication and softmax, built from common hardware instructions and memory operations. Calibrating these building blocks and recording compiler effects creates knowledge that can apply beyond the model that produced it.
The system retains calibrations and successful and failed experiments across operations, shapes, precisions and layouts. Results keep their compiler choices, measured performance, correctness and device conditions, so the agent can judge where each finding applies.
This shared knowledge is designed to help the agent match new workloads with relevant findings from other models, avoid repeated unsuccessful experiments and refine hardware models where predictions differ from measurements.
Hardware calibrations and successful & failed experiments
Calibrated costs and resource limits
Compiler choices and measured effects
With shapes, precisions, layouts and device conditions
Applies relevant findings to new models
Shared operations make reuse possible. Hardware and execution conditions determine where findings apply.
Shared scheduling keeps devices testing one agent’s changes while other agents write code or analyze results. This reduces idle time and runs more experiments with the same hardware.
On-device experiments check performance and output correctness under controlled conditions. Results link to the implementation, inputs and device settings, so agents can compare changes and check the reasoning engine’s predictions.
Write code · analyze results
Test changes while other agents code or analyze.
Measured performance + output correctness
More experiments with the available hardware, with less idle time.
Help customers get more from your silicon. Bring their models onto your platform faster, broaden your model zoo, and improve the software behind it.
Book A Call (opens in a new tab)Port and optimize customer models faster without hiring more support engineers.
Add and maintain support for more models without hiring a large internal optimization team.
Use insights from RunLocal to improve your compiler, libraries, and developer tools.
See how it works for the workflow, and read the docs for more detail.
The challenge is helping the agent decide which implementation changes to try, understand their combined effects, and know how much performance remains. RunLocal adds compiler control, calibrated hardware models, reusable experiment knowledge and infrastructure for running and validating the search. See our secret sauce.
Building it means maintaining compiler integrations, hardware calibrations, reliable device experiments and correctness checks across models, devices and SDK versions. RunLocal provides that system and the hardware knowledge already collected, so your team can focus on its autonomy/perception stack.
We’re also building reusable compiler rules and experiment knowledge across projects, extending the coverage an individual team would otherwise need to develop itself.
For supported workloads, our own MLIR-based compiler provides an alternative path when the standard vendor compilation path restricts useful changes. On Qualcomm, it can produce QNN graphs directly while retaining Qualcomm’s backend and runtime. On NVIDIA, we can generate custom kernels for selected parts of a workload and retain TensorRT where it is useful. The choice depends on the workload and the performance opportunity.
Compiler-level changes include custom kernels, operator fusion, memory layouts, data movement, precision choices and execution scheduling. We consider the full inference workload, including interactions between operations and models. Distillation and broader model architecture changes are outside the current scope.
No. The Request Package describes your stack without requiring weights, training data, or full source code. See the docs.
A Deliverable Package on the RunLocal platform: optimized code or a binary, a performance report, and integration instructions. Your coding agent downloads and integrates it locally. You validate correctness and performance on your target compute.
A subscription gives your team access to RunLocal for a fixed recurring fee. Optimization requests are priced separately. For each request, we agree on performance targets, provide a fixed quote, and estimate the timeline. You only pay for that optimization if the agreed targets are met.
Start by booking a call to discuss your stack and target compute. We can then align on a 1-week pilot to prove RunLocal's optimization capability before you commit to something longer-term.
Tell us about your stack and target hardware. We’ll discuss where RunLocal could help and whether a 1-week pilot is a good fit.
BACKED BY