Skip to content

The agentic ML compiler
for edge compute platforms

Get more performance from your autonomy/perception stack than you thought possible -
with clarity on how much more your onboard compute can actually give.

NVIDIAJetson
QualcommDragonwing
Ambarella

Supported platforms today. More coming soon.

How it works

The steps below let RunLocal optimize your proprietary stack while minimizing IP sharing. For open-source models, just share the repository and target HW/SW versions.

Read the docs
  1. Prepare & submit

    Your coding agent uses the RunLocal Skill to locally prepare a Request Package from your internal model stack. You review it, then your agent uploads it to the RunLocal Platform.

    No weights or training data are uploaded; full source code is not required. See the docs for details.

  2. RunLocal optimizes

    Our agentic ML compiler autonomously optimizes your model stack for your target compute platform (running on RunLocal’s infrastructure).

    It experiments with inference optimizations, including custom kernels, quantization, operator fusion, and scheduling. It doesn’t perform distillation or broader model architecture changes (currently).

  3. Apply & validate

    Optimized code, results, and insights are uploaded to the RunLocal Platform as a Deliverable Package. Your coding agent downloads it and integrates the changes locally. You then validate the results locally.

Where it’s most useful

When standard compilers and libraries don’t deliver the performance you need

Dynamic workloads

Workloads whose shapes or execution paths vary at runtime, such as sparse LiDAR perception with changing occupied-voxel counts.

Complex pipelines

Multiple models sharing hardware or feeding into one another, where coordinating execution and data movement can reduce overall latency.

Unusual operations

Custom and emerging model operations that need more than standard library kernels.

Value to your business

Improve performance, reduce costs, and plan your next move with more confidence.

Better performance

Lower latency

Run your stack faster on the same compute

More capable models

Fit larger models on the same compute

Lower costs

Hire fewer specialists

Do more without hiring more rare ML compiler talent

Reduce onboard compute

Meet performance targets with cheaper compute

More informed planning

What will fit

Better estimate whether models will fit your compute

How much headroom

Better estimate performance headroom before committing

Our Secret Sauce

Helping the agent choose better optimizations, understand the performance headroom remaining, and squeeze every drop out of your target compute.

Hardware Reasoning Engine

Predict latency and remaining headroom

For supported workloads, demand modeling and calibrated device models help estimate how fast your workload could run and where performance headroom remains. These predictions help the agent choose which implementation changes to try and understand their combined effect on latency.

Demand Modeling

Calculates the instructions, data movement and resources an implementation requires. Our MLIR-based compiler gives agents control over kernels, fusion, layouts and scheduling beyond the standard vendor path. Custom dialects connect those choices to their hardware demand through compiler stages.

Calibrated Device Modeling

Benchmarks across opcodes, memory access patterns, parallelism and operating conditions calibrate models of hardware speeds and limits.

Simulator

Combines compiler demand with calibrated device models to predict latency, identify modeled bottlenecks and explain the combined effects of implementation changes.

From workload demand to performance predictions
Demand Modeling

Instructions, data movement and resource demand for a proposed implementation

Calibrated Device Modeling

Hardware speeds and limits, calibrated against real benchmarks

Simulator

Combines workload demand with the device’s calibrated behavior

Workload performance predictions

Latency

Estimated time to run the workload

Compiler trade-offs

How proposed changes affect performance

Remaining headroom

Where further optimization could help

Generic Coding Agent

Uses predictions to choose what to try next

Predictions guide which implementations to test on the device.

Reusable Hardware Knowledge

Shared hardware primitives

Different models share operations such as matrix multiplication and softmax, built from common hardware instructions and memory operations. Calibrating these building blocks and recording compiler effects creates knowledge that can apply beyond the model that produced it.

Results with the context needed for reuse

The system retains calibrations and successful and failed experiments across operations, shapes, precisions and layouts. Results keep their compiler choices, measured performance, correctness and device conditions, so the agent can judge where each finding applies.

Reuse across models and projects

This shared knowledge is designed to help the agent match new workloads with relevant findings from other models, avoid repeated unsuccessful experiments and refine hardware models where predictions differ from measurements.

Experiments across different models
Model A
Model B
Model C

Knowledge of shared primitives

Hardware calibrations and successful & failed experiments

InstructionsMemory operations

Calibrated costs and resource limits

Compiler choices and measured effects

With shapes, precisions, layouts and device conditions

Generic Coding Agent

Applies relevant findings to new models

Shared operations make reuse possible. Hardware and execution conditions determine where findings apply.

Hardware Experiment Engine

Scheduling across a device pool

Shared scheduling keeps devices testing one agent’s changes while other agents write code or analyze results. This reduces idle time and runs more experiments with the same hardware.

Check predictions against measurements

On-device experiments check performance and output correctness under controlled conditions. Results link to the implementation, inputs and device settings, so agents can compare changes and check the reasoning engine’s predictions.

Shared scheduling across devices

Agents

Write code · analyze results

Shared device pool

Test changes while other agents code or analyze.

Measured performance + output correctness

More experiments with the available hardware, with less idle time.

For silicon vendors

Help customers get more from your silicon. Bring their models onto your platform faster, broaden your model zoo, and improve the software behind it.

Book A Call (opens in a new tab)

Resolve customer requests faster

Port and optimize customer models faster without hiring more support engineers.

Grow your model zoo faster

Add and maintain support for more models without hiring a large internal optimization team.

Improve your software stack

Use insights from RunLocal to improve your compiler, libraries, and developer tools.

How do we use RunLocal?

See how it works for the workflow, and read the docs for more detail.

What makes your agent so good?

The challenge is helping the agent decide which implementation changes to try, understand their combined effects, and know how much performance remains. RunLocal adds compiler control, calibrated hardware models, reusable experiment knowledge and infrastructure for running and validating the search. See our secret sauce.

Why shouldn't my team build this?

Building it means maintaining compiler integrations, hardware calibrations, reliable device experiments and correctness checks across models, devices and SDK versions. RunLocal provides that system and the hardware knowledge already collected, so your team can focus on its autonomy/perception stack.

We’re also building reusable compiler rules and experiment knowledge across projects, extending the coverage an individual team would otherwise need to develop itself.

Is RunLocal an alternative to TensorRT or QAIRT?

For supported workloads, our own MLIR-based compiler provides an alternative path when the standard vendor compilation path restricts useful changes. On Qualcomm, it can produce QNN graphs directly while retaining Qualcomm’s backend and runtime. On NVIDIA, we can generate custom kernels for selected parts of a workload and retain TensorRT where it is useful. The choice depends on the workload and the performance opportunity.

What kinds of optimizations does your agent make?

Compiler-level changes include custom kernels, operator fusion, memory layouts, data movement, precision choices and execution scheduling. We consider the full inference workload, including interactions between operations and models. Distillation and broader model architecture changes are outside the current scope.

Do we need to share weights, training data, or full source code?

No. The Request Package describes your stack without requiring weights, training data, or full source code. See the docs.

How do we receive your optimizations?

A Deliverable Package on the RunLocal platform: optimized code or a binary, a performance report, and integration instructions. Your coding agent downloads and integrates it locally. You validate correctness and performance on your target compute.

How does pricing work?

A subscription gives your team access to RunLocal for a fixed recurring fee. Optimization requests are priced separately. For each request, we agree on performance targets, provide a fixed quote, and estimate the timeline. You only pay for that optimization if the agreed targets are met.

Can we try before we buy?

Start by booking a call to discuss your stack and target compute. We can then align on a 1-week pilot to prove RunLocal's optimization capability before you commit to something longer-term.

Get more from your edge compute

Tell us about your stack and target hardware. We’ll discuss where RunLocal could help and whether a 1-week pilot is a good fit.

Book A Call (opens in a new tab)

BACKED BY

468 CapitalY CombinatorRitual Capital