The AI agent for inference optimization that actually understands your target hardware

Minimize onboard latency of your perception/autonomy stack on embedded compute like NVIDIA Jetson. Stop leaving performance on the table and ship faster — without hiring more inference specialists.

Built for ML teams in

Autonomous VehiclesRoboticsSmart CamerasDrones
RunLocal EnvironmentSelf-hosting possible toeliminate IP security concernsRunLocal Web UIFor you to auditthe agent's workYour Trained Model RepoCode · Weights · Validation DataOnly what's required for inferenceoptimization and testingOff-the-shelf coding agentCodex · Claude · GeminiRunLocal CLIFor the agent touse our engineRunLocal EngineOn-Device Benchmark OrchestrationCompute Performance ModelingExperiment LedgerYour Real Target HardwareJetson Orin/Thor · Qualcomm

The RunLocal Engine

Purpose-built systems for on-device benchmarking, modeling what influences performance, and managing experimentation context — enabling generic coding agents to squeeze out model performance on your target compute, and hit your performance targets faster.

Secret Sauce In The RunLocal Engine

Unglamorous infra nobody wants to build: precise experiment tracking, trustworthy on-device benchmarks, and performance data that's refined into actionable insights for coding agents

On-Device Benchmark Orchestration

Enables parallel agents to safely share and benchmark on your target device, with numbers you can actually trust — something a naive SSH script can't do. Jobs live in a durable queue; retries never run anything twice; devices are fenced mid-measurement; crashes are detected and retried automatically.

Compute Performance Modeling

Turns experimentation and on-device benchmarking history into a causal and predictive understanding of how model changes affect on-device performance - clarifying both your model's bottlenecks and performance ceiling for your target compute.

Experiment Ledger

Every measurement is bound to what produced it (code change, software versions, input data, etc.) in an append-only record, published by the system not the agent. So you can audit any claim, reproduce any result, and every experiment compounds into knowledge the agent can learn and iterate on.

How You Use It

Install and run our CLI to launch your coding agent in the RunLocal environment, connect it to your real models and target hardware, and prompt it like you normally would — our infrastructure does its magic under the hood.

1

Integrate

Connect RunLocal to your model repos (code, weights and validation data) and your real target hardware.

Self-hosting our software in your infra is possible to eliminate IP/security concerns.

2

Optimize

Run our CLI to launch your coding agent in the RunLocal environment, then prompt it exactly like you normally would.

Bring your own AI vendor and API keys (e.g. Codex or Claude).

3

Audit

Use our Web UI to track the agent's experimentation and verify its output, in real time as it works.

Value To Your Business

More optimized models, shipped sooner, on cheaper compute onboard, with a leaner team.

Better

More Optimized Models

Lower latency and memory (same accuracy), or better models with less compute.

Faster

Ship Faster

Hit performance targets in days instead of weeks, or hours instead of days.

Cheaper

Less Headcount & Compute

Avoid hiring rare optimization experts. Downgrade your onboard compute.

Coding Agents Alone Aren't Enough

These failure modes don't go away as coding agents get more capable at generic coding — closing them takes specialized infrastructure built for on-device optimization.

Unreliable On-Device Benchmarking

Benchmarking jobs collide, jobs silently crash, and corrupted runs quietly hinder experimentation.

What it takes

Trustworthy on-device benchmarking, which seamlessly deals with many parallel agents, requires a sophisticated benchmarking system.

Shallow Optimization Hypotheses

Basic optimization hypotheses because they don't have a deep understanding of your target hardware.

What it takes

Cause-and-effect must be derived from benchmarking on real hardware with a specific causal analysis system; better generic reasoning isn't a substitute.

Cheating & Non-Trivial Verification

They find ways to hit performance targets by breaking real constraints, and it's non-trivial to verify results.

What it takes

Changes/results must be precisely recorded/validated independently, and a purpose-built UI is needed for inspecting results and artifacts.

Reduce Costly Bottlenecks

Even with today's coding agents, model inference optimization still drags you into the same grind. With RunLocal, the agent actually handles them for you.

Performance Bugs

Poorly supported layers, unexpected issues after quantizing, and other silent-but-deadly surprises you still end up chasing down alongside your agent.

Endless Trial-and-Error

Babysitting the agent through attempt after attempt, re-explaining context, and hand-holding it toward something that actually hits your numbers.

Missed Performance Gains

Not knowing whether you're near the hardware's limit or leaving speed on the table — and no way to tell if another round of optimization is worth it.

A Continuous Optimization Loop

Your models and validation data go in. The agent hypothesizes, transforms, compiles and benchmarks on real hardware — learning each round until it hits your on-device performance targets.

PyTorch
ONNX
+ Validation Data & App Code
(e.g. Pre/Post-Processing)
Hypothesize
Transform
Compile
Benchmark
Learn
RunLocal
NVIDIA
Jetson OrinJetson Thor
Qualcomm
+ Ambarella & TI soon

Backed By

468 Capital
Y Combinator
Ritual Capital

and more

Frequently Asked Questions

Things you might want to know before trying RunLocal