
Here’s where Blackbox is headed.
You give us three things: your workload, the performance you need, and the quality bar the model has to maintain. A team of agents then tunes the inference engine continuously until it reaches your target.
Most inference deployments eventually arrive at the same problem. The model is fixed. The hardware is fixed. What remains is the engine: batching, scheduling, caching, speculative decoding, precision, and the kernels underneath it all.

The right configuration depends heavily on the workload. Settings that perform well for one traffic pattern can underperform badly on another. Today, finding the right combination often means engineers running experiments manually for days or weeks.
We think that work belongs to agents.
Three inputs
A workload profile. A dataset that reflects your real traffic: the kinds of requests you send, their lengths, concurrency patterns, and arrival distribution. The agents optimize for your workload, not a generic benchmark.
Target performance metrics. The numbers that matter to you: throughput, time to first token, latency, cost, or a combination of them.
Quality benchmarks. The evals the model must continue to pass. Every optimization is measured against the original model, so speed never comes at the expense of quality.

One loop until the target is reached
Once the inputs are defined, the agents take over:
1. Propose. An agent suggests a change to the engine and records a hypothesis for why it should help.
2. Measure. A fixed benchmark harness deploys the change, replays your workload, and scores it against the performance targets and quality benchmarks.
3. Keep or discard. If the change clearly beats the current best, it becomes the new baseline. If it does not, it gets rolled back.
4. Share. Results are written to a shared log so every agent can build on what the others have already learned.
Several agents run this loop in parallel. They collaborate rather than repeat the same work, starting with configuration-level changes and moving deeper into the engine as the obvious wins disappear.
The loop continues until the target is reached.
The judge is never an agent
Agents can propose changes, but they never grade their own work.
Scoring is fixed. Agents submit changes to a harness they cannot modify.
Wins must reproduce. A new configuration has to beat the current best when both are measured again.
Quality is non-negotiable. A faster result that fails the quality benchmarks does not count.
You stay in control. Deeper engine changes still go through human review before they reach production.

What’s next
This is the direction we’re building toward at Blackbox.
You bring your workload, your targets, and your quality bar. The agents handle the experimentation, you review what they discover, and the winning configuration runs on your dedicated deployment.
Ready to serve your first token?
Tell us the workload and the controls it has to satisfy. We come back with a deployment plan and a per-token commit.