AI work that can wait should be cheap.
Otium runs your batch AI jobs off-peak, on open models, when GPUs are cheap. Give it a deadline instead of an answer-this-second and the same work costs a lot less. If your job is a multi-stage pipeline, send the whole graph and one deadline covers all of it.
No credit card, no spam — one email when your invite is ready.
or read the case for batch first
Different from the frontier on purpose
Otium isn’t trying to answer you in 200 milliseconds. It is built for the large slice of AI work that only has to be done by a deadline. That work is cheaper to run, and being off the frontier lets us be stricter about what happens to your data.
Cheaper because it can wait
Interactive APIs charge you for sub-second readiness your backlog never uses. Tell us when you need it by, and the work runs when compute is cheap — overnight, off-peak, on capacity that would otherwise sit idle.
Learn more →Nobody can revoke your model
Hosted frontier models get deprecated, restricted, and pulled from general availability — on someone else’s schedule, and increasingly not even the vendor’s call. Otium runs open-weight models. The weights already exist in the world, so nobody can switch them off, and your pipeline doesn’t die on someone else’s decision.
Learn more →Not retained, not trained on
Open models on infrastructure we run. No proprietary vendor sees your data, inputs and outputs expire after delivery, nothing is used for training, and each customer is encrypted under its own key.
Learn more →You can read the code
The API surface, the capacity interface, and the code that encrypts your data are open source today, each with a conformance suite. The privacy and cost claims on this site are things you can check rather than take on trust.
Learn more →We will not take your model away
Building on a hosted frontier model means building on something someone else can deprecate, restrict, or withdraw from general availability — and lately that call isn’t even the vendor’s to make alone. When it happens, the prompts you tuned, the evals you wrote, and the outputs you validated go with it.
Why this keeps happening →- Otium runs open-weight models. The weights already exist in the world — no vendor, and no regulator, can switch them off.
- You choose a size, not a product: small, medium or large. We keep each tier on current open models, so quality moves forward without a migration project on your side.
- Nothing gets deprecated out from under you, so a working pipeline keeps working without a migration you did not plan for.
How it works
Four steps, two of which are yours: send the work, and name a deadline.
Point your client at us
One base-URL change. It’s the OpenAI-compatible batch API you already write against — same client, same request shape, no rewrite and no new SDK.
Tell us when you need it
An hour, a day, a week. That deadline is the price lever: the more slack you can give us, the less the job costs you. If your work is a pipeline rather than a single call, publish it as a workflow and that one deadline covers every stage.
We run it when compute is cheap
Your job sits in durable storage until the scheduler finds cheap capacity, then runs. If a worker is reclaimed mid-job the remaining items re-queue, and you can poll status the whole way.
Results out, inputs expire
You get your output by the deadline. Inputs and outputs expire on a short TTL — we keep token counts and cost, never your content.
A pipeline should cost one deadline, not seven
Chain seven prompts and you normally pay your whole window per step — a seven-day deadline becomes seven weeks, which quietly prices multi-step work out of the cheapest tier. Describe the graph instead, and one window covers all of it.
Authoring, compiling and costing are live today. Running a published workflow lands with the alpha.
- 1 Dependencies are derived from your data, never declared. The graph is whatever the references say it is, so it cannot drift out of step with the prompts.
- 2 You see the plan before you publish: what runs in parallel, the critical path, where the slack is, and a p50/p95 cost estimate with its guesses labelled.
- 3 Compile it in your browser — the same compiler we run, shipped as WebAssembly. No account, no key, nothing to install.
The deadline sets the price
The same job, priced four ways. A wider deadline gives the scheduler more room to wait for cheap capacity, and that room is where the difference comes from.
Figures are illustrative target economics, not a benchmark. We publish real numbers only alongside the code to reproduce them — see open source.
See how pricing worksBuilt for batch workloads
A drop-in API, durable buffering, and a scheduler that treats your deadline as a budget.
Drop-in OpenAI-compatible API
Point your existing client at our base URL. Batch and job semantics you already know, plus one extension for choosing how long you’re willing to wait.
The deadline sets the price
A wider completion window gives the scheduler more room to wait for a price dip, so it costs less. We launch with the widest tier first, where the difference is largest.
Send a pipeline, not a stage at a time
A seven-stage chain normally pays your whole deadline per step — a seven-day window becomes seven weeks. Send the graph instead and one window covers all of it.
You pick a size, not a model name
Small, medium or large, with a current open model kept behind each. Pooling everyone onto a short list is what makes the batches deep enough to be cheap.
Runs on the compute nobody else wants
Your work waits for cheap capacity instead of reserving expensive capacity — spot and idle GPUs, claimed across regions at the moment the price is right.
Nothing retained, nothing trained on
Inputs and outputs expire on a short TTL after delivery, content never reaches a log, and your data is never used to train or fine-tune anything.
Per-customer encryption
Each customer’s data is encrypted under its own key — unreadable to anyone else, and to a compromised machine. Isolation is cryptographic, not just policy.
A spending cap that stops
Set a monthly ceiling and new work holds at the limit instead of billing straight through it. Prepaid balance, optional auto-refill, no surprise invoice.
Cost per job, measured
Cost per job and cost per million tokens are measured and surfaced directly, so you can see what you paid, on which model, and why.
The interfaces and reference code are open
Anyone can promise they don’t keep your data. We publish the code that handles it. The API surface, the capacity interface, and the payload-encryption format are open today, each with a conformance suite, so you can check the privacy and cost claims instead of trusting them. The worker agent and benchmarks are next.
Our open-source strategy- The OpenAI-compatible API surface and a reference server — backed by a conformance suite that drives it with the official OpenAI SDK.
- The capacity-provider interface and a complete AWS Spot reference implementation, with a conformance suite.
- The payload-encryption format — the code that encrypts your data at rest — so the privacy claim is auditable.
- A dependency-free Hugging Face Hub client for model discovery.
Deferrable work flattens the curve
The grid and the cloud both struggle with peaks, not totals. Shifting batch work into the troughs — overnight, off-peak, onto capacity that would otherwise sit idle — means less hardware kept hot for a spike, and less of the dirtiest peaker power burned to serve it. Cheaper for you, and less wasteful.
The case for batch →Otium · oh-tee-um · Latin · noun
Productive leisure — free time put to use rather than wasted.
The interactive AI market prices every token for the rush. We named the company after the other case: work with no reason to hurry, run at the moment compute is cheap. Idle GPUs and a deadline you can stretch are the two things that make it cheaper.
Its opposite, negotium, is “not-leisure”: business, urgency — and the root of “negotiate.”
Stop paying for speed you don’t need.
Otium is in closed alpha. Leave your email and we’ll send an invite when a spot opens — then point your OpenAI-compatible client at our base URL and send your first batch.
No credit card, no spam — one email when your invite is ready.
Closed alpha — onboarding is gated while we calibrate. Already invited? Sign in.