Private model hosting
A language model that belongs to you alone. It remembers your work, runs on disposable hardware, and no one else can read it. Not an API key on someone else's model — a locker with your model inside.
$ parity --atol 0 --ctx 8192 PASS output identical to full-memory run $ memory full 17.49 GB locker 5.84 GB -66.6% $ throughput --batch 16 aggregate 58.6 tok/s per locker 3.66 tok/s scaling 13.72x
Unedited output. Method below.
Every prompt you send goes to a model someone else controls, trained on terms you didn't set, changed without warning, and retired when it suits them. Your context lives in their logs. Your work improves their product.
A silent update rewrites behaviour you'd tuned around.
Your locker is pinned. The weights are yours. Nothing changes unless you change it.
Everything you've told it sits in someone else's system.
Sealed storage. Your model and memory live in your locker, encrypted at rest.
Self-hosting has meant buying hardware and babysitting it.
Hardware is disposable. We hold your model in a third of the memory, so a rented card runs it for cents. The machine dies; the locker persists.
The difference is where the wall sits: cryptographic, or physical. Both keep your model and memory yours. One shares a machine; one doesn't share anything.
Your model and memory are yours and sealed. The engine serving them is shared with other lockers, isolated cryptographically rather than by machine.
Your own process on your own card. Nothing else runs beside it. The wall is physical, not cryptographic — for work where that distinction is the whole point.
Speeds are measured, not projected. A locker is built for sustained and background work — it is deliberately not the fastest way to get a token, it is the most private.
Genie is new and has no customers yet, so there is nothing honest to quote. Here is what we measured instead, on rented hardware, with the method stated.
Your compressed model produces output identical to the uncompressed one — not close, identical, to the last decimal, verified at both short and long context.
A model needing 17.5 GB runs in 5.8 GB. Savings shrink as context grows, so we quote the worst case, not the best.
Sixteen lockers on one card deliver 13.72× the throughput of one, each losing only 14%. That result is why Shared costs $39.
A locker is not fast. Around 4 tokens per second is a fraction of what a commodity API gives you, and if speed is what you need, buy speed instead. We have no paying customers and no uptime history. Shared isolates cryptographically, not physically — if your threat model rules that out, Sovereign is the only honest answer. And the hardware underneath is rented, so we depend on suppliers we do not own.
Your locker keeps what you put in it — your documents, your corrections, the way you want things phrased — as data that belongs to your model, not to a chat history on someone's server. It carries forward across every session and every machine.
Your weights and memory are encrypted and only decrypted for your requests. Other lockers on the same machine can't address them. It is the same boundary your bank uses between accounts on one system — strong, but a software boundary. Sovereign removes the shared machine entirely, and we'd rather state the difference plainly than blur it.
Because it fits in a third of the memory. We stream your model through cheap hardware instead of parking it on expensive hardware, which is what makes a private model cost $39 rather than $600. The output is identical — it simply arrives at reading pace.
You export the locker and run it yourself. It's an open-weights model plus your own data in a documented format — no proprietary lock. That's what makes it yours rather than rented.
Yes. USDC on Base, per session or per month, no account required for the pay-per-use tier. Cards work too if you'd rather.
Open a locker in under a minute. Export it whenever you like.
No card required to start · Export anytime