Why we built it
The same three objections, on every engagement
Clients in banking, healthcare and the public sector wanted the capability but could not send data to a public API. Their prompts had to stay inside the network, their model versions had to be pinned, and their auditors needed a record of who ran what.
We were rebuilding the same private platform for each of them. AI Refinery is that platform, hardened into a product our clients operate themselves.
Data residency
Inference happens on your hardware, in your cloud, your VPC, or fully air-gapped.
Version control
You pin the model and the quantization. Nothing changes underneath your application.
Auditability
Role-based access, single sign-on and a full audit log of every request and deployment.
What is inside
Six parts, one platform
Benchmarked before you commit hardware
Open-source text, embedding and vision models with security scoring, license documents and every quantization variant, each measured for speed, memory and quality first.
An inference server in minutes
Pick a model, pick a quantization, launch. No GPU configuration, no bespoke scripts and no standing DevOps overhead to keep it alive.
One base URL to change
Every deployed model exposes a standard OpenAI-compatible endpoint, so existing clients and SDKs keep working without a migration project.
What your security team asks for
Role-based access control, single sign-on, model size tiers, approval workflows and an audit log covering every deployment and request.
The right model per request
With several models deployed, the routing layer reads each request and forwards it to the best fit, whether that is text, code, embeddings or vision.
Every server, in real time
Request volume, latency, token usage and errors across all inference servers on a single dashboard, so cost and performance stay visible.
Rolling it out
Installed on your own cluster, typically inside an hourInstall
Deploy from the operations runbook into your cloud, your VPC or an air-gapped environment. The install is reproducible from the runbook alone.
Choose models
Browse the catalog, compare the benchmarks for each quantization, and pull what suits your hardware into your own private storage.
Integrate
Launch the inference server and repoint your application at the new base URL. Existing prompts and SDK calls carry over unchanged.
Operate
Run it with your own team, or bring in our Managed IT practice to carry the pager while your engineers come up to speed.
Offline installers are published for Linux, macOS and Windows, so an air-gapped rollout needs no outbound access at all.
Who it is for
Organisations that have a genuine reason not to send prompts to a public API: regulated data, residency requirements, a procurement process that asks where inference happens, or a cost profile that no longer suits per-token billing.
If a hosted API is fine for your case, we will tell you that instead. This product is not always the right answer.
See it running on a workload of your own.
Plans, trials and full product documentation live on the product site. For a walkthrough against your own case, talk to us directly.
