SoftBuildsSoftBuildsGet in touch
Engineered to connect

Every layer. Working together.

From the interfaces people use to the data platforms underneath. One connected engineering team.

Explore our approach →
AIAI TransformationGenerative AIAgentic AI DevelopmentMultimodal AI & Document IntelligenceContext Engineering & Agent HarnessesMCP IntegrationA2A & Multi-Agent SystemsRAG & Enterprise KnowledgeLLM Evaluation & Observability
Data & AnalyticsData ModernisationAdvanced AnalyticsConnected IntelligenceData Management & GovernanceBusiness Intelligence
CloudCloud Foundations & Landing ZonesCloud Operations & MigrationPlatform EngineeringAI & Data InfrastructureReliability & OperationsCost Engineering & FinOps
SecurityRisk & ComplianceIdentity & AccessCloud Security PostureAI & Agent SecurityData Protection & PrivacyDetection & Response
EngineeringCustom Software DevelopmentAPI & Systems IntegrationAI Product InterfacesDedicated Development TeamsQuality Engineering & TestingERP & CRM EngineeringLegacy Modernisation
DigitalDigital Consulting & StrategyDigital CommerceBusiness ApplicationsEmerging Technologies
OperationsIT Management ConsultingManaged Services & SupportDigital Infrastructure ServicesBusiness Process ServicesStaff Augmentation
Home/AI Refinery
Our productA self-hosted LLM platform, built and maintained by SoftBuilds

AI RefineryPrivate AI infrastructure, productised

Everything we learned building AI systems for clients, packaged as software you run yourself. A curated model catalog, managed inference servers and enterprise governance, with no prompt leaving your network.

Visit airefinery.coRequest a walkthrough
Why we built it

The same three objections, on every engagement

Clients in banking, healthcare and the public sector wanted the capability but could not send data to a public API. Their prompts had to stay inside the network, their model versions had to be pinned, and their auditors needed a record of who ran what.

We were rebuilding the same private platform for each of them. AI Refinery is that platform, hardened into a product our clients operate themselves.

01

Data residency

Inference happens on your hardware, in your cloud, your VPC, or fully air-gapped.

02

Version control

You pin the model and the quantization. Nothing changes underneath your application.

03

Auditability

Role-based access, single sign-on and a full audit log of every request and deployment.

What is inside

Six parts, one platform

Quantization variants, benchmarked
Model catalog

Benchmarked before you commit hardware

Open-source text, embedding and vision models with security scoring, license documents and every quantization variant, each measured for speed, memory and quality first.

Deployment

An inference server in minutes

Pick a model, pick a quantization, launch. No GPU configuration, no bespoke scripts and no standing DevOps overhead to keep it alive.

Integration

One base URL to change

Every deployed model exposes a standard OpenAI-compatible endpoint, so existing clients and SDKs keep working without a migration project.

Governance

What your security team asks for

Role-based access control, single sign-on, model size tiers, approval workflows and an audit log covering every deployment and request.

Routing

The right model per request

With several models deployed, the routing layer reads each request and forwards it to the best fit, whether that is text, code, embeddings or vision.

Observability

Every server, in real time

Request volume, latency, token usage and errors across all inference servers on a single dashboard, so cost and performance stay visible.

One endpoint, every capability

Your applications talk to one address and get the right model behind it

Picking a model per call is friction your team does not need. Point the application at the platform and let the routing layer decide, while you keep the option to pin a specific model where it matters.

Text generationCode completionEmbeddingsVision inferenceImage analysis

Rolling it out

Installed on your own cluster, typically inside an hour
1

Install

Deploy from the operations runbook into your cloud, your VPC or an air-gapped environment. The install is reproducible from the runbook alone.

2

Choose models

Browse the catalog, compare the benchmarks for each quantization, and pull what suits your hardware into your own private storage.

3

Integrate

Launch the inference server and repoint your application at the new base URL. Existing prompts and SDK calls carry over unchanged.

4

Operate

Run it with your own team, or bring in our Managed IT practice to carry the pager while your engineers come up to speed.

Offline installers are published for Linux, macOS and Windows, so an air-gapped rollout needs no outbound access at all.

Who it is for

Organisations that have a genuine reason not to send prompts to a public API: regulated data, residency requirements, a procurement process that asks where inference happens, or a cost profile that no longer suits per-token billing.

If a hosted API is fine for your case, we will tell you that instead. This product is not always the right answer.

Banking & FintechResidency and audit trail
HealthcarePatient data never leaves
Public SectorSovereignty and traceability
High-volume productsPredictable inference cost
AI & Data practice →Industries →

See it running on a workload of your own.

Plans, trials and full product documentation live on the product site. For a walkthrough against your own case, talk to us directly.

Visit airefinery.coRequest a walkthrough