Now booking Q3 installations

Private AI that never leaves your building.

Arkanpute installs a production AI stack on hardware you own — local models, your data wired in over MCP, and an agent harness your team will actually use. One installation fee, then one fixed monthly cost. Never a per-token bill.

No data egressNo per-token billingHardware you own
ark-01.internal · on-premise · isolated vlan
Node status
llama-3.3-70bserving
mistral-small-24bwarm
GPU memory118 / 192 GB
Egress this month
0 bytes left the building
Harness · finance team
Which of our Q2 supplier contracts auto-renew before October?
Four contracts auto-renew before 1 Oct. Two carry a 60-day notice window that closes this Friday
via mcp://fileshare · mcp://postgres · 0 external calls

Runs the open models your team already trusts

NVIDIAOllamaMeta LlamaMistral AI
100%On-premise
0 BData egress
~3 wksTypical install
FixedMonthly cost

The problem

Cloud AI asks you to choose. We think that’s a false trade.

Most businesses that stall on AI aren’t confused about the value. They’re stuck on where the data goes, what it costs at scale, and who signs off.

Your data stops being yours

Every prompt, contract, and customer record sent to a hosted model becomes someone else’s log line — and sometimes someone else’s training data. On-prem, none of it crosses your firewall.

The bill scales with usage, not value

Per-token pricing punishes exactly the teams who adopt AI fastest. A fixed monthly turns an unbounded variable cost into a line item finance can actually approve.

Compliance keeps saying no

HIPAA, GDPR, CJIS, ITAR, and most enterprise MSAs get dramatically easier to satisfy when inference happens on a machine sitting in your own rack.

What gets installed

One installation. The entire stack.

A GPU box on its own is a science project. We deliver the machine, the models, the connections to your data, and the interface your staff open every morning.

The machine

Spec’d, procured, racked, and burned in. Sized to your real workload rather than a sales tier — and yours outright from day one.

  • — GPU node sized to workload
  • — Rack, power & cooling plan
  • — Burn-in + benchmark report

Local models

Open-weight models served on your metal behind an OpenAI-compatible endpoint, so the tools your team already uses point at your box with a one-line change.

  • — Llama, Mistral, Qwen, DeepSeek
  • — OpenAI-compatible API
  • — Optional fine-tune on your data

MCP data connectors

Model Context Protocol servers wired into the systems you already run, so the model answers from your data instead of guessing at it.

  • — File shares, NAS & S3
  • — Postgres, MySQL, SQL Server
  • — Jira, Confluence, SharePoint

The harness

Chat UI, agent runner, evals, audit logging, and per-team access control — the part most on-prem AI projects skip and then quietly abandon.

  • — SSO + per-team permissions
  • — Full prompt audit log
  • — Eval suite for your tasks

The process

How an install actually runs.

No discovery phase that bills for six months. We size it, build it, wire it into your systems, and then run it.

Step 01

Assessment

We audit your workloads, data sources, and compliance constraints, then size the machine against what you actually need.

Week 0 · fixed fee, credited
Step 02

Build

Hardware procured, racked, and burned in. Models deployed, network isolated, backups and monitoring configured.

Weeks 1–3 · on your site
Step 03

Connect

MCP servers wired into your file shares, databases, and ticketing. Access rules set per team. Staff onboarding sessions run.

Week 3 · with your IT
Step 04

Operate

Monitoring, model upgrades, connector maintenance, and support — all covered by the fixed monthly, forever.

Ongoing · fixed monthly

Pricing

One installation fee. One fixed monthly.

Installation covers hardware, deployment, and connecting your data. The monthly covers keeping it healthy. Nothing is metered.

Workgroup

A single team proving it out on real data.

$18,000installation

then $1,400 / month

  • Single-GPU node · 48 GB VRAM
  • Up to 25 staff seats
  • 3 MCP data connectors
  • Chat UI + full audit logging
  • Business-hours support
Book an assessment

Enterprise

Multi-site, regulated, or high-security estates.

Custominstallation

then Custom monthly

  • Multi-node HA cluster
  • Unlimited seats
  • Unlimited MCP connectors
  • Network-isolated deployments
  • Custom fine-tuning pipeline
  • 24/7 support with on-site SLA
Talk to us

The hardware is yours from day one. The monthly covers monitoring, model upgrades, connector maintenance, and support — never usage. You are never billed per token.

Questions

The things procurement asks first.

Do we actually own the hardware?

Yes. The machine is procured on your behalf during installation and is yours outright. If you ever stop the monthly, the hardware stays with you and we hand over runbooks, configurations, and credentials.

Does it depend on an internet connection?

No. Inference runs entirely on your own machine, so staff keep working when the line goes down. Model and security updates ship on a schedule you control rather than being pushed to you.

Which models can we run?

Any open-weight model your hardware can hold: Llama, Mistral, Qwen, DeepSeek, and their fine-tunes. We can also fine-tune a model on your own corpus as part of the installation.

What exactly is an MCP?

Model Context Protocol — an open standard for connecting models to external systems. Instead of your staff pasting documents into a chat box, the model queries your file shares, databases, and ticketing systems directly, under permissions you define.

What does the monthly cost actually cover?

Monitoring and alerting, model version upgrades, connector maintenance as your source systems change, security patching, and support. It is fixed — it does not move with how heavily you use the machine.

How long until staff are actually using it?

Most installations are in production three to five weeks after a signed assessment, onboarding sessions included. Procurement lead time on GPUs is usually the long pole.

Can we start small and grow later?

Yes. Additional nodes can be added to an existing installation and the monthly moves to the matching tier. Nothing gets re-platformed and no data is migrated off-site.

Get started

Find out what your install would look like.

Tell us what you’d want the model to answer and where that data lives today. We’ll come back with a machine spec, a connector list, and a fixed quote.

hello@arkanpute.com

Straight to the engineers who do the installs.

What happens next

A 30-minute call, then a written assessment with hardware spec, connector plan, and fixed pricing.

We reply within two business days. No newsletter, no drip sequence.