Your data stops being yours
Every prompt, contract, and customer record sent to a hosted model becomes someone else’s log line — and sometimes someone else’s training data. On-prem, none of it crosses your firewall.
Arkanpute installs a production AI stack on hardware you own — local models, your data wired in over MCP, and an agent harness your team will actually use. One installation fee, then one fixed monthly cost. Never a per-token bill.
Runs the open models your team already trusts
The problem
Most businesses that stall on AI aren’t confused about the value. They’re stuck on where the data goes, what it costs at scale, and who signs off.
Every prompt, contract, and customer record sent to a hosted model becomes someone else’s log line — and sometimes someone else’s training data. On-prem, none of it crosses your firewall.
Per-token pricing punishes exactly the teams who adopt AI fastest. A fixed monthly turns an unbounded variable cost into a line item finance can actually approve.
HIPAA, GDPR, CJIS, ITAR, and most enterprise MSAs get dramatically easier to satisfy when inference happens on a machine sitting in your own rack.
What gets installed
A GPU box on its own is a science project. We deliver the machine, the models, the connections to your data, and the interface your staff open every morning.
Spec’d, procured, racked, and burned in. Sized to your real workload rather than a sales tier — and yours outright from day one.
Open-weight models served on your metal behind an OpenAI-compatible endpoint, so the tools your team already uses point at your box with a one-line change.
Model Context Protocol servers wired into the systems you already run, so the model answers from your data instead of guessing at it.
Chat UI, agent runner, evals, audit logging, and per-team access control — the part most on-prem AI projects skip and then quietly abandon.
The process
No discovery phase that bills for six months. We size it, build it, wire it into your systems, and then run it.
We audit your workloads, data sources, and compliance constraints, then size the machine against what you actually need.
Week 0 · fixed fee, creditedHardware procured, racked, and burned in. Models deployed, network isolated, backups and monitoring configured.
Weeks 1–3 · on your siteMCP servers wired into your file shares, databases, and ticketing. Access rules set per team. Staff onboarding sessions run.
Week 3 · with your ITMonitoring, model upgrades, connector maintenance, and support — all covered by the fixed monthly, forever.
Ongoing · fixed monthlyPricing
Installation covers hardware, deployment, and connecting your data. The monthly covers keeping it healthy. Nothing is metered.
A single team proving it out on real data.
then $1,400 / month
Company-wide rollout across departments.
then $3,200 / month
Multi-site, regulated, or high-security estates.
then Custom monthly
The hardware is yours from day one. The monthly covers monitoring, model upgrades, connector maintenance, and support — never usage. You are never billed per token.
Questions
Yes. The machine is procured on your behalf during installation and is yours outright. If you ever stop the monthly, the hardware stays with you and we hand over runbooks, configurations, and credentials.
No. Inference runs entirely on your own machine, so staff keep working when the line goes down. Model and security updates ship on a schedule you control rather than being pushed to you.
Any open-weight model your hardware can hold: Llama, Mistral, Qwen, DeepSeek, and their fine-tunes. We can also fine-tune a model on your own corpus as part of the installation.
Model Context Protocol — an open standard for connecting models to external systems. Instead of your staff pasting documents into a chat box, the model queries your file shares, databases, and ticketing systems directly, under permissions you define.
Monitoring and alerting, model version upgrades, connector maintenance as your source systems change, security patching, and support. It is fixed — it does not move with how heavily you use the machine.
Most installations are in production three to five weeks after a signed assessment, onboarding sessions included. Procurement lead time on GPUs is usually the long pole.
Yes. Additional nodes can be added to an existing installation and the monthly moves to the matching tier. Nothing gets re-platformed and no data is migrated off-site.
Get started
Tell us what you’d want the model to answer and where that data lives today. We’ll come back with a machine spec, a connector list, and a fixed quote.
Straight to the engineers who do the installs.
A 30-minute call, then a written assessment with hardware spec, connector plan, and fixed pricing.