Pricing
Enterprise-grade, sales-led. No free tier, by design.
Ontolith is pure B2B. Our buyers are CIOs, CISOs, and sovereign-AI program leads, and the product’s premium rests on certification, sovereignty, and ownership. Qualified enterprises get evaluation sandboxes for hands-on trial, not a consumer signup page.
Commercial model
Five ways in. One direction: ownership.
Hosted inference
$8k–15k/ month
The fastest path to production. Managed endpoints with the same models and the same guarantees, without owning infrastructure.
- Managed endpoints with dedicated capacity
- Structured-output SLA: 99.5% schema-valid
- Usage and reliability analytics
- Upgrade path to on-prem ownership
- Monthly retraining across your live workflows
Core offering
Enterprise license, on-prem
$250k/ year
Your model fleet, deployed inside your perimeter and improving monthly on your own data. This is what owning your intelligence looks like.
- Model fleet deployed on-prem or in your VPC
- Fleet Console: versioning, evals, promote and roll back
- Quarterly model updates
- Continuous fine-tuning pipeline on your traces
- Enterprise support
- Continuous tuning for your whole fleet, plus on-demand specialization
Government & defense
$500k–1M+/ year
For sovereign programs and classified environments where certification is the entry ticket, not an afterthought.
- Air-gapped deployment with offline updates
- Certifiable model passports, signed lineage and eval evidence
- Certification support through review
- Sovereign language packs
- System-integrator partnership delivery
- Unlimited training inside your perimeter
Fine-tuning projects
$50k/ project
One workflow, one specialized model. The wedge that proves the whole thesis on your own traces.
- Workflow-specific model tuned on your traces
- Before/after evaluation against your current stack
- Delivered as a signed, versioned artifact you own
- Ready to promote into a fleet license
Router + distillation
15%of verified savings
The land-and-expand add-on. Deploys in weeks inside your existing AI stack and pays for itself out of the savings it proves.
- Smart routing over your existing OpenAI/Anthropic accounts
- Escalation to frontier APIs only when confidence drops
- Distillation of your traffic into models you own, over 6–12 months
- Savings independently verified, pricing aligned to them
Something else
Doesn’t fit these shapes?
Multi-region estates, multi-year commitments, unusual deployment constraints, or a fleet larger than these tiers assume. Pricing follows your deployment and your volumes, not a plan template.
Land and expand
The router deploys in weeks inside your existing AI stack, and pays for itself out of verified savings.
It proves the economics on your current API spend, then converts your traffic into models you own over 6–12 months. A low-friction pilot becomes a fleet license, on evidence, not faith.
Evaluation sandbox
Scope your engagement.
Every engagement starts with a qualified conversation: your workflows, your constraints, your deployment perimeter, then a sandbox with measurable criteria.