Platform
Agent Foundation Models, built for execution.
Most enterprise agents don’t need 400B parameters. They need reliability, privacy, ontology reasoning, and deployment sovereignty. Ontolith models are 3–7B parameters, trained specifically for agentic execution, and shipped with the tooling to own them: deployment, continuous fine-tuning, evaluation, and certification.
Core capabilities
Trained for the work agents actually do.
Agent-native training
Trained on tool-call traces, planning trees, and multi-step execution data, not conversation data. The pretraining objective is the job, not the chat.
Guaranteed structured outputs
Schema-validated JSON and function calls with published reliability rates, 99.5% valid, backed by SLA. Reliability is a contract, not a benchmark claim.
Constrained decoding engine
Validity is enforced at the decoder level. An Ontolith model cannot emit a malformed action, an out-of-schema field, or a call to a tool that doesn’t exist.
Failure-aware planning
Native confidence scores and fallback plans in every output, so your orchestrator knows when to escalate to a human, or to a larger model.
Continuous fine-tuning
A managed pipeline tunes each model monthly on your own traces, with before/after evaluation evidence. Promote or roll back versions with confidence.
Sovereign language packs
First-class performance in your operating languages, including low-resource government languages frontier models treat as an afterthought.
The model family
One fleet. Three builds.
Representative configurations. Every deployed model is specialized to a customer workflow and delivered as a signed, versioned artifact.
| Model | Parameters | Designed for | Deployment targets |
|---|---|---|---|
| Ontolith Edge | 3B | High-volume single workflows: extraction, triage, classification, routing | Edge NPUs, industrial hardware, quantized on-prem |
| Ontolith Core | 5B | Tool-calling and multi-step workflow execution with ontology grounding | On-prem GPU, VPC, air-gapped clusters |
| Ontolith Apex | 7B | Complex planning, long-horizon agent workflows, fleet-lead orchestration | On-prem GPU, VPC, hosted inference |
The front door
Smart router & frontier distillation.
The router sends each request to the smallest capable model and escalates to a frontier API only when confidence drops. You keep your existing OpenAI or Anthropic accounts. Ontolith becomes the front door for 90%+ of traffic.
Behind it, the distillation pipeline turns your frontier-model usage into your own small models over time. Every API call you make today becomes training signal for a model you will own tomorrow.
In the console
Versioned, accountable, self-healing.
Promotion and rollback are one action away and every version carries its evidence. When a model drifts, it diagnoses the cause and proposes its own retraining — you approve.
Beyond the models
The expansion set.
Designed to make the category impossible to enter without competing with us on every front.
Agent memory module
A pluggable long-term memory component, episodic and semantic, purpose-built for the models, so fleets retain workflow experience across sessions.
Execution-reliability benchmark
We publish and maintain the industry benchmark for agentic execution reliability. Owning the scoreboard defines the category and sets the metric competitors must meet.
Edge & NPU builds
Optimized builds for edge NPUs and industrial hardware, opening factory-floor, vehicle, and robotics deployments no cloud API can reach.
Evaluation sandbox
See a specialized model beat a frontier model on your workflow.
The evaluation sandbox runs your traces against an Ontolith model and a frontier baseline, with reliability and cost measured side by side.