Ontolith

Platform

Agent Foundation Models, built for execution.

Most enterprise agents don’t need 400B parameters. They need reliability, privacy, ontology reasoning, and deployment sovereignty. Ontolith models are 3–7B parameters, trained specifically for agentic execution, and shipped with the tooling to own them: deployment, continuous fine-tuning, evaluation, and certification.

Core capabilities

Trained for the work agents actually do.

Agent-native training

Trained on tool-call traces, planning trees, and multi-step execution data, not conversation data. The pretraining objective is the job, not the chat.

Guaranteed structured outputs

Schema-validated JSON and function calls with published reliability rates, 99.5% valid, backed by SLA. Reliability is a contract, not a benchmark claim.

Constrained decoding engine

Validity is enforced at the decoder level. An Ontolith model cannot emit a malformed action, an out-of-schema field, or a call to a tool that doesn’t exist.

Failure-aware planning

Native confidence scores and fallback plans in every output, so your orchestrator knows when to escalate to a human, or to a larger model.

Continuous fine-tuning

A managed pipeline tunes each model monthly on your own traces, with before/after evaluation evidence. Promote or roll back versions with confidence.

Sovereign language packs

First-class performance in your operating languages, including low-resource government languages frontier models treat as an afterthought.

The model family

One fleet. Three builds.

Representative configurations. Every deployed model is specialized to a customer workflow and delivered as a signed, versioned artifact.

Model Parameters Designed for Deployment targets
Ontolith Edge 3B High-volume single workflows: extraction, triage, classification, routing Edge NPUs, industrial hardware, quantized on-prem
Ontolith Core 5B Tool-calling and multi-step workflow execution with ontology grounding On-prem GPU, VPC, air-gapped clusters
Ontolith Apex 7B Complex planning, long-horizon agent workflows, fleet-lead orchestration On-prem GPU, VPC, hosted inference

The front door

Smart router & frontier distillation.

The router sends each request to the smallest capable model and escalates to a frontier API only when confidence drops. You keep your existing OpenAI or Anthropic accounts. Ontolith becomes the front door for 90%+ of traffic.

Behind it, the distillation pipeline turns your frontier-model usage into your own small models over time. Every API call you make today becomes training signal for a model you will own tomorrow.

The gateway OpenAI-compatible endpoint: change one base_url, keep your stack

In the console

Versioned, accountable, self-healing.

Promotion and rollback are one action away and every version carries its evidence. When a model drifts, it diagnoses the cause and proposes its own retraining — you approve.

Model detail Live metrics, version lineage and before/after evaluation evidence
Self-healing fleet Models diagnose their own drift and propose retraining; humans approve

Beyond the models

The expansion set.

Designed to make the category impossible to enter without competing with us on every front.

Agent memory module

A pluggable long-term memory component, episodic and semantic, purpose-built for the models, so fleets retain workflow experience across sessions.

Execution-reliability benchmark

We publish and maintain the industry benchmark for agentic execution reliability. Owning the scoreboard defines the category and sets the metric competitors must meet.

Edge & NPU builds

Optimized builds for edge NPUs and industrial hardware, opening factory-floor, vehicle, and robotics deployments no cloud API can reach.

Evaluation sandbox

See a specialized model beat a frontier model on your workflow.

The evaluation sandbox runs your traces against an Ontolith model and a frontier baseline, with reliability and cost measured side by side.