Agent Foundation Models, built for execution.

Most enterprise agents don’t need 400B parameters. They need reliability, privacy, ontology reasoning, and deployment sovereignty. Ontolith models are 3-7B parameters, trained specifically for agentic execution, and shipped with the tooling to own them: deployment, continuous fine-tuning, evaluation, and certification.

Trained for the work agents actually do.

Eight capabilities that separate a model built to execute from one built to chat, from the training objective down to the decoder that emits every action.

Agent-native training

Trained on execution traces, not chat.

Guaranteed structured outputs

99.5% schema-valid, backed by SLA.

Constrained decoding engine

Invalid actions are impossible.

Failure-aware planning

Confidence and fallbacks in every output.

Continuous fine-tuning

Monthly tuning on your own traces.

Sovereign language packs

First-class in your operating languages.

From a description to a dataset

No labeled data to start.

Reinforced on verifiable rewards

Reinforced on checks you already trust.

One fleet. Three builds.

Representative configurations. Every deployed model is specialized to a customer workflow and delivered as a signed, versioned artifact.

ModelParametersDesigned forDeployment targets
Ontolith Edge3BHigh-volume single workflows: extraction, triage, classification, routingEdge NPUs, industrial hardware, quantized on-prem
Ontolith Core5BTool-calling and multi-step workflow execution with ontology groundingOn-prem GPU, VPC, air-gapped clusters
Ontolith Apex7BComplex planning, long-horizon agent workflows, fleet-lead orchestrationOn-prem GPU, VPC, hosted inference

Smart router & frontier distillation.

The router sends each request to the smallest capable model and escalates to a frontier API only when confidence drops. You keep your existing OpenAI or Anthropic accounts, and Ontolith becomes the front door for 90%+ of traffic.

The gatewayOpenAI-compatible endpoint, screened inbound and outbound, with a semantic cache

Versioned, accountable, self-healing.

Promotion and rollback are one action away, and every version carries its evidence.

Model detailLive metrics, version lineage, trajectory and golden-set evals, and before/after evidence

Incidents that propose their own fix.

When a model’s structured-output validity drifts below its SLA, the fleet opens an incident, diagnoses the likely cause, and drafts the exact retraining to correct it.

Self-healing fleetModels diagnose their own drift and propose retraining; humans approve

The expansion set.

Designed to make the category impossible to enter without competing with us on every front.

Agent memory module

A pluggable long-term memory component, episodic and semantic, purpose-built for the models, so fleets retain workflow experience across sessions.

Execution-reliability benchmark

We publish and maintain the industry benchmark for agentic execution reliability. Owning the scoreboard defines the category and sets the metric competitors must meet.

Edge & NPU builds

Optimized builds for edge NPUs and industrial hardware, opening factory-floor, vehicle, and robotics deployments no cloud API can reach.

See a specialized model beat a frontier model on your workflow.

The evaluation sandbox runs your traces against an Ontolith model and a frontier baseline, with reliability and cost measured side by side.

Eight capabilities separate a model built to execute from one built to chat. They run from the pretraining objective, through the decoder, to how each model is tuned and reinforced on your own work.

Schema-valid outputs99.5%

Published reliability, backed by SLA, enforced at the decoder.

Agent-native training

Trained on tool-call traces, planning trees, and multi-step execution data, not conversation data. The pretraining objective is the job, not the chat.

Guaranteed structured outputs

Schema-validated JSON and function calls with published reliability rates, 99.5% valid, backed by SLA. Reliability is a contract, not a benchmark claim.

Constrained decoding engine

Validity is enforced at the decoder level. An Ontolith model cannot emit a malformed action, an out-of-schema field, or a call to a tool that doesn’t exist.

Failure-aware planning

Native confidence scores and fallback plans in every output, so your orchestrator knows when to escalate to a human, or to a larger model.

Continuous fine-tuning

A managed pipeline tunes each model monthly on your own traces, with before/after evaluation evidence. Promote or roll back versions with confidence.

Sovereign language packs

First-class performance in your operating languages, including low-resource government languages frontier models treat as an afterthought.

From a description to a dataset

No labeled data to start. Describe the agent, its tools and a couple of examples; the advisor picks the recipe and synthesizes the training set, filling any decision branch your examples miss with reward-verified cases. You bring the intent, not a dataset.

Reinforced on verifiable rewards

Beyond imitation: models are reinforced toward outputs that verifiably pass, schema, your enforced policies, and argument correctness, scored by deterministic checks you already trust. No human preference labels, no reward model to game.

Three representative builds cover the range of enterprise agent work, from high-volume single tasks to long-horizon orchestration. Every deployed model is specialized to a customer workflow and delivered as a signed, versioned artifact.

Ontolith Edge3BEdge / on-prem
Ontolith Core5BAir-gapped
Ontolith Apex7BHosted / VPC

Ontolith Edge, 3B

High-volume single workflows: extraction, triage, classification, routing. Runs on edge NPUs, industrial hardware and quantized on-prem targets.

Ontolith Core, 5B

Tool-calling and multi-step workflow execution with ontology grounding. Runs on on-prem GPU, VPC and air-gapped clusters.

Ontolith Apex, 7B

Complex planning, long-horizon agent workflows and fleet-lead orchestration. Runs on on-prem GPU, VPC and hosted inference.

Delivery
Specialisation
Per customer workflow
Artifact
Signed and versioned
Range
3B to 7B

The router sits in front of your existing providers and becomes the front door for 90%+ of traffic, without you migrating anything.

Requestany agent
Routersmallest capable
Escalateonly on low conf.

Route

Each request goes to the smallest capable model and escalates to a frontier API only when confidence drops. You keep your existing OpenAI or Anthropic accounts.

Distil

The distillation pipeline turns your frontier-model usage into your own small models over time. Every API call you make today becomes training signal for a model you will own tomorrow.

Safety and savings gateway

The same gateway screens every request for prompt-injection on the way in and for data loss, PII and secrets on the way out. Repeat or near-duplicate calls are served from a semantic cache, savings on top of routing, before a GPU ever runs.

Gateway
Endpoint
OpenAI-compatible
Inbound
Prompt-injection screen
Outbound
PII, secrets, data-loss
Cache
Semantic, pre-GPU

Model ownership runs as an operational workflow, not a research project. Promotion and rollback are one action away, and every version carries its evidence.

Promote v12 → v13Evals passed
Version lineagev1 → v13Tracked
RollbackOne action

Evidence on every version

Live metrics, version lineage, trajectory and golden-set evaluations, and before/after evidence sit on each model, so a promotion is a decision with proof, not a guess.

Reversible by design

If a promoted version regresses, rollback returns instantly to a known-good model. Nothing changes in production without a record of why.

Operations
Promote
Validated versions only
Roll back
Instant, known-good
Evidence
Evals + lineage

When a model’s structured-output validity drifts below its SLA, the fleet does not just alert, it drafts the correction and waits for a human to approve it.

Driftbelow SLA
Diagnoselikely cause
Approvehuman, by role

Diagnose and draft

The fleet opens an incident, diagnoses the likely cause of the drift, and drafts the exact retraining needed to correct it.

Human in the loop

A human with the right role approves. No autonomous change reaches production, so self-healing never means self-deploying.

Incident
Trigger
Validity below SLA
Output
Drafted retraining
Gate
Role-based approval

The models are the core, but the surrounding set is designed to make the category impossible to enter without competing with us on every front.

Agent memoryReliability benchmarkEdge & NPU builds

Agent memory module

A pluggable long-term memory component, episodic and semantic, purpose-built for the models, so fleets retain workflow experience across sessions.

Execution-reliability benchmark

We publish and maintain the industry benchmark for agentic execution reliability. Owning the scoreboard defines the category and sets the metric competitors must meet.

Edge and NPU builds

Optimized builds for edge NPUs and industrial hardware, opening factory-floor, vehicle, and robotics deployments no cloud API can reach.

Set
Memory
Episodic + semantic
Benchmark
Industry scoreboard
Edge
NPU + industrial