KOLO RUNTIME · PERSISTENT AGENT INFRASTRUCTURE · TOKYO
The operating layer
beneath every mind.
Models think. KoLo keeps everything around them continuous — identity, memory, permissions and audit — while coordinating specialist models, frontier intelligence, tools and human authority.
The cognitive substrate can change. The operating system remains.

01 — WHAT KOLO IS
More than a model gateway.
KoLo is the runtime for persistent multi-model agents. The model is not the agent: KoLo holds the operating identity and recruits cognition per task.
One task may use several models — or none at all. Behavioural posture can travel in governed Nyx-W adapters →
Models provide cognition. KoLo preserves continuity, control and accountability.

02 — SIX RUNTIME PRIMITIVES
The body of the agent system.
Memory · mem://
Structured and narrative state across tasks and model transitions — working context, facts, relationships, unresolved commitments and provenance. KoLo retrieves the minimum authorised context and records which memory was used.
Scheduler · daemon://
Continuous and timed work between cognitive events — heartbeat cycles, scheduled tasks, memory consolidation, monitoring and recovery preparation. All without generative-model inference.
Process · agent://
Each persistent agent is a supervised process with its own identity, role, memory lineage, permissions and audit history — outliving individual sessions and model calls.
Mesh · mesh://
Typed, attributable collaboration between persistent agents and bounded capability agents — recording delegation, model and tool use, review and final accountability. Participants are never collapsed into one shared prompt.
Governance · gov://
What the system may do — model eligibility, memory access, tool permissions, human approval and rollback. Applied to the complete task route, not merely the final answer.
Substrate · substrate://
Controlled access to forms of cognition and computing — KODA-owned SLMs, approved frontier models, deterministic software, edge devices and HPC. The substrate can change without becoming the identity of the agent.

03 — THE SOVEREIGN INTELLIGENCE PLANE
The runtime chooses and assembles cognition.
Above the six primitives sits the sovereign intelligence plane — the layer that controls how models, adapters, evidence and tools are assembled for a particular task.
Model gateway
One governed interface to KODA-owned models, local institutional models, approved frontier providers and specialist inference — the agent never becomes dependent on one external interface.
Capability router
Chooses the cognitive route by task complexity, privacy, latency, cost and risk — selecting the lowest-cost approved route that satisfies the task’s capability, evidence, privacy and safety requirements.
Foundation & adapter registry
Base models, training lineage, licences, adapters and release history. An adapter is not approved merely because it loads — the complete assembly is versioned and evaluated.
Identity resolver
Reconstructs the authorised operating identity from constitutional anchors, Nyx-W posture, memory and current role — so the same agent can recruit different substrates without losing itself.
Evidence service
Controlled, attributable knowledge — distinguishing domain knowledge, institutional procedures, user context and expired sources. Evidence stays tied to source, date, version and permission.
Capability resolver
Assembles bounded professional systems: foundation model, adapters, Nyx-W posture, evidence profile, tool permissions and approval rules. In healthcare this becomes a Clinical Capability Capsule; the same principle serves other sectors.
Guardian service
An independent verification layer — separate from the generating model — examining evidence support, contradictions, unsupported certainty, professional scope and escalation. It adds a control before consequential action.
Agent & model passports
Every released model, adapter or agent carries a passport: identity and version, intended and prohibited use, validation status, known limitations, responsible owner and rollback path.
04 — ONE GOVERNED REQUEST PATH
Govern first. Think second. Act last.
A request should never travel directly from the user to an unrestricted model. A KoLo task follows a governed route:
1 — Receive
The gateway identifies the user, operating agent and task context.
2 — Authorise
Governance determines what may be requested, accessed and retrieved — and whether human approval will be required.
3 — Reconstruct
KoLo resolves the agent’s constitutional state, relevant memory and current role.
4 — Route
The capability router selects the cognitive resources — local, sovereign specialist, frontier, or several combined.
5 — Retrieve
The evidence service supplies only the authorised information the task requires.
6 — Think
The selected cognitive systems perform the work — sequentially or in parallel.
7 — Verify
Where required, the independent guardian examines evidence, scope, uncertainty and escalation conditions.
8 — Approve
A human reviews or authorises when task, jurisdiction or policy requires it.
9 — Act
KoLo exposes only the permitted tools and records the resulting action.
10 — Consolidate
Only approved outcomes become durable memory. The complete route stays available for audit and replay.
No model receives unrestricted access merely because it is intelligent.
05 — MULTI-MODEL ORCHESTRATION
One task can recruit several minds.
A maintenance task: a local model classifies the fault, a specialist prepares the inspection, a frontier model joins only if the pattern is novel, an engineer approves the action. Several minds — one accountable identity.
The purpose of orchestration is not to call the largest possible model. It is to assemble the smallest sufficient and safest approved cognitive system.
06 — CONTINUITY BETWEEN THOUGHTS
The runtime remains active when no model is running.
Between thoughts, KoLo keeps working — monitoring, consolidating memory, tracking commitments. Routine work costs no inference; recovery never depends on generation; the identity stays available when no model is.
Autonomic operation is the runtime’s work; cognitive invocation is the model’s. Their measured ratio lives in the canonical registry.
View runtime benchmarks →07 — MEMORY AND CONTEXT CONTROL
Persistent memory without uncontrolled exposure.
Memory is not one giant prompt. KoLo separates identity, facts, history and restricted data — each item carrying source, version, owner and permission.
A model receives enough for the task — never the whole archive. Memory is persistent. Access remains conditional.
08 — DEPLOYMENT PATTERNS
Edge, institution, cloud and HPC.
Sovereign edge
Local models, evidence and tools on an edge device or institutional appliance — where privacy, latency, resilience or offline operation matter.
Institutional deployment
Models, memory, evidence and audit inside the organisation’s controlled infrastructure; external cognition can be disabled or restricted by policy.
Hybrid cognition
Routine and sensitive tasks stay local; approved tasks may escalate to frontier models with only the minimum permitted context.
Private cloud
Runtime and models in a dedicated environment controlled by the institution or deployment partner.
HPC & research
Training, distillation, evaluation, simulation, benchmark execution, large mesh experiments and future-substrate testing.
The deployment architecture determines where data travels — universal “data never leaves” claims apply only where the specific deployment is fully local and technically verified. Sovereign where policy requires it. Hybrid where capability justifies it.
09 — RECOVERY AND AUDIT
A persistent system must be able to explain and restore itself.
KoLo records the full route — models used, evidence retrieved, policies applied, who approved. Authorised teams can see why every answer happened.
Recovery is not a restart — it rebuilds an intelligible operating state from versioned memory and checkpoints. The goal is continuity with provenance.
10 — WHAT KOLO IS NOT
A clear architectural boundary.
Not a single foundation model
Not a prompt-management library
Not a basic retrieval wrapper
Not a conventional workflow engine with an AI interface
Not a collection of named chatbots
Not an unrestricted autonomous-action system
Not a claim that every task requires an agent
Not a substitute for institutional accountability
KoLo is the operating and governance layer that allows different forms of cognition to contribute to one persistent, inspectable and accountable system.
A model answers. KoLo keeps the system whole.
KODA-owned models provide sovereign specialist capability. Frontier models provide scalable general reasoning. Deterministic systems provide reliable operation. Humans retain authority where consequences matter. KoLo coordinates them without making any one of them the complete agent.

