What is an autonomous application platform?

An application declared as data, run by a platform that can observe it, test it and redeploy it. What that means concretely, and where the autonomy actually comes from.

A system can be defined as a set of elements standing in interrelations.
Ludwig von Bertalanffy, General System Theory (1968)

"Autonomous" is a word that has been spent badly. It usually means an agent with a large toolbox and a long prompt, allowed to keep going until something looks finished. That is autonomy in the sense that a car with no steering wheel is autonomous.

We mean something narrower and, we think, more useful. A Kitsoki application is declared as data — a definition, an ontology, and a set of typed effects — and the platform is what runs it. Because the application is data rather than a codebase wrapped around a process, the platform can do things to it that a runtime normally cannot: read it, reason about it, test it, revise it, redeploy it, and keep a record of having done so. The autonomy is a property of that arrangement, not of a model that was told to be careful.

This post is the introduction. Two companion posts go deeper on where the thing runs — the data plane and the delivery plane.

The application is data

In a conventional stack, "the application" is a distribution of facts across artifacts that do not know about each other. The schema lives in migration files. The workflow lives in controller code and a queue consumer. The permissions live partly in middleware and partly in a policy service. The deployment lives in YAML in a different repository. Nothing can answer a question about the whole, because there is no whole — only parts that happen to be deployed together and a shared belief that they line up.

A drawn map of four great houses - schema, workflow, permissions, deployment - divided by a river, a mountain range and a desert, with one dotted route wandering between them, fording the river and turning back at the pass.

Which is why every change is an expedition. Adding one field to a customer record means a migration, a model, a serializer, a permission rule, a form, a fixture and a release note — separate territories, each with its own dialect, and nothing anywhere that says they are one change. There is no road, so you chart the way yourself: from memory, from grep, from the colleague who did something like it last spring. The next person charts a different route, and neither route is written down, because a route is not an artifact.

An incident is the same journey run in the dark and against the clock. The knowledge that makes it survivable — ah, you also have to bump the cache key — is real skill and the most expensive asset a team can own. It is unwritten, it leaves when people leave, and it cannot be handed to anyone. That last part now includes handing it to an agent.

Kitsoki starts from the other end. An application is a document plus the things it references:

  • a definition — the pages, the state machine, the actions, the data it is allowed to project into a view;
  • an ontology — the typed entities, their relations, and the vocabulary of the domain, versioned and pinned by revision;
  • typed effects — every side effect the application may cause, declared by name, with a schema on its input and a schema on its result.

That last one is where most of the leverage is. An effect is not a function call the application happens to make; it is a capability the application was granted, validated at the boundary, authorized before it runs, and recorded after. The application cannot reach past its declared effects, because there is no ambient runtime under it to reach through.

One pueblo in daylight: definition, ontology and typed effects as three joined wings of the same stepped complex, linked by walkways above and a continuous wall below, with a straight road running past all three.

The practical consequence is that the platform can answer questions about an application without running it. Which effects does this thing perform? Which of them touch a customer's data? What changed between this revision and the last one? Those are lookups against a document, not an expedition across four territories.

The runtime is in charge, not the model

Most LLM systems put the model on top. It holds the plan, the runtime exposes tools, and every decision — which tool, which arguments, which order, whether to ask first — is a moment of model judgment. When something goes wrong the honest answer is usually that nobody knows where.

Kitsoki inverts that. The application's state machine is in charge: it knows every state the process can be in, every intent valid in each state, every transition out, and every effect that fires along it. When it reaches something it cannot resolve deterministically, it calls the model for that specific sub-task, with the maximum relevant context and the smallest reasonable set of tools, and takes the result.

flowchart LR
  Runtime[Application state machine] -->|resolve this sub-task| Model[Model, narrow domain, scoped tools]
  Model -->|named intent or typed payload| Runtime
  Runtime --> Effect[Declared typed effect]
  Runtime --> Trace[Trace and receipt]

The arrow that matters is the return arrow. The model produces a named intent, a typed payload, or a finished artifact. It does not write state, choose the next transition, fire effects, or call hosts — those happen only along edges someone declared in advance.

This is the part a structured-output wrapper cannot reach. Schema validation proves that one response had the right shape. It does not prove that the model was allowed to make only this decision, that the next transition was declared before the call was made, or that the same run can be replayed later with no live model at all.

Determinism is a direction, not a starting position

Nobody writes a correct deterministic workflow first. Useful processes start as a description and a hunch, and the ones that matter are exactly the ones nobody fully understands yet. So the platform is built around a loop rather than a specification.

You begin with as much model latitude as the problem deserves — including, at the extreme, a single state whose whole job is to hand the task to an agent. Even then it runs inside a declared boundary: a write-mode gate, a typed close-out, and a trace. Then you read the trace, and the trace shows you where the model kept making the same judgment. Each one of those is a candidate to be promoted into a deterministic edge.

flowchart LR
  Idea[Idea or description] --> Prove[Prove it end to end, model-heavy]
  Prove --> Inspect[Read the trace for repeated judgment]
  Inspect --> Convert[Promote it to a declared transition or effect]
  Convert --> Better[More deterministic, cheaper, testable]
  Better -.-> Inspect

Each promotion is small and reviewable: a diff against a declared document, a measurable change in trace shape, a measurable drop in cost and latency. A prompt instruction becomes a state with two transitions. A free-form tool call becomes a typed effect with a declared contract. A model asked to judge whether a bug is real becomes a model asked to write the failing test that proves it — strictly better, because a test can be re-run by anyone, forever, and stays behind as a regression guard. The direction reverses too: when a declared path turns out to be wrong, the trace shows that as well, and the surface widens to match.

On a mature application, roughly four turns in five route with no model call at all. That is not a cost optimisation that happened to work out; it is the same mechanism that makes the application testable, viewed from the billing side.

Where the autonomy comes from

Put those together and a few things stop being aspirational.

It can test itself. Because every model call is a bounded sub-task with a typed output, it can be replaced by a recording. Flow tests run the real state machine against recorded model responses and recorded provider responses, deterministically, at zero model cost, and fail on regression. The parts of a system that are usually untestable — "what happens when the provider rate-limits us mid-rollout" — become ordinary fixtures.

It can observe itself. A run is not a pile of log lines; it is a trace of typed events. Which intent was selected, which guard matched, which effect ran, what it returned, what the model received and emitted. That is the raw material the improvement loop above consumes, and it is also the audit record, because they turn out to be the same artifact.

It can deploy itself. This is the one that surprises people, and it gets its own post. Because every resource on our infrastructure is an API call rather than a machine, a deployment is just another application: declared states, typed effects, gates that must be green, and a receipt naming the exact source revision each step ran against. Which means a deployment can be tested like an application and, given the trace, any single step can be replayed in isolation.

And it can be held to account. Autonomy without evidence is just an unsupervised process. Every mutation passes an authorization boundary; every step leaves a typed, durable receipt of what ran, against what, with what result. The question "what exactly happened, and would it still pass?" has an answer rather than an investigation.

What it is not, and what it costs

It is not a general workflow engine. The graph is shaped for processes with a human or an agent in them — typed turns, escapes, mid-flight clarification — not for the data-pipeline shapes that Temporal and Airflow are built around.

It is not free of authorship. Declaring an application is real design work, and the platform is deliberately unhelpful about the alternative: there is no escape hatch where you drop into a service of your own and do it the old way. If a story cannot express something, the capability is missing from the platform, and adding it there is the work. That constraint is the reason any of the above holds, and it is also the thing you will most want to break in week one.

And it is young. We run our own delivery on it, which is the only recommendation worth much at this stage.

The bet underneath all of it is Bertalanffy's, and it is not really a software argument: you cannot understand the behaviour of the whole by studying the parts in isolation. Most stacks are an attempt to do exactly that. This one starts by writing the whole down.

Diagram