The AI Platform Map, part 2 of 6

The Work Has to Outlive the Agent

Your agent is not where the work should live.

Slide titled The work has to outlive the agent. Workflow run #1182: intent captured, plan and step 1 done, step 2 agent crashed, a new agent resumed step 2, human approval paused three days. State is saved after every step and the agent is disposable. What to do now: write intent as an artifact, keep state outside the agent, log every step, run the kill-the-agent test.

Most agent setups get this backwards. The task, the context, every decision made along the way, and the only copy of what the human actually asked for all live inside one session. Close the tab, hit a context limit, or let the process crash, and the work is gone. So is the record of why.

That is fine for a demo. It is not fine for anything you have to operate, audit, or hand to someone else.

This is the first two layers of the nine-layer platform map: human intent and the durable workflow.

Intent is the one thing only a human can supply

Not the prompt. The intent. What outcome we want. What done means. Which constraints apply. What must never be true afterward.

That belongs in a written artifact, versioned and owned by a named person, like a ticket written as a contract. Not scattered across a chat history.

The reason is not tidiness. It is cost. Every check downstream is checking the work against something: the tests, the AI reviewer, the human who handles the exception. If the intent was never written down, each of them has to reconstruct it before they can judge anything. That reconstruction is a large part of why AI output is expensive to trust.

Ten minutes of clear intent is the cheapest verification you will ever buy.

The workflow owns the job. The agent is a step in it

Distributed systems solved a version of this years ago. Long-running work does not live in one process. It lives in a durable workflow that records state, knows which steps finished, retries the ones that failed, and waits as long as it needs to.

Agent work needs the same thing, more urgently, because agents fail in more ways than a normal service does. They time out. They wander. They hit rate limits. They get halfway and lose the thread.

A durable workflow means:

The test

Kill your agent halfway through a real task.

Does the work resume where it left off? Start over from nothing? Or disappear, along with any record of what it was doing?

If it disappears, you do not have an AI platform. You have a very fast intern with no notebook.

The agent is the worker. The workflow is the job. Build it so the job outlives every worker you put on it.

The question for the platform owners reading this

Where does the state of your agent work live today?