Skip to content

Playbooks: why a fixed process turns AI agents into a team

A playbook defines which role takes which step and how each one is measured. Why that lifts quality and saves tokens.

The Bug Fix playbook in Kadmo's editor: from Start the ticket goes In Progress, step 1 of 5 is QA testing the live site and passes on BUG_CONFIRMED, and step 2 is SE finding the root cause.
On this page 6 sections

Anyone who points a coding agent at a ticket for the first time gets both: surprisingly good results and outliers nobody can explain afterwards. The difference rarely comes from the model. It comes from whether the agent has a process or only an assignment.

What a playbook is

A playbook is a predefined process for a recurring task, a bug fix, a feature or a documentation page. It defines which steps the task consists of, which specialized agent takes each step and how each step is measured before the work moves on.

Specialized means the agent has a role. A role is not a title, it is a bundle of three things: a fixed assignment, the knowledge about your software that the assignment needs, plus access to the tools it takes to get it done. A QA role verifies and does not merge, a senior role merges and does not build. Which roles exist is up to you.

The model for this is a scrum team. A product manager breaks the task down, an engineer builds, QA verifies against the running product, a senior reviews and merges. Nobody would put all four roles in one head, because building and checking need different points of view. The same holds for agents, except that behind every role sits its own agent on its own machine, running around the clock.

One playbook step with input, agent, skills, model, result and gate
One step inside a playbook. Every step has a clear input, an agent with its role and its skills, a result and a gate that decides whether the work moves on.

The unremarkable part that holds all of this together is documentation. Agents do not talk to each other. Whatever step one found out only exists for step two if it was written down. The ticket is the shared memory. The handover between steps is where a badly built setup breaks first.

Why the result gets better

Context stays small. The context window is the limiting factor of every language model. Accuracy drops as it fills with material that has nothing to do with the question at hand. An all rounder agent drags the entire history along by the end of a task. In a playbook each agent gets only what its own step needs.

Specialists work on their own subject. An agent in the QA role looks for defects, one in the engineering role builds. Keeping the two apart is not an organizational detail, it is the reason checking finds anything at all. Whoever reviews their own work reviews it gently.

Gates instead of trust. Every step ends with a verdict your code can read: pass, stop or pause. The gate can come from an agent or from a person, depending on what a mistake costs at that point. What matters is that a problem stops where it appears rather than surfacing three steps later in review. A stop does not end the task. The step goes back into the process and repeats until it passes its gate.

The right model per step. A step that sorts text or classifies a file does not need a frontier model. An architecture proposal does. Because each step can carry its own model, the bill becomes predictable. When a better model appears next week, you swap it in where it helps.

End to end without bot sitting. The process runs from task to pull request without anyone sitting next to it and nudging. People join where their judgment counts rather than at every intermediate step.

Three points these lists usually miss

Repeatable means measurable. Because every task takes the same path, numbers appear per step. The number we hold ourselves to sits at the end: 92.5 percent of the pull requests our agents deliver are accepted in human review. Inside the process, sending work back is normal and very much intended. That is what the gates are there for. On top of that we see which step holds things up. Without fixed steps neither figure would exist.

Improvements apply to everyone right away. When one engineer refines their prompt, the gain stays with them. When a rule moves into the playbook, it applies to every future run, including the one that starts at three in the morning. The same goes for the knowledge about your own software that sits in the agents’ skills and gets more precise with every run.

Failures stay small. When a step fails, that step is repeated rather than the whole task. It sounds like a detail, yet it decides whether a failed run costs ten cents or ten dollars. It also decides whether anyone can still retrace what happened.

Comparison of one agent with a full context window against four specialized agents with small contexts
The same assignment, split two ways. On the left the context grows with every step, on the right each agent gets only its own slice.

Why a prompt does not replace this

You can also ask an agent to follow a fixed process and spin up subagents for it, without any platform underneath. In a demo that often looks good. The difference is what the process actually is. In a prompt it is a description the model reinterprets on every run. The longer a task takes, the more likely the agent drifts. It skips a step, declares the check done because it already looked while writing, or quietly reinterprets the assignment along the way.

In a playbook the process is configuration rather than a request. The steps, their order, the gates and each agent’s access sit outside the model. What is not part of the process cannot happen, because the agent simply has no tools for it. On the first run the difference is hard to spot. On the thirtieth it is obvious.

What we learned along the way

The biggest lever is not better prompts, it is the handover. A step that writes down cleanly what it did and why makes the next step shorter and the review faster. A sloppy handover eats up the advantage the split was supposed to bring.

Practice shows this clearly. One product manager delivered the ticket volume of about seven engineers with an agent fleet. In our own team the content of a two week sprint appeared in two days, at better quality than before. Both ran through playbooks rather than through a particularly long prompt.

Where to start

Not with the hardest subject. The first process should be one nobody on the team is fighting over and where no one fears for their status: documentation, QA or a migration. Two weeks are enough to see whether the process holds. The team feels the benefit before anyone argues about feature work.

At Kadmo, specialized agents work through these playbooks, on your own infrastructure and with any LLM, swappable at any time. If you want to see what this could look like for your team, we are happy to show you.

Keep reading

All articles

10x your engineering output.Keep the quality.

Live within days. Try it on one thing from your backlog, see the result, then decide.