Your agent is running.
Is it still answering correctly?
A workflow stops when something goes wrong. An AI agent simply keeps answering β even after a fallback kicked in, a model changed or a prompt quietly lost its effect. That is why agent operations does not just watch whether it runs, but checks monthly whether it still does the right thing.
An agent does not break. It drifts.
Four building blocks in operation.
What runs continuously to keep an agent reliable.
Availability & errors
Availability, error rates and aborted runs are monitored continuously. If the agent goes down, you hear it from me β not from your customers.
Monthly quality sample
A fixed set of test cases runs against the live agent every month. It catches exactly what no error log shows: convincingly worded wrong answers. From the "Operations" tier.
Models & deprecations
Providers deprecate models and change answer behaviour between versions. I track the announcements and migrate in good time instead of on shutdown day. From "Operations" the migration is included within the hours budget.
Cost watch
Token and API cost runs against a cap agreed up front. If it is about to be exceeded, you get a heads-up beforehand β not later on the invoice.
Three tiers β cancel monthly.
Bookable independently of the delivery retainers, including for agents someone else built. No minimum term.
Notice when something fails or gets expensive.
Prove monthly that the agent still answers correctly.
For deployments where someone wants oversight on paper.
What is in β and what is not.
What counts as one agent
One language-model-driven use case, one channel. A WhatsApp support assistant and an automated invoice-processing pipeline are two agents β even if both use the same model. The same use case additionally on Telegram is another channel, and therefore another agent.
Not included
- Token and API cost from the model providers β passed through, with a cap agreed up front and a notice before it is exceeded
- Building new agents (quoted as fixed-price work)
- On-call outside working days
- Changes to the business process itself
From handover to monthly report.
Including agents that someone else built.
Handover
We walk through the existing agent: channels, models, interfaces, cost frame β and where in the process it actually decides anything.
Define test cases
Together we write down what a correct answer is β with an expected result per case. From then on this set is the yardstick measured against every month.
Ongoing operation
Monitoring runs, the cost cap is in place, model announcements are tracked. On an outage I respond within the agreed window on working days.
Monthly report
What ran, what stood out, what was re-tuned β from "Operations" with the quality sample result, from "Operations & Evidence" with a documented model and version state.
Frequently asked.
Scope, limits and what is explicitly not promised.
What counts as "one agent"?
Is there an SLA or 24/7 on-call?
Are token and API costs included?
Does tier 3 make our company AI Act compliant?
Do you also maintain agents built by someone else?
How is this different from the "Automation" retainer?
More on governance and obligations.
What the EU AI Act asks of deployers β and how agents stay manageable day to day.
Is your agent working β or just guessing well?
In a free intro call we go through your agent: where it stands today, what a test-case set would need to cover and which tier fits.




