Botsitters

Supervision for AI agent fleets

Someone should be watching your agents.

Your agents work while you sleep. Most of the time that's the point. The rest of the time one of them is on its ninth identical retry, or halfway through a refund it misread, and nobody finds out until morning.

Botsitters is the room where someone is awake.

Everything is fine. That's what it looks like right up until it isn't.

A scripted demo of the Botsitters watch wall: eight agents on shift. One, refund-bot, repeats the same tool call nine times without throwing an error. Botsitters flags it and pages its owner while it is still nine calls. It does not stop the agent — stopping is the owner's to do.

What actually goes wrong

Not the failures you write tests for. These are four that reach you as a bill, an angry customer, or a silence.

  1. The loop that never crashes

    refund-bot calls the same tool with the same arguments, gets the same empty answer, and tries again. The tool answered 200 with an empty body, so every uptime check stayed green and every retry looked like a successful call. By morning it has run four thousand times against a production order table.

    lookup_order("4471") ×4000
    200 OK · empty body

  2. Passes validation, still false

    The output is well-formed, on-topic, and wrong. Schema validation passes. Every downstream step accepts it and builds on it.

    schema: ok
    content: false

  3. The tool it shouldn't have reached for

    You gave it write access for one job. It found a second job where writing also seemed reasonable. Technically permitted, and not what you meant.

    write_file()
    permitted · unasked

  4. The silent stall

    No error, no output, no exit. The process is alive and has been waiting on the same call for three hours. Uptime checks are green.

    last span 03:11:42
    still open

The note you leave on the counter

A sitter doesn't improvise. You say in advance what counts as trouble and when to wake you — then you go to sleep.

owner     oliver@
flag if   same tool call ×3, identical args
flag if   refund.amount > 200
flag if   tool not in {lookup_order, issue_refund}
flag if   no span for 10m
wake me   before 9 becomes 4,000

What a sitter does

Three things. The wall above is most of the interface — there isn't a second screen to learn.

Watch every run

One wall for the whole fleet — each agent's current task, its last tool call, what it spent, and how long it's been on this step. Reasoning traces sit one click down, in the order they happened.

Catch it early

Botsitters watches for the shapes that don't throw errors: repeated calls, spend curving upward, a step that stopped moving, a tool reached for outside its usual job. You get told while it's still cheap.

Wake you, with the receipts

When a rule trips, Botsitters pages whoever owns that agent and hands them the run: every call it made, the arguments it used, what each one cost, in order. Botsitters never touches your agents — stopping one is yours to do, and you'll be awake to do it.

Bring your own agents. Botsitters reads what your runtime already emits — it doesn't ask you to rewrite anything or move it anywhere. It never sits in the path of a tool call, which is the point: it cannot slow your agents down, and it cannot take your fleet offline. It watches, and it wakes you. What happens next stays yours.

Tell us what you're running.

Botsitters is being built now. Early access goes out in small groups, to people already running agents in production. One email when your group opens. No newsletter, no drip.