← All posts
HiveOperationsAgentsCost

One morning in Hive: the ops brief, the planning room, and the PR that waits for you

Two days after we wrote up Routines, we moved a real daily workflow into Hive: a verified AWS and GitHub brief at 08:30, a feature-planning room, a weekly dependency bump that opens a PR and waits for a human, and budgets that stop themselves. Here is what ran, what broke, and the fix — a dead model alias was killing rooms quietly.

On 18 September we wrote down what a Routine is: one recurring job, a producer on a cheap model, a checker on a strong one, arbitration only on dispute, a budget that pauses itself. Two days later we did the obvious next thing and moved a real morning — the operator’s own — into Hive, the way a small company would run it. Not a demo. The AWS account is the real account, the GitHub PRs are the real PRs, and the numbers below are what the brief said on 20 September 2026.

This post is the log of that morning: four jobs that ran, one thing that broke, and the fix. The broken part is the point, so we’ve kept it in.

08:30 — the ops brief

The first Routine is a room called “Reveal · Ops sáng” (Reveal is the product the team ships; ops sáng is Vietnamese for morning ops). It runs on a schedule at 08:30 and has three seats.

The producer runs on grok-4.6 — a cheap, fast model — with exactly two tools: aws_status and daily_brief. It writes one file, outputs/OPS.md. The checker does not read the producer’s prose and nod. It re-calls the same two tools itself and verifies every number in the file against what the tools return. A lead seat exists but only speaks when the checker and producer disagree. That is the conditional-arbitration shape from the research: no debate unless there is a dispute.

aws_status is read-only. It pulls the CloudWatch alarm list, the RDS and EC2 inventory, month-to-date cost by service — including the Amazon Bedrock lines, because that is where “hidden” AI spend hides in an AWS bill — and the Budgets API. daily_brief covers GitHub: PRs waiting for the operator’s review, team review requests, the operator’s own open PRs, and assigned issues, across two GitHub accounts.

What the brief said this morning, and what the checker confirmed:

The brief also reports on Hive’s own bill. AI spend in Hive this month is about $134. Of that, $125 is a worst-case estimate for 25 legacy runs that predate per-run pricing — we price unpriced runs at the ceiling rather than at zero, so the number errs high on purpose. The caps are unchanged: $20 per day for the workspace, $5 per run.

The planning room

The second job is not on a schedule. “Reveal · Kế hoạch tính năng” (feature planning) is a standing room with four seats: a Product Lead who leads, a Tech Lead and a Cursor Engineer who do the work, and Claude Opus via the Brain as the checker.

You type /plan. The lead plans, the workers run in parallel waves, the checker reviews, and the room writes outputs/REPORT.md with one row per feature: feature · value · effort · risk · owner · week. It is deliberately a table and not an essay. The value of the room is that the operator reads a table over coffee instead of running three chat sessions and merging them by hand.

This morning’s run: the Tech Lead assessed 8 features (12–18 person-days), the Cursor Engineer marked 6 of 8 as automatable by a Cursor Cloud agent, and Claude’s critique cut the scope to 5 features (about 13 person-days including review), deferring the Lambda and multi-model work. The whole room cost $0.59.

The PR that waits for you

The third job is new today: an Automation of kind “Origin task.” Origin is Hive’s autonomous coding agent — it clones a repo, edits, runs tests, opens a PR. Until this morning you launched it by hand. Now it can be scheduled.

The weekly job runs against ampd-reveal. It checks whether pipecat-ai and google-genai have newer releases, bumps them if so, runs the test suite, and opens a PR. The configuration is three header lines at the top of the prompt:

repo: https://github.com/luonghongthuan/ampd-reveal
base: main
approval: push

approval: push is the whole point. The run does everything up to the push and then stops. Nothing lands without a human clicking approve. We did not build a new gate; Origin already had plan and push approvals. The scheduled task just defaults to the push gate. A dependency PR opened by an agent on a Sunday and waiting for Monday’s click is exactly the size of autonomy a small team should be comfortable with.

Budgets that stop themselves

The fourth piece is the least visible and the one we would sell first. Every layer of this morning’s work has a ceiling:

Hitting any ceiling pauses the thing that hit it and makes no model call. The ledger reserves before a call and settles after it, so the pause happens before the spend, not after the invoice. This is the Paperclip lesson from the research — the agent pauses at 100 % — applied at four heights instead of one.

What broke

Now the honest part.

Yesterday the chief-of-staff rooms started failing with “Model error 502.” Nothing in the logs pointed at a room; the rooms themselves looked fine. The root cause was a fallback. When a seat had no model set, the code picked the first model in the pool list. The first model in the pool list was xp-claude-opus-5, an alias on a prepaid gateway whose upstream now answers upstream_unsupported. Every room that relied on the default was quietly routing to a model that no longer existed, and the 502 was the gateway’s way of saying so.

We fixed it in three moves:

  1. One default_model() function that reads HIVE_DEFAULT_MODEL, falling back to the pool list. It replaced every hardcoded gpt-5.5 fallback in flows.py, rooms.py, evals.py and routes.py. There were more of those than we’d like to admit.
  2. The pool list reordered, grok-4.6 first, because the default should be the cheap model that works, not the expensive one that might.
  3. Dead aliases removed: the xp-* family, gpt-5.5, gemini-2.5-pro (removed upstream), and grok-3-mini-fast (out of credits).

The audit turned up more than the 502:

None of these are bugs in the agents. They are the boring failure modes of running real integrations on real accounts: a token expires, a balance runs out, a vendor removes a model, a gateway route needs a login. The agents did not fail loudly on any of them. That is the lesson.

The lesson

Agents die quietly when a model alias dies. A room doesn’t crash; it returns a 502 and moves on, and if nobody is reading the room it looks like nothing happened. The fix is not better error messages. It is structural: the pool list must be the single source of truth for which models exist, and every fallback must read it. A hardcoded model name anywhere in the code is a future outage with a date we don’t know yet.

This is also why the checker earns its seat. The producer on the ops brief could have been routed to a dead alias too. If it had, the checker — on a different model, re-calling the tools itself — would have refused to sign off on an empty OPS.md, and the lead would have been woken up. Two models on two routes is cheap insurance against one route going dark.

Where this leaves the product

We said on the 18th that we would not build a canvas or a fourth orchestrator, and that the thing to sell is one recurring, verified job: producer on a cheap model, checker on a strong one, arbitration only on dispute, a budget that pauses itself. Two days of running our own morning through it hasn’t changed that. It has made the pitch concrete:

Every morning at 08:30, Hive reads your AWS account and your GitHub, writes a one-page brief, has a second model verify every number, and tells you about the two alarms pointing at a database that doesn’t exist. Once a week it bumps your dependencies and opens a PR that waits for your click. It cannot spend more than you told it to.

That is the job. Everything we fixed today was in service of making it run tomorrow without anyone watching.

Want something like this built for your team?

Get a quote →