1 jailing, 1 gate redo, 1 retry loop, 1 concierge debate, 1 spend spike and 1 scale readout

auto-assembled · drafted mechanically from this run's detected moments

Auto-assembled from the 6 sharpest detected moments of this run. None of them ever turned.

Or grab the whole reel as one image: composite.png · or as a looping animation: reel.webp

Animated reel: 1 jailing, 1 gate redo, 1 retry loop, 1 concierge debate, 1 spend spike and 1 scale readout. Auto-assembled from the 6 sharpest detected moments of this run. None of them ever turned. 6 moment cards with source code and transcript excerpts from the run telemetry.
The animated reel - the cards plus source code and transcript excerpts from the run telemetry. The full-size cards follow below.
  1. The concierge argued a product call for 1 round - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.

    1/6 The concierge debates a product call for 1 rounds. A rule gets locked.

  2. At peak, 4 agents were mid-tool-call in the same second. One run, one job - and a headcount no single-agent tool can log. 2452 agents were hired across 66 hours; 2 different models split the work by role. A dedicated supervision lane spent 3899 model calls doing nothing but checkups on the rest of the team.

    2/6 2452 agents hired for one job - a headcount no single-agent tool can log.

  3. The run's priciest five minutes - 5 times a typical window. Every model call is metered, so a run's spend can be read window by window against its own baseline. For five minutes, 91 model requests from 19 agents landed at once - 5 times the run's typical five-minute spend. The burst bought something: the checkup agent finished its job.

    3/6 The priciest five minutes runs 5 times the typical window - 19 agents burning at once.

  4. Jailed 3 times in one run - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 7 demerits at its checkups, the develop agent was jailed. It had already been handed a written performance plan and kept slipping anyway. The agent still finished its job before the run ended.

    4/6 7 demerits at checkups, then jail. It still finished the job.

  5. The same command failed 12 times in 24 minutes - then passed. Agents run their own build and test commands, and in a team of agents the code can change under a command between runs. One agent ran the exact same command 12 times over 24 minutes; every run failed. The command never changed - the code underneath it did: 7 file writes landed inside the loop window, and the next run passed.

    5/6 The same command fails 12 times; the code changes under it, and the next run passes.

  6. A quality gate said no - the same agent passed it 41 minutes later. Steps end at automated quality gates; work that misses the bar goes back to the agent that produced it, however long the redo takes. The develop agent failed its gate with 11 of 12 criteria met. 41 minutes later the same agent brought the work back, and the same gate passed it.

    6/6 A quality gate says no; the same agent redoes the work and the gate passes it.