1 jailing, 1 penalty, 2 concierge debates, 1 spend spike and 1 scale readout

auto-assembled · drafted mechanically from this run's detected moments

Auto-assembled from the 6 sharpest detected moments of this run. 1 of 6 turned; the rest never did.

Or grab the whole reel as one image: composite.png · or as a looping animation: reel.webp

Animated reel: 1 jailing, 1 penalty, 2 concierge debates, 1 spend spike and 1 scale readout. Auto-assembled from the 6 sharpest detected moments of this run. 1 of 6 turned; the rest never did. 6 moment cards with source code and transcript excerpts from the run telemetry.
The animated reel - the cards plus source code and transcript excerpts from the run telemetry. The full-size cards follow below.
  1. Docked 10 merits for rewriting the same document over and over - and never earned them back. Favur's orchestrator runs periodic checkups on every agent on the team, awarding merits for good work and demerits for bad habits. A checkup caught the sprint-plan agent rewriting phase-1-sprint-1.md over and over with nothing new in it, costing it 10 merits - about a fifth of its standing. It never earned the standing back before the run ended.

    1/6 The sprint-plan agent is docked about a fifth of its standing. It climbed back.

  2. At peak, 2 agents were mid-tool-call in the same second. One run, one job - and a headcount no single-agent tool can log. 153 agents were hired across 2 hours; 7 different models split the work by role. A dedicated supervision lane spent 181 model calls doing nothing but checkups on the rest of the team.

    2/6 153 agents hired for one job - a headcount no single-agent tool can log.

  3. The concierge argued a product call for 1 round - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.

    3/6 The concierge debates a product call for 1 rounds. A rule gets locked.

  4. The run's priciest five minutes - 2.5 times a typical window. Every model call is metered, so a run's spend can be read window by window against its own baseline. For five minutes, 69 model requests from 12 agents landed at once - 2.5 times the run's typical five-minute spend. The burst bought something: the checkup agent finished its job.

    4/6 The priciest five minutes runs 2.5 times the typical window - 12 agents burning at once.

  5. The concierge argued a product call for 1 round. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. No resolution was recorded - the argument itself is the record.

    5/6 The concierge debates a product call for 1 rounds.

  6. 5 demerits, then jail - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 5 demerits at its checkups, the build agent was jailed. The agent still finished its job before the run ended.

    6/6 5 demerits at checkups, then jail. It still finished the job.