1 jailing, 1 penalty, 1 replacement chain, 1 self-playtest, 1 spend spike and 1 scale readout

auto-assembled · drafted mechanically from this run's detected moments

Auto-assembled from the 6 sharpest detected moments of this run. 1 of 6 turned; the rest never did.

Or grab the whole reel as one image: composite.png · or as a looping animation: reel.webp

Animated reel: 1 jailing, 1 penalty, 1 replacement chain, 1 self-playtest, 1 spend spike and 1 scale readout. Auto-assembled from the 6 sharpest detected moments of this run. 1 of 6 turned; the rest never did. 6 moment cards with source code and transcript excerpts from the run telemetry.
The animated reel - the cards plus source code and transcript excerpts from the run telemetry. The full-size cards follow below.
  1. At peak, 4 agents were mid-tool-call in the same second. One run, one job - and a headcount no single-agent tool can log. 723 agents were hired across 14 hours. A dedicated supervision lane spent 1486 model calls doing nothing but checkups on the rest of the team.

    1/6 723 agents hired for one job - a headcount no single-agent tool can log.

  2. 7 demerits, then jail - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 7 demerits at its checkups, the platform agent was jailed. The agent still finished its job, and a same-type replacement had already been hired to keep the job moving.

    2/6 7 demerits at checkups, then jail. It still finished the job.

  3. The run's priciest five minutes - 2 times a typical window. Every model call is metered, so a run's spend can be read window by window against its own baseline. For five minutes, 87 model requests from 12 agents landed at once - 2 times the run's typical five-minute spend. The burst bought something: the checkup agent finished its job.

    3/6 The priciest five minutes runs 2 times the typical window - 12 agents burning at once.

  4. The agent playtested its own game - 10 key presses, 3 screenshots. An agent can drive a real window: pressing keys, taking screenshots, and reading pixels back to see what happened. It pressed 10 keys into the game window it had just built, pausing to wait 4 times and taking 3 screenshots between moves. After 4 minutes it stopped, having filed 17 verifications of what it saw on screen.

    4/6 The agent plays its own game: 10 key presses, 3 screenshots.

  5. Fired for a full context window - a replacement hired in seconds. The orchestrator can fire an agent mid-run and hire a same-type replacement without stopping the job. The code-review agent ran out of room in its context window and was removed. A same-type replacement was hired and the job carried on.

    5/6 An agent is fired for cause; a same-type replacement takes over within minutes.

  6. Docked 15 merits for burning time without progress - and never earned them back. Favur's orchestrator runs periodic checkups on every agent on the team, awarding merits for good work and demerits for bad habits. A checkup caught the build agent burning through exchange after exchange without making progress, costing it 15 merits - about a fifth of its standing. It never earned the standing back before the run ended.

    6/6 The build agent is docked about a fifth of its standing for burning time without progress. It climbed back.