1 penalty, 1 replacement chain, 1 gate redo, 1 network storm, 1 spend spike and 1 scale readout

auto-assembled · drafted mechanically from this run's detected moments

Auto-assembled from the 6 sharpest detected moments of this run. None of them ever turned.

Or grab the whole reel as one image: composite.png · or as a looping animation: reel.webp

Animated reel: 1 penalty, 1 replacement chain, 1 gate redo, 1 network storm, 1 spend spike and 1 scale readout. Auto-assembled from the 6 sharpest detected moments of this run. None of them ever turned. 6 moment cards with source code and transcript excerpts from the run telemetry.
The animated reel - the cards plus source code and transcript excerpts from the run telemetry. The full-size cards follow below.
  1. A provider storm - 185 retries across 62 agents, absorbed unattended. Dozens of agents share the same model providers; when a provider stumbles, every one of them feels it at once. 185 network retries hit 62 agents over 3 hours 36 minutes, all handled by backoff. 1918 model requests still completed inside the storm window. No human touched anything.

    1/6 185 network retries across 62 agents, absorbed with no human in the loop.

  2. A quality gate said no - the same agent passed it 11 minutes later. Steps end at automated quality gates; work that misses the bar goes back to the agent that produced it, however long the redo takes. The develop agent failed its gate with 0 of 7 criteria met. 11 minutes later the same agent brought the work back, and the same gate passed it.

    2/6 A quality gate says no; the same agent redoes the work and the gate passes it.

  3. The run's priciest five minutes - 2.4 times a typical window. Every model call is metered, so a run's spend can be read window by window against its own baseline. For five minutes, 45 model requests from 9 agents landed at once - 2.4 times the run's typical five-minute spend. The burst bought something: the develop agent finished its job.

    3/6 The priciest five minutes runs 2.4 times the typical window - 9 agents burning at once.

  4. Fired for a full context window - a replacement hired in seconds. The orchestrator can fire an agent mid-run and hire a same-type replacement without stopping the job. The code agent ran out of room in its context window and was removed. A same-type replacement was hired and the job carried on.

    4/6 An agent is fired for cause; a same-type replacement takes over within minutes.

  5. At peak, 4 agents were mid-tool-call in the same second. One run, one job - and a headcount no single-agent tool can log. 449 agents were hired across 13 hours. A dedicated supervision lane spent 278 model calls doing nothing but checkups on the rest of the team.

    5/6 449 agents hired for one job - a headcount no single-agent tool can log.

  6. Docked 10 merits for burning time without progress - and never earned them back. Favur's orchestrator runs periodic checkups on every agent on the team, awarding merits for good work and demerits for bad habits. A checkup caught the code agent burning through exchange after exchange without making progress, costing it 10 merits - about a fifth of its standing. It never earned the standing back before the run ended.

    6/6 The code agent is docked about a fifth of its standing for burning time without progress. It never recovered.