1 false claim caught, 3 penalties, 1 spend spike and 1 scale readout

auto-assembled · drafted mechanically from this run's detected moments

Auto-assembled from the 6 sharpest detected moments of this run. 1 of 6 turned; the rest never did.

Or grab the whole reel as one image: composite.png · or as a looping animation: reel.webp

Animated reel: 1 false claim caught, 3 penalties, 1 spend spike and 1 scale readout. Auto-assembled from the 6 sharpest detected moments of this run. 1 of 6 turned; the rest never did. 6 moment cards with source code and transcript excerpts from the run telemetry.
The animated reel - the cards plus source code and transcript excerpts from the run telemetry. The full-size cards follow below.
  1. The run's priciest five minutes - 2.1 times a typical window. Every model call is metered, so a run's spend can be read window by window against its own baseline. For five minutes, 111 model requests from 8 agents landed at once - 2.1 times the run's typical five-minute spend. The burst bought something: the checkup agent finished its job.

    1/6 The priciest five minutes runs 2.1 times the typical window - 8 agents burning at once.

  2. At peak, 2 agents were mid-tool-call in the same second. One run, one job - and a headcount no single-agent tool can log. 57 agents were hired across 1 hours; 2 different models split the work by role. A dedicated supervision lane spent 65 model calls doing nothing but checkups on the rest of the team.

    2/6 57 agents hired for one job - a headcount no single-agent tool can log.

  3. Docked 10 merits for drifting off its task - and never earned them back. Favur's orchestrator runs periodic checkups on every agent on the team, awarding merits for good work and demerits for bad habits. A checkup caught the code agent drifting off the task it was given, costing it 10 merits - about a fifth of its standing. It never earned the standing back before the run ended.

    3/6 The code agent is docked about a fifth of its standing for drifting off its task. It never recovered.

  4. Docked 10 merits for rewriting the same document over and over - earned back 5 by fixing it. Favur's orchestrator runs periodic checkups on every agent on the team, awarding merits for good work and demerits for bad habits. A checkup caught the code agent rewriting test_main.py over and over with nothing new in it, costing it 10 merits - about a fifth of its standing. The agent corrected course and earned 5 merits back before the run ended.

    4/6 The code agent is docked about a fifth of its standing for rewriting the same document over and over. It climbed back.

  5. Caught claiming finished work - docked 20 merits. Agents report their own progress; at every checkup the orchestrator reads the records and the files for itself. The build agent claimed its work was done; the checkup read the record and found otherwise, costing it 20 merits. The demerit stood on the record for the rest of the run.

    5/6 Caught claiming finished work; docked 20 merits at a checkup.

  6. Docked 20 merits for a command that kept failing - and never earned them back. Favur's orchestrator runs periodic checkups on every agent on the team, awarding merits for good work and demerits for bad habits. A checkup caught the build agent pressing on after a command kept failing, costing it 20 merits - over a third of its standing. That dropped it into the warning tier. It never earned the standing back before the run ended.

    6/6 The build agent is docked over a third of its standing for a command that kept failing. It never recovered.