Quick Draw - a circle-drawing game

Run of , staffed by a mixed roster (google + x-ai + anthropic + openai + qwen + meta). Combines Gpt Terra (primary) with minority vendors for cost savings - composite 79.03.

The verdict.

79.03 composite · rubric v15

duration
116 min
requests
814
failed requests
0 (0%)
cost
$19.03

Scores come from the open scoring pipeline - methodology and leaderboards at Favur Evals ↗

Per-subject scores.

subjectscore
code quality 74.77
cost efficiency 56.95
deliverables 55.94
effort efficiency 60.79
process discipline 51.18
test quality 50
tool discipline 91
velocity 64.11

Who staffed it.

rolemodel
codeopenai/gpt-5.6-terra
buildmeta/muse-spark-1.1
scoutopenai/gpt-5.6-luna
developgoogle/gemini-3.1-flash-lite
conciergegoogle/gemini-3.1-flash-lite
pseudocodeanthropic/claude-sonnet-4.6
code-reviewqwen/qwen3.7-plus
sprint-planx-ai/grok-4.5
orchestratorgoogle/gemini-3.1-flash-lite
sprint-reviewqwen/qwen3.7-plus

What shipped.

tests
20/20 passing
source
129 lines
test code
150 lines
coverage
52.87%

Detected moments.

  • At peak, 2 agents were mid-tool-call in the same second. One run, one job - and a headcount no single-agent tool can log. 153 agents were hired across 2 hours; 7 different models split the work by role. A dedicated supervision lane spent 181 model calls doing nothing but checkups on the rest of the team.
  • 5 demerits, then jail - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 5 demerits at its checkups, the build agent was jailed. The agent still finished its job before the run ended.
  • The concierge argued a product call for 1 round - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • The run's priciest five minutes - 2.5 times a typical window. Every model call is metered, so a run's spend can be read window by window against its own baseline. For five minutes, 69 model requests from 12 agents landed at once - 2.5 times the run's typical five-minute spend. The burst bought something: the checkup agent finished its job.
  • The concierge argued a product call for 1 round. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. No resolution was recorded - the argument itself is the record.
  • Docked 10 merits for rewriting the same document over and over - and never earned them back. Favur's orchestrator runs periodic checkups on every agent on the team, awarding merits for good work and demerits for bad habits. A checkup caught the sprint-plan agent rewriting phase-1-sprint-1.md over and over with nothing new in it, costing it 10 merits - about a fifth of its standing. It never earned the standing back before the run ended.
  • Docked 10 merits for rewriting the same document over and over - earned back 5 by fixing it. Favur's orchestrator runs periodic checkups on every agent on the team, awarding merits for good work and demerits for bad habits. A checkup caught the code-review agent rewriting phase-1-sprint-1-task-1-review.md over and over with nothing new in it, costing it 10 merits - about a fifth of its standing. The agent corrected course and earned 5 merits back before the run ended.
  • Docked 10 merits for misusing its tools - and never earned them back. Favur's orchestrator runs periodic checkups on every agent on the team, awarding merits for good work and demerits for bad habits. A checkup caught the code-review agent misusing its tools, costing it 10 merits - about a fifth of its standing. It never earned the standing back before the run ended.