Evolution Arena - an evolutionary simulation

Run of , staffed by a mixed roster (google + meta + openai). Combines Muse Spark (primary) with minority vendors for cost savings - composite 69.99.

The verdict.

69.99 composite · rubric v15

duration
193 min
requests
13838
failed requests
96 (0.69%)
cost
$724.01

Scores come from the open scoring pipeline - methodology and leaderboards at Favur Evals ↗

Per-subject scores.

subjectscore
code quality 34.62
cost efficiency 71.81
deliverables 87.16
test quality 61.48
tool discipline 84.34
velocity 76.82
effort efficiency 52.18
process discipline 54.98

Who staffed it.

rolemodel
codemeta/muse-spark-1.2
specmeta/muse-spark-1.2
testopenai/gpt-5.6-luna
buildopenai/gpt-5.6-luna
scoutmeta/muse-spark-1.2
developgoogle/gemini-3.6-flash
directormeta/muse-spark-1.2
architectmeta/muse-spark-1.2
conciergeopenai/gpt-5.6-luna
pseudocodemeta/muse-spark-1.2
code-reviewmeta/muse-spark-1.2
sprint-planmeta/muse-spark-1.2
orchestratorgoogle/gemini-3.6-flash
sprint-reviewmeta/muse-spark-1.2

What shipped.

tests
535/536 passing
source
7011 lines
test code
7217 lines
coverage
78.56%

Detected moments.

  • At peak, 5 agents were mid-tool-call in the same second. One run, one job - and a headcount no single-agent tool can log. 1861 agents were hired across 28 hours; 3 different models split the work by role. A dedicated supervision lane spent 2859 model calls doing nothing but checkups on the rest of the team.
  • 7 demerits, then jail - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 7 demerits at its checkups, the sprint-review agent was jailed. The agent still finished its job, and a same-type replacement had already been hired to keep the job moving.
  • 5 demerits, then jail - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 5 demerits at its checkups, the code agent was jailed. The agent still finished its job, and a same-type replacement had already been hired to keep the job moving.
  • The concierge argued a product call for 1 round - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • The same command failed 3 times in 15 minutes - then passed. Agents run their own build and test commands, and in a team of agents the code can change under a command between runs. One agent ran the same command, `poetry run python benchmark_calibration.py`, 3 times over 15 minutes; every run failed. The command never changed - the code underneath it did: 1 commit and 3 file writes landed inside the loop window, and the next run passed.
  • 2 failing tests, one commit, 296 passing. The suite runs after every change, so a failing run points straight at the commits that follow it. The suite came back with 2 failing tests. One code-touching commit landed, from the develop agent - the only change in the window. The suite ran green 84 minutes after the red run - 296 passing, 252 of them new tests landing in the same window.
  • The concierge argued a product call for 1 round - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • The concierge argued a product call for 1 round - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • The run's priciest five minutes - 2.8 times a typical window. Every model call is metered, so a run's spend can be read window by window against its own baseline. For five minutes, 103 model requests from 13 agents landed at once - 2.8 times the run's typical five-minute spend. The burst bought something: the checkup agent finished its job.
  • A quality gate said no - the same agent passed it 13 minutes later. Steps end at automated quality gates; work that misses the bar goes back to the agent that produced it, however long the redo takes. The develop agent failed its gate with 3 of 7 criteria met. 13 minutes later the same agent brought the work back, and the same gate passed it.
  • The same command failed 4 times in 6 minutes - then passed. Agents run their own build and test commands, and in a team of agents the code can change under a command between runs. One agent ran the same command, `poetry run pytest tests/test_variation.py -v`, 4 times over 6 minutes; every run failed. The command never changed - the code underneath it did: 3 file writes landed inside the loop window, and the next run passed.
  • A quality gate said no - the same agent passed it 6 minutes later. Steps end at automated quality gates; work that misses the bar goes back to the agent that produced it, however long the redo takes. The develop agent failed its gate with 0 of 1 criteria met. 6 minutes later the same agent brought the work back, and the same gate passed it.
  • Fired for spinning in circles - a replacement hired in seconds. The orchestrator can fire an agent mid-run and hire a same-type replacement without stopping the job. The code agent was pulled after spinning through the same cycle over and over and was removed. A same-type replacement was hired and the job carried on.
  • Fired for spinning in circles - a replacement hired in seconds. The orchestrator can fire an agent mid-run and hire a same-type replacement without stopping the job. The code agent was pulled after spinning through the same cycle over and over and was removed. A same-type replacement was hired and the job carried on.
  • Caught claiming finished work - docked 5 merits. Agents report their own progress; at every checkup the orchestrator reads the records and the files for itself. The code-review agent claimed its work was done; the checkup read the record and found otherwise, costing it 5 merits. The demerit stood on the record for the rest of the run.
  • Caught claiming finished work - docked 5 merits. Agents report their own progress; at every checkup the orchestrator reads the records and the files for itself. The code agent claimed its work was done; the checkup read the record and found otherwise, costing it 5 merits. The demerit stood on the record for the rest of the run.
  • The reviewer held the work back over 1 critical and 2 high-severity problems - 3 edits later it passed. Before anything ships, a code-review agent on the same team grades its teammates' work. It required fixes before it would sign the change off - 1 critical and 2 high-severity problems in simulation.py. The team fixed it: 3 edits later, the review passed it.
  • Sprint 1 came back rejected - the rework passed. At the end of each sprint, a sprint-review agent judges the whole sprint's output at once. It sent sprint 1 back - the work did not meet the bar. The team redid the work and the next review signed off; the sprint closed with 2 of its 4 planned tasks complete.
  • Sprint 3 came back rejected - and never came back for approval. At the end of each sprint, a sprint-review agent judges the whole sprint's output at once. It sent sprint 3 back - the work did not meet the bar. No later approval is on record for that sprint.
  • Sprint 1 came back rejected - the rework passed. At the end of each sprint, a sprint-review agent judges the whole sprint's output at once. It sent sprint 1 back: 1 CRITICAL + 2 HIGH code findings with 2 failed acceptance criteria. The team redid the work and the next review signed off; the sprint closed with 1 of its 6 planned tasks complete.
  • 2 failing tests went green - 296 passing after 44 edits. The team runs the whole test suite after every change. 2 of 296 tests came back failing. 44 edits to evolution.py later, the suite ran green - 296 passing.
  • The build broke - 88 edits later it came back clean. Nothing counts as done until the build runs clean. The build failed with 1 error. 88 edits to __init__.py and ast_nodes.py later, it came back clean.