Solar System - an accurate 2D simulation

Run of , staffed by a mixed roster (google + meta + openai). Combines Muse Spark (primary) with minority vendors for cost savings - composite 57.99.

The verdict.

57.99 composite · rubric v15

duration
24 min
requests
11135
failed requests
19 (0.17%)
cost
$573.29

Scores come from the open scoring pipeline - methodology and leaderboards at Favur Evals ↗

Per-subject scores.

subjectscore
code quality 33.98
cost efficiency 38.47
deliverables 83.21
effort efficiency 33.84
process discipline 52.8
test quality 59.93
tool discipline 91.98
velocity 74.39

Who staffed it.

rolemodel
codemeta/muse-spark-1.2
specmeta/muse-spark-1.2
testopenai/gpt-5.6-luna
buildopenai/gpt-5.6-luna
scoutmeta/muse-spark-1.2
developgoogle/gemini-3.6-flash
directormeta/muse-spark-1.2
platformmeta/muse-spark-1.2
architectmeta/muse-spark-1.2
conciergeopenai/gpt-5.6-luna
pseudocodemeta/muse-spark-1.2
code-reviewmeta/muse-spark-1.2
sprint-planmeta/muse-spark-1.2
orchestratorgoogle/gemini-3.6-flash
sprint-reviewmeta/muse-spark-1.2

What shipped.

tests
275/277 passing
source
3741 lines
test code
4426 lines
coverage
79.74%

Detected moments.

  • The same command failed 19 times in 52 minutes - then passed. Agents run their own build and test commands, and in a team of agents the code can change under a command between runs. One agent ran the same command, `poetry run pytest tests/test_camera.py -v`, 19 times over 52 minutes; every run failed. The command never changed - the code underneath it did: 2 commits and 40 file writes landed inside the loop window, and the next run passed.
  • At peak, 4 agents were mid-tool-call in the same second. One run, one job - and a headcount no single-agent tool can log. 1466 agents were hired across 22 hours; 4 different models split the work by role. A dedicated supervision lane spent 2310 model calls doing nothing but checkups on the rest of the team.
  • 7 demerits, then jail - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 7 demerits at its checkups, the sprint-review agent was jailed. It had already been handed a written performance plan and kept slipping anyway. The agent still finished its job, and a same-type replacement had already been hired to keep the job moving.
  • 9 demerits, then jail - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 9 demerits at its checkups, the code agent was jailed. It had already been handed a written performance plan and kept slipping anyway. The agent still finished its job, and a same-type replacement had already been hired to keep the job moving.
  • The same command failed 11 times in 9 minutes - then passed. Agents run their own build and test commands, and in a team of agents the code can change under a command between runs. One agent ran the same command, `pre-commit run --all-files`, 11 times over 9 minutes; every run failed. The command never changed - the code underneath it did: 1 file write landed inside the loop window, and the next run passed.
  • One agent generated 25 images in 7 minutes. An agent can call an image model mid-run when the job needs art it cannot draw itself. The platform agent produced 25 generated images in 7 minutes, one after another. The set landed in the project as files - a coherent batch of assets built to order in one sitting.
  • 3 demerits, then jail - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 3 demerits at its checkups, the sprint-review agent was jailed. The agent still finished its job, and a same-type replacement had already been hired to keep the job moving.
  • One agent generated 22 images in 10 minutes. An agent can call an image model mid-run when the job needs art it cannot draw itself. The code agent produced 22 generated images in 10 minutes, one after another. The set landed in the project as files - a coherent batch of assets built to order in one sitting.
  • The concierge argued a product call for 1 round - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • The concierge argued a product call for 2 rounds - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 2 rounds, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • The concierge argued a product call for 1 round - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • The run's priciest five minutes - 3.6 times a typical window. Every model call is metered, so a run's spend can be read window by window against its own baseline. For five minutes, 116 model requests from 14 agents landed at once - 3.6 times the run's typical five-minute spend. The burst bought something: the pseudocode agent finished its job.
  • A quality gate said no - the same agent passed it 9 minutes later. Steps end at automated quality gates; work that misses the bar goes back to the agent that produced it, however long the redo takes. The develop agent failed its gate with 0 of 1 criteria met. 9 minutes later the same agent brought the work back, and the same gate passed it.
  • Fired for spinning in circles - a replacement hired in seconds. The orchestrator can fire an agent mid-run and hire a same-type replacement without stopping the job. The sprint-review agent was pulled after spinning through the same cycle over and over and was removed. A same-type replacement was hired and the job carried on.
  • Fired for spinning in circles - a replacement hired in seconds. The orchestrator can fire an agent mid-run and hire a same-type replacement without stopping the job. The sprint-plan agent was pulled after spinning through the same cycle over and over and was removed. A same-type replacement was hired and the job carried on.
  • Fired for spinning in circles - a replacement hired in seconds. The orchestrator can fire an agent mid-run and hire a same-type replacement without stopping the job. The code agent was pulled after spinning through the same cycle over and over and was removed. A same-type replacement was hired and the job carried on.
  • 1 failing test, one commit, 140 passing. The suite runs after every change, so a failing run points straight at the commits that follow it. The suite came back with 1 failing test. One code-touching commit landed, from the develop agent - the only change in the window. The suite ran green 45 minutes after the red run - 140 passing, 23 of them new tests landing in the same window.
  • Sprint 1 came back rejected - the rework passed. At the end of each sprint, a sprint-review agent judges the whole sprint's output at once. It sent sprint 1 back - the work did not meet the bar. The team redid the work and the next review signed off; the sprint closed with 3 of its 4 planned tasks complete.
  • The build broke - 80 edits later it came back clean. Nothing counts as done until the build runs clean. The build failed with 1 error. 80 edits to bodies.py and scaling.py later, it came back clean.