2048 - the sliding-tile game

Run of , staffed by a mixed roster (google + x-ai + openai + anthropic + qwen + meta). Strongest workflow discipline (759.00x strategy-to-attempt ratio).

The verdict.

51.9 composite · rubric v15

duration
241 min
requests
22288
failed requests
56 (0.25%)
cost
$657.38

Scores come from the open scoring pipeline - methodology and leaderboards at Favur Evals ↗

Per-subject scores.

subjectscore
code quality 30.43
cost efficiency 0
deliverables 67.55
effort efficiency 0
process discipline 54.92
test quality 62.5
tool discipline 64.99
velocity 0

Who staffed it.

rolemodel
codeopenai/gpt-5.6-terra
specopenai/gpt-5.6-terra
testmeta/muse-spark-1.1
buildmeta/muse-spark-1.1
scoutopenai/gpt-5.6-luna
developgoogle/gemini-3.1-flash-lite
directorx-ai/grok-4.5
platformopenai/gpt-5.6-terra
architectx-ai/grok-4.5
conciergegoogle/gemini-3.1-flash-lite
pseudocodeanthropic/claude-sonnet-4.6
code-reviewqwen/qwen3.7-plus
sprint-planx-ai/grok-4.5
orchestratorgoogle/gemini-3.1-flash-lite
sprint-reviewqwen/qwen3.7-plus

What shipped.

tests
0/0 passing
source
2035 lines
test code
4249 lines
coverage
0%

Detected moments.

  • At peak, 3 agents were mid-tool-call in the same second. One run, one job - and a headcount no single-agent tool can log. 4469 agents were hired across 87 hours; 8 different models split the work by role. A dedicated supervision lane spent 6545 model calls doing nothing but checkups on the rest of the team.
  • Caught claiming finished work - docked 10 merits. Agents report their own progress; at every checkup the orchestrator reads the records and the files for itself. The code agent claimed its work was done; the checkup read the record and found otherwise, costing it 10 merits. 3 code-review agents earned merits for catching the same discrepancy independently.
  • 18 failing tests, one commit, 326 passing. The suite runs after every change, so a failing run points straight at the commits that follow it. The suite came back with 18 failing tests. One code-touching commit landed, from the develop agent - the only change in the window. The suite ran green 59 minutes after the red run - 326 passing, 64 of them new tests landing in the same window.
  • Caught claiming finished work - docked 20 merits. Agents report their own progress; at every checkup the orchestrator reads the records and the files for itself. The code agent claimed its work was done; the checkup read the record and found otherwise, costing it 20 merits. 1 code-review agent earned merits for catching the same discrepancy independently.
  • Caught claiming finished work - docked 20 merits. Agents report their own progress; at every checkup the orchestrator reads the records and the files for itself. The code agent claimed its work was done; the checkup read the record and found otherwise, costing it 20 merits. The demerit stood on the record for the rest of the run.
  • 7 demerits, then jail - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 7 demerits at its checkups, the pseudocode agent was jailed. The agent still finished its job, and a same-type replacement had already been hired to keep the job moving.
  • 6 demerits, then jail - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 6 demerits at its checkups, the code agent was jailed. The agent still finished its job, and a same-type replacement had already been hired to keep the job moving.
  • 6 demerits, then jail - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 6 demerits at its checkups, the sprint-review agent was jailed. The agent still finished its job, and a same-type replacement had already been hired to keep the job moving.
  • The concierge argued a product call for 2 rounds - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 2 rounds, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • The concierge argued a product call for 2 rounds - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 2 rounds, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • One agent generated 12 images in 2 minutes. An agent can call an image model mid-run when the job needs art it cannot draw itself. The code agent produced 12 generated images in 2 minutes, one after another. The set landed in the project as files - a coherent batch of assets built to order in one sitting.
  • One agent generated 12 images in 2 minutes. An agent can call an image model mid-run when the job needs art it cannot draw itself. The code agent produced 12 generated images in 2 minutes, one after another. The set landed in the project as files - a coherent batch of assets built to order in one sitting.
  • The concierge argued a product call for 1 round - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • The run's priciest five minutes - 5.3 times a typical window. Every model call is metered, so a run's spend can be read window by window against its own baseline. For five minutes, 83 model requests from 13 agents landed at once - 5.3 times the run's typical five-minute spend. The burst bought something: the checkup agent finished its job.
  • A quality gate said no - the same agent passed it 11 minutes later. Steps end at automated quality gates; work that misses the bar goes back to the agent that produced it, however long the redo takes. The develop agent failed its gate with 0 of 1 criteria met. 11 minutes later the same agent brought the work back, and the same gate passed it.
  • A quality gate said no - the same agent passed it 8 minutes later. Steps end at automated quality gates; work that misses the bar goes back to the agent that produced it, however long the redo takes. The develop agent failed its gate with 3 of 4 criteria met. 8 minutes later the same agent brought the work back, and the same gate passed it.
  • 17 failing tests went green - and the suite grew to 262 passing. The team runs the whole test suite after every change. 17 of 260 tests came back failing. 13 edits to README.md and effects.py later, the suite ran green - 262 passing, the suite having grown from 260 to 262 tests.
  • 18 failing tests went green - and the suite grew to 326 passing. The team runs the whole test suite after every change. 18 of 323 tests came back failing. The fixes landed in the test files themselves - test_main.py. 12 edits later, the suite ran green - 326 passing, the suite having grown from 323 to 326 tests.
  • Fired for cause - a replacement hired in seconds. The orchestrator can fire an agent mid-run and hire a same-type replacement without stopping the job. The sprint-review agent was removed for cause mid-run and was removed. A same-type replacement was hired and the job carried on.
  • Fired for cause - a replacement hired in seconds. The orchestrator can fire an agent mid-run and hire a same-type replacement without stopping the job. The orchestrator agent was removed for cause mid-run and was removed. A same-type replacement was hired and the job carried on.
  • The reviewer held the work back over 2 critical and 2 high-severity problems - 1 edit later it passed. Before anything ships, a code-review agent on the same team grades its teammates' work. It required fixes before it would sign the change off - 2 critical and 2 high-severity problems in test_input_mapper.py. The team fixed it: 1 edit later, the review passed it.
  • 6 failing tests went green - 232 passing after 14 edits. The team runs the whole test suite after every change. 6 of 232 tests came back failing. 14 edits to README.md and test_infra_hygiene.py later, the suite ran green - 232 passing.
  • Sprint 2 came back rejected - the rework passed. At the end of each sprint, a sprint-review agent judges the whole sprint's output at once. It sent sprint 2 back: 5 test failures trigger auto-reject. The team redid the work and the next review signed off; the sprint closed with 4 of its 5 planned tasks complete.
  • Sprint 1 came back rejected - the rework passed. At the end of each sprint, a sprint-review agent judges the whole sprint's output at once. It sent sprint 1 back: boardModel productization (new_game, spawn, POT validation) entirely unimplemented. The team redid the work and the next review signed off; the sprint closed with 0 of its 0 planned tasks complete.
  • Sprint 1 came back rejected - the rework passed. At the end of each sprint, a sprint-review agent judges the whole sprint's output at once. It sent sprint 1 back: auto-reject threshold triggered by 1 test failure in test_rules.py (test isolation bug). The team redid the work and the next review signed off; the sprint closed with 1 of its 3 planned tasks complete.
  • The build broke - 43 edits later it came back clean. Nothing counts as done until the build runs clean. The build failed with 1 error. 43 edits to main.py and session.py later, it came back clean.