2048 - the sliding-tile game

Run of , staffed by a mixed roster (xiaomi). Combines Mimo Pro (primary) with minority vendors for cost savings - composite 63.30.

The verdict.

62.35 composite · rubric v15

duration
225 min
requests
20281
failed requests
1059 (5.22%)
cost
$224.63

Scores come from the open scoring pipeline - methodology and leaderboards at Favur Evals ↗

Per-subject scores.

subjectscore
code quality 42.81
cost efficiency 71.52
deliverables 99.69
effort efficiency 53.49
test quality 82.43
tool discipline 30
velocity 34.44
process discipline 57.57

Who staffed it.

rolemodel
codexiaomi/mimo-v2.5-pro
specxiaomi/mimo-v2.5
testxiaomi/mimo-v2.5-pro
buildxiaomi/mimo-v2.5-pro
scoutxiaomi/mimo-v2.5-pro
developxiaomi/mimo-v2.5
directorxiaomi/mimo-v2.5
platformxiaomi/mimo-v2.5-pro
architectxiaomi/mimo-v2.5
conciergexiaomi/mimo-v2.5-pro
pseudocodexiaomi/mimo-v2.5-pro
code-reviewxiaomi/mimo-v2.5
sprint-planxiaomi/mimo-v2.5
orchestratorxiaomi/mimo-v2.5
sprint-reviewxiaomi/mimo-v2.5

What shipped.

tests
422/430 passing
source
3125 lines
test code
8090 lines
coverage
90.82%

Detected moments.

  • Jailed 4 times in one run - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 8 demerits at its checkups, the pseudocode agent was jailed. It had already been handed a written performance plan and kept slipping anyway. The agent still finished its job, and a same-type replacement had already been hired to keep the job moving.
  • At peak, 4 agents were mid-tool-call in the same second. One run, one job - and a headcount no single-agent tool can log. 3611 agents were hired across 94 hours; 3 different models split the work by role. A dedicated supervision lane spent 5008 model calls doing nothing but checkups on the rest of the team.
  • Jailed 3 times in one run - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 10 demerits at its checkups, the sprint-review agent was jailed. It had already been handed a written performance plan and kept slipping anyway. The agent still finished its job before the run ended.
  • 5 agents of one type fired into the same trap - each replaced in minutes. The orchestrator can fire an agent mid-run and hire a same-type replacement without stopping the job. The code agent was assigned work its model could not see - it cannot read images and was removed. A replacement was hired at once - and 5 same-type agents went down the same way inside 44 minutes.
  • A quality gate said no - the same agent passed it 2 hours 11 minutes later. Steps end at automated quality gates; work that misses the bar goes back to the agent that produced it, however long the redo takes. The develop agent failed its gate with 0 of 4 criteria met. 2 hours 11 minutes later the same agent brought the work back, and the same gate passed it.
  • The run's priciest five minutes - 11.6 times a typical window. Every model call is metered, so a run's spend can be read window by window against its own baseline. For five minutes, 58 model requests from 7 agents landed at once - 11.6 times the run's typical five-minute spend. The burst bought something: the checkup agent finished its job.
  • 5 demerits, then jail - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 5 demerits at its checkups, the sprint-plan agent was jailed. The agent still finished its job, and a same-type replacement had already been hired to keep the job moving.
  • 2 agents of one type fired into the same trap - each replaced in minutes. The orchestrator can fire an agent mid-run and hire a same-type replacement without stopping the job. The code agent was assigned work its model could not see - it cannot read images and was removed. A replacement was hired at once - and 2 same-type agents went down the same way inside 9 minutes.
  • Fell 25 merits, then climbed 20 back. Every agent carries a merit standing, marked up and down at checkups as its work lands or slips. The sprint-review agent slid from a standing of 65 down to 40 across a string of demerits. It corrected course and climbed back to 60 before the run ended - no jail, just the fall and the recovery.
  • The agent playtested its own game - 46 key presses, 5 screenshots. An agent can drive a real window: pressing keys, taking screenshots, and reading pixels back to see what happened. It pressed 46 keys into the game window it had just built, pausing to wait 46 times and taking 5 screenshots between moves. After 5 minutes it stopped, having filed 99 verifications of what it saw on screen.
  • One agent generated 13 images in 2 minutes. An agent can call an image model mid-run when the job needs art it cannot draw itself. The spec agent produced 13 generated images in 2 minutes, one after another. The set landed in the project as files - a coherent batch of assets built to order in one sitting.
  • One agent generated 11 images in 3 minutes. An agent can call an image model mid-run when the job needs art it cannot draw itself. The spec agent produced 11 generated images in 3 minutes, one after another. The set landed in the project as files - a coherent batch of assets built to order in one sitting.
  • The concierge argued a product call for 1 round - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • A quality gate said no - the same agent passed it 11 minutes later. Steps end at automated quality gates; work that misses the bar goes back to the agent that produced it, however long the redo takes. The develop agent failed its gate with 1 of 2 criteria met. 11 minutes later the same agent brought the work back, and the same gate passed it.
  • The same command failed 3 times in 4 minutes - then passed. Agents run their own build and test commands, and in a team of agents the code can change under a command between runs. One agent ran the same command, `poetry run mypy src/ --strict`, 3 times over 4 minutes; every run failed. The command never changed - the code underneath it did: 6 file writes landed inside the loop window, and the next run passed.
  • The same command failed 3 times in 5 minutes - then passed. Agents run their own build and test commands, and in a team of agents the code can change under a command between runs. One agent ran the same command, `dist/the2048.exe`, 3 times over 5 minutes; every run failed. The command never changed - the code underneath it did: 1 file write landed inside the loop window, and the next run passed.
  • Caught claiming finished work - docked 10 merits. Agents report their own progress; at every checkup the orchestrator reads the records and the files for itself. The code agent claimed its work was done; the checkup read the record and found otherwise, costing it 10 merits. The demerit stood on the record for the rest of the run.
  • 2 agents of one type fired into the same trap - each replaced in minutes. The orchestrator can fire an agent mid-run and hire a same-type replacement without stopping the job. The code agent was assigned work its model could not see - it cannot read images and was removed. A replacement was hired at once - and 2 same-type agents went down the same way inside 7 minutes.
  • The concierge argued a product call for 3 rounds. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 3 rounds. No resolution was recorded - the argument itself is the record.
  • Caught claiming finished work - docked 5 merits. Agents report their own progress; at every checkup the orchestrator reads the records and the files for itself. The code agent claimed its work was done; the checkup read the record and found otherwise, costing it 5 merits. The demerit stood on the record for the rest of the run.
  • Sprint 2 came back rejected - the rework passed. At the end of each sprint, a sprint-review agent judges the whole sprint's output at once. It sent sprint 2 back - the work did not meet the bar. The team redid the work and the next review signed off; the sprint closed with 3 of its 3 planned tasks complete.
  • Sprint 1 came back rejected - the rework passed. At the end of each sprint, a sprint-review agent judges the whole sprint's output at once. It sent sprint 1 back - the work did not meet the bar. The team redid the work and the next review signed off; the sprint closed with 2 of its 3 planned tasks complete.
  • Sprint 2 came back rejected - the rework passed. At the end of each sprint, a sprint-review agent judges the whole sprint's output at once. It sent sprint 2 back - the work did not meet the bar. The team redid the work and the next review signed off; the sprint closed with 2 of its 0 planned tasks complete.