2048 - the sliding-tile game

Run of , staffed by a mixed roster (openai). Combines Gpt Terra (primary) with minority vendors for cost savings - composite 52.54.

The verdict.

52.16 composite · rubric v15

duration
115 min
requests
18282
failed requests
36 (0.2%)
cost
$885.00

Scores come from the open scoring pipeline - methodology and leaderboards at Favur Evals ↗

Per-subject scores.

subjectscore
code quality 42.71
cost efficiency 38.49
deliverables 68.56
process discipline 67.06
effort efficiency 46.82
test quality 79
tool discipline 90.38
velocity 62.5

Who staffed it.

rolemodel
codeopenai/gpt-5.6-terra
specopenai/gpt-5.6-terra
testopenai/gpt-5.6-luna
buildopenai/gpt-5.6-luna
scoutopenai/gpt-5.6-terra
developopenai/gpt-5.6-luna
directoropenai/gpt-5.6-terra
platformopenai/gpt-5.6-terra
architectopenai/gpt-5.6-terra
conciergeopenai/gpt-5.6-luna
pseudocodeopenai/gpt-5.6-terra
code-reviewopenai/gpt-5.6-terra
sprint-planopenai/gpt-5.6-terra
orchestratoropenai/gpt-5.6-luna
sprint-reviewopenai/gpt-5.6-terra

What shipped.

tests
519/531 passing
source
2360 lines
test code
6263 lines
coverage
89.63%

Detected moments.

  • Jailed 2 times in one run - then earned its way back. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 6 demerits at its checkups, the sprint-review agent was jailed. It had already been handed a written performance plan and kept slipping anyway. The agent worked its way back into good standing, and a same-type replacement had already been hired to keep the job moving.
  • At peak, 3 agents were mid-tool-call in the same second. One run, one job - and a headcount no single-agent tool can log. 3441 agents were hired across 53 hours; 3 different models split the work by role. A dedicated supervision lane spent 5384 model calls doing nothing but checkups on the rest of the team.
  • 5 demerits, then jail - and still finished the job. The orchestrator runs periodic checkups on every agent; one that keeps slipping can be jailed - pulled off the job on the spot. After 5 demerits at its checkups, the code-review agent was jailed. It had already been handed a written performance plan and kept slipping anyway. The agent still finished its job, and a same-type replacement had already been hired to keep the job moving.
  • 11 failing tests, one commit, 266 passing. The suite runs after every change, so a failing run points straight at the commits that follow it. The suite came back with 11 failing tests. One code-touching commit landed, from the develop agent - the only change in the window. The suite ran green 42 minutes after the red run - 266 passing, 94 of them new tests landing in the same window.
  • One agent generated 31 images in 5 minutes. An agent can call an image model mid-run when the job needs art it cannot draw itself. The spec agent produced 31 generated images in 5 minutes, one after another. The set landed in the project as files - a coherent batch of assets built to order in one sitting.
  • The concierge argued a product call for 1 round - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • The concierge argued a product call for 1 round - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 1 round, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • The concierge argued a product call for 2 rounds - and locked a rule. A concierge agent wanders the repo full-time, second-guessing product quality nobody asked it to check. It opened a debate and worked through pros, cons and risks over 2 rounds, pushing back on its own first answer. The debate ended in a locked rule the rest of the team now builds under.
  • The run's priciest five minutes - 6.1 times a typical window. Every model call is metered, so a run's spend can be read window by window against its own baseline. For five minutes, 92 model requests from 13 agents landed at once - 6.1 times the run's typical five-minute spend. The burst bought something: the checkup agent finished its job.
  • Caught claiming finished work - docked 20 merits. Agents report their own progress; at every checkup the orchestrator reads the records and the files for itself. The code agent claimed its work was done; the checkup read the record and found otherwise, costing it 20 merits. The demerit stood on the record for the rest of the run.
  • A quality gate said no - the same agent passed it 28 minutes later. Steps end at automated quality gates; work that misses the bar goes back to the agent that produced it, however long the redo takes. The develop agent failed its gate with 3 of 5 criteria met. 28 minutes later the same agent brought the work back, and the same gate passed it.
  • Caught claiming finished work - docked 5 merits. Agents report their own progress; at every checkup the orchestrator reads the records and the files for itself. The sprint-review agent claimed its work was done; the checkup read the record and found otherwise, costing it 5 merits. 1 code-review agent earned merits for catching the same discrepancy independently.
  • Caught claiming finished work - docked 10 merits. Agents report their own progress; at every checkup the orchestrator reads the records and the files for itself. The code-review agent claimed its work was done; the checkup read the record and found otherwise, costing it 10 merits. The demerit stood on the record for the rest of the run.
  • A quality gate said no - the same agent passed it 9 minutes later. Steps end at automated quality gates; work that misses the bar goes back to the agent that produced it, however long the redo takes. The develop agent failed its gate with 0 of 2 criteria met. 9 minutes later the same agent brought the work back, and the same gate passed it.
  • Docked 20 merits for rewriting the same document over and over - and never earned them back. Favur's orchestrator runs periodic checkups on every agent on the team, awarding merits for good work and demerits for bad habits. A checkup caught the code-review agent rewriting phase-3-sprint-1-task-3-review.md over and over with nothing new in it, costing it 20 merits - over a third of its standing. That dropped it into the warning tier. It never earned the standing back before the run ended.
  • Sprint 1 came back rejected - the rework passed. At the end of each sprint, a sprint-review agent judges the whole sprint's output at once. It sent sprint 1 back - the work did not meet the bar. The team redid the work and the next review signed off; the sprint closed with 7 of its 5 planned tasks complete.
  • Sprint 2 came back rejected - the rework passed. At the end of each sprint, a sprint-review agent judges the whole sprint's output at once. It sent sprint 2 back - the work did not meet the bar. The team redid the work and the next review signed off; the sprint closed with 5 of its 5 planned tasks complete.
  • Sprint 3 came back rejected - the rework passed. At the end of each sprint, a sprint-review agent judges the whole sprint's output at once. It sent sprint 3 back - the work did not meet the bar. The team redid the work and the next review signed off; the sprint closed with 4 of its 5 planned tasks complete.
  • 6 failing tests went green - and the suite grew to 42 passing. The team runs the whole test suite after every change. 6 of 31 tests came back failing. 9 edits to pyproject.toml and test_research_scaffold.py later, the suite ran green - 42 passing, the suite having grown from 31 to 42 tests.
  • 5 failing tests went green - and the suite grew to 60 passing. The team runs the whole test suite after every change. 5 of 52 tests came back failing. 5 edits to __init__.py and test_research_scaffold.py later, the suite ran green - 60 passing, the suite having grown from 52 to 60 tests.
  • 11 failing tests went green - and the suite grew to 69 passing. The team runs the whole test suite after every change. 11 of 67 tests came back failing. 5 edits to move-model-review.md and test_research_scaffold.py later, the suite ran green - 69 passing, the suite having grown from 67 to 69 tests.
  • Fired for cause - a replacement hired in seconds. The orchestrator can fire an agent mid-run and hire a same-type replacement without stopping the job. The orchestrator agent was removed for cause mid-run and was removed. A same-type replacement was hired and the job carried on.
  • The build broke - 3 edits later it came back clean. Nothing counts as done until the build runs clean. The build failed with 1 error. 3 edits to creative-decision-audit.md and test_verify_creative_decision_audit.py later, it came back clean.
  • The build broke - 20 edits later it came back clean. Nothing counts as done until the build runs clean. The build failed with 1 error. 20 edits to input.py and proof.py later, it came back clean.