2048 - the sliding-tile game
Run of , staffed by a mixed roster (openai). Combines Gpt Terra (primary) with minority vendors for cost savings - composite 52.54.
The verdict.
52.16 composite · rubric v15
- duration
- 115 min
- requests
- 18282
- failed requests
- 36 (0.2%)
- cost
- $885.00
Scores come from the open scoring pipeline - methodology and leaderboards at Favur Evals ↗
Per-subject scores.
| subject | score |
|---|---|
| code quality | 42.71 |
| cost efficiency | 38.49 |
| deliverables | 68.56 |
| process discipline | 67.06 |
| effort efficiency | 46.82 |
| test quality | 79 |
| tool discipline | 90.38 |
| velocity | 62.5 |
Who staffed it.
| role | model |
|---|---|
| code | openai/gpt-5.6-terra |
| spec | openai/gpt-5.6-terra |
| test | openai/gpt-5.6-luna |
| build | openai/gpt-5.6-luna |
| scout | openai/gpt-5.6-terra |
| develop | openai/gpt-5.6-luna |
| director | openai/gpt-5.6-terra |
| platform | openai/gpt-5.6-terra |
| architect | openai/gpt-5.6-terra |
| concierge | openai/gpt-5.6-luna |
| pseudocode | openai/gpt-5.6-terra |
| code-review | openai/gpt-5.6-terra |
| sprint-plan | openai/gpt-5.6-terra |
| orchestrator | openai/gpt-5.6-luna |
| sprint-review | openai/gpt-5.6-terra |
What shipped.
- tests
- 519/531 passing
- source
- 2360 lines
- test code
- 6263 lines
- coverage
- 89.63%