Quick Draw - a circle-drawing game
Run of , staffed by a mixed roster (google + x-ai + anthropic + openai + qwen + meta). Combines Gpt Terra (primary) with minority vendors for cost savings - composite 79.03.
The verdict.
79.03 composite · rubric v15
- duration
- 116 min
- requests
- 814
- failed requests
- 0 (0%)
- cost
- $19.03
Scores come from the open scoring pipeline - methodology and leaderboards at Favur Evals ↗
Per-subject scores.
| subject | score |
|---|---|
| code quality | 74.77 |
| cost efficiency | 56.95 |
| deliverables | 55.94 |
| effort efficiency | 60.79 |
| process discipline | 51.18 |
| test quality | 50 |
| tool discipline | 91 |
| velocity | 64.11 |
Who staffed it.
| role | model |
|---|---|
| code | openai/gpt-5.6-terra |
| build | meta/muse-spark-1.1 |
| scout | openai/gpt-5.6-luna |
| develop | google/gemini-3.1-flash-lite |
| concierge | google/gemini-3.1-flash-lite |
| pseudocode | anthropic/claude-sonnet-4.6 |
| code-review | qwen/qwen3.7-plus |
| sprint-plan | x-ai/grok-4.5 |
| orchestrator | google/gemini-3.1-flash-lite |
| sprint-review | qwen/qwen3.7-plus |
What shipped.
- tests
- 20/20 passing
- source
- 129 lines
- test code
- 150 lines
- coverage
- 52.87%