Evidence before a leaderboard.
Keep published scores, independent tests, and your own workload separate. This preview has not run a live model benchmark.
No invented scores
We will publish scores with the exact model, prompt set, settings, date, and method. Until then, use the source links and this evaluation checklist.
| Task | Measure | Reference |
|---|---|---|
| Code review | Runnable fixes; false positives | Tests and expected behavior |
| Document answers | Correctness and grounded evidence | Source passages |
| Image understanding | Label and object accuracy | Annotated images |
| Response efficiency | First text, total time, token cost | Server usage records |