Large 4 or Large 3: choose by task
A reproducible evaluation plan rather than an unsupported winner claim.
On this page
Start with verifiable work
Build a test set from a code bug with a known fix, a question with a source passage, and an image with labels you can check. Use identical prompts and settings for both models.
Record quality, latency, and cost
Save model ID, date, prompt, output limit, and reasoning setting. Score against a reference. Measure first-text latency separately from total time.
| Task | What to check |
|---|---|
| Code | Diagnosis, runnable patch, scope |
| Documents | Evidence from the source |
| Images | Labels and uncertainty |
| JSON | Schema validity |
Upgrade with evidence
Higher benchmark scores do not guarantee better results on your work. Compare failures too. This build has not run a live comparison and publishes no invented winner.