Preview — services not connectedAbout this workspace

Large 4 or Large 3: choose by task

A reproducible evaluation plan rather than an unsupported winner claim.

On this page

Start with verifiable work

Build a test set from a code bug with a known fix, a question with a source passage, and an image with labels you can check. Use identical prompts and settings for both models.

Record quality, latency, and cost

Save model ID, date, prompt, output limit, and reasoning setting. Score against a reference. Measure first-text latency separately from total time.

TaskWhat to check
CodeDiagnosis, runnable patch, scope
DocumentsEvidence from the source
ImagesLabels and uncertainty
JSONSchema validity

Upgrade with evidence

Higher benchmark scores do not guarantee better results on your work. Compare failures too. This build has not run a live comparison and publishes no invented winner.

Try a task from your own work.

Bring the source material and check the result.

Open workspace