Plan a local deployment
Check weights, license, memory, and serving requirements before choosing hardware.
On this page
Verify the release first
Read the current model card and weight repository. An announcement, an available API, and downloadable weights are different milestones. We do not host model weights.
Estimate memory from actual files
Check precision, shard sizes, and quantization support. KV cache, activations, runtime buffers, and concurrency need memory beyond the checkpoint itself.
Validate your serving stack
Confirm architecture, vision, reasoning, and checkpoint compatibility. Test a small workload before estimating production throughput.
Compare operating cost
Include hardware, energy, maintenance, and utilization. Occasional work may be simpler through an API. Review data policies before processing private documents.