Choose your GPU server for local artificial intelligence
Your model, your volume: the memory needed and the hardware that is enough to run an AI in-house.
- Size
- active
- Capability index
Your usage
| Hardware | Memory | Units | Theoretical maximum throughput |
|---|
Calculation: model weights (parameters × bytes per parameter) + 20% for context and runtime. Throughput: memory bandwidth ÷ weights read per token; this is a maximum, expect 50 to 70% in practice.
Sources : NVIDIA · Apple · official model cards from the model publishers (size)
Was this information useful?