EN
Sign in Publish

Choose your GPU server for local artificial intelligence

Your model, your volume: the memory needed and the hardware that is enough to run an AI in-house.

Size
active
Capability index

Your usage

HardwareMemoryUnitsTheoretical maximum throughput

Calculation: model weights (parameters × bytes per parameter) + 20% for context and runtime. Throughput: memory bandwidth ÷ weights read per token; this is a maximum, expect 50 to 70% in practice.

Sources : NVIDIA · Apple · official model cards from the model publishers (size)

Was this information useful?