What is an inference config?
What is it?
An inference config is the set of settings needed to run a model on your own machine: the runtime (Ollama, vLLM, llama.cpp), quantization, context size and launch command.
When to use it
- Run an LLM locally, without sending data to the cloud.
- Find the right settings for specific hardware (Mac, GPU, server).
Page fields
| Runtime | Ollama, vLLM, llama.cpp, LM Studio, TGI. |
|---|---|
| Configuration | The Modelfile or settings file. |
| Command | The line to run. |
| Quantization, context, memory | What it takes to fit on the machine. |
How to use it
Copy the configuration, then run the command shown under “Usage”.