Was ist eine Inferenz-Konfiguration?
Was ist das?
Eine Inferenz-Konfiguration bündelt die Einstellungen, um ein Modell auf dem eigenen Rechner zu betreiben: Laufzeit (Ollama, vLLM, llama.cpp), Quantisierung, Kontextgröße und Startbefehl.
Wann verwenden
- Run an LLM locally, without sending data to the cloud.
- Find the right settings for specific hardware (Mac, GPU, server).
Felder der Seite
| Runtime | Ollama, vLLM, llama.cpp, LM Studio, TGI. |
|---|---|
| Configuration | The Modelfile or settings file. |
| Command | The line to run. |
| Quantization, context, memory | What it takes to fit on the machine. |
So nutzt du es
Copy the configuration, then run the command shown under “Usage”.