DE
Anmelden Veröffentlichen

Was ist eine Inferenz-Konfiguration?

Was ist das?

Eine Inferenz-Konfiguration bündelt die Einstellungen, um ein Modell auf dem eigenen Rechner zu betreiben: Laufzeit (Ollama, vLLM, llama.cpp), Quantisierung, Kontextgröße und Startbefehl.

Wann verwenden

  • Run an LLM locally, without sending data to the cloud.
  • Find the right settings for specific hardware (Mac, GPU, server).

Felder der Seite

RuntimeOllama, vLLM, llama.cpp, LM Studio, TGI.
ConfigurationThe Modelfile or settings file.
CommandThe line to run.
Quantization, context, memoryWhat it takes to fit on the machine.

So nutzt du es

Copy the configuration, then run the command shown under “Usage”.

Inferenz-Konfigurationen ansehen →