¿Qué es una configuración de inferencia?
¿Qué es?
Una configuración de inferencia reúne los ajustes para ejecutar un modelo en tu propio equipo: el motor (Ollama, vLLM, llama.cpp), la cuantización, el tamaño del contexto y el comando de lanzamiento.
Cuándo usarlo
- Run an LLM locally, without sending data to the cloud.
- Find the right settings for specific hardware (Mac, GPU, server).
Campos de la página
| Runtime | Ollama, vLLM, llama.cpp, LM Studio, TGI. |
|---|---|
| Configuration | The Modelfile or settings file. |
| Command | The line to run. |
| Quantization, context, memory | What it takes to fit on the machine. |
Cómo usarlo
Copy the configuration, then run the command shown under “Usage”.