Ein KI-Modell selbst bereitstellen
LLMs, Inferenz-Konfigurationen und Stacks, um KI auf eigenen Rechnern zu hosten.
llama.cpp server for a GGUF model
Run any GGUF model with the llama.cpp server: OpenAI-compatible API on port…
Mistral 7B on a 16 GB Mac with Ollama
Ollama Modelfile to run Mistral 7B Instruct locally on a 16 GB Mac: Q4…
Qwen2.5 7B with vLLM on a 24 GB GPU
Serve Qwen2.5 7B Instruct with vLLM on a 24 GB GPU: OpenAI-compatible API, 8k…
Local automations: n8n and Ollama
docker-compose stack to automate with a local AI: n8n (workflows) and Ollama…
Private chat on your documents: Ollama, Qdrant, Open WebUI
docker-compose stack for a private chat on your documents: Ollama (model),…
Minimal private AI chat: Ollama and Open WebUI
The simplest stack for a private AI chat: Ollama and Open WebUI with…
Mistral 7B Instruct v0.3
Fiche de référence du modèle Mistral 7B Instruct v0.3 publié par Mistral AI : 7…
Was heißt ein KI-Modell bereitstellen?
Bereitstellen heißt, ein Modell auf eigenen Rechnern zu betreiben statt eine externe API zu nutzen: Die Daten bleiben bei dir. Man wählt ein Modell, eine Laufzeit (Ollama, vLLM, llama.cpp) und bei Bedarf einen kompletten Stack.