Deploy an AI model on your own machines
LLMs, inference configs and stacks to deploy AI on your own machines, without sending your data elsewhere.
contextetech
Mistral 7B on a 16 GB Mac with Ollama
Ollama Modelfile to run Mistral 7B Instruct locally on a 16 GB Mac: Q4…
contextetech
Qwen2.5 7B with vLLM on a 24 GB GPU
Serve Qwen2.5 7B Instruct with vLLM on a 24 GB GPU: OpenAI-compatible API, 8k…
contextetech
llama.cpp server for a GGUF model
Run any GGUF model with the llama.cpp server: OpenAI-compatible API on port…
What does deploying an AI model mean?
Deploying means running a model on your own machines or servers instead of calling an external API: your data stays with you. You pick a model (LLM, embeddings), a runtime (Ollama, vLLM, llama.cpp) and, if needed, a full docker-compose stack.