Sign in Publish

AI configurations: inference, stacks, embeddings and settings

Technical settings to run or adapt a model: inference, stacks, embeddings, LoRA, RAG…

contextetech

Mistral 7B on a 16 GB Mac with Ollama

Ollama Modelfile to run Mistral 7B Instruct locally on a 16 GB Mac: Q4…

ConfigurationOllamaQ4_K_M 6 days
contextetech

Qwen2.5 7B with vLLM on a 24 GB GPU

Serve Qwen2.5 7B Instruct with vLLM on a 24 GB GPU: OpenAI-compatible API, 8k…

ConfigurationvLLMbf16 6 days
contextetech

llama.cpp server for a GGUF model

Run any GGUF model with the llama.cpp server: OpenAI-compatible API on port…

Configurationllama.cppQ4_K_M 6 days
contextetech

Private chat on your documents: Ollama, Qdrant, Open WebUI

docker-compose stack for a private chat on your documents: Ollama (model),…

Configuration3 service(s) 6 days
contextetech

Minimal private AI chat: Ollama and Open WebUI

The simplest stack for a private AI chat: Ollama and Open WebUI with…

Configuration2 service(s) 6 days
contextetech

Local automations: n8n and Ollama

docker-compose stack to automate with a local AI: n8n (workflows) and Ollama…

Configuration2 service(s) 6 days
contextetech

nomic-embed-text embeddings locally with Ollama

nomic-embed-text for a 100% local English RAG with Ollama: 768 dimensions, long…

Configuration768 dim. 6 days
contextetech

BGE-M3 embeddings for RAG over long documents

BGE-M3 for RAG over long documents: 1024 dimensions, up to 8,192 tokens, 100+…

Configuration1024 dim. 6 days

What is a configuration for AI?

A configuration gathers the technical settings that run or adapt a model: LoRA training parameters, a RAG pipeline, an embedding model, local inference, a stack of services, a reward function or a synthetic data recipe. You copy it and run it again to get the same result.