llama.cpp server for a GGUF model
Run any GGUF model with the llama.cpp server: OpenAI-compatible API on port 8080, works without a GPU too.
| Model | model.gguf |
|---|---|
| Tested hardware | CPU only, or a GPU (use -ngl to offload layers) |
Launch command
llama-server -m model.gguf -c 8192 --host 127.0.0.1 --port 8080
Estimated cost
Input tokens per call : ≈ 135
| Model | Per call | Per 1,000 calls |
|---|---|---|
| Mistral Large · Mistral AI | €0.00006 | €0.06 |
| DeepSeek Flash · DeepSeek | €0.00004 | €0.04 |
| DeepSeek V4 Pro · DeepSeek | €0.00016 | €0.16 |
| GPT-6 Luna · OpenAI | €0.00001 | €0.01 |
| GPT-6 Sol · OpenAI | €0.00024 | €0.24 |
| Claude Sonnet 5.5 · Anthropic | €0.00024 | €0.24 |
| Claude Opus 5.5 · Anthropic | €0.00048 | €0.48 |
| Self-hosted open-source model (Ollama, vLLM) | €0 in API fees (server cost only) | |
Input cost only, excluding the model's answer. Providers' public standard prices (Mistral AI, DeepSeek, OpenAI, Anthropic) converted from USD to euros at the ECB rate of September 29, 2026. Estimate: 1 token ≈ 3.6 characters.
Common to all 3 regions
Created by the author, no outside source.
Commercial use allowed · credit the author · changes allowed.
Text or data only: no access to files, network or commands.
Hosted in France (Scaleway, Paris). No transfer outside the EU.
European Union · 5/5
No personal data.
No training data: no summary required.
United States · 5/5
No personal data.
No training data: no documentation to publish.
China · 5/5
No personal data.
No training data; AI-generated content published in China must be labelled (2025).
Indicative summary as of 30/09/2026, not a legal certification. Method and sources →
Statistics
Likes : 0 · Comments : 0
Click a counter to show or hide it, hover the chart for details. One view per visitor per day, bots excluded; counting started on 30 September 2026.
- Identifier
- contextetech--llama-cpp-server-gguf
- Type
- Inference configs
- File
- llama-cpp-server-gguf.inference.json
- Format
- JSON (application/json)
- Size
- 521 bytes
- Encoding
- UTF-8
- Estimated tokens
- ≈ 145
- Language
- English
- License
- MIT
- Commercial use
- Allowed
- Personal data
- None
- Version
- 1.0
- Published on
- October 2, 2026 at 9:54 PM
- Updated on
- October 2, 2026 at 9:54 PM
- Storage
- in database
- Uses
- 0
- Likes
- 0
llama-cpp-server-gguf.inference.json
{
"format": "contextetech/inference/v1",
"id": "contextetech/llama-cpp-server-gguf",
"description": "Run any GGUF model with the llama.cpp server: OpenAI-compatible API on port 8080, works without a GPU too.",
"tags": [
"llama-cpp",
"gguf",
"cpu"
],
"license": "MIT",
"language": "en",
"runtime": "llamacpp",
"model": "model.gguf",
"config": "",
"command": "llama-server -m model.gguf -c 8192 --host 127.0.0.1 --port 8080",
"quant": "Q4_K_M",
"contextLength": 8192,
"ramGb": null
}Launch command
llama-server -m model.gguf -c 8192 --host 127.0.0.1 --port 8080
Community
No comments yet. Share your feedback, it will help the next person.