EN
Sign in Publish
contextetech /

llama.cpp server for a GGUF model

v1
English License: MIT Published on updated 1 hour ago 0 uses

Run any GGUF model with the llama.cpp server: OpenAI-compatible API on port 8080, works without a GPU too.

Modelmodel.gguf
Tested hardwareCPU only, or a GPU (use -ngl to offload layers)

Launch command

llama-server -m model.gguf -c 8192 --host 127.0.0.1 --port 8080

Community

No comments yet. Share your feedback, it will help the next person.