A local MCP server that runs Llama models entirely on your machine. No API keys, no cloud costs, 100% private and offline-capable.
Claim it to get a verified publisher badge, a free copy of our full audit findings, and direct contact for any high-priority issues we find.
Install from
M8ven verifies MCPs across every public registry — install directly from whichever one you prefer.
process.env. You'll be asked to provide them before it can run.MODEL_PATH— C:\path\to\your\model.ggufN_THREADS— Number of CPU threads 4N_GPU_LAYERS— GPU layers (use -1 for all, 0 for CPU only) 0CONTEXT_SIZE— Maximum context window size 2048SESSION_HISTORY_DIR— Directory for storing conversation history historySESSION_MAX_MESSAGES— Maximum messages per session (older messages trimmed) 40SESSION_MAX_FILE_BYTES— Maximum size per session file (bytes) 2097152 (~2MB)SESSION_AUTO_TRIM— Automatically trim history when limits exceeded trueSTREAMING_ENABLED— Enable streaming responses (tokens sent incrementally) falseSTREAMING_CHUNK_SIZE— Approximate chunk size for streaming (characters) 50MODELS_DIRPORT[](https://m8ven.ai/mcp/marcel-msc-local-llm-mcp-tool-13tr52)