May 4, 2026·1 min read

KoboldCpp — Single-File Local LLM Inference Engine

KoboldCpp is a self-contained local LLM inference engine that runs GGUF models with GPU acceleration on consumer hardware, providing an OpenAI-compatible API and built-in web UI without requiring Python or complex setup.

Agent ready

Ready-to-run agent install

This asset can be installed after the agent chooses its runtime, checks the plan, and runs the matching command.

Native · 98/100Policy: allow
Agent surface
Any MCP/CLI agent
Kind
Skill
Install
Single
Trust
Trust: Community
Entrypoint
KoboldCpp LLM Engine
Direct install command
npx -y tokrepo@latest install f0ec1009-4771-11f1-9bc6-00163e2b0d79 --target codex

Run after dry-run confirms the install plan.

KoboldCpp is a self-contained local LLM inference engine that runs GGUF models with GPU acceleration on consumer hardware, providing an OpenAI-compatible API and built-in web UI without requiring Python or complex setup.

Discussion

Sign in to join the discussion.
No comments yet. Be the first to share your thoughts.