ScriptsSep 11, 2026·3 min read

Tabby — Self-Hosted AI Code Completion

An open-source, self-hosted AI coding assistant that provides IDE-integrated code suggestions and chat without sending data to third-party services.

Agent ready

Ready-to-run agent install

This asset can be installed after the agent chooses its runtime, checks the plan, and runs the matching command.

Native · 98/100Policy: allow
Agent surface
Any MCP/CLI agent
Kind
Skill
Install
Single
Trust
Trust: Established
Entrypoint
Tabby Overview
Direct install command
npx -y tokrepo@latest install 0f6772bd-ae2e-11f1-9bc6-00163e2b0d79 --target codex

Run after dry-run confirms the install plan.

Introduction

Tabby is a self-hosted AI code completion server that runs entirely on your own infrastructure. It lets teams use LLM-powered coding assistance while keeping source code private and under full organizational control.

What Tabby Does

  • Provides real-time code completions inside VS Code, JetBrains, Vim, and Emacs
  • Runs a built-in chat interface for code Q&A and generation
  • Indexes your repository to deliver context-aware suggestions
  • Supports multiple open-weight models including StarCoder, CodeLlama, and DeepSeek
  • Works with NVIDIA CUDA, Apple Metal, and CPU-only setups

Architecture Overview

Tabby consists of a Rust-based HTTP server that loads a quantized code model and exposes a completions API. IDE extensions communicate with this server over HTTP. A background indexer crawls configured Git repositories to build a retrieval-augmented context window, improving suggestion relevance for private codebases.

Self-Hosting & Configuration

  • Deploy via Docker, Docker Compose, or bare-metal binary
  • GPU acceleration with CUDA or Metal for fast inference
  • Configure models through the admin web UI at port 8080
  • Add Git repositories for codebase-aware completions
  • Set up team accounts with LDAP or OAuth integration

Key Features

  • Full data privacy with no code leaving your network
  • Sub-200ms latency on modern GPUs
  • Repository-level code context through RAG indexing
  • Multi-user support with admin dashboard and usage analytics
  • Answer Engine for codebase-aware chat and code search

Comparison with Similar Tools

  • GitHub Copilot — cloud-hosted, requires sending code to Microsoft servers
  • Cody by Sourcegraph — strong context engine but partially SaaS-dependent
  • Continue — open-source IDE extension but requires separate model hosting
  • Codeium — free tier available but proprietary backend
  • FauxPilot — early self-hosted alternative, now largely unmaintained

FAQ

Q: What hardware do I need? A: A machine with an NVIDIA GPU (8 GB+ VRAM) is recommended. Smaller models like StarCoder-1B run on 4 GB VRAM or even CPU-only, though with higher latency.

Q: Which IDEs are supported? A: VS Code, JetBrains (IntelliJ, PyCharm, etc.), Vim/Neovim, and Emacs all have official extensions.

Q: Can I use my own fine-tuned model? A: Yes. Tabby supports GGML and HuggingFace model formats, so you can serve a custom fine-tuned checkpoint.

Q: Is Tabby free? A: The core server and all IDE extensions are open-source under the Apache 2.0 license. An optional Tabby Cloud offering exists for managed hosting.

Sources

Discussion

Sign in to join the discussion.
No comments yet. Be the first to share your thoughts.

Related Assets