How to Self-Host an LLM on a VPS? – Ollama, DeepSeek & LLaMA Installation Guide

Rajdeep Singh

Last Updated:

hero-image

AI has moved fast, and so have the expectations around control, privacy, and cost. Relying on third-party APIs may work at the start, but over time, limits begin to show. Usage costs rise, data handling becomes a concern, and flexibility starts to feel restricted. Many developers now look for a better way to run AI without depending entirely on external platforms.

A self-hosted LLM offers that shift. It allows you to run powerful language models on your own VPS, manage your data within your own environment, and build systems that fit your exact needs. With tools like Ollama, setting up models such as DeepSeek or LLaMA has become far more practical than it once was. In this guide, you will learn how to set up, run, and manage your own AI server step by step.

Key Takeaways from the Article

  • A self-hosted LLM eliminates reliance on external AI APIs.

  • VPS environments can handle real AI workloads with proper configuration.

  • Ollama simplifies deployment and API access.

  • DeepSeek excels at reasoning, while LLaMA works well for general tasks.

  • NVMe storage directly improves model load speed and response time.

  • Security setup is essential before exposing any API.

  • A KVM-based VPS offers consistent performance for AI workloads.

What is a Self-Hosted LLM?

A self-hosted LLM is a language model running on infrastructure you control, such as a VPS, instead of relying on external AI services. In other words, the model lives on your own server, and your data stays within your environment. You decide how it runs, how it is accessed, and how it fits into your workflow.

In real use, your server becomes a private AI system. It handles prompts, generates responses, and connects directly with your applications through a local API. That setup reduces delays, removes ongoing API costs, and gives you the freedom to adjust how the model behaves based on your specific needs.

Privacy and Data Control

The best case of self-hosting is data privacy. Any request made via an external API will traverse networks and can be logged, analysed, or stored. That is dangerous, particularly with sensitive information such as internal documents, client files, or internal code. Using a dedicated AI server keeps all the information within your environment. It is up to you to decide what gets stored, logged, or deleted. Such a degree of control is adequate regarding compliance and enterprise-level security practices.

Cost Savings Over API Services

API pricing models depend on token usage. As usage increases, costs scale linearly. High-frequency workloads such as chatbots, automation pipelines, or analytics tools quickly become expensive. A VPS introduces a fixed monthly cost. Once the infrastructure is running, you can process an unlimited number of requests within the resource limits. Over time, the cost advantage becomes significant, especially for continuous workloads.

Full Control and Customization

Self-hosting allows deeper control over how models behave. You can switch models, adjust prompts, fine-tune behaviour, and integrate with internal systems without restrictions. That flexibility matters in production environments. Instead of adapting your workflow to match an API, you shape the AI system to fit your workflow.

Hardware Requirements for Self-Hosting an LLM

RAM, CPU, and Storage Breakdown

For a self-hosted LLM, the most crucial resource is RAM. The model should be easily loaded into memory to operate; otherwise, you may encounter crashes or extremely slow responses. Smaller models, such as 3B or 7B, can operate using 8-16 GB of RAM, but you can definitely notice the difference when you have more memory, particularly when you deal with longer prompts or requests concurrently.

CPU performance affects how quickly the model responds. More cores and better processing power help generate outputs faster, even without a GPU. Storage also plays its part. Since model files are quite large, using NVMe storage helps load them quickly and reduces delays when switching between models.

Model Size

Minimum RAM

Recommended RAM

CPU

Storage (NVMe)

3B – 7B models

8 GB

16 GB

2–4 cores

40–60 GB

7B – 13B models

16 GB

32 GB

4–8 cores

80–120 GB

13B+ models

32 GB

64 GB+

8+ cores

150 GB+

 

What is Quantisation?

Quantisation is a technique that reduces the precision of a model's internal calculations, which in turn shrinks the file size and lowers the amount of RAM needed to run it. A full-precision 7B model might need 14–16 GB of RAM, but a quantised version of the same model (referred to as Q4 or Q5, depending on the level of compression) can run in 6–8 GB. The trade-off is a small reduction in output quality, but for most practical tasks, the difference is negligible.

Why NVMe Storage Matters?

Self-hosted LLM model files are typically large, ranging from a few gigabytes to tens of gigabytes, depending on the model size. Whenever a model is started, they have to be loaded into RAM from disk. In the event of slow storage, that process causes perceivable delays until the model can be used. NVMe drives have much higher read speeds than traditional SSDs, and load times are also much lower, making the system feel more responsive.

Faster storage also improves day-to-day usability. Switching between models, restarting services, or handling frequent requests becomes smoother with NVMe. In setups where multiple models run on the same VPS, quick data access prevents bottlenecks and keeps performance consistent. As a result, NVMe storage plays a key role in delivering a seamless and efficient AI workflow.

DeepSeek vs LLaMA vs Other LLM Models – Which is the Best Model?

DeepSeek R1

DeepSeek R1 is an open-source AI model that was created recently, with an emphasis on reasoning and technical accuracy. It comes from the DeepSeek team, and it has won publicity for how effectively it is at solving problems that require explicit reasoning, like coding or other systematic problem-solving. The replies are usually more sequential than conversational. It is not the easiest model to run on a basic VPS. You require a good RAM and a smooth-running CPU. There are lighter options, but they may not feel as sharp when it comes to detailed outputs.

LLaMA 3

LLaMA 3 is the general-purpose language model developed by Meta to perform generalized everyday tasks. It is also good in writing, summarizing, and general queries, hence most developers consider it a starting point. It does not try to specialise too much, and that actually works in its favour. The ease with which it may be run on a VPS is another strength. It performs reasonably and remains stable even with average resources. If you want something reliable without spending too much time tweaking settings, LLaMA 3 is usually a safe choice.

Mistral 7B

Mistral 7B is a lightweight model that is made to run without requiring any excessive input from the system. It is constructed by Mistral AI and mostly utilised when you have limited resources, but you still want a usable setup. It is fast and reacts to basic tasks; nevertheless, it has its own limits. In more complex queries or when more reasoning is required, it may not deliver the same level of detail as larger models. On simple applications, however, it is effective and does not make things difficult.

Qwen

Qwen is a model family developed by Alibaba Cloud, with a strong focus on handling multiple languages smoothly. It works well in setups where different languages come into play, and it manages context across them without much confusion. Along with that, it handles general tasks like writing or summaries without issues. It does not feel too heavy on system resources either, so running it on a VPS is manageable. If your project needs flexibility across languages, Qwen is a solid option.

 

Model

Parameters

Core Strength

RAM Needed (Quantised)

Best For

DeepSeek R1

7B

Reasoning, code, math

~8 GB

Dev tools, structured analysis

LLaMA 3

8B

General-purpose, conversation

~8 GB

Chatbots, content, Q&A

Mistral 7B

7B

Fast inference, lightweight

~6 GB

Low-resource VPS setups

Qwen

7B

Multilingual (Chinese + English)

~8 GB

Non-English workloads

 

Step-by-Step Process to Install Ollama and Run an LLM on Your VPS

1. Connect to your VPS via SSH  

2. Update system packages  

3. Install Ollama  

4. Verify installation  

5. Start service  

Step 1: Connect to Your VPS via SSH

Before you begin installation, you need secure access to your VPS. SSH allows you to remotely control your server through the command line. Make sure you have your server IP address and credentials ready. Using SSH keys instead of passwords is strongly recommended for better security and stability.

ssh username@your-server-ip

Replace username with your actual user (typically root on a fresh VPS) and your-server-ip with the server’s public IP address. If you are using a non-standard SSH port, add the -p flag:

ssh -p 2222 username@your-server-ip

Step 2: Update System Packages

Keeping your system updated ensures that all existing packages are secure and compatible with new installations. Outdated libraries can cause unexpected errors during setup. A quick update at the start helps avoid troubleshooting issues later in the process.

sudo apt update && sudo apt upgrade -y

On CentOS or AlmaLinux, replace apt with dnf:

sudo dnf update -y

Step 3: Install Ollama

Ollama simplifies the entire process of running a self-hosted LLM by handling model management and execution. Instead of manually configuring dependencies, you can install everything using a single command. The installation script automatically sets up the required environment.

curl -fsSL https://ollama.com/install.sh | sh

The script detects your OS and CPU architecture. It installs the Ollama binary to /usr/local/bin/ollama and registers a systemd service called ollama that starts on boot. The whole process takes under a minute on a decent connection.

Step 4: Verify the Installation

After installation, it is important to confirm that Ollama is working correctly. A quick version check ensures that the tool is properly installed and accessible from your system. If the command returns a version number, you are ready to proceed.

ollama --version

You should see output like ollama version 0.x.x (the exact number depends on the latest release). If the command is not found, your shell may need to be refreshed, run source ~/.bashrc,`or start a new SSH session.

Step 5: Start the Ollama Service

Now you need to start the Ollama service, which runs a local server to handle model requests. This service acts as the backend for your AI system. Once started, your VPS can run and serve LLM responses.

The install script typically starts the service automatically. Check its status:

sudo systemctl status ollama

You should see active (running) in the output. If the service is not running, start it manually:

sudo systemctl start ollama

To make sure Ollama starts automatically after every reboot:

sudo systemctl enable ollama

Process to Run DeepSeek R1 on Your VPS with Ollama

Step 1: Pull the DeepSeek R1 Model

Before using DeepSeek, you need to download the model files to your VPS. Ollama handles this process efficiently by pulling the model from its repository. Ensure your server has enough storage space, as model files can be quite large.

ollama pull deepseek-r1:7b

The 7B variant is approximately 4.7 GB, so download time depends on your server’s bandwidth.

Step 2: Run and Test DeepSeek

Once the model is downloaded, you can run it directly from the terminal. This step allows you to interact with the model and verify that it is working as expected. You can test prompts and observe how the model responds in real time.

ollama run deepseek-r1:7b

Step 3: Test API Endpoint

After confirming the model works in the terminal, the next step is to test the API. Ollama exposes a local endpoint that allows applications to send requests and receive responses. This is essential for integrating your model into real-world systems.

curl http://localhost:11434/api/generate -d '{"model": "deepseek-r1:7b", "prompt": "What is a VPS?", "stream": false}'

You can also check which models are available on your server at any time:

ollama list

This displays all pulled models, along with their sizes and modification dates. You should see deepseek-r1:7b in the list.

How to Run LLaMA 3 on Your VPS with Ollama?

Running LLaMA follows the same structure as DeepSeek. Ollama handles all model management uniformly, so switching between models is just a matter of pulling the weights and specifying the model name.

Pull the LLaMA 3 8B model:

ollama pull llama3:8b

This downloads Meta’s LLaMA 3 in its quantised 8B variant. The file size is similar to DeepSeek R1 7B — roughly 4.7 GB.

Start an interactive session:

Ollama run llama3:8b

Test it with a prompt:

>>> Write a short summary of how DNS resolution works.

LLaMA 3 tends to produce more conversational, clearly structured responses compared to DeepSeek R1, which leans more toward analytical and code-oriented responses. Both models coexist on the same server without any conflict. Switch between them freely using ollama run deepseek-r1:7b or ollama run llama3:8b, and the API endpoint supports both simultaneously — just change the model parameter in your request.

Test LLaMA 3 via the API:

curl http://localhost:11434/api/generate -d '{"model": "llama3:8b", "prompt": "Explain load balancing in one paragraph.", "stream": false}'

Steps to Add a Browser Interface with Open WebUI

Step 1: Install Docker on Your VPS

If Docker is not already on your VPS, run the following:

sudo apt install -y apt-transport-https ca-certificates curl software-properties-common

curl -fsSL https://get.docker.com | sh

Add your user to the Docker group so you can run commands without sudo:

sudo usermod -aG docker $USER

Log out and back in (or run newgrp docker) for the group change to take effect. Verify Docker is working:

docker --version

Step 2: Run Open WebUI as a Docker Container

The --network=host flag lets the container talk directly to Ollama on localhost:11434:

docker run -d --network=host -v open-webui:/app/backend/data -e OLLAMA_BASE_URL=http://127.0.0.1:11434 --name open-webui --restart always ghcr.io/open-webui/open-webui:main

This pulls the latest Open WebUI image, starts the container in the background, and creates a persistent Docker volume for your chat history and settings. The --restart-always flag ensures the container automatically restarts after a crash or server reboot.

Once the container is running, access the interface in your browser:

http://your-server-ip:8080

Step 3: Switch Between Models in the Interface

Visit:

http://your-server-ip:8080

You can switch models, test prompts, and manage sessions visually.

VPS vs GPU Server – When is a VPS Enough?

A VPS can handle many real-world AI tasks effectively.

When a VPS Works Well

  • Chatbots and assistants

  • Internal tools

  • API-based automation

  • Development environments

When You Need a GPU Server

  • Large-scale inference

  • Real-time applications

  • High concurrency workloads

A VPS offers a strong balance between cost and performance. For most developers, it provides enough power to deploy and test AI systems without heavy infrastructure investment. If you want a stable environment with fast storage, explore a KVM NVMe VPS.

Tips to Secure Your Self-Hosted LLM

Bind Ollama to Localhost Only

By default, services may accept external connections, which can expose your model API to the public internet. Restricting Ollama to localhost ensures that only internal processes or authorised services can access it. This simple step significantly reduces the risk of unauthorised access.

ss -tlnp | grep 11434

Set Up a Reverse Proxy with Nginx

A reverse proxy acts as a controlled gateway between your server and external traffic. It allows you to add SSL encryption, authentication, and request filtering before traffic reaches your LLM service. This adds an important security layer while also improving traffic management.

sudo apt install nginx -y

Create a server block at /etc/nginx/sites-available/llm:

server { listen 80; server_name llm.yourdomain.com;  location / {     proxy_pass http://127.0.0.1:8080;     proxy_set_header Host $host;     proxy_set_header X-Real-IP $remote_addr;     proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;     proxy_set_header X-Forwarded-Proto $scheme;     proxy_http_version 1.1;     proxy_set_header Upgrade $http_upgrade;     proxy_set_header Connection "upgrade"; } }

Enable the site and restart Nginx:

sudo ln -s /etc/nginx/sites-available/llm /etc/nginx/sites-enabled/ sudo nginx -t sudo systemctl restart nginx

Firewall and SSH Hardening

Basic firewall rules help block unnecessary incoming traffic and limit access to only essential services. At the same time, securing SSH access prevents brute-force attacks and unauthorised logins. Using SSH keys instead of passwords adds an extra layer of protection.

sudo ufw allow 22/tcp sudo ufw allow 80/tcp sudo ufw allow 443/tcp sudo ufw enable

The Bottom Line

Running your own LLM changes how you deal with AI daily. You are not tied to external APIs, your data stays where it should, and costs stop jumping around every time usage grows. Tools like Ollama have made the setup far less complicated than it used to be, and models like DeepSeek or LLaMA are now practical options even on a VPS.

If you are planning to set up your own system, the server you choose matters more than people expect. A stable VPS with fast storage makes everything smoother, from loading models to handling requests. For setups that require consistent performance, a KVM NVMe VPS from HostSailor can be considered as a practical option.

Frequently Asked Questions About Self-Hosted LLMs

Can you run an LLM on a VPS without a GPU?

Yes, you can run a self-hosted LLM on a VPS without a GPU, especially smaller or quantised models. A well-configured CPU setup with enough RAM can handle most tasks, making it possible to host LLM locally without expensive hardware.

What is Ollama and how does it work?

Ollama is a tool that helps you set up and run a local AI server with minimal effort. It handles model downloads, execution, and API access, which makes the ollama server setup simple and practical for most VPS environments.

Is a self-hosted LLM private?

A self-hosted LLM works as a private AI server, where all data stays within your own infrastructure. No external service processes your requests, which gives you full control over how data is handled and stored.

How does DeepSeek compare to LLaMA for self-hosting?

DeepSeek focuses more on reasoning and technical tasks, while LLaMA provides balanced performance across general use cases. For a local AI server, DeepSeek suits structured workflows, whereas LLaMA works better for everyday interactions and content tasks.

 

Reliable Hosting You Can Trust

Experience lightning-fast, secure hosting that easily scales as your business grows, empowering you to succeed online effortlessly.

Start Hosting Now

Join Our Newsletter

Your information will never be Shared with third parties, and you can unsubscribe from our updates at any time.