Skip to content

Repository files navigation

PeronAI Logo

🚀 LiteLLM On-Premise (PeronAI Edition)

Docker

A professional, self-hosted deployment of LiteLLM Proxy with a PostgreSQL database and custom branding. This setup allows you to consolidate multiple LLM providers (OpenAI, Anthropic, Gemini, Azure, etc.) into a single, OpenAI-compatible API gateway.


📑 Table of Contents


✨ Features

  • Unified API: Access any LLM (100+ providers) using the standard OpenAI SDK format.
  • Admin UI: Beautiful dashboard to manage models, track usage, and manage keys.
  • Database Integrated: Includes PostgreSQL for persistent logs, usage tracking, and API key management.
  • Custom Branding: Pre-configured with PeronAI branding.
  • Cross-Platform: Automated setup scripts for both Windows and Linux/macOS.
  • Cost Tracking: Monitor spend across different models and providers in real-time.

🏗️ Architecture & Flow

System Overview

This diagram illustrates how the LiteLLM Proxy acts as a central gateway between your applications and various LLM providers.

graph TD
    subgraph "External Providers"
        OpenAI[OpenAI]
        Anthropic[Anthropic]
        Gemini[Google Gemini]
        Azure[Azure OpenAI]
    end

    subgraph "Your Lab / Infrastructure"
        Apps[User Applications] -->|OpenAI SDK| Proxy[LiteLLM Proxy]
        Proxy <--> Config[config.yaml]
        Proxy <--> DB[(PostgreSQL)]
        AdminUI[Admin Dashboard] <--> Proxy
    end

    Proxy -->|API Request| OpenAI
    Proxy -->|API Request| Anthropic
    Proxy -->|API Request| Gemini
    Proxy -->|API Request| Azure
Loading

Request Lifecycle

A quick look at how a single request is handled by the system.

sequenceDiagram
    participant App as User Application
    participant Proxy as LiteLLM Proxy
    participant DB as PostgreSQL
    participant LLM as AI Provider (e.g. GPT-4)

    App->>Proxy: Chat Request (with Key)
    Proxy->>DB: Validate Key & Permissions
    DB-->>Proxy: Authorized
    Proxy->>LLM: Forwarded Request (Mapped)
    LLM-->>Proxy: Model Response
    Proxy->>DB: Log Usage & Costing
    Proxy-->>App: Return Formatted Response
Loading

🛠️ Prerequisites


🚀 Quick Start

1. Initialize Configuration

Clone this repository and run the setup script for your operating system. This will generate your .env and initialize the environment.

Windows (PowerShell):

./setup.ps1

Linux/macOS (Bash):

chmod +x setup.sh run.sh
./setup.sh

2. Configure Your Environment

Open the newly created .env file:

  1. LITELLM_MASTER_KEY: Set this to a secure string (e.g., sk-my-secret-1234). You will need this to log in to the UI.
  2. API Keys: Add your provider API keys (e.g., OPENAI_API_KEY, ANTHROPIC_API_KEY).

3. Define Models

Edit config.yaml to define the models you want to expose through the proxy. Refer to the commented examples in the file.

4. Launch the Proxy

Start the stack using the provided run scripts:

Windows:

./run.ps1

Linux/macOS:

./run.sh

📊 Admin Dashboard

Once running, open your browser to: 👉 http://localhost:4000/ui

Log in using the Master Key defined in your .env.

In the UI, you can:

  • Monitor real-time usage and costs.
  • Generate unique API keys for different applications.
  • Set rate limits, budgets, and model permissions per key.

💻 Usage Example

Point your OpenAI client to your local proxy:

import openai

client = openai.OpenAI(
    api_key="your-proxy-key", # Use your Generated Key or Master Key
    base_url="http://localhost:4000"
)

response = client.chat.completions.create(
    model="gpt-4",
    messages=[{"role": "user", "content": "Hello from PeronAI!"}]
)

print(response.choices[0].message.content)

📂 Project Structure

  • config.yaml: Model definitions and proxy settings.
  • docker-compose.yml: Docker orchestration for LiteLLM and PostgreSQL.
  • .env: Secret environment variables (ignored by git).
  • peronai_logo.png: Branding asset.
  • setup.*: Cross-platform initialization scripts.
  • run.*: Execution scripts to simplify Docker commands.

🛠️ Management Commands

Action Command
View Logs docker compose logs -f
Stop Services docker compose down
Update Images docker compose pull
Reset DB docker compose down -v (Warning: Deletes all logs/keys)

📄 License

This project follows the LiteLLM license. For more details, visit the LiteLLM Documentation.

About

LiteLLM On-Prem — A lightweight, self-hosted LLM gateway for running and managing large language models entirely on-premises. Provides unified OpenAI-compatible APIs, secure model routing, usage tracking, and enterprise-grade control without external dependencies.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages