Skip to content

Repository files navigation

AI Voice Agent 🎤🤖

A real-time AI voice agent that can listen, reason, and respond using agentic decision-making. Built with FastAPI, LangGraph, and Next.js.

✨ Features

  • Voice Interaction: Speak naturally and get voice responses
  • Intent Classification: Automatically understands user intent (chat, question, task, unclear)
  • Tool Calling: Uses tools like search, notes, summarizer, and time
  • Memory: Remembers conversation context
  • LangGraph Workflow: Agentic decision-making with LangGraph

🛠️ Tech Stack

Backend

  • FastAPI - Web framework
  • LangChain - LLM framework
  • LangGraph - Agent workflow orchestration
  • Google Gemini - LLM (FREE tier available!)
  • OpenAI Whisper - Speech-to-text
  • Google Text-to-Speech (gTTS) - Text-to-speech
  • Python 3.11

Frontend

  • Next.js 14 - React framework
  • Web Audio API - Audio recording/playback
  • Tailwind CSS - Styling
  • TypeScript - Type safety

🚀 Quick Start

Prerequisites

  • Python 3.11+
  • Node.js 18+
  • Google Gemini API key (FREE - get it at https://aistudio.google.com/app/apikey)
  • All services are FREE!:
    • LLM: Google Gemini (free tier) or rule-based fallback
    • TTS: Google Text-to-Speech (gTTS) - completely free
    • Search: DuckDuckGo - completely free, no API key needed
    • STT: OpenAI Whisper - runs locally, free

Backend Setup

  1. Create virtual environment:
cd backend
python -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate
  1. Install dependencies:
pip install -r ../requirements.txt
  1. Set up environment variables: Create a .env file in the project root:
# Get your FREE API key from: https://aistudio.google.com/app/apikey
GOOGLE_API_KEY=your_google_api_key_here
GEMINI_MODEL=gemini-pro

To get a Google Gemini API key (FREE!):

  1. Go to https://aistudio.google.com/app/apikey
  2. Sign in with your Google account
  3. Click "Create API Key"
  4. Copy the key and add it to your .env file
  5. See GEMINI_SETUP.md for detailed instructions

Note: The system works without API key too! It will use rule-based responses as fallback.

  1. Run the backend:
cd backend
python main.py

The API will be available at http://localhost:8000

Frontend Setup

  1. Install dependencies:
npm install
  1. Run the development server:
npm run dev

The frontend will be available at http://localhost:3000

📖 Usage

  1. Open http://localhost:3000 in your browser
  2. Click and hold the microphone button
  3. Speak your message
  4. Release the button to send
  5. The agent will process your request and respond with voice

Example Interactions

  • Casual Chat: "Hello, how are you?"
  • Question: "What's the weather in Lahore?"
  • Task with Tool: "Search for information about LangGraph"
  • Notes: "Save a note: Meeting at 3pm tomorrow"
  • Time: "What time is it?"
  • Memory: "What did we discuss earlier?"

🧩 Agent Workflow

START
 ↓
Speech Input Node
 ↓
Intent Classification Node
 ↓
 ├─ Simple Chat → Response Node
 ├─ Tool Required → Tool Node
 ├─ Unclear → Clarification Node
 ↓
Memory Update Node
 ↓
TTS Node
 ↓
END

🔧 API Endpoints

POST /api/transcribe

Convert speech to text

  • Body: audio (multipart/form-data)
  • Returns: Transcribed text

POST /api/chat

Process text through the agent

  • Body: { "text": "...", "session_id": "..." }
  • Returns: { "response": "...", "intent": "...", "tools_used": [...] }

POST /api/voice-chat

End-to-end voice chat (STT → Agent → TTS)

  • Body: audio (multipart/form-data), session_id (form data)
  • Returns: Audio stream (MPEG)

GET /api/memory/{session_id}

Get conversation memory

  • Returns: { "session_id": "...", "memory": [...] }

DELETE /api/memory/{session_id}

Clear conversation memory

🤖 Agent Tools

  1. Search Tool: Web search using DuckDuckGo (free, no API key needed)
  2. Note Tool: Store, retrieve, and manage notes (in-memory)
  3. Summarizer Tool: Summarize text content
  4. Time Tool: Get current time and date

📝 Project Structure

ai-voice-agent/
├── backend/
│   ├── main.py              # FastAPI app
│   ├── agent/
│   │   ├── agent.py         # LangGraph agent
│   │   └── tools.py         # Agent tools
│   └── services/
│       ├── stt.py           # Speech-to-text
│       ├── tts.py           # Text-to-speech
│       └── memory.py        # Memory management
├── app/
│   ├── page.tsx             # Main page
│   ├── layout.tsx           # Root layout
│   ├── globals.css          # Global styles
│   └── components/
│       └── VoiceChat.tsx    # Voice chat component
├── requirements.txt         # Python dependencies
├── package.json             # Node dependencies
└── README.md

🎯 Intent Types

Intent Action
casual_chat Direct reply
question Reason + answer
task Call tool
unclear Ask follow-up

🔐 Environment Variables

All services are FREE and work without API keys!

Optional:

Free Services Used:

  • LLM: Hugging Face (free tier) or rule-based fallback
  • TTS: Google Text-to-Speech (gTTS) - no API key needed
  • Search: DuckDuckGo - no API key needed
  • STT: OpenAI Whisper - runs locally, completely free

🚧 Future Enhancements

  • Redis integration for persistent memory
  • Multiple voice options
  • Conversation export
  • Custom tool creation
  • Multi-language support
  • Streaming responses

📄 License

MIT

🙏 Acknowledgments

  • LangChain & LangGraph teams
  • OpenAI for Whisper
  • ElevenLabs for TTS
  • Tavily for search

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages