A real-time AI voice agent that can listen, reason, and respond using agentic decision-making. Built with FastAPI, LangGraph, and Next.js.
- Voice Interaction: Speak naturally and get voice responses
- Intent Classification: Automatically understands user intent (chat, question, task, unclear)
- Tool Calling: Uses tools like search, notes, summarizer, and time
- Memory: Remembers conversation context
- LangGraph Workflow: Agentic decision-making with LangGraph
- FastAPI - Web framework
- LangChain - LLM framework
- LangGraph - Agent workflow orchestration
- Google Gemini - LLM (FREE tier available!)
- OpenAI Whisper - Speech-to-text
- Google Text-to-Speech (gTTS) - Text-to-speech
- Python 3.11
- Next.js 14 - React framework
- Web Audio API - Audio recording/playback
- Tailwind CSS - Styling
- TypeScript - Type safety
- Python 3.11+
- Node.js 18+
- Google Gemini API key (FREE - get it at https://aistudio.google.com/app/apikey)
- All services are FREE!:
- LLM: Google Gemini (free tier) or rule-based fallback
- TTS: Google Text-to-Speech (gTTS) - completely free
- Search: DuckDuckGo - completely free, no API key needed
- STT: OpenAI Whisper - runs locally, free
- Create virtual environment:
cd backend
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate- Install dependencies:
pip install -r ../requirements.txt- Set up environment variables:
Create a
.envfile in the project root:
# Get your FREE API key from: https://aistudio.google.com/app/apikey
GOOGLE_API_KEY=your_google_api_key_here
GEMINI_MODEL=gemini-proTo get a Google Gemini API key (FREE!):
- Go to https://aistudio.google.com/app/apikey
- Sign in with your Google account
- Click "Create API Key"
- Copy the key and add it to your
.envfile - See
GEMINI_SETUP.mdfor detailed instructions
Note: The system works without API key too! It will use rule-based responses as fallback.
- Run the backend:
cd backend
python main.pyThe API will be available at http://localhost:8000
- Install dependencies:
npm install- Run the development server:
npm run devThe frontend will be available at http://localhost:3000
- Open
http://localhost:3000in your browser - Click and hold the microphone button
- Speak your message
- Release the button to send
- The agent will process your request and respond with voice
- Casual Chat: "Hello, how are you?"
- Question: "What's the weather in Lahore?"
- Task with Tool: "Search for information about LangGraph"
- Notes: "Save a note: Meeting at 3pm tomorrow"
- Time: "What time is it?"
- Memory: "What did we discuss earlier?"
START
↓
Speech Input Node
↓
Intent Classification Node
↓
├─ Simple Chat → Response Node
├─ Tool Required → Tool Node
├─ Unclear → Clarification Node
↓
Memory Update Node
↓
TTS Node
↓
END
Convert speech to text
- Body:
audio(multipart/form-data) - Returns: Transcribed text
Process text through the agent
- Body:
{ "text": "...", "session_id": "..." } - Returns:
{ "response": "...", "intent": "...", "tools_used": [...] }
End-to-end voice chat (STT → Agent → TTS)
- Body:
audio(multipart/form-data),session_id(form data) - Returns: Audio stream (MPEG)
Get conversation memory
- Returns:
{ "session_id": "...", "memory": [...] }
Clear conversation memory
- Search Tool: Web search using DuckDuckGo (free, no API key needed)
- Note Tool: Store, retrieve, and manage notes (in-memory)
- Summarizer Tool: Summarize text content
- Time Tool: Get current time and date
ai-voice-agent/
├── backend/
│ ├── main.py # FastAPI app
│ ├── agent/
│ │ ├── agent.py # LangGraph agent
│ │ └── tools.py # Agent tools
│ └── services/
│ ├── stt.py # Speech-to-text
│ ├── tts.py # Text-to-speech
│ └── memory.py # Memory management
├── app/
│ ├── page.tsx # Main page
│ ├── layout.tsx # Root layout
│ ├── globals.css # Global styles
│ └── components/
│ └── VoiceChat.tsx # Voice chat component
├── requirements.txt # Python dependencies
├── package.json # Node dependencies
└── README.md
| Intent | Action |
|---|---|
casual_chat |
Direct reply |
question |
Reason + answer |
task |
Call tool |
unclear |
Ask follow-up |
All services are FREE and work without API keys!
Optional:
HUGGINGFACE_API_KEY- For better LLM responses (free tier available at https://huggingface.co/settings/tokens)- If not provided, the system uses intelligent rule-based responses
Free Services Used:
- LLM: Hugging Face (free tier) or rule-based fallback
- TTS: Google Text-to-Speech (gTTS) - no API key needed
- Search: DuckDuckGo - no API key needed
- STT: OpenAI Whisper - runs locally, completely free
- Redis integration for persistent memory
- Multiple voice options
- Conversation export
- Custom tool creation
- Multi-language support
- Streaming responses
MIT
- LangChain & LangGraph teams
- OpenAI for Whisper
- ElevenLabs for TTS
- Tavily for search