Turn every conversation into actionable intelligence.
MeetMind AI is an AI-powered conversation and video intelligence platform that transforms meetings, lectures, interviews, and other recorded conversations into structured, searchable knowledge.
By automating the tedious process of transcribing and summarizing, MeetMind AI helps you focus on what actually matters in your sessions.
Capabilities:
- Media transcription (Audio/Video formats)
- English transcription with Whisper
- Hinglish transcription with Sarvam AI
- AI-generated summaries
- Action item extraction
- Decision extraction
- Open question extraction
- Searchable transcript
- RAG-powered "Ask MeetMind"
- PDF & TXT export
- Session history
- Session-isolated knowledge retrieval
🎙️ Multi-format Ingestion Supports common audio and video formats including MP3, MP4, MOV, MKV, WAV, and more.
📝 AI Transcription Powered by OpenAI's Whisper for English and Sarvam AI for Hinglish capabilities.
🧠 Conversation Intelligence Automatically extracts high-level summaries, action items, key decisions, and open questions.
🔎 Ask MeetMind Ask questions against the current session using Retrieval-Augmented Generation (RAG).
🔐 Session Isolation Knowledge retrieval is scoped strictly to the individual session being analyzed.
📄 Export Export your analyzed sessions cleanly as TXT or PDF documents for easy sharing.
📊 Session Workspace A premium dashboard to view your transcript, insights, and AI conversation history interactively.
Frontend
- Next.js (App Router)
- React
- TypeScript
- Tailwind CSS & Framer Motion
Backend
- FastAPI
- SQLAlchemy
- SQLite
AI Pipeline
- OpenAI Whisper
- Sarvam AI
- Mistral / Codestral
Knowledge Retrieval
- Chroma (Vector DB)
- HuggingFace embeddings
Media Processing
- FFmpeg
- Pydub
flowchart TD
User([User]) --> Frontend(Next.js Frontend)
Frontend --> Backend(FastAPI)
subgraph Session Processing
Backend --> Media[FFmpeg / Pydub]
Media --> Transcribe[Whisper / Sarvam]
Transcribe --> Infer[Mistral / Codestral]
Transcribe --> Index[Chroma]
end
Infer --> SQLite[(SQLite Session Knowledge)]
Index --> SQLite
SQLite --> Chat[Ask MeetMind]
Chat --> User
meetmind-ai/
├── frontend/ # Next.js Application
│ ├── app/ # Routes (Dashboard, Sessions, Upload)
│ ├── components/ # Reusable UI Components
│ └── lib/ # API and Utils
├── backend/ # FastAPI Application
│ ├── api/ # API Route Handlers
│ ├── database/ # SQLAlchemy Models & Connection
│ └── services/ # Background Processing Logic
├── core/ # AI Intelligence Core
│ ├── transcriber.py # Whisper & Sarvam Logic
│ ├── summarizer.py # Summary Generation
│ ├── extractor.py # Insights Extraction
│ ├── vector_store.py # ChromaDB Indexing
│ └── rag_engine.py # Session Chat
├── utils/ # Audio Processing & Helpers
├── downloads/ # Runtime Media Downloads (Ignored)
├── temp_uploads/ # Runtime Local Uploads (Ignored)
└── vector_db/ # Runtime ChromaDB Storage (Ignored)
- Python 3.12+
- Node.js & npm (v20+)
- FFmpeg installed and added to your system PATH.
Create a .env file in the root directory (do not commit this):
MISTRAL_API_KEY=your_mistral_key_here
SARVAM_API_KEY=your_sarvam_key_here
WHISPER_MODEL=small
SARVAM_STT_MODEL=saaras:v2.5Create a .env.local file inside the frontend/ directory:
NEXT_PUBLIC_API_URL=http://localhost:8000Backend: Open PowerShell or your preferred terminal:
python -m venv .venv
.\.venv\Scripts\activate
pip install -r requirements.txtFrontend: Open a new terminal window:
cd frontend
npm installStart the Backend:
# Ensure your virtual environment is active
python -m uvicorn backend.main:app --reload --port 8000Start the Frontend:
cd frontend
npm run devThe application will be available at http://localhost:3000.
- Local Media:
.mp4,.mov,.mkv,.avi,.webm,.mp3,.wav,.m4a(Max 200MB) - URLs: Public YouTube URLs
- Upload: User creates a session by uploading local media or providing a YouTube URL.
- Media Prep: FFmpeg extracts and formats the audio into standardized WAV chunks.
- Transcription: The chunks are processed by Whisper (English) or Sarvam AI (Hinglish).
- Analysis: The AI core analyzes the raw transcript to generate summaries and extract tasks/decisions.
- Knowledge Base: The transcript is embedded into ChromaDB for semantic search.
- Interaction: The user inspects the insights on the dashboard and asks follow-up questions using the "Ask MeetMind" chat.
- Export: The user can export the analyzed session as a polished PDF or TXT file.
The FastAPI backend exposes the following documented endpoints:
System
GET /api/health
Sessions
POST /api/sessions- Create a new empty sessionGET /api/sessions- List all sessionsGET /api/sessions/{session_id}- Retrieve session details and insightsGET /api/sessions/{session_id}/status- Check background processing status
Ingestion
POST /api/sessions/{session_id}/upload- Upload local mediaPOST /api/sessions/{session_id}/youtube- Submit YouTube URL
Interactive Chat
GET /api/sessions/{session_id}/chat- Get session chat historyPOST /api/sessions/{session_id}/chat- Send a RAG query to the session
Export
GET /api/sessions/{session_id}/export/txt- Download session as TXTGET /api/sessions/{session_id}/export/pdf- Download session as PDF
- API Keys: Stored securely in
.envvariables and never exposed to the frontend. - Session Isolation: The knowledge base (Chroma DB) strictly scopes vector retrieval using the
session_id. Users cannot query data across different sessions. - Upload Validation: Uploaded files are strictly validated against size limits and supported MIME types before processing.
- FFmpeg not found: Ensure FFmpeg is installed and added to your system's Environment Variables (
PATH). Restart your terminal after installing. - Backend not running: Ensure you activated the virtual environment before starting Uvicorn. Check that your port
8000is free. - Frontend cannot reach backend: Ensure
NEXT_PUBLIC_API_URLis properly set infrontend/.env.local. - Whisper model downloading slowly: The first run will securely download the model weights (e.g.,
small.pt) to your cache. This may take a few minutes depending on your internet connection. - PDF Export failing: Ensure the session has reached the "completed" status. Exports cannot be generated for failed or processing sessions.
This is an active development/portfolio project. It is intended for local usage and demonstration purposes and is not yet configured for production cloud deployment.
MeetMind AI is an independently evolved project built from an existing open-source AI video assistant foundation (originally the AI Video Assistant project). The project has been substantially reworked and extended with new product identity, input workflows, UX, and advanced capabilities, while preserving the MIT license terms of the original foundation.