Skip to content

About

AI-powered meeting & video intelligence platform that transcribes conversations, extracts insights, and enables RAG-powered Q&A with session-isolated knowledge.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

MeetMind AI

Turn every conversation into actionable intelligence.

Overview

MeetMind AI is an AI-powered conversation and video intelligence platform that transforms meetings, lectures, interviews, and other recorded conversations into structured, searchable knowledge.

By automating the tedious process of transcribing and summarizing, MeetMind AI helps you focus on what actually matters in your sessions.

Capabilities:

  • Media transcription (Audio/Video formats)
  • English transcription with Whisper
  • Hinglish transcription with Sarvam AI
  • AI-generated summaries
  • Action item extraction
  • Decision extraction
  • Open question extraction
  • Searchable transcript
  • RAG-powered "Ask MeetMind"
  • PDF & TXT export
  • Session history
  • Session-isolated knowledge retrieval

Features

🎙️ Multi-format Ingestion Supports common audio and video formats including MP3, MP4, MOV, MKV, WAV, and more.

▶️ YouTube Ingestion Analyze and extract knowledge directly from supported YouTube URLs.

📝 AI Transcription Powered by OpenAI's Whisper for English and Sarvam AI for Hinglish capabilities.

🧠 Conversation Intelligence Automatically extracts high-level summaries, action items, key decisions, and open questions.

🔎 Ask MeetMind Ask questions against the current session using Retrieval-Augmented Generation (RAG).

🔐 Session Isolation Knowledge retrieval is scoped strictly to the individual session being analyzed.

📄 Export Export your analyzed sessions cleanly as TXT or PDF documents for easy sharing.

📊 Session Workspace A premium dashboard to view your transcript, insights, and AI conversation history interactively.


Architecture

Frontend

  • Next.js (App Router)
  • React
  • TypeScript
  • Tailwind CSS & Framer Motion

Backend

  • FastAPI
  • SQLAlchemy
  • SQLite

AI Pipeline

  • OpenAI Whisper
  • Sarvam AI
  • Mistral / Codestral

Knowledge Retrieval

  • Chroma (Vector DB)
  • HuggingFace embeddings

Media Processing

  • FFmpeg
  • Pydub
flowchart TD
    User([User]) --> Frontend(Next.js Frontend)
    Frontend --> Backend(FastAPI)
    
    subgraph Session Processing
        Backend --> Media[FFmpeg / Pydub]
        Media --> Transcribe[Whisper / Sarvam]
        Transcribe --> Infer[Mistral / Codestral]
        Transcribe --> Index[Chroma]
    end
    
    Infer --> SQLite[(SQLite Session Knowledge)]
    Index --> SQLite
    
    SQLite --> Chat[Ask MeetMind]
    Chat --> User
Loading

Project Structure

meetmind-ai/
├── frontend/             # Next.js Application
│   ├── app/              # Routes (Dashboard, Sessions, Upload)
│   ├── components/       # Reusable UI Components
│   └── lib/              # API and Utils
├── backend/              # FastAPI Application
│   ├── api/              # API Route Handlers
│   ├── database/         # SQLAlchemy Models & Connection
│   └── services/         # Background Processing Logic
├── core/                 # AI Intelligence Core
│   ├── transcriber.py    # Whisper & Sarvam Logic
│   ├── summarizer.py     # Summary Generation
│   ├── extractor.py      # Insights Extraction
│   ├── vector_store.py   # ChromaDB Indexing
│   └── rag_engine.py     # Session Chat
├── utils/                # Audio Processing & Helpers
├── downloads/            # Runtime Media Downloads (Ignored)
├── temp_uploads/         # Runtime Local Uploads (Ignored)
└── vector_db/            # Runtime ChromaDB Storage (Ignored)

Setup & Installation

Prerequisites

  • Python 3.12+
  • Node.js & npm (v20+)
  • FFmpeg installed and added to your system PATH.

1. Backend Environment Setup

Create a .env file in the root directory (do not commit this):

MISTRAL_API_KEY=your_mistral_key_here
SARVAM_API_KEY=your_sarvam_key_here
WHISPER_MODEL=small
SARVAM_STT_MODEL=saaras:v2.5

2. Frontend Environment Setup

Create a .env.local file inside the frontend/ directory:

NEXT_PUBLIC_API_URL=http://localhost:8000

3. Installation

Backend: Open PowerShell or your preferred terminal:

python -m venv .venv
.\.venv\Scripts\activate
pip install -r requirements.txt

Frontend: Open a new terminal window:

cd frontend
npm install

Running the Application

Start the Backend:

# Ensure your virtual environment is active
python -m uvicorn backend.main:app --reload --port 8000

Start the Frontend:

cd frontend
npm run dev

The application will be available at http://localhost:3000.


Supported Inputs

  • Local Media: .mp4, .mov, .mkv, .avi, .webm, .mp3, .wav, .m4a (Max 200MB)
  • URLs: Public YouTube URLs

How It Works

  1. Upload: User creates a session by uploading local media or providing a YouTube URL.
  2. Media Prep: FFmpeg extracts and formats the audio into standardized WAV chunks.
  3. Transcription: The chunks are processed by Whisper (English) or Sarvam AI (Hinglish).
  4. Analysis: The AI core analyzes the raw transcript to generate summaries and extract tasks/decisions.
  5. Knowledge Base: The transcript is embedded into ChromaDB for semantic search.
  6. Interaction: The user inspects the insights on the dashboard and asks follow-up questions using the "Ask MeetMind" chat.
  7. Export: The user can export the analyzed session as a polished PDF or TXT file.

API Overview

The FastAPI backend exposes the following documented endpoints:

System

  • GET /api/health

Sessions

  • POST /api/sessions - Create a new empty session
  • GET /api/sessions - List all sessions
  • GET /api/sessions/{session_id} - Retrieve session details and insights
  • GET /api/sessions/{session_id}/status - Check background processing status

Ingestion

  • POST /api/sessions/{session_id}/upload - Upload local media
  • POST /api/sessions/{session_id}/youtube - Submit YouTube URL

Interactive Chat

  • GET /api/sessions/{session_id}/chat - Get session chat history
  • POST /api/sessions/{session_id}/chat - Send a RAG query to the session

Export

  • GET /api/sessions/{session_id}/export/txt - Download session as TXT
  • GET /api/sessions/{session_id}/export/pdf - Download session as PDF

Security & Privacy

  • API Keys: Stored securely in .env variables and never exposed to the frontend.
  • Session Isolation: The knowledge base (Chroma DB) strictly scopes vector retrieval using the session_id. Users cannot query data across different sessions.
  • Upload Validation: Uploaded files are strictly validated against size limits and supported MIME types before processing.

Troubleshooting

  • FFmpeg not found: Ensure FFmpeg is installed and added to your system's Environment Variables (PATH). Restart your terminal after installing.
  • Backend not running: Ensure you activated the virtual environment before starting Uvicorn. Check that your port 8000 is free.
  • Frontend cannot reach backend: Ensure NEXT_PUBLIC_API_URL is properly set in frontend/.env.local.
  • Whisper model downloading slowly: The first run will securely download the model weights (e.g., small.pt) to your cache. This may take a few minutes depending on your internet connection.
  • PDF Export failing: Ensure the session has reached the "completed" status. Exports cannot be generated for failed or processing sessions.

Project Status

This is an active development/portfolio project. It is intended for local usage and demonstration purposes and is not yet configured for production cloud deployment.

Attribution & License

MeetMind AI is an independently evolved project built from an existing open-source AI video assistant foundation (originally the AI Video Assistant project). The project has been substantially reworked and extended with new product identity, input workflows, UX, and advanced capabilities, while preserving the MIT license terms of the original foundation.

About

AI-powered meeting & video intelligence platform that transcribes conversations, extracts insights, and enables RAG-powered Q&A with session-isolated knowledge.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages