Skip to content

About

design and implement an Agentic AI system that autonomously profiles, cleans, and preprocesses raw structured or semi-structured data by dynamically detecting data quality issues, selecting optimal transformation strategies, executing corrective actions, and validating the improved dataset for downstream analytics or machine learning tasks.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

4 Commits

Folders and files

Repository files navigation

Agentic Data System

Overview

The Agentic Data System is an autonomous AI system designed for comprehensive data ingestion, profiling, cleaning, preprocessing, validation, and insight generation. It operates without human intervention, leveraging a multi-agent reasoning and reflection loop to ensure high data quality, explainability, and reliability. This system is domain-agnostic and built for industry-grade scalability.

Features

• Multi-Agent Architecture: Composed of specialized agents for each stage of the data pipeline.

• Autonomous Data Processing: Automatically detects and fixes data quality issues, applies optimal preprocessing strategies, and validates improvements.

• Broad File Format Support: Handles CSV, Excel (XLS/XLSX), JSON, XML, SQL/Database Tables, TXT (Plain Text), LOG Files, and PDF (structured, semi-structured, and unstructured content).

• Detailed Reporting: Generates a comprehensive Markdown report detailing all actions, measurable quality metrics, and generated insights.

• Streamlit Web Interface: Provides an intuitive web application for easy file uploads, real-time processing feedback, and cleaned data downloads.

• Reflection & Optimization Loop: Iteratively refines data cleaning strategies to achieve a target data quality score.

Agentic Architecture

The system is built upon a robust multi-agent architecture:

Data Ingestion Agent: Automatically detects file type and structure, extracts data, and converts all inputs into a unified internal representation (tabular or document-based).

Data Profiling Agent: Infers schema and data types, detects missing values, duplicates, outliers, and inconsistencies, validates ranges, constraints, and formats, and generates statistical summaries.

Issue Detection Agent: Identifies and categorizes missing/null values, duplicate records, invalid ranges/formats, incorrect data types, inconsistent labels/encodings, noise, corrupted/irrelevant rows, and domain-independent anomalies.

Planning Agent: Autonomously decides optimal cleaning strategies, choosing between imputation, removal, normalization, standardization, transformation, or preservation, minimizing unnecessary data loss.

Execution Agent: Applies transformations such as missing value imputation, deduplication, data type correction, normalization, standardization, text cleaning, schema alignment, and outlier handling.

Validation Agent: Re-profiles the cleaned dataset, compares before vs. after quality metrics, verifies schema integrity, checks statistical consistency, and flags unresolved or residual issues.

Reflection & Optimization Loop: Evaluates the effectiveness of applied strategies, refines techniques if quality thresholds are not met, and iterates until acceptable data quality is achieved, logging all decisions transparently.

Installation

To set up the Agentic Data System, follow these steps:

Clone the repository (or download the project files):

Bash

git clone <repository_url> cd agentic-data-system

Create a virtual environment (recommended):

Bash

python3 -m venv venv source venv/bin/activate

Install dependencies:

Bash

pip install -r requirements.txt

Usage

Command-Line Interface (CLI)

YouYou can run the core agentic system directly from the command line:

Bash

python3 agentic_data_system.py <path_to_your_dataset>

If no dataset path is provided, the system will generate and process a sample CSV file for demonstration.

Streamlit Web Application

For an interactive experience, use the Streamlit web interface:

Start the Streamlit application:

Bash

streamlit run app.py

Access the application: Open your web browser and navigate to the URL provided by Streamlit (usually http://localhost:8501 ).

Upload and Process: Use the web interface to upload your dataset, adjust the target quality score, and initiate the data processing. View the report, cleaned data preview, and download the results directly from the dashboard.

Project Structure

• agentic_data_system.py: Contains the core logic for all the agents and the orchestrator.

• app.py: The Streamlit web application interface.

• requirements.txt: Lists all Python dependencies required for the project.

• README.md: This documentation file.

Contributing

Contributions are welcome! Please feel free to fork the repository, open issues, or submit pull requests.

.

Generated by Manus AI

About

design and implement an Agentic AI system that autonomously profiles, cleans, and preprocesses raw structured or semi-structured data by dynamically detecting data quality issues, selecting optimal transformation strategies, executing corrective actions, and validating the improved dataset for downstream analytics or machine learning tasks.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages