Skip to content
View hassaanch23's full-sized avatar

Block or report hassaanch23

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
hassaanch23/README.md

Hi, I'm Muhammad Hassaan

AI Research Engineer & Full Stack Developer

I build reinforcement learning evaluation environments for frontier models, and the backend systems around them: APIs, agentic pipelines, document intelligence, and full-stack applications. I also contribute upstream to the AI evaluation and data infrastructure projects I depend on.

Profile views LinkedIn Portfolio


Currently

  • AI Research Engineer (RL) @ Tensium (UK) — building reinforcement learning evaluation environments for frontier AI models; designing tasks, verifiers, and sandboxed eval pipelines
  • AI & Full Stack Engineer (Contract) @ Atlast — Barcelona (Remote)

Focus Areas

  • RL evaluation environments for frontier models
  • Agentic systems (LangChain / LangGraph, tool-calling agents)
  • Scalable backend systems and production LLM pipelines
  • Document intelligence, RAG, and OCR

Open Source Contributions

I contribute to AI evaluation and data infrastructure projects, focusing on correctness bugs that fail silently — code that returns the wrong answer without raising an error.

Merged

  • Apache DataFusion #25402 — Order-insensitive aggregates carried ordering fields into their partial state schema, so min(v ORDER BY k) in a grouped query failed with an Arrow schema mismatch. Fixed at the builder so the inconsistent state never exists.
  • DeepEval #2916 — Three silent benchmark mis-scoring bugs. MathQA's answer schema only allowed a–d while the dataset has five options, making ~20% of it unscoreable even for a perfect model. Also a DROP comma-delimiter corruption, and a BigBenchHard batch path truncating (A) to (A.
  • DeepEval #2849 — UnboundLocalError crash when a document chunked to zero pieces, masking the real error message.
  • InsForge #1786 — Concurrency bug: a credential read racing a cache invalidation repopulated the cache it was meant to clear.
  • InsForge #1749 — OpenAI spec compliance: the gateway rejected valid assistant messages carrying tool_calls without content, breaking multi-turn agent tool loops.
  • InsForge #1740 — Token usage reporting for streaming chat completions.
  • Opik (Comet ML) #8120 — Removed an orphaned CI workflow that was broken on the project's default branch.

Co-authored — landed inside maintainers' release PRs

  • Instructor #2597 — Docs lint failures, 34 down to 18. Shipped in Instructor 1.17.0.
  • TraceRoot #1593 — Anthropic model pricing: fast-mode rate cards and dot-notation model IDs.

Open — PRs in review across Langfuse, LiteLLM, EleutherAI's lm-evaluation-harness, Hugging Face evaluate, Future AGI, Parea, Ragas, and Graphify.


Tech Stack

Backend & AI

Python FastAPI TypeScript NestJS Node.js LangChain Jest

Frontend

React Next.js Vue.js Nuxt Tailwind CSS Vite Vitest

Data & Infra

PostgreSQL MongoDB Prisma Redis Docker Terraform AWS Sentry


Experience

AI Research Engineer (RL) — Tensium (UK)

  • Build RL evaluation environments for frontier models
  • Design tasks, verifiers, and sandboxed eval pipelines
  • Tested against Claude and GPT on the HUD platform

AI & Full Stack Engineer (Contract) — Atlast, Barcelona (Remote)

Aug 2026 – Present

AI Developer — Techfy, Lahore, Pakistan (Remote)

  • Built a financial document intelligence platform with Azure Document Intelligence, FastAPI, OCR, and structured extraction
  • Integrated LLM chatbots, AWS Cognito authentication, and Stripe payments

Software Engineer — Mercor

  • Collaborated with global AI teams including Anthropic on AI model training and evaluation
  • Designed scalable backend services, database schemas, and FastAPI-based APIs for production AI infrastructure

Software Engineer — AfterQuery

  • Evaluated and optimized AI models

Education

BS Software Engineering — FAST National University of Computer and Emerging Sciences, Lahore, Pakistan


GitHub Analytics

GitHub stats Top languages GitHub trophies

Connect

Open to collaborations in AI, backend, and full-stack development.

Pinned Loading

  1. zohaibfast99/MedQuick-App zohaibfast99/MedQuick-App Public

    Cross-platform AI-powered healthcare assistance and medicine delivery system

    Dart 1

  2. iron-fit iron-fit Public

    TypeScript

  3. RecallAI RecallAI Public

    Swift

  4. BRICKnCLICK BRICKnCLICK Public

    A property listing platform

    JavaScript

  5. ai-voice-agent ai-voice-agent Public

    Python

  6. project-sentiment-analysis-of-custumer-reviews project-sentiment-analysis-of-custumer-reviews Public

    Jupyter Notebook