I build reinforcement learning evaluation environments for frontier models, and the backend systems around them: APIs, agentic pipelines, document intelligence, and full-stack applications. I also contribute upstream to the AI evaluation and data infrastructure projects I depend on.
- AI Research Engineer (RL) @ Tensium (UK) — building reinforcement learning evaluation environments for frontier AI models; designing tasks, verifiers, and sandboxed eval pipelines
- AI & Full Stack Engineer (Contract) @ Atlast — Barcelona (Remote)
- RL evaluation environments for frontier models
- Agentic systems (LangChain / LangGraph, tool-calling agents)
- Scalable backend systems and production LLM pipelines
- Document intelligence, RAG, and OCR
I contribute to AI evaluation and data infrastructure projects, focusing on correctness bugs that fail silently — code that returns the wrong answer without raising an error.
Merged
- Apache DataFusion #25402 — Order-insensitive aggregates carried ordering fields into their partial state schema, so
min(v ORDER BY k)in a grouped query failed with an Arrow schema mismatch. Fixed at the builder so the inconsistent state never exists. - DeepEval #2916 — Three silent benchmark mis-scoring bugs. MathQA's answer schema only allowed
a–dwhile the dataset has five options, making ~20% of it unscoreable even for a perfect model. Also a DROP comma-delimiter corruption, and a BigBenchHard batch path truncating(A)to(A. - DeepEval #2849 —
UnboundLocalErrorcrash when a document chunked to zero pieces, masking the real error message. - InsForge #1786 — Concurrency bug: a credential read racing a cache invalidation repopulated the cache it was meant to clear.
- InsForge #1749 — OpenAI spec compliance: the gateway rejected valid assistant messages carrying
tool_callswithoutcontent, breaking multi-turn agent tool loops. - InsForge #1740 — Token usage reporting for streaming chat completions.
- Opik (Comet ML) #8120 — Removed an orphaned CI workflow that was broken on the project's default branch.
Co-authored — landed inside maintainers' release PRs
- Instructor #2597 — Docs lint failures, 34 down to 18. Shipped in Instructor 1.17.0.
- TraceRoot #1593 — Anthropic model pricing: fast-mode rate cards and dot-notation model IDs.
Open — PRs in review across Langfuse, LiteLLM, EleutherAI's lm-evaluation-harness, Hugging Face evaluate, Future AGI, Parea, Ragas, and Graphify.
Backend & AI
Frontend
Data & Infra
- Build RL evaluation environments for frontier models
- Design tasks, verifiers, and sandboxed eval pipelines
- Tested against Claude and GPT on the HUD platform
Aug 2026 – Present
- Built a financial document intelligence platform with Azure Document Intelligence, FastAPI, OCR, and structured extraction
- Integrated LLM chatbots, AWS Cognito authentication, and Stripe payments
- Collaborated with global AI teams including Anthropic on AI model training and evaluation
- Designed scalable backend services, database schemas, and FastAPI-based APIs for production AI infrastructure
- Evaluated and optimized AI models
BS Software Engineering — FAST National University of Computer and Emerging Sciences, Lahore, Pakistan
- Portfolio: https://imhassaan04.vercel.app
- LinkedIn: https://linkedin.com/in/imhassaan04
- GitHub: https://github.com/hassaanch23
- Email: imhassaan04@gmail.com
Open to collaborations in AI, backend, and full-stack development.


