Shipping AI is easy.
Knowing it works is hard.
I help AI teams build evaluation systems, agent benchmarks, and quality gates before broken workflows reach customers.
Verified proof points
Verified receipts
lm-evaluation-harness (EleutherAI) — fail-fast guard against silent MoE regressions #3734
CL-Bench (Tencent, published) — benchmark-task delivery arXiv:2602.03587 arXiv:2604.27043 case study
matplotlib · pytest · Sphinx · astropy — 8 merged PRs GitHub
Every claim links to the merged PR or write-up. All receipts →
Three ways I can help
Productized offers with clear scope, timelines, and deliverables. No vague "AI consulting." You know exactly what you get.
What you get
- Failure mode analysis across your AI workflows
- Custom evaluation dataset with judge rubrics
- Pass/fail scoring logic and regression baseline
- Repeatable quality scorecard you can run on every release
Teams about to launch an AI feature and need to know if it actually works.
What you get
- Scenario-based benchmark suite (multi-step, tool-use)
- MCP server or API harness with deterministic validation
- LLM-as-a-Judge rubric scoring for open-ended tasks
- CI-compatible regression testing and comparison dashboard
Teams shipping AI agents with tool-use, multi-step reasoning, or complex orchestration.
What you get
- Production-ready AI-powered application
- Built-in evaluation and quality monitoring
- High test coverage (94%+ target)
- Deployed with CI/CD and observability
Solo founders and small teams who need speed without sacrificing quality.

Fazeel Usmani
AI Evaluation Engineer with 8+ years of experience building production systems. I got my start introducing accessibility features to Google Maps (impacting 52M+ users, invited to Google HQ twice), then spent years scaling enterprise platforms and shipping evaluation infrastructure for frontier AI models.
Today I lead benchmark and evaluation work at Turing, where I managed the team that delivered tasks for Tencent's published CL-bench (arXiv:2602.03587 and CL-Bench Life) and solo-built MBPI, a 25K-line platform that surfaces failure modes across GPT-5.1, Claude 4.5, Gemini 3 Pro, and Hunyuan at up to 128K API calls/day.
What I specialize in
AI Evaluation & Benchmarking
Built LLM-as-a-Judge systems processing 100K+ evaluations/day. Designed SWE-Bench+ extensions and multi-verifier evaluation pipelines.
Agent Testing & MCP Servers
Architected a 43-tool MCP server with 78 benchmarks for SAP S/4HANA. Expertise in tool-use evaluation, multi-step orchestration testing.
Open-Source Contributor
Merged PRs into matplotlib, pytest, Sphinx, and astropy. Arctic Code Vault Contributor. Active in evaluation-related repos.
Production Engineering
Scaled WorkSpan search 10,000x (2K to 20M users). Built BookAFarm.com to 10K+ MAU with 94.3% test coverage. Ships in 14-day sprints.
Career highlights
Delivery Engineering Manager
Turing · Jun 2024 – Present
Leading 6 teams (30+ members). Shipped benchmark tasks into Tencent’s CL-bench. Built MBPI for automated LLM quality assurance.
Senior Software Engineer
WorkSpan · Jul 2021 – Dec 2023
Led Salesforce AppExchange integrations, onboarding 84+ enterprises. Scaled search 10,000x.
Google Maps Contributor
Google · 2016 – 2019
Introduced accessibility features impacting 52M+ users. Invited to Google HQ twice to collaborate with Maps engineering.
Tech stack
Education
B.E. in Information Technology · Osmania University · 2014–2018
Open Source Contributions
lm-evaluation-harness
EleutherAI · 12.5k★
AI evaluation framework
- #3734: Fail-fast guard preventing a silent ∼17× MoE accuracy regression
Python scientific stack
8 PRs
- matplotlib: #30746, #30756, #30760, #30795
- pytest: #13930, #13954
- Sphinx: #14046
- astropy: #18861
Common questions
Things founders usually ask before we start working together.
Can you work async with US teams?
Yes. I’m based in Hyderabad (IST) and regularly overlap with US Eastern and Pacific time zones. Most of my current work is fully async with weekly syncs.
Do you sign NDAs?
Absolutely. I sign NDAs before every engagement. I’m also happy to work under your standard contractor agreement.
Can you work with our existing codebase?
Yes. I’ve integrated evaluation harnesses into existing Python, TypeScript, and Go codebases. I adapt to your stack, not the other way around.
Can you help us compare GPT / Claude / Gemini?
That’s exactly what MBPI does. I can build a multi-model benchmarking harness with per-model Pass@1 metrics, side-by-side comparison, and failure-mode analysis.
How long does a typical sprint take?
An Evaluation Sprint is 10–14 days. An Agent Benchmark Harness build is 4–8 weeks. A Rapid MVP is 2–4 weeks. We scope it together on the first call.
What if I’m not sure what I need?
That’s what the free call is for. I’ll listen to what you’re building and give you an honest recommendation — even if it’s "you don’t need me yet."