Shipping AI is easy.
Knowing it works is hard.

I help AI teams build evaluation systems, agent benchmarks, and quality gates before broken workflows reach customers.

Verified proof points

Contributed benchmark-task delivery for Tencent’s published CL-bench
Built a 43-tool MCP benchmark system with 78 agent scenarios
Solo-built MBPI — 25K-line platform handling 128K API calls/day
Open-source contributor to matplotlib, pytest, Sphinx & astropy
Scaled WorkSpan search from 2K to 20M users (10,000x)

Three ways I can help

Productized offers with clear scope, timelines, and deliverables. No vague "AI consulting." You know exactly what you get.

From $7,500
AI Evaluation Sprint
10–14 days
Define failure modes, build your first evaluation set, create rubrics, and leave with a repeatable scorecard for your AI product.

What you get

  • Failure mode analysis across your AI workflows
  • Custom evaluation dataset with judge rubrics
  • Pass/fail scoring logic and regression baseline
  • Repeatable quality scorecard you can run on every release

Teams about to launch an AI feature and need to know if it actually works.

From $15,000
Agent Benchmark Harness
4–8 weeks
Custom benchmark workflows, MCP/tool-use scenarios, pass/fail logic, judge rubrics, and a dashboard for regression testing your AI agents.

What you get

  • Scenario-based benchmark suite (multi-step, tool-use)
  • MCP server or API harness with deterministic validation
  • LLM-as-a-Judge rubric scoring for open-ended tasks
  • CI-compatible regression testing and comparison dashboard

Teams shipping AI agents with tool-use, multi-step reasoning, or complex orchestration.

From $5,000
Rapid AI MVP
2–4 weeks
For founders who need to ship a working AI product fast — but with testing and quality instrumentation built in from day one.

What you get

  • Production-ready AI-powered application
  • Built-in evaluation and quality monitoring
  • High test coverage (94%+ target)
  • Deployed with CI/CD and observability

Solo founders and small teams who need speed without sacrificing quality.

Not sure which offer fits?

Book a free call. I'll listen to what you're building and tell you honestly whether I can help.

About the Founder
Fazeel Usmani

Fazeel Usmani

AI Evaluation Engineer with 8+ years of experience building production systems. I got my start introducing accessibility features to Google Maps (impacting 52M+ users, invited to Google HQ twice), then spent years scaling enterprise platforms and shipping evaluation infrastructure for frontier AI models.

Today I lead benchmark and evaluation work at Turing, where I managed the team that delivered tasks for Tencent's published CL-bench (arXiv:2602.03587 and CL-Bench Life) and solo-built MBPI, a 25K-line platform that surfaces failure modes across GPT-5.1, Claude 4.5, Gemini 3 Pro, and Hunyuan at up to 128K API calls/day.

Hyderabad, India
US timezone overlap available

What I specialize in

AI Evaluation & Benchmarking

Built LLM-as-a-Judge systems processing 100K+ evaluations/day. Designed SWE-Bench+ extensions and multi-verifier evaluation pipelines.

Agent Testing & MCP Servers

Architected a 43-tool MCP server with 78 benchmarks for SAP S/4HANA. Expertise in tool-use evaluation, multi-step orchestration testing.

Open-Source Contributor

Merged PRs into matplotlib, pytest, Sphinx, and astropy. Arctic Code Vault Contributor. Active in evaluation-related repos.

Production Engineering

Scaled WorkSpan search 10,000x (2K to 20M users). Built BookAFarm.com to 10K+ MAU with 94.3% test coverage. Ships in 14-day sprints.

Career highlights

Delivery Engineering Manager

Turing · Jun 2024 – Present

Leading 6 teams (30+ members). Shipped benchmark tasks into Tencent’s CL-bench. Built MBPI for automated LLM quality assurance.

Senior Software Engineer

WorkSpan · Jul 2021 – Dec 2023

Led Salesforce AppExchange integrations, onboarding 84+ enterprises. Scaled search 10,000x.

Google Maps Contributor

Google · 2016 – 2019

Introduced accessibility features impacting 52M+ users. Invited to Google HQ twice to collaborate with Maps engineering.

Tech stack

Python
TypeScript
Rust
Go
PyTorch
OpenAI API
Anthropic Claude
Gemini
Docker
Kubernetes
AWS
GCP
PostgreSQL
Redis
Kafka

Education

B.E. in Information Technology · Osmania University · 2014–2018

Open Source Contributions

inspect_ai

UK AI Safety Institute · 2k★ · 5 PRs, hand-merged by JJ Allaire

AI evaluation framework

  • #3668: vLLM provider restart-after-close fix (system failure analysis)
  • #3690: Judge-prompt injection hardening (22 tests)
  • #3722: Anthropic Opus 4.7 sampling-parameter guard
  • #3794: Pluggable URI resolver registry for custom scheme handling
  • #3836: HuggingFace dataset retry on transient errors

lm-evaluation-harness

EleutherAI · 12.5k★

AI evaluation framework

  • #3734: Fail-fast guard preventing a silent ∼17× MoE accuracy regression

Python scientific stack

8 PRs

Common questions

Things founders usually ask before we start working together.

Can you work async with US teams?

Yes. I’m based in Hyderabad (IST) and regularly overlap with US Eastern and Pacific time zones. Most of my current work is fully async with weekly syncs.

Do you sign NDAs?

Absolutely. I sign NDAs before every engagement. I’m also happy to work under your standard contractor agreement.

Can you work with our existing codebase?

Yes. I’ve integrated evaluation harnesses into existing Python, TypeScript, and Go codebases. I adapt to your stack, not the other way around.

Can you help us compare GPT / Claude / Gemini?

That’s exactly what MBPI does. I can build a multi-model benchmarking harness with per-model Pass@1 metrics, side-by-side comparison, and failure-mode analysis.

How long does a typical sprint take?

An Evaluation Sprint is 10–14 days. An Agent Benchmark Harness build is 4–8 weeks. A Rapid MVP is 2–4 weeks. We scope it together on the first call.

What if I’m not sure what I need?

That’s what the free call is for. I’ll listen to what you’re building and give you an honest recommendation — even if it’s "you don’t need me yet."

Ready to know if your AI actually works?

Book a free call. No pitch deck, no sales script. Just a conversation about what you're building and whether I can help.

fazeelusmani18@gmail.com · Hyderabad, India · US timezone overlap available