
Fazeel Usmani
AI Evaluation Engineer with 8+ years of experience building production systems. I got my start introducing accessibility features to Google Maps (impacting 52M+ users, invited to Google HQ twice), then spent years scaling enterprise platforms and shipping evaluation infrastructure for frontier AI models.
Today I lead benchmark and evaluation work at Turing, where I managed the team that delivered tasks for Tencent's published CL-bench (arXiv:2602.03587 and CL-Bench Life) and solo-built MBPI, a 25K-line platform that surfaces failure modes across GPT-5.1, Claude 4.5, Gemini 3 Pro, and Hunyuan at up to 128K API calls/day.
What I specialize in
AI Evaluation & Benchmarking
Built LLM-as-a-Judge systems processing 100K+ evaluations/day. Designed SWE-Bench+ extensions and multi-verifier evaluation pipelines.
Agent Testing & MCP Servers
Architected a 43-tool MCP server with 78 benchmarks for SAP S/4HANA. Expertise in tool-use evaluation, multi-step orchestration testing.
Open-Source Contributor
Merged PRs into matplotlib, pytest, Sphinx, and astropy. Arctic Code Vault Contributor. Active in evaluation-related repos.
Production Engineering
Scaled WorkSpan search 10,000x (2K to 20M users). Built BookAFarm.com to 10K+ MAU with 94.3% test coverage. Ships in 14-day sprints.
Career highlights
Delivery Engineering Manager
Turing · Jun 2024 – Present
Leading 6 teams (30+ members). Shipped benchmark tasks into Tencent’s CL-bench. Built MBPI for automated LLM quality assurance.
Senior Software Engineer
WorkSpan · Jul 2021 – Dec 2023
Led Salesforce AppExchange integrations, onboarding 84+ enterprises. Scaled search 10,000x.
Google Maps Contributor
Google · 2016 – 2019
Introduced accessibility features impacting 52M+ users. Invited to Google HQ twice to collaborate with Maps engineering.
Tech stack
Education
B.E. in Information Technology · Osmania University · 2014–2018
Open Source Contributions
lm-evaluation-harness
EleutherAI · 12.5k★
AI evaluation framework
- #3734: Fail-fast guard preventing a silent ∼17× MoE accuracy regression
Python scientific stack
8 PRs
- matplotlib: #30746, #30756, #30760, #30795
- pytest: #13930, #13954
- Sphinx: #14046
- astropy: #18861