About the Founder
Fazeel Usmani

Fazeel Usmani

AI Evaluation Engineer with 8+ years of experience building production systems. I got my start introducing accessibility features to Google Maps (impacting 52M+ users, invited to Google HQ twice), then spent years scaling enterprise platforms and shipping evaluation infrastructure for frontier AI models.

Today I lead benchmark and evaluation work at Turing, where I managed the team that delivered tasks for Tencent's published CL-bench (arXiv:2602.03587 and CL-Bench Life) and solo-built MBPI, a 25K-line platform that surfaces failure modes across GPT-5.1, Claude 4.5, Gemini 3 Pro, and Hunyuan at up to 128K API calls/day.

Hyderabad, India
US timezone overlap available

What I specialize in

AI Evaluation & Benchmarking

Built LLM-as-a-Judge systems processing 100K+ evaluations/day. Designed SWE-Bench+ extensions and multi-verifier evaluation pipelines.

Agent Testing & MCP Servers

Architected a 43-tool MCP server with 78 benchmarks for SAP S/4HANA. Expertise in tool-use evaluation, multi-step orchestration testing.

Open-Source Contributor

Merged PRs into matplotlib, pytest, Sphinx, and astropy. Arctic Code Vault Contributor. Active in evaluation-related repos.

Production Engineering

Scaled WorkSpan search 10,000x (2K to 20M users). Built BookAFarm.com to 10K+ MAU with 94.3% test coverage. Ships in 14-day sprints.

Career highlights

Delivery Engineering Manager

Turing · Jun 2024 – Present

Leading 6 teams (30+ members). Shipped benchmark tasks into Tencent’s CL-bench. Built MBPI for automated LLM quality assurance.

Senior Software Engineer

WorkSpan · Jul 2021 – Dec 2023

Led Salesforce AppExchange integrations, onboarding 84+ enterprises. Scaled search 10,000x.

Google Maps Contributor

Google · 2016 – 2019

Introduced accessibility features impacting 52M+ users. Invited to Google HQ twice to collaborate with Maps engineering.

Tech stack

Python
TypeScript
Rust
Go
PyTorch
OpenAI API
Anthropic Claude
Gemini
Docker
Kubernetes
AWS
GCP
PostgreSQL
Redis
Kafka

Education

B.E. in Information Technology · Osmania University · 2014–2018

Open Source Contributions

inspect_ai

UK AI Safety Institute · 2k★ · 5 PRs, hand-merged by JJ Allaire

AI evaluation framework

  • #3668: vLLM provider restart-after-close fix (system failure analysis)
  • #3690: Judge-prompt injection hardening (22 tests)
  • #3722: Anthropic Opus 4.7 sampling-parameter guard
  • #3794: Pluggable URI resolver registry for custom scheme handling
  • #3836: HuggingFace dataset retry on transient errors

lm-evaluation-harness

EleutherAI · 12.5k★

AI evaluation framework

  • #3734: Fail-fast guard preventing a silent ∼17× MoE accuracy regression

Python scientific stack

8 PRs