Skip to content
LogoSamir Sawarkar

AI Reliability Engineer · LLM Evaluation · Agent Infrastructure

I build AI systems that fail loudly before they fail in production.

I design evaluation harnesses, deterministic safety gates, fault-injection systems, and trace/replay infrastructure for agents and AI workflows that touch real systems of record.

Pune, India · Open to high-ownership AI reliability, evaluation, and agent infrastructure roles

302

tests in the FAULTLINE reliability harness

400+

real invoices validated before production cutover

50,000+

users on a platform operated end to end

<5 min

incident response, reduced from 45 minutes

Current research

v0.2 · Pre-pilot

Capability–Reliability Substitution in Enterprise Agent Systems

Testing whether recovery engineering can make a smaller model as safely reliable as a frontier model, at materially lower token cost—and measuring the workflow depth where that substitution fails.

Read the experimental specification
Core test
Smaller model+ recovery system
≈?
Frontier modelbare system

Correctness · duplicate safety · total-token cost · dependency depth

Selected evidence

Systems with consequences

View all projects

Production AI

Invoice → ERP validation pipeline

LLM extraction backed by a deterministic validation gate before every write to a live financial system of record. Ambiguous cases escalate instead of becoming confident guesses.

  • 400+ invoices validated
  • 300+ invoices/day
  • Two reviewers reduced to one

Published research

EBOM → MBOM SafeGate

A deterministic validation-first pipeline with BLOCK / REVIEW / PASS conflict handling, bounded transformation, and a formal safety gate before Odoo execution.

  • ~30 min → under 2 min
  • Zero silent merges by design
  • Rollback-safe ERP push

Production operations

ScholarNote platform

Sole technical operator for a large education platform, owning reliability, monitoring, incident response, query performance, and infrastructure.

  • 50,000+ users
  • 99%+ uptime over 18 months
  • 3.2s → 0.8s page load

Technical judgment

Writing

View all