Two years at Genpact turning tangled operations into numbers you could act on — then
bringing the same discipline to AI, where almost nobody measures anything. I build
agentic RAG, multi-agent reasoning and LLM evaluation, each backed by
tests and CI, and I’m currently pursuing a PGP in Applied AI and Agentic
Systems at Masters’ Union.
0Faster turnaround
0Fewer defects
0Productivity uplift
0Less reporting effort
Selected work
Things I built, and why they exist.
01Python · FastAPI · Ollama/Groq · pytest
FirstLine — An AI First Line of Support
I worked in an ops queue for two years. FirstLine is that job rebuilt as an
agent: it triages a live incident queue — root causes, SLA clocks,
escalation judgment, RCA drafts — on a simulated ops floor that knows the
truth behind every ticket it generates.
Which means the agent can be graded, not vibed about. Every shift
ends in a scorecard — triage accuracy, missed P0s, escalation precision —
because the simulator holds hidden ground truth the agent never sees.
FirstLine · eval real capture
$ firstline run --scenario fog --seed 735 tickets · severities lie · shift 480m rules triage 1.000 · esc-prec 1.000 · 100.0 llama3.2 triage 0.857 · esc-prec 0.529 · 81.4the 3B cried wolf — 17 escalations, 8 needless0 missed P0s on either backend
A client pays late, and the whole deal lives in a WhatsApp chat. Day 46 reads
the chat and the invoice, builds a timeline where every date opens the message
behind it, and fills a payment reminder, a demand letter and a case file for an
advocate, which the freelancer checks, signs and sends themselves.
The AI only points: it marks the messages that record an order, a
delivery or a promise to pay, and an entry survives only if its quote is word
for word in the message it cites. It supplies no dates, no amounts and no
wording. Code reads those from the messages, and the interest comes from an
engine that matches a blind rewrite to the paisa on 2,500 cases.
day46 · accuracy real run · 24 Sep
GET /accuracy · 12 unseen chats262 events, answers fixed first opus-5 found 99% · right 94% sonnet-5 found 97% · right 95% haiku-4.5 found 78% · right 95% sonnet-5 · one-line prompt found 92% · right 85%0 of 1,043 quotes hallucinated
1,707-test CIBlind rewrite, to the paisaMasked before the AI sees it
You tweak a prompt to fix one bad case, it works, you ship it — and three other
cases broke silently. PromptDrift snapshots prompt behaviour into a committed
baseline and fails CI when a change makes things worse.
It refuses to answer “is this prompt good?”, which is
unanswerable, and answers “is it worse than yesterday?” —
the question that actually blocks a merge. Cheap deterministic checks run first
and short-circuit the paid LLM-as-judge calls.
BM25 and FAISS fused through Reciprocal Rank Fusion with LLM listwise reranking,
feeding a four-stage agent loop of planner, retriever, synthesizer and critic.
The critic is the interesting part: it grounds every answer in retrieved evidence
and flags its own unsupported claims.
synapse · critic real capture · llama3.2
query what changed in Q3 policy?draft retention now 365 days [1] applies to all regions [1]critic✗ not supported — [1]says 180, not 365✗ "all regions" — no excerpt→ sent back to draft[1] policy_v3.pdf [2] notes.md
one answer, four passes — the critic can send it back
A full-stack finance tracker with dashboards, budgets and CSV import/export, plus
AI categorisation using concurrent LLM calls with a keyword fallback that works
with no network at all.
Its assistant calls tools to query real transaction data rather than trusting an
LLM to do arithmetic.
ledgr · tool call real capture
user dining spend last month?tool get_totals( start="2026-07-01", end="2026-07-31", category="Food & Dining")→ expenses 79.32 · 3 rowsreply 79.32 on dining in July the database did the maths
An RF prototype that pulls remotely transmitted temperature signals out of the air
and decodes them — the full software-defined radio chain of capture, filtering,
demodulation and decoding, validated against a reference sensor.
The waterfall behind the title of this page is a spectrogram, the display this
project lives in. It seemed a fitting thing to open with.
Analysed high-volume workflows and drove process redesign, contributing to an 18% reduction in turnaround time.
Ran structured root-cause analysis of recurring defects and SLA deviations, contributing to a 30% reduction in operational defects.
Automated management reporting in Excel and Power BI, cutting recurring effort by 25%.
Partnered with Risk, Compliance and Quality on controls and workflow improvements, contributing to a 22% productivity uplift.
Documented process gaps, tracked issues through Buganizer and validated corrective actions.
Internships
Where I learned the fundamentals.
2023Jul — Aug 2023
Bhagwan Parshuram Institute of Technology, GGSIPU
Data Analytics Intern · New Delhi
Analysed institutional datasets in Python and Excel, building KPI dashboards and automated reports that cut repetitive manual effort by 25%.
2022Jul — Sep 2022
Defence Research and Development Organisation (DRDO)
Research and Development Intern · India
Supported development, validation and signal analysis of an embedded wireless sensing prototype, improving reliability and reducing errors.
2020Sep — Oct 2020
AGIXURY
Web Development Intern · Remote
Built responsive interfaces from business requirements and improved usability across desktop and mobile.
Education
Electronics, then AI.
2026Present
Masters’ Union Current
PGP in Applied AI and Agentic Systems · Gurugram
2020To 2024
B.Tech, Electronics and Communication Engineering
Bhagwan Parshuram Institute of Technology, GGSIPU · CGPA 8.53 / 10
2018To 2020
Faith Academy Senior Secondary School, CBSE
Class XII 92.8% · Class X 94.2% · New Delhi
Certified
Snowflake Generative AI·Microsoft Azure AI Essentials·Become an AI Engineer·Advanced LLMs with RAG·Vector Databases and RAG·GenAIOps Foundations·Fine-Tuning for LLMs·Generative AI in Cloud Computing·Astronomer Data Engineering·AI Ecosystem for Developers·Snowflake Generative AI·Microsoft Azure AI Essentials·Become an AI Engineer·Advanced LLMs with RAG·Vector Databases and RAG·GenAIOps Foundations·Fine-Tuning for LLMs·Generative AI in Cloud Computing·Astronomer Data Engineering·AI Ecosystem for Developers·