AI systems that show their work and refuse to fake it

Rohit Agrawal — principal engineer. From understanding the requirement to what's actually running in production — and feeding what I learn back into the next release. Fourteen years deciding whether software was safe to release; now I build AI systems that make that decision about themselves.

Experience
14+ yrs — Oracle, Amazon, LimeRoad, Mobileum, Snapdeal, Subex
Most recent
Oracle Principal MTS · 2019–2026
Since April 2026
Building independent AI systems, full time
Led
11+ quality engineers and SDETs, globally
Seeking
Senior / Principal — AI platform engineering, LLMOps, forward-deployed AI · AI quality engineering, test automation and agentic AI

Fourteen years I can tell you about. Six practices you can check yourself.

How I build · enforced in CI, not in reviewCheckingReplay of the last run
An evaluation gate, before promotion
A retrieval suite on the candidate index decides; overrides are audited · CiteVyn (52 / 52 cases)
CheckingEnforced
Cite it, or refuse it
No source, no answer. Refusal is a valid outcome, not a failure · CiteVyn (1.0 hit-rate, 26 answerable)
CheckingEnforced
Floors in config, not in good intentions
The build fails below the coverage floor; the number lives in the repo · Quorum‑AI (88%, enforced in config)
CheckingEnforced
Uncertainty returned, never averaged away
Disagreement is its own field; a skipped round is recorded as skipped · Quorum‑AI (32 ADRs, each named)
CheckingEnforced
Provenance on every reading
Live, cached, or no reading. Degraded modes are disclosed, not hidden · SaafSaans (19 live, 2 known gaps)
CheckingEnforced
A readiness review that is allowed to say no
One of my own systems is not deployed, because its own review refused it · NarraTwin AI (Phase 1, No-Go)
CheckingBlocking

The last row is the one that costs me something. It stays.

Each figure read from the named repo at a pinned commit
Practices, enforced6
Systems of my own6
Blocked by its own gate1
Every figureCounted, or marked approximate

Two ledgers, deliberately kept apart.

Employer outcomes are reported and marked approximate; the independent systems are evidence you can inspect. The two are never averaged into a single number.

Employer outcomes

Approximate · Fourteen years, six employers

Manual test design
~40% less effort
Automation coverage
65% → 95%
Production incidents
~20% fewer
Payment defects in production
~25% fewer
Release validation
~35% less manual effort

Independent StackClimb systems

Inspectable evidence · public code

StackClimb is where Rohit Agrawal builds independent AI systems — outside any employer.

CiteVyn
Answers only what it can cite, and refuses the rest.
Quorum-AI
Four models answer; a moderated critique maps their disagreement.
SaafSaans
Your air, not the city average — readings labelled live, cached, or no reading.
NarraTwin AI
Cited walkthroughs; every claim is checked against its source.
EvalAxis
A regression gate that fails the build when quality drops — private, in progress.

Not claimed — users, customers, revenue, production scale, awards or partnerships for any independent system.

What you can use, and what is still being built

Use them now

  • CiteVynLive — cold-startsAnswers from official docs and shows its sources — or says it can’t.Golden suite — 52 / 52 cases
  • Quorum‑AILiveAsks four AI models the same question and shows where they disagree.Coverage floor — 88%, enforced in config
  • SaafSaansDeployed — sleeps when idleScores your air risk for today — yours, not the city average.Feeds — 19 live · 2 known gaps

Being built

  • Aegis ContractsIn progress · closedRules for what one AI system may promise another.
  • EvalAxisIn progress · closedBlocks a release when answer quality drops, the way a failing test would.12,978 lines · 388 test functions
  • NarraTwin AIPhase 1 — No-GoTurns project knowledge into cited walkthroughs; its own gate says not yet.Surface parity — 25 / 25 agree
Live — cold-starts

CiteVyn

“Can I trust this answer — and trace every claim to its source?”

Citation-grounded Q&A over official AI documentation. Answers quote their sources verbatim; where no source supports an answer, CiteVyn refuses instead of guessing. Index updates reach production only through an evaluation gate.

Twenty-six answerable questions. It found the right source for every one.

GateA 15-case retrieval suite on the candidate index gates promotion; any override is audited

CiteVyn answering a question about Claude Code, with two numbered citations to real docs.claude.com pages

Committed visual baseline, citevyn at df8cfc3. Both citations are real documentation URLs.

Retrieval1.0 hit-rate · 26 answerable
Judge4.63 / 5 · 54 judged
Golden suite52 / 52 cases
Tests1,036 backend · 125 e2e
Not claimed
Corpus scale · adoption · SLA
Live

Quorum‑AI

“What happens when four models disagree about your question?”

One question runs against four models in parallel. A separate moderator pass — its own call, with its own configured model — reads all four answers and critiques them over two rounds; a round dropped for time is recorded as skipped, not hidden. The four never read each other. A synthesis then returns consensus, disagreement, source support, uncertainty, and a recommendation.

Disagreement and uncertainty come back as their own fields, not folded into one confident answer.

GateCost approved before anything runs; fallbacks disclosed

Quorum-AI's verdict panel reading '4 of 4 models aligned, 3 revised their position', with a note that only 13% of claims carried citations

SAMPLE — an end-to-end acceptance fixture rendering canned data (e2e/fixtures/golden-run.ts), not a measured run. Live execution is off by default. The revision count is inferred from position movement, and the interface says so.

Tests2,095 python · 358 e2e
Coverage floor88%, enforced in config
Decisions32 ADRs, each named
Limits16 concurrent · 180s deadline
Not claimed
Latency · accuracy · adoption
Deployed — sleeps when idle

SaafSaans

“Is it safe for me to go outside right now — and if not, when?”

A Delhi-NCR air-quality companion that scores your risk — age, condition, planned activity — rather than the city’s average, across a repo-stated 21 stations, and answers questions with cited health guidance. The Hindi draft ships behind a banner saying no Hindi speaker has reviewed it yet.

Two people, one sky. The adult with asthma scores 76. The healthy adult, same plans, scores 64.

RuleEvery reading labelled live, cached, or no reading

SaafSaans showing a risk score of 76 out of 100 for an adult with asthma beside 64 for a healthy adult on the same plans

Captured live 11 Aug 2026, 5:00 PM. Cropped above a time-window strip showing a defect under repair.

Risk delta76 vs 64 · identical plans
Feeds19 live · 2 known gaps
Tests767 functions · 16,072 lines @10f4213
Guard39 patterns · EN, Hindi, Hinglish
Not claimed
Medical advice · uptime
Phase 1 — No-Go

NarraTwin AI

“Can project knowledge become a walkthrough without inventing a claim?”

Grounded walkthrough generation with citations, claim evaluation, consent checks, and release gates that run before anything is generated. Its own release-readiness review currently reads No-Go — so it is not deployed, and this page says so.

Twenty-five languages, six script classes. Every one proved to agree across five surfaces against a pinned fixture.

GateIts own readiness review says No-Go, so it is not deployed

NarraTwin interface design: a simulated host context panel, a consent checkbox before any presenter is generated, and external web disabled by policy

SAMPLE — interface design, not a running capture. No screenshot of the built system exists yet, so this stands in and is labelled rather than implied.

StatePhase 1 — No-Go · not deployed
Languages25 · 6 script classes
Surface parity25 / 25 agree
Code26,646 lines · 1,743 tests @a022862
Not claimed
Deployment · video · avatar Q&A
In progress — closed

Carried, not shown

Two systems are still being built. They stay closed until they can be judged on finished work rather than on intent. They are named because a record that discloses its gaps should also say what exists, and the question each one is built to answer costs nothing to state.

EvalAxis

“Why can a failing test stop a release, when a measured drop in answer quality cannot?”

Evaluates LLM, RAG, and agent changes with evidence, and blocks CI on a quality regression.

In progress · closed

Aegis Contracts

“What should one AI system be allowed to promise another — and who checks?”

Early-stage work on contract-shaped guarantees between AI systems.

In progress · closed
Two zipped garment bags on a rail, each carrying a name tag.EVALAXISAEGISCONTRACTS

One message away

Open to senior and principal roles in AI platform engineering, LLMOps and forward-deployed AI, and to AI quality engineering, test automation and agentic AI leadership. Fourteen years of evaluation and release governance sits underneath all of them.

Based
Bengaluru, India · IST (UTC+5:30)
Open to
Relocation worldwide · international travel
Work authorisation
India · United States — H-1B approved