Skip to content
All work

AGENTREADY COMMERCE·AI SALES AGENT·HUMAN-IN-THE-LOOP

The model recommends. The rules approve.

AgentReady Commerce is a merchant-specific sales agent. A customer types "something for the gym under ₹2,000" and the agent picks a product, explains why, and completes a Razorpay payment. But the LLM only narrows the catalog. Hard-coded rules — price ceilings, inventory, variant match, and merchant policy — decide whether money moves.

Role

Sole author

Language

TypeScript, Python

Model

GPT-4o-mini, structured output

Payments

Razorpay test mode, x402/Solana rail

AgentReady Commerce

Architecture

How the system fits together

Intent → candidate products → approval binding → Razorpay / x402 settlement

The problem

Somebody has to say which product matches the intent

A customer messages a sports store: "I need something for the gym under ₹2,000." The catalog has whey protein at ₹1,899, a shaker at ₹349, lifting gloves at ₹599, and shoes at ₹4,499. The agent has to turn that fuzzy request into one exact SKU before any payment can happen.

The naive version gives the model a catalog and a checkout API and hopes for the best. The safe version splits the work: the model proposes, code binds the proposal to exact terms, and rules approve only the bound terms.

Competitive gap

What existing checkout flows get wrong

Traditional chatbots answer questions but stop before payment. Conversational commerce demos take payment but hide the rules inside the model. Neither is safe for real money.

AgentReady sits in the middle: the chat is natural, but the charge is governed by merchant policy that the merchant can read and change.

ApproachWhere it fails
Rule-based chatbotCannot handle fuzzy intent; every question needs a hand-coded branch.
LLM + checkout APILooks smart, but the model can charge the wrong product or invent a price.
AgentReadyModel handles language; rules handle money. Auditable and tunable.

User context

Who this is for

The system is built for small merchants who sell through WhatsApp or Instagram and do not have engineering teams to build safe automation. They need conversions without the risk of a wrong charge.

UserNeedWhat AgentReady does
Merchant ownerMore sales, no chargebacksSets policy once; rules enforce it every time.
Store operatorLess time on repetitive questionsCommon requests auto-bind and auto-approve.
CustomerFast, clear checkoutSees exact product, price, and policy before paying.

What a binding order looks like

The shape of a safe sale

Before any payment call, the system needs a product SKU, a variant, a unit price, the available inventory, and the merchant policy for that category. If any piece is missing, the order stays at clarification.

For example, "Whey Protein, 1 kg, Chocolate, ₹1,899, 12 in stock, refundable" is a bound term. "Whey protein" is not. The model is good at getting from the message to the bound term; it is not trusted to decide whether ₹1,899 is an acceptable charge.

FieldComes fromWhy it matters
SKU + variantLive catalog lookupStops the model from inventing a product that does not exist.
Unit priceCatalog, not the LLMThe price on checkout must match the catalog price exactly.
Inventory countCatalogPrevents selling something that is out of stock.
Merchant policyMerchant configCategory-level rules: max auto-approval value, return window, restricted items.
Customer consentExplicit confirmationThe human says yes to the exact bound terms before pay.

Three passes

How I got there

I built it in three passes because each pass exposed a different failure mode.

First, let the model narrow the catalog and explain its reasoning. This worked for product discovery but hallucinated variants and prices.

Second, bind every candidate to exact catalog terms before showing it to the user. This stopped hallucinations but made the system too cautious: almost everything needed human approval because the binding was too strict.

Third, add merchant rules that approve bound terms automatically when they pass policy. That is where the automation rate came from.

The router

One order, from intent to payment

The flow has three branches. If the customer message is ambiguous, the agent asks a clarifying question. If it is exact, the system binds the product and checks rules. Only bound terms that pass rules reach payment.

One order, from intent to payment

What the model is not allowed to do

Three rules I did not break

One: the model never sees Razorpay credentials or the payment API. It returns structured output. The payment call is server-side code.

Two: the model never computes money. It may read prices from the catalog, but the checkout amount is calculated by code from the bound terms.

Three: every transaction writes an audit log with the raw customer message, the model recommendation, the bound terms, the rule outcome, and the payment ID. A merchant can reconstruct why any charge happened.

Guardrails

The thresholds, in one file

Keeping the thresholds together means a merchant can tune the automation curve in one place. A higher auto-approval ceiling means more automation and more risk; a lower one means more human checks.

src/rules/merchant-policy.ts
export const POLICY = {
  maxAutoApprovePaise: rupees(2_000),
  restrictedCategories: ['prescription', 'bulk-order'],
  requireConfirmationAbove: rupees(500),
  maxQuantityPerOrder: 3,
};

export function canAutoApprove(order: BoundOrder): Decision {
  if (order.category in POLICY.restrictedCategories)
    return { ok: false, reason: 'Restricted category' };
  if (order.totalPaise > POLICY.maxAutoApprovePaise)
    return { ok: false, reason: 'Above auto-approval ceiling' };
  if (order.inventory < order.quantity)
    return { ok: false, reason: 'Insufficient inventory' };
  return { ok: true };
}

Results

What it does on the held-out set

I generated 120 synthetic conversations against a catalog with known correct answers. The model could recommend anything; the rules decided what was allowed to become a charge.

0

Wrong auto-approvals

NEVER ALLOWED TO MOVE

78.3%

Handled without a human

94 OF 120

₹0

Disputed test transactions

RAZORPAY TEST MODE

Source · synthetic conversations with answer key, 120 records

What each layer bought

Turning the pieces on one at a time

I ran the same 120 conversations with one more layer switched on each time. The numbers show where the real value is.

ConfigurationWhat it bought
Model only · 91% automated · 7 wrongThe obvious build. Quietly approves seven payments it should not. High automation is the warning sign.
+ binding · 0% automated · 0 wrongNothing gets approved because every request lacks exact terms. Correct, and useless.
+ rules · 78.3% automated · 0 wrongThe single biggest jump. Binding without rules is powerless; rules without binding is blind.
+ audit log · 78.3% automated · 0 wrongThree points of trust: the customer, the merchant, and the engineer debugging at 2 AM.

Source · same held-out set for every row

Lessons

What building it taught me

The layer everyone would build first — the chat — is the one that does the least damage control. The layer with the model on it is worth the least if it can move money.

The interesting question was never how to automate more orders. It was what an order has to survive before it is allowed to count.

Trust in an agent comes from what it cannot do, not from how eloquently it explains what it can.

Still open

Where it is still rough

Synthetic catalog

The test data contains the messes I thought to generate. A real merchant catalog has stale prices, missing images, and variants that are not actually in stock.

Small held-out set

120 conversations is enough to catch obvious mistakes, not enough to claim a fraud rate. Zero wrong approvals at this size means nothing went wrong, not that nothing would.

Policy belongs to the merchant

The thresholds are sensible defaults, not derived from real transactions. A merchant with actual sales data should own them.

In short

What AgentReady came down to

01

Recommend, never decide

The model chooses between options that code has already priced; rules decide whether the choice becomes an action.

02

Bind before comparing

Fuzzy intent becomes a concrete SKU, variant, and price before any rule or payment runs.

03

Measure what got approved wrongly

Automation rate is the wrong metric. Wrong approvals are the right one.

More work