AGENTREADY COMMERCE·AI SALES AGENT·HUMAN-IN-THE-LOOP
The model recommends. The rules approve.
AgentReady Commerce is a merchant-specific sales agent. A customer types "something for the gym under ₹2,000" and the agent picks a product, explains why, and completes a Razorpay payment. But the LLM only narrows the catalog. Hard-coded rules — price ceilings, inventory, variant match, and merchant policy — decide whether money moves.
Role
Sole author
Language
TypeScript, Python
Model
GPT-4o-mini, structured output
Payments
Razorpay test mode, x402/Solana rail

Architecture
How the system fits together
Intent → candidate products → approval binding → Razorpay / x402 settlement
The problem
Somebody has to say which product matches the intent
A customer messages a sports store: "I need something for the gym under ₹2,000." The catalog has whey protein at ₹1,899, a shaker at ₹349, lifting gloves at ₹599, and shoes at ₹4,499. The agent has to turn that fuzzy request into one exact SKU before any payment can happen.
The naive version gives the model a catalog and a checkout API and hopes for the best. The safe version splits the work: the model proposes, code binds the proposal to exact terms, and rules approve only the bound terms.
Competitive gap
What existing checkout flows get wrong
Traditional chatbots answer questions but stop before payment. Conversational commerce demos take payment but hide the rules inside the model. Neither is safe for real money.
AgentReady sits in the middle: the chat is natural, but the charge is governed by merchant policy that the merchant can read and change.
User context
Who this is for
The system is built for small merchants who sell through WhatsApp or Instagram and do not have engineering teams to build safe automation. They need conversions without the risk of a wrong charge.
What a binding order looks like
The shape of a safe sale
Before any payment call, the system needs a product SKU, a variant, a unit price, the available inventory, and the merchant policy for that category. If any piece is missing, the order stays at clarification.
For example, "Whey Protein, 1 kg, Chocolate, ₹1,899, 12 in stock, refundable" is a bound term. "Whey protein" is not. The model is good at getting from the message to the bound term; it is not trusted to decide whether ₹1,899 is an acceptable charge.
Three passes
How I got there
I built it in three passes because each pass exposed a different failure mode.
First, let the model narrow the catalog and explain its reasoning. This worked for product discovery but hallucinated variants and prices.
Second, bind every candidate to exact catalog terms before showing it to the user. This stopped hallucinations but made the system too cautious: almost everything needed human approval because the binding was too strict.
Third, add merchant rules that approve bound terms automatically when they pass policy. That is where the automation rate came from.
The router
One order, from intent to payment
The flow has three branches. If the customer message is ambiguous, the agent asks a clarifying question. If it is exact, the system binds the product and checks rules. Only bound terms that pass rules reach payment.
What the model is not allowed to do
Three rules I did not break
One: the model never sees Razorpay credentials or the payment API. It returns structured output. The payment call is server-side code.
Two: the model never computes money. It may read prices from the catalog, but the checkout amount is calculated by code from the bound terms.
Three: every transaction writes an audit log with the raw customer message, the model recommendation, the bound terms, the rule outcome, and the payment ID. A merchant can reconstruct why any charge happened.
Guardrails
The thresholds, in one file
Keeping the thresholds together means a merchant can tune the automation curve in one place. A higher auto-approval ceiling means more automation and more risk; a lower one means more human checks.
export const POLICY = {
maxAutoApprovePaise: rupees(2_000),
restrictedCategories: ['prescription', 'bulk-order'],
requireConfirmationAbove: rupees(500),
maxQuantityPerOrder: 3,
};
export function canAutoApprove(order: BoundOrder): Decision {
if (order.category in POLICY.restrictedCategories)
return { ok: false, reason: 'Restricted category' };
if (order.totalPaise > POLICY.maxAutoApprovePaise)
return { ok: false, reason: 'Above auto-approval ceiling' };
if (order.inventory < order.quantity)
return { ok: false, reason: 'Insufficient inventory' };
return { ok: true };
}Results
What it does on the held-out set
I generated 120 synthetic conversations against a catalog with known correct answers. The model could recommend anything; the rules decided what was allowed to become a charge.
0
Wrong auto-approvals
NEVER ALLOWED TO MOVE
78.3%
Handled without a human
94 OF 120
₹0
Disputed test transactions
RAZORPAY TEST MODE
Source · synthetic conversations with answer key, 120 records
What each layer bought
Turning the pieces on one at a time
I ran the same 120 conversations with one more layer switched on each time. The numbers show where the real value is.
Source · same held-out set for every row
Lessons
What building it taught me
The layer everyone would build first — the chat — is the one that does the least damage control. The layer with the model on it is worth the least if it can move money.
The interesting question was never how to automate more orders. It was what an order has to survive before it is allowed to count.
Trust in an agent comes from what it cannot do, not from how eloquently it explains what it can.
Still open
Where it is still rough
Synthetic catalog
The test data contains the messes I thought to generate. A real merchant catalog has stale prices, missing images, and variants that are not actually in stock.
Small held-out set
120 conversations is enough to catch obvious mistakes, not enough to claim a fraud rate. Zero wrong approvals at this size means nothing went wrong, not that nothing would.
Policy belongs to the merchant
The thresholds are sensible defaults, not derived from real transactions. A merchant with actual sales data should own them.
In short
What AgentReady came down to
01
Recommend, never decide
The model chooses between options that code has already priced; rules decide whether the choice becomes an action.
02
Bind before comparing
Fuzzy intent becomes a concrete SKU, variant, and price before any rule or payment runs.
03
Measure what got approved wrongly
Automation rate is the wrong metric. Wrong approvals are the right one.


