Key takeaway
Hiding credentials does not fully protect user intent. This guide constrains merchant, amount, time, product, and executor in delegated purchasing, with revocation, exception, and audit paths.
Executing constrained authority
Separate credential protection from permission to buy
Imagine a shopper allowing an agent to buy one box of household supplies from a specified store. The agent receives a payment token, then encounters a product page describing an upgraded bundle as better value. Even if the card number never appears, the agent can exceed the intended scope. Hiding credentials addresses one problem; buying the right item from the right merchant within budget requires additional rules. Token use alone does not establish end-to-end safety.
Stripe’s guide describes shared payment tokens constrained by seller, amount, and time window. OpenAI’s delegated-payment concepts also describe amount and expiry limits. This guide does not assume identical fields or capabilities across providers. Establish who validates each restriction, where enforcement occurs, and what happens if validation is unavailable. A prompt saying not to exceed the budget is not equivalent to a payment system enforcing it.
Identify the specific user intent the token serves
Before creating executable authority, record the task, permitted products, target merchant, maximum quantity, budget, and stopping conditions. If the shopper asks to buy the same as last time, establish whether that order remains a suitable reference instead of inferring ongoing permission from vague conversation. Stock, composition, address, and price may have changed. The record should refer to current confirmation rather than historical preference alone.
One intent can involve many searches without requiring payment power for each. Separate research from execution: research reads products and quotes, while execution obtains constrained authority after conditions are established. The agent need not hold charging capability while visiting unfamiliar pages. A business-state transition should trigger the switch, supported by user participation or prior authorization evidence, rather than the model deciding for itself.
Use a permission matrix rather than a master switch
Treat catalog reading, quoting, order creation, payment confirmation, cancellation, refund requests, and payout-detail changes as separate permissions. A shopping agent should not gain merchant refund, withdrawal, or account-administration powers merely because it needs to pay. Distinguish production and test environments, merchant accounts, and employee roles. Administrator credentials usually exceed the needs of one purchase and should not be handed to a general reasoning flow.
For each permission, identify its authority source and validation point. Shopper permission to buy is not merchant permission to change prices, and order lookup access is not access to every customer’s orders. Enforce checks server-side using trusted identity and ownership. An agent’s explanation can inform the shopper but cannot justify unauthorized backend actions. Test similar roles with different entitlements to reveal accidental approval.
Bind the transaction object as well as the amount
A payment below its ceiling can still be wrong. If the shopper requests unscented cleaning supplies and the agent substitutes a scented version at the same price, the budget passes while the requirement fails. Bind business authorization to the product or explicit substitution conditions, quantity, merchant, and delivery constraints. The payment layer may not express every product condition, so the application must check them before charging and link results to the order snapshot.
Identify application-only conditions as outside the token’s guarantee. The token handler may check amount and payee, the merchant cart version, and the agent interface confirmation. Explicit responsibilities prevent every participant from assuming someone else checked the product. The same total is not the same order. Substitution policies should remain explainable and testable.
Match validity to the task
Limit authority to the task’s required time rather than issuing indefinite permission for an immediate purchase. Delayed tasks need stopping conditions and status lookup. After expiry, stop using the authority and follow supported processes for renewed confirmation or a new constrained grant; do not silently extend it. Cancelling a task should affect unexecuted charging capability, not merely hide the task in chat.
Short validity can disrupt legitimate checkout; long validity expands exposure. Choose settings from observed duration and recovery behavior, testing absence, extra verification, and delay. There is no universal number of minutes. Explain why the window supports the task and what remains allowed afterward. Existing-order lookup normally should not require reopening charging authority.
Revocation must account for queued actions
When the shopper cancels, queued tasks, calls in progress, and completed transactions may coexist. Stop new execution, then determine existing stages. Unsent tasks can be cancelled, unknown calls need checks, and completed orders enter applicable cancellation or support. Revocation is not guaranteed recovery of a completed funds movement.
Record revocation time and task version, with another check immediately before payment execution. Creation-time checks alone let old queued actions execute after revocation. For requests already beyond the boundary, explain established facts and what is being verified. Test queue delays, multiple executors, and disconnected networks for windows in which an unauthorized action still charges.
External content must not rewrite permissions
Descriptions, reviews, email, and tool results may contain unrelated instructions. Treat them as information, not authority to change budgets, merchant allowlists, or approval rules. Authority should come from trusted user consent and system policy with independent deterministic checks. OWASP’s excessive-agency material supplies background; implementation still needs merchant-specific trust boundaries.
Test a description instructing the agent to ignore its budget and pay immediately. It should remain page content, without expanding permission or becoming approval evidence. Recognition of malicious wording alone is insufficient because instructions can be rephrased. Require the executor to accept only structured parameters within original authority, regardless of external text.
Reduce exposure while retaining diagnostics
Only services needing the token should receive its value. Other logs and interfaces should use linkable nonsensitive identifiers. Investigation needs to know which authorization, order, and executor did what, not necessarily see usable payment material. Avoid public URLs, long-retained chat history, and uncontrolled error reports. Determine sensitivity and storage requirements from provider documentation and merchant practices.
Recording nothing prevents investigation. Preserve constraint summaries, validation outcomes, times, associations, and execution results without unnecessary credentials. Distinguish audit access from payment-execution roles. Debugging needs boundaries too: do not copy real charge-capable material onto personal devices or shared documents for convenience. Reproduce equivalent conditions with test materials.
Do not amplify authority through multiple agents
A shopping system may have research, comparison, and execution agents. Give each only the capability needed for its step. Research passes candidates and evidence, not a self-expanded budget. Internal origin is no reason to skip consent checks. Messages should trace to original intent and currently permitted actions, otherwise more agents make responsibility harder to locate.
For necessary delegation chains, define who may redelegate, what can be passed on, and when authority ends. Arbitrary redistribution turns a limited task into invisible permission spread. Test upstream restarts, repeated downstream execution, role termination, and revocation for consistent authority. Identify the final executor and its basis rather than recording only that AI performed the action.
Define behavior when authorization services fail
Unavailable checks create conversion pressure. Decide beforehand which actions may continue, such as reading public products, and which must wait, such as charging or expanding orders. Do not improvise fail-open checks during incidents. Explain temporary purchasing unavailability and preserve selections instead of creating an improperly authorized order.
Recovery should not blindly execute all waiting tasks: intent, stock, and quotes may have changed. Recheck task validity, authority, and current quote. Recover original operations rather than creating transactions from retry counts. Observe availability together with incorrect approvals; a service that always permits seldom blocks but offers no meaningful protection.
Prove constraints through out-of-scope tests
Test excess amount, incorrect currency or merchant, substituted products, expiry, revocation, changed executor, and token reuse. Expected outcomes include no extra effective charge, understandable states, continued lookup, and an audit explanation, not merely an API error. Provider reuse rules differ and need specification-specific tests.
Pair restrictions with valid transactions to prevent indiscriminate refusal. Identify enforcement at the token provider, merchant service, or confirmation interface. Another controlled layer can supply a missing restriction, but do not present application checks as network guarantees. Rerun tests when model, tool, or adapter upgrades change payment paths.
Maintain authorization as a product
Let users inspect active tasks, limits, and cancellation options, while merchants inspect exceptional authorizations and executions. Reassess boundaries for new categories, methods, and entry points. An initial review does not make later permission expansion routine: changes can alter financial consequences and need owners and validation.
Review legitimate tasks wrongly blocked, out-of-scope tasks approved, post-revocation execution, and unexplained authority. Repair boundaries rather than adding vague prompt reminders. Tokens are useful within evidenced, constrained, revocable, recoverable transaction chains. They protect certain stages without replacing product, intent, and fulfillment judgments.
FAQ
- Can a token without a card number be stored freely in chat history?
- No. It may still carry transaction capability. Restrict transmission and storage according to provider requirements, using nonsensitive correlation identifiers for diagnostics where possible.
- Does cancelling a task automatically refund a completed purchase?
- No. Revocation stops actions that should no longer execute. Completed transactions require cancellation or refund handling according to order state and applicable conditions.
- Is model judgment alone enough?
- No. A model can help interpret intent, but financially consequential actions need independent permission and business checks with reviewable results. Prompts complement controls, not replace them.
Sources & further reading
- Agentic Commerce: A Getting Started Guide · Stripe
- Key concepts – Agentic Commerce · OpenAI
- LLM06:2025 Excessive Agency · OWASP
AI-assisted original research. Scenarios are hypothetical; rely on the primary sources listed for facts. Not investment or legal advice.
