AI Agents Begin Paying Each Other, But Dispute Resolution Remains a Challenge

iconTechFlow
Share
AI summary iconSummary
AI and crypto news reveals that AI agents are now managing microtransactions, but disputes over task quality remain unresolved. Blockchain news highlights GenLayer’s new system, which uses AI validators and rules to assess outcomes. The platform targets markets such as Daydreams TaskMarket, aiming to clarify machine-to-machine transactions where payment is secure but interpretation is not.

Author: Thejaswini M A

Compiled by Deep潮 TechFlow

Shenchao Summary: Token Dispatch applies the human dispute resolution systems of PayPal, FINRA, and the International Chamber of Commerce to machine payments as small as fractions of a cent—funds can be escrowed, identities can be on-chain, but “whether the work qualifies” still requires human judgment. The author zeroes in on GenLayer’s Intelligent Contract, which uses multi-verifier systems and large models to adjudicate against contract rules, slotting it into the gaps between UMA, Kleros, Google AP2, and ERC-8004: the evidence layer is in place; what’s missing is the interpretation layer.

It seems we have built civilization on the premise that evidence and acceptance are not the same thing.

Researchers are expected to use authentic, direct primary sources. Designers are expected to understand that “clean” doesn’t mean handing in a blank page. Developers are required to deliver a functional checkout page that doesn’t take four minutes to load. Humans are adept at carrying these unwritten contextual expectations because we share an intuitive understanding of what constitutes reasonable work. Software rarely needed this in the past, because the instructions we gave it were typically narrow enough to be executed literally.

If you've seen the joke from The Orville, you probably know what I'm talking about.

AI agents are entering a messier world: when instructions are incomplete, “do a good job”究竟包含什么?一旦资金与那段缺失的上下文绑定,解释本身便成为交易的一部分。

If someone hires a freelancer for $15 to research 20 companies with a reasonably clear brief, what do they expect? Find the companies, explain what each one does, use independent sources, and submit the report by the deadline. Once the task is delivered on time with a complete list of 20 companies, the freelancer will feel they’ve finished the job.

Buyers have a different perspective. Several citations refer back to the company’s own blog, and two passages are nearly identical to the original text.

When both sides are involved, the next step is arguing, sharing screenshots, going back to the original brief, and going back and forth. If they still can’t agree, someone else must ultimately decide: whether this work qualifies for payment.

What if both sides are AI agents?

Money itself isn't difficult.

The escrow contract can hold the $15 until the work is submitted. The blockchain can show when the funds were deposited and when the deliverables were submitted. The identity system can help confirm which agent performed the work. Signature records preserve the instructions both parties originally agreed to.

What remains unresolved is that someone still needs to read the brief, review the deliverables, and determine whether the seller has truly completed the work.

Businesses have long had systems in place to handle such issues—we just notice them when problems arise. When a PayPal buyer claims they didn’t receive the item or that it’s significantly different from the description, both parties first attempt to resolve the issue directly; if they can’t agree, either side can escalate it to a claim, prompting PayPal to review the evidence and make a ruling. PayPal states that most decisions are made within 14 days, though some may take 30 days or longer.

Larger commercial disputes have their own machinery. U.S. broker-dealers have the self-regulatory organization FINRA, which operates its own arbitration forum: in 2025, it received 2,597 new cases, with an average resolution time of 13.4 months. The nonprofit American Arbitration Association reported receiving 580,000 cases in 2025, with B2B claims and counterclaims totaling over $29 billion. The International Chamber of Commerce received 881 new arbitrations that year, with the total value of pending cases reaching $299 billion. These disputes are clearly much larger than those the agent is likely to encounter. But they illustrate this: as transaction volumes increase, disputes become frequent enough that an entire industry emerges around “reviewing evidence and deciding who should get paid.”

AI agents are beginning to route similar issues to the other end: trades that are extremely small and occur at very high frequencies make traditional dispute processes simply not worthwhile.

Visa and Artemis examined two emerging payment protocols designed for machine-to-machine commerce and found that, from its launch in May 2025 through April 2026, x402 processed approximately 109.6 million adjusted transactions totaling around $15 million. On the protocols studied by Visa, the average payment amount was less than one cent. Human customer service teams could not reasonably spend half an hour investigating each dispute worth just a few cents.

I went to find out who was solving this problem and came across GenLayer—it aims to make this into infrastructure.

GenLayer is building its own blockchain specifically to handle decisions that traditional smart contracts cannot manage. Traditional smart contracts excel at tasks where all machines run the same mathematical logic and arrive at the same result: whether funds have been received or whether a deadline has passed. Problems arise when there’s a need to “read and interpret”—such as reviewing research reports, browsing web pages, understanding contract clauses, or determining whether an image aligns with a creative brief. AI models never produce identical responses; different models can reach different conclusions from the same input, thereby breaking the assumption that every node must produce exactly the same output.

GenLayer features an "Intelligent Contract"—an advanced version of a smart contract. Instead of requiring every validator to reproduce identical computations, it allows multiple validators to assess whether a proposed result is acceptable, based on rules encoded in the contract. Validators can leverage large language models and publicly available web data during this process.

At first glance, this might seem like another missing layer in the agent tech stack. But several components are already covered by existing systems, so GenLayer isn’t being built from scratch.

Its more specific role is to determine what the evidence actually means when the parties remain in dispute.

Google's Agent Payments Protocol (AP2) can now generate cryptographically signed records: what the agent was authorized to purchase, what price the merchant quoted, and which payment was approved. Its specification explicitly states that these authorizations and receipts can later be combined as evidence in disputes. However, AP2 itself does not determine who is right or who should receive the money—it only preserves a trustworthy record of "what happened," leaving the interpretation of that record to others.

ERC-8004 is a standard on Ethereum for proxy identities and reputation, also featuring a Validation Registry. Proxies can send tasks to another system for verification, and the results are recorded on-chain as part of the proxy's history.

Blockchain can prove whether money has moved, so pure payment disputes are easy to resolve; when external information is needed, oracles bring that data on-chain. Now the evidence is complete, but one link is still missing: someone must interpret what the evidence means—GenLayer fills exactly this gap.

To see how much real work falls into each category, my agent on Slack—which I blindly trust because it’s our shared view of the world—collected over 300 online tasks on September 24 from Freelancer.com, Fiverr, Upwork, Scale AI’s Outlier, Prolific, and the on-chain task marketplace Daydreams TaskMarket. My agent reported that approximately 23% of tasks could be easily verified with a clear yes/no check; another 45% primarily relied on personal judgment; and the remaining 32% provided evidence but still required judgment on whether the work was “good enough.”

This isn't a market size estimate—the classification itself is also being determined by AI. But you get the point. And not every task involving AI requires an AI court.

For example, “design the most beautiful restaurant logo”: several validators might agree that one design looks better, but agreement doesn’t turn taste into an objective standard. Unless the buyer has provided an extremely detailed creative brief in advance, disputes still stem from personal preference. Marketing campaigns may deliver the promised engagement metrics, yet still spark arguments over how many interactions come from real users. In each case, evidence is available—but explanations are still needed before payment is made.

The crypto community has been experimenting with decentralized dispute resolution for years. These systems handle disagreements in very different ways.

UMA's Optimistic Oracle assumes the proposed claim is valid unless challenged. The proposer posts a deposit upfront; if no one challenges during the waiting period, the answer is accepted. If challenged, UMA token holders vote on the outcome. This keeps costs low—most claims don’t require full review, and expensive voting only occurs when a challenge is made.

Kleros is a decentralized dispute resolution system on Ethereum. Instead of company employees or court judges, jurors are randomly selected from individuals who have staked PNK tokens. If one party appeals, the next round brings in a larger jury. This approach is useful for cases that truly require human judgment, but it also means it can be time-consuming and costly. A three-juror round in General Court costs approximately 0.015 ETH (not including appeal fees), making this model impractical when the dispute itself is worth only a few dollars.

GenLayer replaces Kleros’s human jurors’ “judgment calls” with AI-augmented verifiers: one verifier proposes an answer, others review it, and if enough agree, it moves to final approval; if challenged, a new set of verifiers re-evaluates, and if the dispute persists, the scale expands. Moreover, both UMA and Kleros are also moving toward AI-augmented adjudication, so I don’t want to frame GenLayer as an “AI replacement for two entirely human systems.”

Rules must be specific so that validators have concrete criteria to verify. If a contract merely states that work should be “insightful” or “original,” agreement among several models will not make the standard any clearer.

GenLayer addresses this using what it calls the Equivalence Principle: telling validators how similar the answers must be for the network to consider them the same decision. Some contracts require exact matches; others allow different wording or reasoning, as long as the key outcomes are consistent. GenLayer’s documentation repeatedly encourages developers to narrow down these decisions.

As a result, significant power rests in the hands of those who write the contracts—they determine what evidence is valid, what constitutes success, and how much variation is allowed between answers. GenLayer can adjudicate according to these rules, but it cannot retroactively fix ambiguous standards once a dispute has begun.

Rulings only matter if they can genuinely alter the transaction. In practice, this means funds or other valuable outcomes must be tied to the decision. Disputes can be resolved through disbursement, refund, or updating the agent’s reputation.

For small-scale agency transactions, the goal may simply be to ensure the money ends up on the correct side, rather than obtaining an enforceable legal judgment.

Where will it actually be used?

The agency market has increasingly tied together delegation, payments, and on-chain identity, making this one of the most noticeable trends.

Taking Daydreams TaskMarket as an example: it runs on Base and allows humans or AI agents to post tasks backed by USDC. Different task formats clearly reveal where adjudication is useful versus unnecessary. Benchmark tasks enable workers to compete on measurable metrics such as accuracy or latency; if tasks are based on clear scores—say, one agent at 96% and another at 91%—the platform can simply compare the numbers and select the better result.

However, multiple participants may submit different types of deliverables for a bounty task, and the requester must determine which one is the best. When there’s no simple metric or rule to determine the superior submission, someone must review the results and make a decision.

The same applies to software and security bounties: tasks requiring an agent to pass a fixed test suite can be mechanically settled; vulnerability bounties, however, may involve genuine disagreements over whether an exploit is within scope, whether the severity is accurately described, or whether the same underlying issue has already been reported.

In performance marketing, counting 50,000 social interactions isn’t difficult. But are they real users or bot accounts? GenLayer has demonstrated Rally—a performance marketing app that uses its validators to analyze social activity and assess the authenticity of interactions.

A Service Level Agreement (SLA) is essentially a commitment between a service provider and a customer regarding the minimum level of service. For example, a cloud provider may guarantee 99.9% availability or a response time within two hours; if these standards are not met, the customer may be eligible for a refund or service credit.

If the protocol explicitly specifies which logs, monitoring services, and exclusions constitute evidence, the scope of what the automated arbiter must address will be significantly narrower.

Research is an example of what GenLayer can achieve at this stage. Some aspects of research are easy to verify: Has the report answered all questions? Are the sources authentic? Validators can check these.

But determining whether an insight is sufficiently original is much harder—there may be no clear rules or evidence to definitively settle the matter.

This type of work cannot be judged by simple yes/no scripts, but there is still sufficient evidence and structure to allow different verifiers to reach reasonable decisions. If a task were almost entirely based on taste or ambiguous quality, the system would have far fewer handles to work with.

Thinking that using several AI judges automatically makes it reliable—doesn't work.

If five validators use similar models, they may make the same mistakes: misinterpreting the same instruction, trusting the same bad source, or being deceived by manipulated content. GenLayer aims to reduce this risk by having validators independently verify results and replacing validators when an appeal is filed.

The evidence itself may also be unreliable. Web pages change, APIs provide different answers at different times, and documents may contain instructions specifically designed to mislead AI reading them. Therefore, the system only works well when verifiers can properly examine the same evidence. Sometimes there is no clear answer—GenLayer can return "Undetermined" rather than forcing a yes or no.

The design of GenLayer relies on the Condorcet Jury Theorem: a group of independent reasoners outperforms any single individual—provided the reasoners are truly independent; however, AI models often learn from overlapping data. Kleros researchers conducted near-real tests on 99 actual customer disputes from a Latin American crypto exchange: ChatGPT sided with the customer 3 times, while Claude Opus sided with the customer 14 times; when rerun with the next version of Claude, the platform’s win rate increased from 86% to 95%. Mixing models is useful when biases are random; but if an entire generation of models drifts in the same direction, a mixed jury will drift along with it.

How large is the market right now? It’s hard to estimate—only a portion of agency transactions ultimately become disputes. But existing systems have already shown that demand for rulings is inherently high.

Merchants processed approximately 261 million chargebacks in 2025, amounting to roughly $33 billion at an average dispute resolution cost of $128 per case. On the other hand, GenLayer’s current testnet handles about 25,800 decisions per day; even at $1 per decision, this amounts to only about $9.4 million annually. In comparison, eBay’s peak volume reached approximately 60 million disputes per year (about 164,000 per day): only when machine commerce reaches that scale could GenLayer become a major business.

To make this system work, the contract needs clear rules, the evidence must be reliable, the validators need to be sufficiently independent so that their consensus truly matters; and appeals must remain inexpensive.

Human commerce has long had courts, arbitration, and dispute resolution teams to handle situations where two parties cannot agree on the same transaction. If agents begin trading more frequently with each other, they may need a faster, cheaper version of the same—because machines always require inexpensive, high-speed versions of everything.

Perhaps it would be beneficial to teach machines how to determine what is fair before teaching them to do everything else.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.