Anthropic just ran an internal marketplace where AI agents negotiated real deals on behalf of employees. The experiment, called Project Deal, produced 186 completed transactions worth over $4,000. It involved 69 participants, real goods, and zero human intervention during the trading phase. This is not a product demo or a chatbot upgrade. It is the first credible signal that agent-to-agent commerce is moving from whitepaper speculation to working prototype.
What Project Deal Actually Did
The setup was straightforward. Anthropic recruited 69 employees and had each one sit through a structured interview with Claude. The agent learned what each person wanted to sell, what they wanted to buy, their price ranges, and their negotiation style (aggressive, flexible, firm on price but open on timing, etc.). These preferences were compiled into individualized system prompts for each participant’s agent.
Then the agents were dropped into Slack marketplace channels. They posted listings, browsed offers from other agents, initiated negotiations, counter-offered, and closed deals. All autonomously. When the experiment ended, humans exchanged the actual items based on agreements their agents had reached.
Anthropic ran four parallel market configurations: some populated entirely by Opus-level agents (their strongest model), others mixing in Haiku-level agents (their lightest). This created a controlled comparison of what happens when agents of different capability levels compete in the same market.
The findings were blunt. Opus agents closed more deals and secured better prices for their principals. Users paired with weaker models got worse outcomes on average, accepting lower sell prices and paying more as buyers. And the participants running weaker agents did not realize they were getting outperformed.
That last point matters more than it sounds. In a future where businesses delegate procurement and sales to AI agents, the quality of your agent becomes a competitive variable. Not just your information advantage or your relationships, but your proxy’s negotiation skill. The old question was “are you good at haggling?” The new question is “is your AI good at haggling?” And unlike human negotiation skill, which people roughly self-assess, agent performance is invisible to most users unless they run explicit benchmarks against alternatives.
Why This Is a Commercial Entry Point Shift, Not a Feature Update
If you read Project Deal as “Anthropic helped employees swap used bikes,” you are underselling it by several orders of magnitude.
The structural pattern here is: humans express preferences and constraints, agents handle discovery, comparison, negotiation, and execution. Today the items are secondhand office equipment. Tomorrow it is SaaS subscriptions, cloud capacity, logistics contracts, and API access deals. The mechanism is identical. Only the catalog and the stakes change.
Consider what each agent actually did in the experiment: it scanned available listings, matched them against its principal’s stated preferences, calculated acceptable price ranges, initiated outreach to relevant counterparties, handled multi-round negotiation (including concessions and walk-aways), and finalized binding agreements. That is not autocomplete. That is a full transaction lifecycle executed without human involvement at any step.
This raises a question that should concern anyone in e-commerce or digital marketing: when both buyer and seller are represented by agents, who needs a homepage? Who needs a banner ad? Who needs a recommendation engine optimized for impulse purchases? Who needs an influencer unboxing video when the buyer has no eyes?
The entire consumer internet was built around human browsing behavior. Agent commerce does not browse. It queries, compares, negotiates, and transacts. Twenty years of conversion funnel optimization may need rethinking when the “visitor” is a piece of software that cannot be emotionally triggered by a countdown timer or a “only 2 left in stock” warning. Agents do not experience FOMO. They do not respond to social proof. They evaluate structured data against defined criteria and either transact or move on.
The Infrastructure Stack That Does Not Exist Yet
Project Deal worked because it was a closed system. Identity was known (all participants were Anthropic employees). Payment was handled offline. Fraud, chargebacks, disputes, and cross-platform discovery were not factors.
In the real world, agent commerce requires infrastructure layers that are still being built:
| Layer | Function | Current State |
|---|---|---|
| , , , – | , , , , , | , , , , , , , – |
| Tool Connection (MCP) | Lets agents interact with external systems, APIs, and data sources | Open-sourced by Anthropic in 2024; growing adoption |
| Agent Communication (A2A) | Enables agents to discover each other, exchange tasks, negotiate | Early protocol stage; no dominant standard |
| Payment & Settlement | Agents need to authorize payments, enforce limits, and create audit trails | Fenwick identifies agentic payments as a key 2026 trend |
| Identity & Trust | Platforms must verify who authorized an agent, its spending limits, and liability chain | Conceptual; no production-grade system deployed |
| Dispute Resolution | Post-transaction auditing, chargebacks, and accountability | Undefined for agent-to-agent scenarios |
The real signal from Project Deal is not “Claude can buy things.” It is that the agent transaction stack is transitioning from concept to pilot. MCP gives agents system access. Payment protocols give agents spending capability. Identity layers prevent agents from becoming unaccountable autonomous spenders.
Without all four layers working together, agent commerce stays in the lab. With them, it scales to enterprise procurement.
B2B Procurement Will Move First
Consumer shopping will change, but the first major wave of agent commerce will not be “AI orders your coffee.” Humans still manage to buy coffee without digital delegation.
The high-value, near-term opportunity sits in B2B scenarios:
- Software subscription procurement and renewal negotiation
- Cloud resource and API usage optimization
- Supplier quote aggregation and comparison
- Logistics, inventory, and restocking negotiations
- Small-ticket purchases for advertising, data services, and outsourced work
These scenarios share common traits: they are rule-heavy, repetitive, auditable, and operate within defined budgets. They are perfect for agents working within human-set boundaries.
A practical example: a corporate procurement agent with a $5,000 monthly budget, restricted to whitelisted vendors, required to collect three quotes before any purchase, authorized to auto-approve below $500, and escalating to a human above that threshold. The human never reads 17 quote emails. They review exceptions only.
Another example: a mid-size company renewing its annual SaaS stack. Today, this involves a procurement manager spending two weeks emailing vendors, comparing tiers, checking contract terms, and negotiating discounts. An agent could run the same process in hours: query each vendor’s API for current pricing, compare against usage data from the past 12 months, identify underutilized subscriptions, flag contracts approaching auto-renewal, and negotiate volume discounts with counterparty agents on the vendor side. The human approves a summary table instead of managing dozens of threads.
In these contexts, agents do not need taste or intuition. They need to stay within budget, avoid being exploited, and explain their decisions. Modest requirements with hard commercial value.
The New Competitive Moat: Agent Readability
If agents become the buyer’s entry point, what merchants optimize for will shift.
Historically, websites courted humans: beautiful layouts, compelling stories, prominent CTAs, social proof, urgency signals. In an agent-mediated world, websites also need to serve machines: structured product data, explicit pricing rules, available APIs, reliable inventory signals, and machine-parseable terms of service.
Think of it as the next layer beyond SEO. Call it AEO (Agent Engine Optimization): making your business discoverable, evaluable, and transactable by procurement agents. Not just ranking in search results, but earning a spot on an agent’s shortlist.
This reshuffles competitive positions. A brand with beautiful design but messy data might lose to a plain-looking supplier whose API, inventory feeds, pricing, and delivery terms are clean and structured. Agents will not be swayed by “flash sale” banners. They ask: what is the total cost? What is the delivery probability? What happens on failure?
The startup opportunities that follow from this are less glamorous but more durable:
- Agent-readable product and service catalogs
- Quoting and settlement APIs designed for machine clients
- Enterprise-grade authorization and spending limit systems
- Agent identity verification and reputation scoring
- Post-transaction audit trails, dispute handling, and insurance products
These are infrastructure plays. They are where capital accumulates in platform shifts, regardless of which consumer-facing agent wins the popularity contest. The pattern repeats from every prior platform shift: the picks-and-shovels businesses build durable value while the consumer-facing layer remains volatile and winner-take-most.
The Risk: Weaker Agents as a New Source of Commercial Disadvantage
The most concerning finding from Project Deal is the asymmetry created by agent quality differences.
In human markets, you have some awareness of your own negotiation weaknesses. You know if you are conflict-averse, impatient, or lazy about comparison shopping. You accept the trade-offs. In agent markets, the dynamic is more opaque: you assume your AI is performing well, but it might simply be losing negotiations at a consistent, invisible rate.
Anthropic’s data showed that weaker-model participants did not notice their disadvantage. Scale this to enterprise procurement and the implications are serious. If different companies deploy agents of different capability tiers, “agent quality” itself becomes a business resource. Strong agents help large companies extract better terms. Weak agents lock smaller companies into worse contracts. The gap compounds.
This means agent commerce cannot be evaluated purely on efficiency gains. Transparency requirements will follow:
- What comparisons did the agent run?
- Why did it accept this specific price?
- Did it miss cheaper or better-fit alternatives?
- Did a negotiation fail because of market conditions or agent limitations?
- Is the platform’s ranking algorithm favoring its own services?
Regulators and enterprise procurement teams will both focus on these questions. Letting AI spend money autonomously is acceptable. Letting it spend money with no explainability is not.
For SaaS vendors selling to enterprises, this creates a new requirement: your product needs to be evaluable by agents representing the buyer. If a competing vendor provides structured comparison data and your product requires a 45-minute sales call to get pricing, you will lose deals before a human ever enters the conversation. The sales process itself gets disintermediated when the buyer’s agent cannot parse your offering.
Agent Commerce Will Break Out Through Delegated Transactions
Over the past year, most agent products have been stuck in an awkward position: impressive in demos, fragmented in daily use. Automation looks elegant until every edge case requires human intervention. The core problem is that these products try to take over entire workflows at once, and reality pushes back at every integration point.
Transaction scenarios are different. They have clear objectives, quantifiable outcomes, and natural boundary conditions. Did the purchase happen? At what price? Was the budget exceeded? Was a better option available? All of these are recordable, comparable, and auditable.
This is why delegated transactions will be the first agent use case to reach production scale, ahead of the “universal AI employee” vision. The universal employee sounds like science fiction. Delegated procurement sounds like an expense management upgrade. Less exciting, more shippable.
Project Deal was small: 69 people, $4,000 in transactions. In the context of global commerce, that is statistically invisible. But it proved one thing: when humans hand over preferences, budgets, and constraints to agents, those agents can complete real market transactions without further human involvement.
The real competition ahead is not about which model writes better emails. It is about who can connect agents to live commercial systems: service discovery, identity verification, payment authorization, audit logging, and accountability when things go wrong.
The next phase of AI agents is less likely to be “writes your weekly status report” and more likely to be “negotiates your vendor contracts, files the paperwork, and flags anomalies.” That is less romantic than artificial general intelligence. But commerce has never been romantic. Whoever spends money more accurately wins the round.
For technology leaders evaluating where to invest: the agent commerce stack is where middleware was in 2005 and cloud infrastructure was in 2010. Unsexy, essential, and about to absorb a lot of enterprise spending. The companies that build reliable agent transaction infrastructure in the next 18 months will own the rails that everyone else runs on.



