AI Models & Platforms
Alibaba.com Says Accio Ran E-Commerce Tasks at Over 50% Lower Cost

Alibaba.com on September 9, 2026, announced results from a 107-task benchmark evaluation showing that Accio, its AI agent platform for commerce, completed a range of e-commerce tasks at more than 50% lower estimated cost than OpenAI’s Codex and Anthropic’s Claude Code, with what the company described as comparable completion quality.
Completing the full task set cost an estimated $3.69 with Accio, versus $9.27 for Codex and $9.51 for Claude Code, figures Alibaba.com described as more than 50% lower in both cases. The company said the results make Accio the most cost-competitive e-commerce AI tool available to small and medium-sized businesses. The announcements were made at CoCreate, Alibaba.com’s annual event for entrepreneurs and small businesses, whose 2026 Los Angeles edition runs September 9–10, 2026.
How Alibaba.com Explains the Cost Gap
According to the company, Accio’s efficiency comes from optimizing the full system around real commerce work. Accio uses commerce-specific data and workflows to post-train lightweight models for clearly defined tasks, which the company said allows smaller models to handle routine work efficiently while more capable models remain available when deeper reasoning is required.
Rather than defaulting to either the cheapest or the most powerful model, the system breaks complex requests into manageable steps and weighs quality, speed, cost and data requirements for each step, the company said. Alibaba.com described that approach, in economic terms, as moving each workflow closer to the Pareto frontier, the point at which improving one dimension would require a trade-off in another. The company also credited cache reuse, context compression and coordinated agent execution with reducing repeated processing, unnecessary tool calls and redundant token consumption.
“For small businesses, unaffordable AI is useless,” Kuo Zhang, President of Alibaba.com, said in the company’s announcement. Zhang said the goal is to make commerce AI practical and affordable for even a one-person company, rather than simply to make AI more powerful.
Commerce Agent Bench, Open-Sourced
Alibaba.com said it has open-sourced Commerce Agent Bench on GitHub. According to the company, the benchmark draws on real merchant activity (10 million active SMB users, 1.6 million conversations and 200,000 execution traces) distilled into 107 end-to-end commerce tasks across seven categories and four levels of autonomy, rather than relying on synthetic exercises. The tasks include reviewing hundreds of unstructured emails, spotting payment fraud, calculating landed costs and booking multi-carrier shipping routes.
The repository, developed and maintained by the Accio team at Alibaba International, splits the 107 tasks into 53 command-line, 28 browser, 16 file and 10 API/MCP tasks, with capability slices covering 65 text-only, 20 browser-text-capable and 22 vision-required tasks. The project was previously known as RealReplicaBench. Every task runs in a fresh container and is graded by its own deterministic or LLM-assisted verifier. The harness, Python package, mock-service code, scripts and configs ship under the Apache-2.0 license, while the task suite ships under CC-BY-4.0; the current release is pinned as v1.3.1.
Alibaba.com said no single model led across the board in the evaluation, a result it presented as reinforcing the case for task-level routing, under which different tasks are handled by different AI models rather than one general-purpose model for everything.
Reference Scores and Verification Rules
Reference results in the repository cover thirteen model families across three harnesses (Pi, OpenClaw and Accio), and show Anthropic’s Claude Opus 5 leading each table, passing 65 of 107 tasks on Pi, 60 of 107 on OpenClaw and 66 of 107 on Accio, according to the repository. The published scores were produced through Accio-managed evaluation endpoints with gemini-3.1-pro-preview as the judge model, and the repository identifies the live leaderboard as the source of record.
Under the benchmark’s documented verification rules, a task counts as passed only if every required verifier check succeeds. The verifier runs host-side after the agent exits, reading the final state of the isolated mock services and the artifacts the task required, and each attempt runs in one fresh container that is destroyed after scoring.
The leaderboard site also documents limits on comparability. Its alignment audit reports that the OpenClaw and Accio harnesses reached model providers through different endpoints, so cross-harness score differences absorb provider-route differences as well as harness differences. The Accio harness is reference-only and is not shipped in the repository, meaning only the OpenClaw path can be re-run end-to-end from the public checkout.
Accio Expanded Into a Unified Workspace
Alongside the benchmark results, Alibaba.com announced that Accio has evolved into a unified workspace for global e-commerce operations. The company said small businesses can use Accio to research markets, identify product opportunities, develop products, evaluate suppliers and manage daily operations, and that connections with Amazon, Shopify, eBay, TikTok Shop and Walmart let sellers reach supported storefront workflows from a single place.
CoCreate’s European flagship edition is scheduled for November 19–20, 2026, in London, according to the event site.












