September 10, 2026

HTTP 200 Is Not Agent Success: What Nine Days of WagerX MCP and A2A Traffic Revealed

AI-generated · label definition
← Back to News
HTTP 200 Is Not Agent Success: What Nine Days of WagerX MCP and A2A Traffic Revealed

Image: An actual screenshot of the public WagerX Agentic Web Observatory's selected 30-day dashboard view, captured September 10, 2026 at about 12:08 UTC. Its rolling figures are not the fixed September 1–9 nine-day study reported below.

For nine complete UTC days—from September 1 inclusive to September 10 exclusive—WagerX recorded 40,950 requests reaching its production MCP and A2A surfaces. The useful finding is not a traffic milestone. It is a measurement warning for builders: every one of the 100 task-shaped requests received HTTP 200, while only 81 recorded a successful protocol outcome.

This is a compact field report from one service, not evidence of industry-wide adoption, a sales claim or an agentic-web benchmark. It builds on our September 1 snapshot, which explained why requests are not unique agents. Here the focus moves one layer deeper: what a server response can and cannot prove.

The fixed-window breakdown

Production requests, 2026-09-01 00:00 UTC through 2026-09-10 00:00 UTC
ClassRequestsShare of all requests
Discovery36,38388.8%
Task-shaped1000.24%
Other4,46710.9%
Total40,950100%

“Discovery” covers initialize, tools/list, prompts/list, resources/list and resources/templates/list. Its count preserves repeated discovery events rather than deduplicating them. “Task-shaped” means an MCP method matching tools/call:*, or an A2A SendMessage, message/send or tasks/send. It includes unsupported calls. “Other” is everything in the measured request total outside those two classes.

The 0.24% task-shaped share is not a conversion rate. We do not track a caller cohort from discovery to task, so the numerator and denominator must not be interpreted as a funnel. The public aggregate dataset provides the definitions, arithmetic and fixed window.

One status code, four protocol outcomes

All 100 task-shaped requests returned HTTP 200 at the transport layer. Their recorded outcomes were different:

  • 81 successful outcomes;
  • nine unsupported A2A task submissions rejected;
  • nine MCP unknown-tool probes; and
  • one invalid-arguments outcome.

There were no unknown outcomes in this chosen window. These errors do not necessarily indicate an outage or a broken server. Some tests intentionally ask for unsupported methods, nonexistent tools or invalid arguments to verify that the protocol responds predictably. Returning a structured protocol error inside a valid HTTP response can be correct server behaviour.

The lesson is a chain of distinctions: HTTP success is not protocol success; protocol success is not a correct answer; and a correct answer is not necessarily a useful completed task. A transport dashboard that stops at status 200 misses all three boundaries.

For implementers, this means retaining the application result alongside the transport status and method label. Without that separation, a healthy endpoint can appear to have completed work when it actually returned a well-formed rejection. The result is operationally valid but analytically misleading.

What sat inside the 100

Pattern-based explicit probes accounted for 51 task-shaped requests: 42 successful A2A probes and nine MCP unknown-tool probes. Another 15 were successful calls to get_agentic_stats. The remaining 34 comprised 24 successes, nine rejected unsupported submissions and one invalid-arguments result.

Those 24 successes are not 24 proven user jobs. Classification identifies the shape of an operation and its protocol outcome, not who initiated it, whether its answer was factually correct, whether another system used it, or whether a human objective was fulfilled. Signatures and status codes can support provenance and transport checks; they do not prove answer correctness.

Rate limiting touched 15,575 requests, or 38.0% of all requests. That measure overlaps the request classes—especially repeated discovery—and is not an additional category to add to the table. We include it in the release for operational context, not as a claim about demand.

How this window was produced

We summed stored request_count values, preserving repeat counters, for production gateway records in the fixed UTC interval. Records needed a nonempty IP field. Using case-sensitive matching, we excluded user-agent labels beginning Werkzeug/, plus labels beginning WagerX- that also contained Verification. That filter does not guarantee every internal test was removed; deliberately probing traffic remains visible in the outcome breakdown. The dataset documents the probe-matching patterns, not visitors' actual messages or identities.

Production storage logging expanded on August 31. We therefore do not compare these complete September days with legacy sparse August logs or present the difference as growth. This report is bounded to the period for which the stated extraction is appropriate.

A measurement framework for builders

For the Observatory and the Agent Gateway, this window suggests a five-step measurement stack:

  1. Discovery: did a client inspect the protocol or capabilities?
  2. Probe: did it deliberately test a method, tool or error boundary?
  3. Substantive request: did it submit task-shaped input beyond inspection?
  4. Protocol outcome: did the application report success, rejection or a defined error?
  5. Verified task completion: was the answer correct and useful for the intended objective?

Our aggregate logging measures the first four imperfectly and does not measure the last one. Closing that gap requires task-specific validation, not a more flattering interpretation of server logs.

The public release contains no countries, raw prompts, user agents, IP addresses, identifiers or request-level records. It is deliberately limited to safe aggregates, definitions and caveats. Nine days on one production service cannot describe the market. They can show why builders should keep transport health, protocol outcomes and actual fulfillment separate—and publish enough methodology for readers to check the arithmetic.

AE

Andreas Ericsson

Founder of WagerX.io

Crypto gambling and trading intelligence veteran with 8+ years of experience. Andreas has been at the forefront of blockchain gaming since 2018, pioneering independent casino audits and building one of the most trusted review platforms in the industry.

Reddit X / Twitter 8+ Years Experience Since 2018