# The Data — Every First-Party Number We've Published, With Provenance

The first-party dataset behind this site: 76 tools declared by one Claude Code session, 285,814 metered tool calls, a $148.16 working day, a 13-model agent benchmark, 7,522 skills audited. Method, date, and source for each.

# The data

Most claims about AI agents are estimates. The numbers on this page aren't — each one comes from something we ran, captured, metered, or scanned ourselves. This is the canonical index: the number, what it means, how it was produced, when, and the write-up with the full detail. If you cite one (please do), cite the method with it.

Two ground rules we hold ourselves to: cost and traffic figures are from **our own workspaces — dogfood, not customer data** — and every number links to the post where the methodology is spelled out, including its limitations.

## Tool surfaces (captured from live traffic) {#tool-surfaces}

| Number | What it is | Method & date | Source |
|---|---|---|---|
| **76 tools** | Declared by one real Claude Code session (v2.1, Chrome extension + connectors available): 35 core harness, 21 browser control, 20 connectors | Captured from live API request bodies — the `tools` array the harness sends on every call. 2026-07 | [Tool Surface Index](/tool-surfaces) · [the argued posture](/blog/which-claude-code-tools-to-deny-out-of-the-box) |
| **64 of 76** | Tools declared but never invoked in that session — standing capability, not used capability | Same capture, declaration vs invocation log. 2026-07 | [ACP for coding agents](/for-coding-agents) |
| **17 tools** | Declared by OpenAI Codex CLI (v0.142) via the Responses API | Same capture method. 2026-07 | [Tool Surface Index](/tool-surfaces) |
| **+21 tools, mid-session** | One session's surface grew by 21 tools partway through the day — deferred tools loaded, no prompt, no changelog | Declaration diffing across requests in captured traffic. 2026-07 | [Which tools to deny out of the box](/blog/which-claude-code-tools-to-deny-out-of-the-box) |

*Cite this section:* `https://agenticcontrolplane.com/data#tool-surfaces`

## Metered cost (priced per call, at API rates) {#metered-cost}

| Number | What it is | Method & date | Source |
|---|---|---|---|
| **984,335 calls** | Every tool and model call recorded by the gateway across all 162 workspaces, ours and external, since launch | Firestore count over every workspace's log collection, re-run 2026-09-02 (`mine-acp.mjs`); this is the homepage counter | [Homepage](/) |
| **285,814 tool calls** | The July research corpus: governed tool calls metered across 96 of our own workspaces, the dataset the cost posts below are built on | Every call through the ACP gateway logged with model, tokens, and estimated cost | [What 285,000 agent tool calls actually cost](/blog/what-280k-agent-tool-calls-look-like) (2026-04, updated 2026-07) |
| **97%** | Share of total spend that is the orchestration **loop** (the model re-reading context to pick the next step), not the leaf work | Every call tagged `callKind: loop` or `leaf`; spend split by tag. 2026-07 snapshot | [The loop tax](/blog/the-loop-tax) |
| **77% of spend, 30% of calls** | One frontier model's share of the bill vs its share of call volume — roughly **467×** the per-call cost of the cheapest model in the workload | Per-call cost attribution across the same 285,814 calls | [The teardown](/blog/what-280k-agent-tool-calls-look-like) |
| **85% reads** | Share of sampled tool calls that are read operations (`read_file`, `grep`, `cd`, …) | Tool-name classification over the ~10,200-call sample (2026-06) | [The teardown](/blog/what-280k-agent-tool-calls-look-like) |
| **2.7 seconds** | Average duration of an orchestration step (`chat.completion`) — the loop is the expensive part | Metered latency on governed calls. 2026-07 | [The loop tax](/blog/the-loop-tax) |
| **$148.16** | One full working day of Claude Code on a Max subscription, priced at API rates: 276 model calls, 1,697 tool calls | Model traffic routed through the ACP cost proxy; each call priced at list rates while the subscription passes through untouched. 2026-06 | [ACP for coding agents](/for-coding-agents) · [Claude Code cost tracking](/blog/claude-code-cost-tracking) |
| **100% loop tax, 72M tokens** | That $148 session's spend was entirely loop — 72M tokens of context re-read at a 100% cache hit rate (the only reason it wasn't ~10× more); the last 28 turns bought 10% of the output | Turn-by-turn session X-ray on the same proxy data | [Claude Code cost tracking](/blog/claude-code-cost-tracking) |
| **90% on one step** | Share of our agent-builder's model bill spent on a single step — deliberately, because per-call attribution let us route each step to the model it needs | Per-step cost attribution on our own production agent. 2026-06 | [One step is 90% of our agent's model bill](/blog/one-step-90-percent-of-our-agent-bill) |

*Cite this section:* `https://agenticcontrolplane.com/data#metered-cost`

## Model benchmark (agents, not leaderboards) {#model-benchmark}

| Number | What it is | Method & date | Source |
|---|---|---|---|
| **13 models, two ways** | Flagships from Anthropic, OpenAI, Google + open models (Llama, DeepSeek, Qwen, GLM), tested as isolated tool calls and as full agent loops | Deterministic grading where possible, a 3-judge model panel for prose; agent runs scored on completion, 3 runs per scenario with spread reported; cost = live pricing × actual tokens. 2026-06 | [We benchmarked 13 models on real agent runs](/blog/we-benchmarked-14-models-on-real-agent-runs) |
| **0.83–0.95 vs 0.06–0.77** | The whole field ties on isolated tool calls; the same models spread by more than 10× on completing a real agent loop | Same benchmark, both test modes | [The benchmark](/blog/we-benchmarked-14-models-on-real-agent-runs) |
| **0.83 → 0.06** | DeepSeek V3.2's isolated score vs its agent-completion score — perfect calls in a vacuum, cannot drive a loop to a finish | Same benchmark | [The benchmark](/blog/we-benchmarked-14-models-on-real-agent-runs) |

*Cite this section:* `https://agenticcontrolplane.com/data#model-benchmark`

## Ecosystem scans (security research) {#ecosystem-scans}

| Number | What it is | Method & date | Source |
|---|---|---|---|
| **7,522 skills scanned** | Every skill on the ClawHub registry, statically analyzed: 4,931 findings across 746 skills, ~61% estimated false-positive rate after triage | 40 regex patterns from published research (Snyk, Cisco, Kaspersky), run airgapped in Docker with `--network=none`; static analysis — a floor, not a ceiling | [I audited 7,522 AI agent skills](/blog/i-audited-7522-ai-agent-skills) (2026-03) |
| **8,216 MCP servers** | Public MCP servers scanned for input-validation posture (7,840 tools) and classified by auth appropriateness | Registry-wide static scans; methodology and caveats in each post | [Input validation](/blog/mcp-input-validation-attack-surface) · [auth appropriateness](/blog/mcp-auth-appropriateness-audit) (2026-03) |

*Cite this section:* `https://agenticcontrolplane.com/data#ecosystem-scans`

## Using these numbers

Cite freely with attribution — "per Agentic Control Plane's metered data" plus a link is ideal. The canonical URL for this page is `https://agenticcontrolplane.com/data`, and each section above has a stable anchor (`#tool-surfaces`, `#metered-cost`, `#model-benchmark`, `#ecosystem-scans`) if you're citing one group of numbers. Prefer linking the source post where you can, because each post carries the caveats that keep the number honest (sample sizes, dogfood-not-customer scope, static-analysis limits). If a number here disagrees with a post, the post is canonical and this page needs an update — [tell us](/community).

The captures and meters that produce this data run continuously. To point them at your own agents:

```bash
curl -sf https://agenticcontrolplane.com/install.sh | bash
```

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Dataset",
  "name": "Agentic Control Plane first-party agent data",
  "url": "https://agenticcontrolplane.com/data",
  "description": "First-party measurements of AI agent behavior and cost: 76 tools declared by one Claude Code session (17 by Codex CLI), 285,814 metered tool calls across 96 workspaces with 97% of spend in the orchestration loop, a $148.16 Claude Code working day priced at API rates, a 13-model agent benchmark, and static scans of 7,522 agent skills and 8,216 MCP servers.",
  "creator": { "@type": "Organization", "name": "Agentic Control Plane", "url": "https://agenticcontrolplane.com" },
  "temporalCoverage": "2026-03/2026-07",
  "isAccessibleForFree": true,
  "license": "https://creativecommons.org/licenses/by/4.0/"
}
</script>
